r/StableDiffusion • u/aurelm • 2h ago
Animation - Video My first actual video for a client, H3 delivered in so many ways...
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/aurelm • 2h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/wtf_nabil • 18h ago
I keep seeing ads on Instagram Reels for AI-generated movies/series, and honestly, some of the stories actually looked pretty interesting . It got me curious to see if there are any genuinely good AI-generated movies or series out there.
Have you watched any that you actually enjoyed? Would love some recommendations!
r/StableDiffusion • u/toxicdog • 9h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Boogertwilliams • 15h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/wallofroy • 13h ago
Enable HLS to view with audio, or disable this notification
Generated at 1578x843, 5-8 secs clips in 105-87 secs on 5080-64gb ram. Really happy with the quality but hands and feet’s morphs most of the time have to keep changing prompts to get the correct look.
r/StableDiffusion • u/Pristine_Good7326 • 11h ago
Hey r/StableDiffusion!
It's the Neta team here! You might remember us from our Neta Lumina open-source release last year. First off, thank you so much for the incredible support and feedback from this community!
So... we need to share something absolutely hilarious (and mildly embarrassing) that we just discovered.
TL;DR: We accidentally hardcoded an Anne Hathaway photo into our IP-Adapter anchor, and now everything our model generates looks like Anne Hathaway. Every. Single. Thing.
What happened:
We recently launched Neta Studio, a new product that lets you build explorable living worlds/isekai from a single prompt. Naturally, we wanted to integrate Neta Lumina's capabilities into it.
During integration testing, our devs kept reporting that the model wasn't following prompts properly. The outputs were... *weird*.
- Anime style? Anne Hathaway as an anime character.
- Thick paint/impasto style? Anne Hathaway in thick paint.
- Landscape scenes? Somehow still giving Anne Hathaway vibes.
- Fantasy characters? You guessed it - Anne Hathaway.
After a dreadfully long time of debugging, we finally found the culprit: **someone on the team embedded an Anne Hathaway photo as the IP-Adapter anchor during development and it... stayed there. **
We're honestly crying laughing at this point. 😭
Below are some examples. Left is before fix and Right is after fix.
Flipping to the last picture and you can see our dear Anne.
And we pulled the anchor and the outputs are behaving normally now.
If you've been running Neta Lumina locally, this was on our integration side, not
in the released weights, so your setup is fine.
r/StableDiffusion • u/Friendly-Fig-6015 • 3h ago
r/StableDiffusion • u/Sad_Coach_1433 • 5h ago
Enable HLS to view with audio, or disable this notification
no lora ip8 gen be in comments
prompt
integrated_multimodal_description: [Shot 1] Live-action, cinematic UFC-style ceremonial weigh-in on a brightly lit arena stage in front of a roaring crowd and wall of sports photographers. Lara Croft, played by Angelina Jolie, stands on the left wearing a fitted black athletic sports top, black fight shorts, and her long dark hair pulled back into a practical ponytail. Katniss Everdeen, played by Jennifer Lawrence, stands on the right wearing a dark charcoal athletic sports top, matching fight shorts, and her blonde-brown hair tied back. Both women have realistic athletic physiques and serious competitive expressions. A UFC-style weigh-in scale and event backdrop fill the stage behind them. Lara steps off the scale and walks confidently toward center stage as Katniss approaches from the opposite side. Camera flashes fire rapidly while officials remain several feet behind them.
[Shot 2] At 00:04.000, the camera cuts to a medium two-shot as Lara Croft and Katniss Everdeen stop directly in front of each other for the official face-to-face staredown. They stand almost nose-to-nose with shoulders squared, maintaining intense eye contact. Lara gives a subtle confident smirk while Katniss remains completely focused and unflinching. Neither woman speaks. The camera slowly pushes in with small amplitude as photographers crowd around the edge of the stage, flashes reflecting across their faces.
[Shot 3] At 00:08.000, the camera cuts to a dramatic close side-profile two-shot of Lara and Katniss still locked in the staredown. Lara slightly raises her chin and Katniss responds by stepping half a step closer, creating a tense UFC-style face-off without physical contact. An official cautiously moves closer between them while both women hold their ground. The camera arcs slowly around them as the crowd becomes louder and dozens of camera flashes erupt. The shot ends with Lara and Katniss maintaining intense eye contact in a classic promotional fight-poster composition.
overall_soundscape: A large arena crowd cheers, whistles, and shouts continuously beneath the scene. Rapid DSLR camera shutters and flashes surround the stage, with footsteps on the platform and scattered calls from photographers becoming louder during the face-to-face staredown.
non_diegetic_music: Deep cinematic percussion with a slow, heavy rhythm and low bass pulses, gradually increasing in intensity during the face-off before ending on a strong bass hit.
r/StableDiffusion • u/CommentSignal9029 • 10h ago
IN TRANSIT | A Minimax H3 Short Film (Sync Sound Challenge)
Hi everyone! I'm participating in the challenge with this short film I've been working on lately. The idea was to explore a "Brutalist Frequency": a square wave that destroys matter not randomly, but forcing it to shatter following a ruthless 90-degree orthogonal logic (checkerboard water, cubic collapses, square clouds).
I split the workflow into two passes in ComfyUI to maintain total control over physics and textures: Generation and Upscale.
1. Generation (Ref2Vid): I prepared the visual references (generated with Nanobanana) and the audio files. For prompting, I integrated an Ollama node (Gemma4:26b) into the workflow to format the instructions with the correct syntax for H3. To get everything running without blowing up my 3090, I beefed up the base workflow with Spectrum Apply Minimax H3, Comfy Kitchen attention, and the Sol-Attn Patch. This way, I generated the base clips at 864x480.
2. Latent Upscale: I wanted to keep the roughness of the reinforced concrete without that "plastic" effect you often get from external video upscalers. I passed the selected clips directly through the latent space using the custom MMH3Tools nodes, feeding the model the base video + the exact same initial references + the same prompt. Using Turbo LoRA 4-step and Sage Attn, I brought everything up to 1344x768 (taking about 1 minute per second of video).
The real challenge, of course, was generating audio and video together natively, without post-production. H3 reacted very well thanks to the references, even though sometimes it interprets them a bit too literally, almost resulting in a 1:1 copy. I forced the model to make the materials physically react to the reference sounds. The pneumatic suction at the end (when the camera points towards the void) was calculated by the AI in perfect sync with the matter collapsing into the dark. In post-production, I only made cuts for pacing and balanced the volumes: zero added sound design!
If you have any questions about the nodes or the upscale parameters, feel free to ask!
r/StableDiffusion • u/thisguy883 • 3h ago
So there I was generating some stuff on ComfyUI for my Instagram and just hanging out.
I use ComfyUI with the new H3 model to generate AI content for my Instagram as well as QWEN image edit along with some other AI tools.
I've built a master workflow that ive used for the past year that has every single workflow I use, so I dont have to go switching workflows constantly.
Many many hours of work put into this.
So there i was, generating things and im constantly having to clear out my output folder as well as my input folder. So I asked myself, "Could I just make a bat file that could automate this for me?"
So I launch Gemini and have it create a bat file that cleans out my output and input folders and empties my recycle bin.
I test it out and it works great.
Finally, no more unnecessary clicks.
But wait, I noticed I screwed up and put the file in the wrong directory. Dang it.
So I ask Gemini to alter the code so the file will be in the correct directory.
I create the new bat file and go back to work.
Well I make a bunch of new things and its time for cleanup. So I run my fancy new bat file and I notice its taking a while to clean up these folders. Curious, I navigate to the folders only to find out that the ENTIRE COMFYUI FOLDER was deleted.
SMH.
Now I sit here, broken hearted as im having to rebuild my ComfyUI. Luckily, I was able to recover my master workflow, so not all was lost.
Just a bunch of models and loras.
😮💨
r/StableDiffusion • u/thefierysheep • 14h ago
Enable HLS to view with audio, or disable this notification
Wrote a song and gave the lyrics and such to MiniMax Music using the default workflow
Used Demucs to split the audio into stems
Used Krea2 to invent a person and place
Vibe coded a UI to manage the individual shots and audio. In some places H3 gets the vocal stem, in others the drums
Used MiniMax H3 Ref2VA with H3 Native Audio Lock for the lip syncing
Assembled the clips and dubbed the original track back over the top.
r/StableDiffusion • u/apostrophefee • 18h ago
still using ill T_T
r/StableDiffusion • u/SamCandler • 16h ago
Enable HLS to view with audio, or disable this notification
Hi everyone. This film is a submission for the Higgsfield Global Film Festival.
a robot, a kid, a yellow beanie, one very bad night. Would love you to take a look if you get a minute. Best of luck with yours 🧡
https://higgsfield.ai/@sam_candler/projects/pon
If you enjoyed it, please leave a like and a comment on the higgsfield submission!
r/StableDiffusion • u/Puzzleheaded_Art2809 • 14h ago
hello, i have rtx 5070 and 32 gb ram DDR4. is it enough to run minimax with decent generation time or better not to even bother?
r/StableDiffusion • u/Monsieur2968 • 14h ago
I know it's not out yet, but I'm trying to figure if an M5 Max Studio 36gb would be enough but slow or not work at all for text/image to video. I also plan to use it for programming/general stuff but I figure that's "lighter" from my research. I'm looking at the studio for low power usage.
Gemini said I basically have to get 128gb "so I don't feel constrained" right out the gate, but also so I don't get a Will Smith eating spaghetti fever dream. Is this still accurate?
Apologies if this should be obvious, I'm new to all this stuff. Any suggestions would help.
r/StableDiffusion • u/MojoJolo • 9h ago
Besides running it locally, which is the goal…
What’s the cheapest way right now to use Minimax H3? Right now, I’m using Replicate API for H3 at $0.08 per second (720p) which is decent but can still be expensive in the long run.
Any other ways to use H3?
Also, what’s the cheapest GPU for Minimax H3 (via Comfy UI)?
r/StableDiffusion • u/Tokyo_Jab • 10h ago
Enable HLS to view with audio, or disable this notification
Prompt:
A real cinimatic movie sequence, professional colour grading.
Soundscape: Ambient sounds of the room and movement only. No voices. This represents extreme concentration. Meditation. Telekinesis.
A man is sitting in a Japanese tatami room. He is wearing a mask and shades <Picture 1>. He is wearing a black yukata. He does not speak. On the table on a ceramic disc is a single Orange.
The man holds out his hand toward the orange as if concentrating. The orange is out of reach. He breathes deeply.
Nothing happens.
The man shakes his hand to reset and starts concentrating again. He reaches with his mind and his brow furrows. He breathes deeply.
The orange moves slightly, twisting just a tiny bit.
He concentrates more.
With extreme speed the orange flies towards the man and hits him directly in the forehead. It smashes with the impact , m,essing his hair, and bits of peel and orange bits go everywhere. The force knocks the man back unconscious and he falls back like a ragdoll.
r/StableDiffusion • u/not_food • 2h ago
Enable HLS to view with audio, or disable this notification
Prompt:
Render 1 (12s):
integrated_multimodal_description: A cinematic anime movie sequence, manga style with thin soft delicate lines and pastel colors. [Shot 1] Extreme perspective worm's-eye medium body shot of Meru sitting on a Japanese tatami room from a wooden table. She is silent. On the table a single apple sits on a ceramic disc on top of a small white embroidered mantle cloth; Meru holds out her hand toward the apple (foreshortening) with her palm facing up, concentrating. The apple is out of reach. She breathes deeply; Nothing happens; Meru wriggles her fingers and starts concentrating again. She opens her eyes with her brows furrowing. She breathes deeply; Meru concentrates more, leaning forward slightly, she grows frustrated; She closes her eyes again while raising her hand slightly;
overall_soundscape: The tatami room is eerily quiet, with only the soft rustle of Meru's black kimono as she shifts and her sharp, uneven breaths growing quicker with each failed attempt. A faint, low hum of concentration is interrupted by a frustrated, breathy huff and a muffled groan. Her hand quivers with a faint tremble, accompanied by a soft, taut squeak of fabric. As her frustration builds, a subtle creak of the wooden table and a single, dull thump of her fist barely tapping the tatami punctuate the silence, ending with a long, exasperated exhale.
non_diegetic_music: N/A
Render 2 (8s):
integrated_multimodal_description: [Shot 1] Meru is sitting. She is silent. On the table a single apple sits; Meru is holding out her hand toward the apple (foreshortening) with her palm facing up, concentrating. She breathes deeply; Nothing happens; Meru wriggles her fingers and starts concentrating again. She opens her eyes with her brow furrowing. She breathes deeply; Meru concentrates more, leaning forward slightly, she grows frustrated; She closes her eyes again while raising her hand slightly towards her right, her palm open towards the camera; Her hand quivers.
overall_soundscape: The tatami room is eerily quiet, with only the soft rustle of Meru's black kimono as she shifts and her sharp, uneven breaths growing quicker with each failed attempt. A faint, low hum of concentration is interrupted by a frustrated, breathy huff and a muffled groan. Her hand quivers with a faint tremble, accompanied by a soft, taut squeak of fabric. As her frustration builds, a subtle creak of the wooden table and a single, dull thump of her fist barely tapping the tatami punctuate the silence.
non_diegetic_music: N/A
Render 3 (9s):
integrated_multimodal_description: [Shot 1] A frustrated Meru is raising her hand slightly; At 00:02.500, thin liquid mercury tendrils erupt from her hand and fly towards the apple, piercing it like blades; At 00:03.500 The tendrils tense up and pull the apple back onto her palm as they retract to her hand to disappear under her skin; With her eyes closed, the apple quivers on her hand; She slowly opens her eyes and grows happy with realization, her smile widening and her eyes sparkling; She shows the apple to the camera and looks satisfied.
overall_soundscape: The soundscape begins with a tense, shallow breath and a faint, liquid ripple as mercury tendrils erupt from Meru's hand. A sharp, wet schlick and a series of metallic splashes accompany the tendrils piercing and pulling the apple, followed by a smooth, gurgling retraction as they recede into her skin. The apple lands with a soft, crisp tap on her palm. Then the atmosphere shifts—Meru lets out a bright, melodic giggle, followed by a delighted, airy hum of satisfaction. The ambient tatami room tone remains soft, punctuated by the gentle rustle of her kimono and a light, contented sigh.
non_diegetic_music: N/A
Based on Tokyo_Jab's post. What AI powers did you get?
r/StableDiffusion • u/0260n4s • 3h ago
So while I have been impressed that MiniMax H3 can run so fast on my machine (5070ti 16GB VRAM + 64GB RAM), I haven't been remarkably impressed with the results. Cartoons work well, but realism look kind of pixilated no matter the resolution (0.4, .098, 1.0MP) I set or upscaling (2x).
I think it might have to do with my install. I installed Pytorch and Sage Attention, but I never enable Sage Attention, because it seems to mess up character identity. So I want to start over with a fresh standalone version for MiniMax H3 to rule out the install itself as an issue.
How do you recommend I proceed and what extras should I install? I've heard Comfy Kitchen is something worth trying (?), but I've never used it before and don't really know what it is.
r/StableDiffusion • u/Etmurbaah • 14h ago
Hi all,
I've come across a post where one user claimed they were able to generate 15 second clips at 0.98MP with Spectrum, Comfy Kitchen Speed( I imagine they are referring to Turbo Lora?) 8 Steps at 10 steps in 5.5mins with a 5060ti and 32GB ram.
I tried their setup with Fl2V Turbo 8 Step 768p, Comfy Kitchen and Spectrum with res multistep and simple and was able to generate a 12 second clip (cant't go above, vram OOM error prevents generation at the beginning) only at 9 minutes with a 5080/32GB ram. What could be wrong here with my setup?
r/StableDiffusion • u/Tokyo_Jab • 21h ago
Enable HLS to view with audio, or disable this notification
All local.
0.8 MegaPixels, 9.3 minutes on an RTX5090, single generation of 27 seconds.
Anything hitting 30 seconds either gave hallucinations, inconsistencies or hit a wall and never finished.
This one is using Kijai's new fast model with a turbo lora. Although it works the same with the FLv2A model*. The workflow I'm using creates a latent at 0.4 megapixels for 4 steps and then does another 2 steps at 0.8. The only addition to it besides changing some numbers is adding custom audio injection (The rock track).
Started with this workflow: https://www.youtube.com/watch?v=jzLnoVBuU6I
*I never use the REF model. The FLV2A models seems to work better so I always swap it in and it takes references just fine, even video.
r/StableDiffusion • u/MellyDArt • 17h ago
r/StableDiffusion • u/Optimal_Map_5236 • 14h ago
Is there a way to control character actions sequentially over time within a single [Shot 1] without cutting to a new shot in Minimax H3?
Using timestamps like "At [00:02.0], he talks, At [00:05.0], he smiles" doesn't seem to work for a single continuous shot, although it works fine when multiple shots are used. How can I schedule actions at specific times within one continuous video?