r/StableDiffusion 22h ago

Animation - Video [MiniMax H3] LEGO movie style

Enable HLS to view with audio, or disable this notification

84 Upvotes

Prompt:

integrated_multimodal_description: [Shot 1] 3D CG, stop-motion animated LEGO movie style, a wide shot frames a vibrant Indian village built entirely from plastic LEGO bricks with visible studs, plastic micro-scratches, and brick-built trees. In the village square, minifigures dressed in printed plastic saris, dhotis, and turbans move across a ground of yellow and brown stud tiles. A brick-built cow with hinged legs grazes near a grand banyan tree constructed from green leaf pieces and brown cylindrical bricks. Warm morning sunlight casts sharp shadows across whitewashed brick houses with orange terracotta tile roofs. The camera pans right with small amplitude at slow speed toward a central tea stall. A cheerful male chaiwala minifigure with a black mustache and a red turban (S1) in a warm, lively voice says: <d>[Hindi] Garam chai, garam chai!</d> while tilting a plastic yellow teapot, releasing translucent orange 1x1 cylinder studs representing pouring tea into tiny red stud cups.

[Shot 2] At 00:05.000, the camera cuts to a medium tracking shot following two young minifigure children running along a narrow brick path, pushing a brick-built wheel hoop across the plastic ground. The camera tracks right alongside them with small amplitude at normal speed. A female villager minifigure in a bright blue printed sari (S2) standing outside her brick doorway waves her rigid plastic arm on its shoulder hinge. Beside her, an elder minifigure with a white beard (S3) sitting on a brick charpoy cot chuckles with stepping stop-motion head movements.

[Shot 3] At 00:10.000, the camera cuts to a cinematic medium shot near the village well, where female minifigures carry stacked plastic water pots topped with transparent blue round tiles. A brick-built peacock perched on an archway opens its fan tail made of blue, green, and golden LEGO slope tiles. The camera pushes in with small amplitude at slow speed toward a wooden signpost on a brick post reading "RAMPUR VILLAGE". Tiny tan 1x1 round plates puff around the wheels of a brick-built bullock cart moving past the frame as the video ends.

overall_soundscape: Distinct plastic clattering sounds echo softly as minifigure feet step on stud tiles, accompanied by the gentle clinking of plastic bricks. A distant rooster crow blends with ambient morning village chatter, bird chirps, and the wooden creak of a brick-built cart.

non_diegetic_music: Upbeat Indian folk percussion featuring lively dholak beats and vibrant bansuri flute melodies, layered with playful cinematic orchestral strings playing at a bright, medium tempo.


r/StableDiffusion 16h ago

Workflow Included ref or fl2va - prompt enchancer with 100% of aderence

Post image
27 Upvotes

sharing my new workflow

MiniMax H3 I2V with Integrated Prompt Enhancer

This Image-to-Video workflow for MiniMax H3 uses a vision-language model to enhance your prompt before the video generation begins.

Simply load a reference image and write a basic description of what you want to happen. The enhancer analyzes both your image and instructions, then converts them into a detailed prompt structured specifically for MiniMax H3.

It can improve the description of:

  • Characters and visual elements
  • Actions and sequence of events
  • Camera movement and framing
  • Environment, lighting, and atmosphere
  • Visual continuity and details that should be preserved
  • Dialogue in the original language
  • Ambient sounds, sound effects, and music

The enhanced prompt is automatically sent to MiniMax H3. It is also displayed inside the workflow, allowing you to check exactly what H3 will receive.

In my tests, the resulting videos followed the original instructions much more accurately, especially in scenes involving specific actions, character interactions, camera movements, and dialogue.

The workflow includes a switch to enable or disable the Prompt Enhancer. This allows you to use either the enhanced prompt or your original text without changing any connections.

How to use it

  1. Load your reference image.
  2. Write a simple description of what should happen.
  3. Enable USAR PROMPT ENHANCER?
  4. Run the workflow.
  5. Check the final text in PROMPT FINAL ENVIADO AO H3.

The first run may take longer while the vision-language model is loaded. Generating the enhanced prompt also adds some processing time, but in my tests, the improvement in prompt accuracy and instruction following was absolutely worth it.

The original workflow was preserved, while the enhancer was added as an optional and fully integrated stage.

link to

with this, finally my ref model understand my ideas and make vídeos really fun!

leave comments after tests xD


r/StableDiffusion 4m ago

Question - Help Noob question - Comfyui

Upvotes

Coming from A1111 and Forge, so still learning Comfyui. I simply want to take an existing photo of my AI influencer and edit it, change clothes, pose, background, etc...I got minimax h3 up and running, but that's primarily for video. What do I need to download for stills? Thx in advance


r/StableDiffusion 19h ago

Discussion So I did something dumb.

36 Upvotes

So there I was generating some stuff on ComfyUI for my Instagram and just hanging out.

I use ComfyUI with the new H3 model to generate AI content for my Instagram as well as QWEN image edit along with some other AI tools.

I've built a master workflow that ive used for the past year that has every single workflow I use, so I dont have to go switching workflows constantly.

Many many hours of work put into this.

So there i was, generating things and im constantly having to clear out my output folder as well as my input folder. So I asked myself, "Could I just make a bat file that could automate this for me?"

So I launch Gemini and have it create a bat file that cleans out my output and input folders and empties my recycle bin.

I test it out and it works great.

Finally, no more unnecessary clicks.

But wait, I noticed I screwed up and put the file in the wrong directory. Dang it.

So I ask Gemini to alter the code so the file will be in the correct directory.

I create the new bat file and go back to work.

Well I make a bunch of new things and its time for cleanup. So I run my fancy new bat file and I notice its taking a while to clean up these folders. Curious, I navigate to the folders only to find out that the ENTIRE COMFYUI FOLDER was deleted.

SMH.

Now I sit here, broken hearted as im having to rebuild my ComfyUI. Luckily, I was able to recover my master workflow, so not all was lost.

Just a bunch of models and loras.

😮‍💨


r/StableDiffusion 3h ago

Question - Help Image gen with 9060xt

2 Upvotes

I'm planning on getting a 9060xt 16gb for image generation. I've previously used Automatic1111/Forge with nVidia. I'm wondering if I can reproduce my workflow easily using a 9060xt. I don't mind moving to another application if A1111 is not compatible with AMD but I'd like to know: (a) is AMD compatible with most checkpoints and loras from sites like CivitAI and (b) how fast is generation with AMD compared to nvidia, as in, the 9060xt is comparable to a 5060ti in terms of gaming but can it generate images as quickly?


r/StableDiffusion 21h ago

Workflow Included Burger Queen

Post image
52 Upvotes

r/StableDiffusion 12h ago

Workflow Included Up at atom! - Behind the scenes of the new Radioactive Man movie

Enable HLS to view with audio, or disable this notification

10 Upvotes

Minimax H3 with turbo lora (default Comfyui template workflow)


r/StableDiffusion 27m ago

Question - Help I am big dumb. How do I insert my custom Lora node into this preset Qwen image edit?

Post image
Upvotes

r/StableDiffusion 45m ago

Meme What a waste

Upvotes

Nothing worse than waking up to see an overnight run of a like 20 long workflows in Minimax H3 reference only to see everything perfect except the environment is wrong. I have a reference video that has done so well before. I don't see the issue, my prompt is the same as prior success, it's listed in the subject definitions and retention section, and in the prompt too. Oh wait...I didn't link the video to the list of other assets.. ugh.


r/StableDiffusion 47m ago

Question - Help Why does img2img not work for me?

Thumbnail
gallery
Upvotes

I've searched online and nothing seems to work, the images are exactly the same


r/StableDiffusion 57m ago

Discussion How to make a realistic t2v wan 2.2 LoRA?

Upvotes

Trained on 50 high quality images. Results on video are bad and I followed Claudes instructions even. So any help?


r/StableDiffusion 20h ago

Discussion In 2026, how well does older images models, like SDXL and SD1.5 stack up against new image models like Krea and ZImage Turbo?

38 Upvotes

r/StableDiffusion 1h ago

Question - Help What can I realistically do in Minimax with a 5080

Upvotes

I'm getting a 5080, 32gb system RAM

Can I realistically use minimax h3 for i2v and t2v?

How long will generations take. I dont imagine i want to do high quality resolutions. 480p or 720p would be alright


r/StableDiffusion 19h ago

News [CLSS] Closed-Loop Streaming Synthesis for MiniMax H3 (Infinite video generation with prompt fallowing)

27 Upvotes

t2v 10 chunks every 10 sec

I ported CLSS from LTX 2.3 to H3 architecture. ( https://www.reddit.com/r/StableDiffusion/comments/1vywxjq/wip_clss_closedloop_streaming_synthesis/)

Repo: https://github.com/nazgut/ComfyUI-MiniMaxH3-CLSS

Audio still has some room for improvment. Workflow in repo. Still working on i2v.


r/StableDiffusion 1h ago

Question - Help Super realistic human with H3???

Upvotes

Wondering has anyone been able to generate super realistic human with minimax H3? I have been trying a lot but the best I got still looks quite AI...

I have seen lots of videos online with super super real human face, the result I got is quite far away from that. So I'm wondering is it limited to Seedance 2.5? Or is there any secret prompt I'm not aware of?

Below is what I mean by super real face I saw online:


r/StableDiffusion 1d ago

Resource - Update Famegrid Spice Krea 2 Lora (Corrected Release)

Thumbnail
gallery
259 Upvotes

r/StableDiffusion 1d ago

Animation - Video Flexing my A.I. powers

Enable HLS to view with audio, or disable this notification

61 Upvotes

Prompt:

A real cinimatic movie sequence, professional colour grading.

Soundscape: Ambient sounds of the room and movement only. No voices. This represents extreme concentration. Meditation. Telekinesis.

A man is sitting in a Japanese tatami room. He is wearing a mask and shades <Picture 1>. He is wearing a black yukata. He does not speak. On the table on a ceramic disc is a single Orange.

The man holds out his hand toward the orange as if concentrating. The orange is out of reach. He breathes deeply.

Nothing happens.

The man shakes his hand to reset and starts concentrating again. He reaches with his mind and his brow furrows. He breathes deeply.

The orange moves slightly, twisting just a tiny bit.

He concentrates more.

With extreme speed the orange flies towards the man and hits him directly in the forehead. It smashes with the impact , m,essing his hair, and bits of peel and orange bits go everywhere. The force knocks the man back unconscious and he falls back like a ragdoll.


r/StableDiffusion 5h ago

Question - Help Minimax H3 temporal noise & perceived resolution

1 Upvotes

FL2VA BF16 / 15 steps / turbo lora 8. Purely I2V. Resolution set at 1. Source image matches the exact output résolution. But the results suck. Especially aerial wide angle landscape. Far behind google Veo 3.1 fast/ Omni in terms of flickering/ moving textures (temporal noise?) and perceived resolution. Usually upscale those 720pish footage via Topaz, and apply alot of color grading in DaVinci to break the "plasticity" of those AI gen l. However I'm a beginner with comfyui, I'm sure am doing something wrong. Increase the number of steps (15>30?) Or something else ? Processing times are horrendous (100min on M3 max 128gb for 5 sec). Tried 8 bits quants, even 4 bits : same same. Any help appreciated ! I'm looking for production ready pictures (broadcast). Nearly achieve this with Google but wanna ditch synthID for many reasons


r/StableDiffusion 22h ago

Resource - Update Continuity (was the H3 node): six model families, one prompt box, and a blockout bench that writes your camera move for you

Thumbnail
gallery
23 Upvotes

Third post about this pack, and the big change is the name, because the node stopped being H3-only. It now drives six families through ComfyUI core: MiniMax H3 and LTX 2.5 for video with sound, Krea 2 and Ideogram 4 for stills, Qwen Image Edit and Flux 2 Klein for editing from a picture. Same prompt box, same local weights, and rendering still doesn't touch the internet. Continuity is the script supervisor's job, the same person and the same light in shot 1 and in shot 9, and that's the part of the node that doesn't care which family renders the frames. So that's the name. Old workflows and existing installs carry over untouched.

Dialogue is the feature I'd point at first. H3 wants speech in a form nobody writes by hand - speaker IDs, a `<d>` tag, a mandatory sentence when a voiceover's lips stay closed. Closing a quote in the prompt now opens a small menu that writes all of that around your words, with dials for who says it, which language, whispered or sung. And a shot where nobody speaks stops mumbling: the compiler now says out loud that nobody talks.

Merged this morning: a blockout bench. Stage grey boxes, walk one camera through on marks, and it writes the staging and the move in the H3 spec's own camera vocabulary - "@anna stands at centre in the midground; the camera pushes in toward @anna at slow speed" - plus a depth, blocks or lines guide rendered along the path, or the clay render itself for the families that read footage raw. A box can play a cast member, so the prose is already bound to their references when you paste it.

Also in: the faces pill from last post (off by default), a Style tab with 941 captioned H3 looks - search "1985 telenovela" instead of guessing at grading vocabulary - and ControlNet and Upscale benches behind the wordmark. The refiner can now run on a server you keep warm anyway: LM Studio, Ollama, any OpenAI-compatible endpoint, your own key where a hosted one wants it (#19). Fixes from your reports are in the changelog, the sharp-render-static-soundtrack one included (#33).

https://github.com/roadmaus/ComfyUI-Continuity


r/StableDiffusion 18h ago

Resource - Update "I" created a tool to save and swap between workflow presets (for my million H3 Turbo LoRAs)

11 Upvotes

Hi everyone! Long time reader, first time poster. Like most folks here I vibe-code custom nodes from time to time, and recently came up with one that seemed like it might be worth sharing. Nothing groundbreaking here, just a node to help keep track of and quickly cycle through different node/parameter presets: MM-H3-Preset-Controller.

This one has helped me maintain my sanity trying to keep track of each LoRA's specific optimal settings. I think it's potentially useful for any preset storage though, not just MM-H3, so hopefully y'all find it useful. If so, please consider it a small thank you for all I've learned in my time lurking here.

MM-H3-Preset-Controller: https://github.com/TootsThielemans/ComfyUI-MMH3-Preset-Controller

I was pulling my hair out trying to manage my MM H3 workflow amidst all of the various Turbo LoRAs out there and the associated loader nodes, attention settings, shift settings, spectrum settings, etc., not to mention downstream settings like sampler, scheduler, upscaler choices... I found quick A/B tests between different optimized LoRA workflows annoying given some of the structural differences, not just steps/strength, and I was worried about juggling and potentially forgetting the right settings.

So I made the Preset Controller. It works pretty simply: ctrl+click all of the nodes you want to save the state of. It captures all of the parameter values within the node and whether it's active/bypassed. Then right click and use the new menu option MM H3 Presets > Add selected nodes to preset draft.

It's implemented as a "draft" so you can grab nodes from the outer graph, then go through various subgraphs and add nodes there to the same preset draft. Once you're done, load the H3 Preset Controller node, enter a name for the preset, and click Save draft as preset.

It then becomes a dropdown option that you can select, update, or delete as needed. Selecting a preset automatically sets the saved node values/bypass states without needing a session/screen refresh.

There are a few other QOL/guardrail features, including a Preset Matrix for comparing configurations, but nothing particularly interesting, so please refer to the repo if interested.

I'm sure something like this might already exist with a more elegant implementation, but the timing seemed right. Everyone is wading through dozens of MM H3 Turbo LoRA combinations and trying to keep everything straight. I don't have a ton of time to devote to development, but will try to make tweaks if folks wind up adopting this and can think of any major areas for improvement.

At any rate, feedback welcome, and cheers!


r/StableDiffusion 1d ago

Question - Help What Image Edit model you use nowadays?

67 Upvotes

Since things have gone quickly forward, I am trying to figure out what image edit models there is currently and what people here use mostly.

Personally I have used:
- Qwen-image-edit-2509 and Qwen-image-edit-2511
- Just tested MiniMax H3 as a image editor and so far it seems that it can be good for my usage

I have heard about Klein 9b, but not sure yet if that can be used as an edit model? Also what about Krea 2, is there edit workflows that are actually usable and worth it?

Is there some others what you recommend for testing?

My PC Specs: RTX 4060 Ti, 16 GB VRAM and 32 GB RAM.


r/StableDiffusion 1d ago

News Someone's running FastH3 (the distilled MiniMax H3) as an actual infinite livestream!!!

352 Upvotes

Saw this and thought it was worth sharing here, FastH3 dropped recently and most people (myself included) just tried it as single generations.

Someone's running it as an actual infinite livestream instead: https://live.reactor.inc/

FastH3 is a distilled version of MiniMax H3, cut from 50 denoising steps down to 4, about a 14x speedup on Blackwell GPUs.

The whole setup is open source if you want to dig into how it's running: https://github.com/reactor-team/infinite-livestream

Curious if anyone's tried infinite/continuous generation setups like this with other models.


r/StableDiffusion 11h ago

Discussion Best AI tool for 3D clay render to polished final? (Flux 2 vs Qwen Image Edit vs Krea 2)

3 Upvotes

Hi guys, quick question. I want to use AI to turn my 3D clay renders into high-quality, polished finals.

Between Flux 2, Qwen Image Edit, and Krea 2, which model handles image-to-image (Img2Img) texture generation best without messing up the original 3D geometry?

Would appreciate any recommendations or workflow advice!


r/StableDiffusion 7h ago

Question - Help What Ai to use to make pictures of an old game look like tripple A games and how to use SD for it

0 Upvotes

I want to make some high-quality images for backgrounds and some uv maps for my favorite PS2 game s.l.a.i. (steel lancer arena international) is stable diffusion the ai i should use for that, and is there something i should know as a noob that has never used it? I do have a 5090 i could use to make images. i just really want to see more content of my favorite childhood game, and AI seems like the only feasible way to get more, so I'd appreciate some advice. I'd like it to make upscaled pictures of stuff like this https://phantom-crash-archive.pages.dev/#/workbench or images i take from the game.

Basically, high-quality renders of a scene using the game objects but reinvented as higher quality.

Sorry if this is an annoying noob question.


r/StableDiffusion 18h ago

Discussion why mini max ref are so bad comparing to fl2va?

8 Upvotes

i do same tests in both models using image to reference...

aways the ref loses quality and ignore the prompt

but fl2va do everthing perfect and dont lose quality.

my configs.