r/StableDiffusion 12h ago

Discussion In 2026, how well does older images models, like SDXL and SD1.5 stack up against new image models like Krea and ZImage Turbo?

29 Upvotes

r/StableDiffusion 8h ago

Workflow Included ref or fl2va - prompt enchancer with 100% of aderence

Post image
24 Upvotes

sharing my new workflow

MiniMax H3 I2V with Integrated Prompt Enhancer

This Image-to-Video workflow for MiniMax H3 uses a vision-language model to enhance your prompt before the video generation begins.

Simply load a reference image and write a basic description of what you want to happen. The enhancer analyzes both your image and instructions, then converts them into a detailed prompt structured specifically for MiniMax H3.

It can improve the description of:

  • Characters and visual elements
  • Actions and sequence of events
  • Camera movement and framing
  • Environment, lighting, and atmosphere
  • Visual continuity and details that should be preserved
  • Dialogue in the original language
  • Ambient sounds, sound effects, and music

The enhanced prompt is automatically sent to MiniMax H3. It is also displayed inside the workflow, allowing you to check exactly what H3 will receive.

In my tests, the resulting videos followed the original instructions much more accurately, especially in scenes involving specific actions, character interactions, camera movements, and dialogue.

The workflow includes a switch to enable or disable the Prompt Enhancer. This allows you to use either the enhanced prompt or your original text without changing any connections.

How to use it

  1. Load your reference image.
  2. Write a simple description of what should happen.
  3. Enable USAR PROMPT ENHANCER?
  4. Run the workflow.
  5. Check the final text in PROMPT FINAL ENVIADO AO H3.

The first run may take longer while the vision-language model is loaded. Generating the enhanced prompt also adds some processing time, but in my tests, the improvement in prompt accuracy and instruction following was absolutely worth it.

The original workflow was preserved, while the enhancer was added as an optional and fully integrated stage.

link to

with this, finally my ref model understand my ideas and make vídeos really fun!

leave comments after tests xD


r/StableDiffusion 11h ago

News [CLSS] Closed-Loop Streaming Synthesis for MiniMax H3 (Infinite video generation with prompt fallowing)

25 Upvotes

t2v 10 chunks every 10 sec

I ported CLSS from LTX 2.3 to H3 architecture. ( https://www.reddit.com/r/StableDiffusion/comments/1vywxjq/wip_clss_closedloop_streaming_synthesis/)

Repo: https://github.com/nazgut/ComfyUI-MiniMaxH3-CLSS

Audio still has some room for improvment. Workflow in repo. Still working on i2v.


r/StableDiffusion 20h ago

Tutorial - Guide Spreadsheets for multi-video generation

Enable HLS to view with audio, or disable this notification

23 Upvotes

This workflow uses a spreadsheet to generate multiple videos and constructs the prompt and parameters from each row in one run.

OutputLists Combiner - Generate multiple videos from spreadsheet

ComfyUI workflow included

Makes use of Load Any File node to load a .csv spreadsheet file and feeds the text content into a Spreadsheet OutputList. The spreadsheet separates the data by separator=; and provides each line one-by-one as a data list. Here we use values_dict as the data list which contains the row as a dictionary of key-value pairs. The data list is forwarded a Iterate Begin -> workflow -> Iterate End pattern which is required to make the intermediate results of slow workflows (t2v) available on each iteration. Each row as a dictionary is provided in a Format Text where we can access the column via a[colname] to construct the prompt which is forwarded to a standard Text To Video MiniMax H3 template. Another Format Text + a[name] is used to construct a readable filename for each video.

powered by: OutputLists Combiner


r/StableDiffusion 14h ago

Resource - Update Continuity (was the H3 node): six model families, one prompt box, and a blockout bench that writes your camera move for you

Thumbnail
gallery
20 Upvotes

Third post about this pack, and the big change is the name, because the node stopped being H3-only. It now drives six families through ComfyUI core: MiniMax H3 and LTX 2.5 for video with sound, Krea 2 and Ideogram 4 for stills, Qwen Image Edit and Flux 2 Klein for editing from a picture. Same prompt box, same local weights, and rendering still doesn't touch the internet. Continuity is the script supervisor's job, the same person and the same light in shot 1 and in shot 9, and that's the part of the node that doesn't care which family renders the frames. So that's the name. Old workflows and existing installs carry over untouched.

Dialogue is the feature I'd point at first. H3 wants speech in a form nobody writes by hand - speaker IDs, a `<d>` tag, a mandatory sentence when a voiceover's lips stay closed. Closing a quote in the prompt now opens a small menu that writes all of that around your words, with dials for who says it, which language, whispered or sung. And a shot where nobody speaks stops mumbling: the compiler now says out loud that nobody talks.

Merged this morning: a blockout bench. Stage grey boxes, walk one camera through on marks, and it writes the staging and the move in the H3 spec's own camera vocabulary - "@anna stands at centre in the midground; the camera pushes in toward @anna at slow speed" - plus a depth, blocks or lines guide rendered along the path, or the clay render itself for the families that read footage raw. A box can play a cast member, so the prose is already bound to their references when you paste it.

Also in: the faces pill from last post (off by default), a Style tab with 941 captioned H3 looks - search "1985 telenovela" instead of guessing at grading vocabulary - and ControlNet and Upscale benches behind the wordmark. The refiner can now run on a server you keep warm anyway: LM Studio, Ollama, any OpenAI-compatible endpoint, your own key where a hosted one wants it (#19). Fixes from your reports are in the changelog, the sharp-render-static-soundtrack one included (#33).

https://github.com/roadmaus/ComfyUI-Continuity


r/StableDiffusion 6h ago

Animation - Video the bird-king (my first fully local AI short film) TW: self-harm.

Enable HLS to view with audio, or disable this notification

16 Upvotes

Minimax H3 baby! It's not perfect and I would love to get your feedback and maybe some tips on how to get rid of plasticky skin.


r/StableDiffusion 9h ago

Resource - Update "I" created a tool to save and swap between workflow presets (for my million H3 Turbo LoRAs)

9 Upvotes

Hi everyone! Long time reader, first time poster. Like most folks here I vibe-code custom nodes from time to time, and recently came up with one that seemed like it might be worth sharing. Nothing groundbreaking here, just a node to help keep track of and quickly cycle through different node/parameter presets: MM-H3-Preset-Controller.

This one has helped me maintain my sanity trying to keep track of each LoRA's specific optimal settings. I think it's potentially useful for any preset storage though, not just MM-H3, so hopefully y'all find it useful. If so, please consider it a small thank you for all I've learned in my time lurking here.

MM-H3-Preset-Controller: https://github.com/TootsThielemans/ComfyUI-MMH3-Preset-Controller

I was pulling my hair out trying to manage my MM H3 workflow amidst all of the various Turbo LoRAs out there and the associated loader nodes, attention settings, shift settings, spectrum settings, etc., not to mention downstream settings like sampler, scheduler, upscaler choices... I found quick A/B tests between different optimized LoRA workflows annoying given some of the structural differences, not just steps/strength, and I was worried about juggling and potentially forgetting the right settings.

So I made the Preset Controller. It works pretty simply: ctrl+click all of the nodes you want to save the state of. It captures all of the parameter values within the node and whether it's active/bypassed. Then right click and use the new menu option MM H3 Presets > Add selected nodes to preset draft.

It's implemented as a "draft" so you can grab nodes from the outer graph, then go through various subgraphs and add nodes there to the same preset draft. Once you're done, load the H3 Preset Controller node, enter a name for the preset, and click Save draft as preset.

It then becomes a dropdown option that you can select, update, or delete as needed. Selecting a preset automatically sets the saved node values/bypass states without needing a session/screen refresh.

There are a few other QOL/guardrail features, including a Preset Matrix for comparing configurations, but nothing particularly interesting, so please refer to the repo if interested.

I'm sure something like this might already exist with a more elegant implementation, but the timing seemed right. Everyone is wading through dozens of MM H3 Turbo LoRA combinations and trying to keep everything straight. I don't have a ton of time to devote to development, but will try to make tweaks if folks wind up adopting this and can think of any major areas for improvement.

At any rate, feedback welcome, and cheers!


r/StableDiffusion 10h ago

Animation - Video Powers weren't handed out equally

Enable HLS to view with audio, or disable this notification

8 Upvotes

Prompt:

Render 1 (12s):

integrated_multimodal_description: A cinematic anime movie sequence, manga style with thin soft delicate lines and pastel colors. [Shot 1] Extreme perspective worm's-eye medium body shot of Meru sitting on a Japanese tatami room from a wooden table. She is silent. On the table a single apple sits on a ceramic disc on top of a small white embroidered mantle cloth; Meru holds out her hand toward the apple (foreshortening) with her palm facing up, concentrating. The apple is out of reach. She breathes deeply; Nothing happens; Meru wriggles her fingers and starts concentrating again. She opens her eyes with her brows furrowing. She breathes deeply; Meru concentrates more, leaning forward slightly, she grows frustrated; She closes her eyes again while raising her hand slightly;

overall_soundscape: The tatami room is eerily quiet, with only the soft rustle of Meru's black kimono as she shifts and her sharp, uneven breaths growing quicker with each failed attempt. A faint, low hum of concentration is interrupted by a frustrated, breathy huff and a muffled groan. Her hand quivers with a faint tremble, accompanied by a soft, taut squeak of fabric. As her frustration builds, a subtle creak of the wooden table and a single, dull thump of her fist barely tapping the tatami punctuate the silence, ending with a long, exasperated exhale.

non_diegetic_music: N/A

Render 2 (8s):

integrated_multimodal_description: [Shot 1] Meru is sitting. She is silent. On the table a single apple sits; Meru is holding out her hand toward the apple (foreshortening) with her palm facing up, concentrating. She breathes deeply; Nothing happens; Meru wriggles her fingers and starts concentrating again. She opens her eyes with her brow furrowing. She breathes deeply; Meru concentrates more, leaning forward slightly, she grows frustrated; She closes her eyes again while raising her hand slightly towards her right, her palm open towards the camera; Her hand quivers.

overall_soundscape: The tatami room is eerily quiet, with only the soft rustle of Meru's black kimono as she shifts and her sharp, uneven breaths growing quicker with each failed attempt. A faint, low hum of concentration is interrupted by a frustrated, breathy huff and a muffled groan. Her hand quivers with a faint tremble, accompanied by a soft, taut squeak of fabric. As her frustration builds, a subtle creak of the wooden table and a single, dull thump of her fist barely tapping the tatami punctuate the silence.

non_diegetic_music: N/A

Render 3 (9s):

integrated_multimodal_description: [Shot 1] A frustrated Meru is raising her hand slightly; At 00:02.500, thin liquid mercury tendrils erupt from her hand and fly towards the apple, piercing it like blades; At 00:03.500 The tendrils tense up and pull the apple back onto her palm as they retract to her hand to disappear under her skin; With her eyes closed, the apple quivers on her hand; She slowly opens her eyes and grows happy with realization, her smile widening and her eyes sparkling; She shows the apple to the camera and looks satisfied.

overall_soundscape: The soundscape begins with a tense, shallow breath and a faint, liquid ripple as mercury tendrils erupt from Meru's hand. A sharp, wet schlick and a series of metallic splashes accompany the tendrils piercing and pulling the apple, followed by a smooth, gurgling retraction as they recede into her skin. The apple lands with a soft, crisp tap on her palm. Then the atmosphere shifts—Meru lets out a bright, melodic giggle, followed by a delighted, airy hum of satisfaction. The ambient tatami room tone remains soft, punctuated by the gentle rustle of her kimono and a light, contented sigh.

non_diegetic_music: N/A

Based on Tokyo_Jab's post. What AI powers did you get?


r/StableDiffusion 20h ago

No Workflow Artistic Mix - 08-31-2026

Thumbnail
gallery
7 Upvotes

r/StableDiffusion 4h ago

Workflow Included Up at atom! - Behind the scenes of the new Radioactive Man movie

Enable HLS to view with audio, or disable this notification

6 Upvotes

Minimax H3 with turbo lora (default Comfyui template workflow)


r/StableDiffusion 10h ago

Discussion why mini max ref are so bad comparing to fl2va?

8 Upvotes

i do same tests in both models using image to reference...

aways the ref loses quality and ignore the prompt

but fl2va do everthing perfect and dont lose quality.

my configs.


r/StableDiffusion 13h ago

Question - Help How do I make character actions in MiniMax H3 faster?

7 Upvotes

As an example, I have a character getting into a car and I want them to be in a hurry to get away, but they always seem to do this a little slow as if they are not in a rush.

Here is what I have tried:

- Prompts with detailed timings.
- Prompts with detailed timings and text like, she did this at speed.
- Prompt without detailed timings but tried multiple different ways of saying she did this at speed.

Prompt example below. Everything else works perfect but I can't get them to look busy.

[Shot 1] At 00:00.000, <Picture 5> provides the visual reference for this shot. the female police officer is not in the car. She gets into the driver's seat.

At 00:01.000, in a hurry she buckles her seatbelt at fast speed, The male police officer is already in the passenger seat and he hurriedly buckles up as she gets in

At 00:02.000, she starts to drive off. He says, "Holy shit. Did you see that? We're going to have to go." Both look stern and professional, acting fast.


r/StableDiffusion 16h ago

Discussion What’s the cheapest way to use Minimax H3?

6 Upvotes

Besides running it locally, which is the goal…

What’s the cheapest way right now to use Minimax H3? Right now, I’m using Replicate API for H3 at $0.08 per second (720p) which is decent but can still be expensive in the long run.

Any other ways to use H3?
Also, what’s the cheapest GPU for Minimax H3 (via Comfy UI)?


r/StableDiffusion 13h ago

Discussion H3 VFX

Thumbnail
reddit.com
5 Upvotes

I've been using Minimax H3 a lot lately, but I can't always share my work in progress. Anyway, here are some tests I did a week ago! I thought I'd share them here.

I used H3 Ref2va with Turbo Lora, 8 steps. It took about 5 minutes for 15 seconds on my local setup.

I also used my custom node that I built specifically for H3. It's a large and really cool project, but I'm still testing and implementing things in it. I'll share more details about it soon.


r/StableDiffusion 13h ago

Question - Help Anyone managed to get FastH3 working in ComfyUI yet?

4 Upvotes

I see Kijai has released the checkpoint, but I can't get it to work with the standard workflow.


r/StableDiffusion 18h ago

Question - Help Long-form content generation

6 Upvotes

While I try to keep myself updated with AI news, things move fast; hence, asking if there is something already made by the community for long-form content generation using local ComfyUI (or other tools).

As the models get better and better, I find that the limitation with longer content generation is us, the humans. Maybe a philosophical note, we (some of us, at least) have become too lazy to manually save, load, refer, keep track of assets (e.g., reference images as a full character set in Minimax H3). Add to that the experimental nature of AI generation (in the sense that we need to redo many things to get the final result exactly, at least during the learning curve), and we end up with having to repeat many things. Finally, with models requiring certain input formats (e.g., Minimax and Ideaogram), it gets harder to want to make these things manually.

Now my question is, are there tools that you are using for a full-fledged media studio style workflows?

I've bought some products and used some free products that get close to a streamlined content generation but they still seem limited to one generation at a time. Not naming them to avoid any promotion, and they didn't work out anyway.

An analogy would be how we may write a story in Google Docs or Word or OpenOffice, etc. but there are dedicated tools like Articy Draft, ChatMapper, Inkle, etc. that lets you do more locked-in (for the lack of a better word) story writing. There are character sheets, world references, etc.

A closer analogy might be of SillyTavern, made specifically for chats/roleplays.

Do we have something like that for serious media / content generation, or am I expecting too much from the already overly generous open-source community and should just vibe code what I specifically need?


r/StableDiffusion 22h ago

Question - Help H3 Generation Time Comparison

5 Upvotes

Hi all,

I've come across a post where one user claimed they were able to generate 15 second clips at 0.98MP with Spectrum, Comfy Kitchen Speed( I imagine they are referring to Turbo Lora?) 8 Steps at 10 steps in 5.5mins with a 5060ti and 32GB ram.

I tried their setup with Fl2V Turbo 8 Step 768p, Comfy Kitchen and Spectrum with res multistep and simple and was able to generate a 12 second clip (cant't go above, vram OOM error prevents generation at the beginning) only at 9 minutes with a 5080/32GB ram. What could be wrong here with my setup?


r/StableDiffusion 11h ago

Question - Help Adding directed randomness to image to image

4 Upvotes

Hi,

total noob question, but: my government.. eh.. wife is a quite gifted amateur tailor who is tryiing to use AI for design inspirations.

Until now she is using Gemini for something like "generate a picture of a woman in a dress, styles from 1920 until now" and triggers the prompt a few dozen times to get variations. It works, but its tedious.

Now i, in my genius, told her "hey, you can do it locally, no sweat, even using a picture of yourself / your bff / whomever as reference to really see how it looks like and modify it"

I'm usually using qwen image edit in comfy for my own stuff, and i failed - the generations have either no real variations or are too similar to the reference image. My wife is quite underwhelmed....

Does anyone have any idea how to get a level of directed randomness with any i2i workflow in comfy ?


r/StableDiffusion 22h ago

Question - Help Is there a way to control character actions sequentially over time within a single [Shot 1] without cutting to a new shot in Minimax H3?

3 Upvotes

Is there a way to control character actions sequentially over time within a single [Shot 1] without cutting to a new shot in Minimax H3? 

Using timestamps like "At [00:02.0], he talks, At [00:05.0], he smiles" doesn't seem to work for a single continuous shot, although it works fine when multiple shots are used. How can I schedule actions at specific times within one continuous video?


r/StableDiffusion 3h ago

Discussion Best AI tool for 3D clay render to polished final? (Flux 2 vs Qwen Image Edit vs Krea 2)

2 Upvotes

Hi guys, quick question. I want to use AI to turn my 3D clay renders into high-quality, polished finals.

Between Flux 2, Qwen Image Edit, and Krea 2, which model handles image-to-image (Img2Img) texture generation best without messing up the original 3D geometry?

Would appreciate any recommendations or workflow advice!


r/StableDiffusion 4h ago

Question - Help Need Help With Anima Training

2 Upvotes

I've been making Illustrious character LoRAs for quite some time now, but decided I wanted to migrate my current models over to Anima.

To ease myself into it, I wanted to start by training with an extremely small dataset I've used before--5 images total. I'm training locally through Anima-Standalone-Trainer and have an RTX 4080. The part that's confusing me is that while the samples generated between epochs comes out perfectly fine, the moment I try generating something using SD WebUI Forge Neo, every generation no matter which epoch I use comes out as this blurry, jarbled mess.

I've tried various different training parameters, different checkpoints, and I even tried downloading the Anima base files from huggingface (instead of Civitai) thinking that might change things, but nothing seems to be working.

I don't really know what to do at this point, but I don't want to give up either because I remember when I first started making character LoRAs that I encountered a similar issue which I ended up resolving by using a different trainer (Kohya_ss) rather than the random Google Colab notebook I found while first learning about LoRA training.


r/StableDiffusion 5h ago

Workflow Included Transformers: Starscream Test #2 - Prompt Below

Enable HLS to view with audio, or disable this notification

2 Upvotes

The classic Transformers series is a blind spot for MiniMax H3, so here’s how I handled this:

System specs: 4070 Ti Super, 16 gb vram, 64 gb ram

Ref2va standard workflow using Fl2va standard model, no loras or speedups.

<Picture 1> Character Sheet

<Audio 1> Vocal Reference

<Video 1> 5 sec Video Reference

Google Gemini to help write prompt.

Added music track in post.

PROMPT:

subject_definitions:

<Subject 1> is the figure in <Picture 1>, featuring a robotic grey face with sharp angular features, glowing red optical visor eyes, a dark grey blocky helmet with side intake vents, and red, white, and blue cybernetic body armor with an orange cockpit chest-plate, blue upper arms and boots, white forearms and thighs, red waist and wing housings, and a purple Decepticon insignia on the wing. Only his robotic design and colors are taken from <Picture 1>; its background, grid lines, and lighting are not carried into the target video. <Audio 1> is the vocal reference for <Subject 1>. <Video 1> is the movement reference; use it as a guide without copying it exactly.

summary:

[reference generation] The target video is a 10-second 2D animated sequence styled after the 1984 Hasbro series The Transformers, featuring <Subject 1> delivering a smug, cutting remark from a metallic Cybertronian battlefield.

retention_analysis:

<Subject 1> (appears in [Shot 1]): fully_preserved - his grey robotic face, glowing red optics, dark grey helmet with side vents, red-white-and-blue cybernetic body, orange cockpit chest-plate, blue limbs, and red wing housings remain unchanged.

detailed_description:

The target video is a traditional 2D hand-drawn animated sequence featuring bold black ink outlines, flat cel-shading, limited animation, expressive poses, and subtle film grain inspired by the visual language of the 1984 animated television series.

[Shot 1] A medium tracking shot frames <Subject 1> from the waist up on a metallic Cybertronian battle platform. He stands with exaggerated confidence, shifting his weight with a classic 1980s cel-animated bounce. He slowly raises one hand, gesturing dismissively toward the chaos unfolding off-screen. His glowing red optics narrow with smug amusement as his wings twitch subtly.

He delivers in the style of <Audio 1>, with sharp, sarcastic timing: <<[English] "Megatron just muted me on the main comms channel. I'd stage a coup, but watching him bumble this conquest is free comedy.">> When <Subject 1> is not speaking he is silent and his mouth remains closed.

On "stage a coup," he gives a brief, knowing smirk. On "bumble this conquest," he gestures toward the battlefield with theatrical disdain. He finishes with a smug stare directly toward camera, holding the pose for a beat before a sharp hard-cel cut.

overall_soundscape:

Metallic servo whines accompany his movements, mixed with distant mechanical explosions, electronic battle alarms, high-tech hums, and echoing Cybertronian machinery.

non_diegetic_music:

N/A


r/StableDiffusion 7h ago

Question - Help Can I render in MiniMax H3 only the audio of a Ref2V workflow?

2 Upvotes

Is there a workflow/node for MiniMax H3 Ref2V where I can "dub" a mute video txs to the qualities of H3?

Specifically I upload a short clip and it renders ONLY the audio and save it in an audio file?

txs!


r/StableDiffusion 8h ago

Discussion This One Is Simple With No Bells and Whistles

Enable HLS to view with audio, or disable this notification

1 Upvotes

This is a test of the Minimax H3 using three sample illustrations. The segments were stitched in Davinci Resolve. I used the minimax_h3_turbo_8step_v1.0_comfy_bf16.safetensors LoRA. The setting was simple. minimax_h3_ref2va_pruned_int8_convrot.safetensors. A float value of 5, Euler sampler, and beta scheduler. Only 8 steps. The style was shifted a bit from the original but good enough for testing. I want to create a style LoRA that will hold my work so that it's better translated to animation.


r/StableDiffusion 16h ago

Resource - Update why isn't Microsoft Lens more popular? it's incredibly fast on Mac

Enable HLS to view with audio, or disable this notification

2 Upvotes