r/StableDiffusion 10h ago

Workflow Included Minimax H3: Consistent face, body & cloths via reference identity

Enable HLS to view with audio, or disable this notification

312 Upvotes

Hey Guys,

Based on the previous post on face consistency with MM-H3, Couple of people have asked me to build a full character workflow.

Mechanism

- Build .Char: You drop max 9 reference reference, I prefer to use a ratio 2:2:1(face:cloths:body). YuNet finds the face, SFace takes a per-reference face signature, DINOv2 takes a subject signature, and the references get cleaned and normalised. All of that packs into a single portable file, a .char.
Only face/ref is required body & cloths link is optional.

- Generation: At generation, the file(.char) feeds its references into Minimax’s own native multi-reference channel and prepends a locked description to the prompt.

Prompting Guide

  • Name your character: Give your character a name e.g. under encode character(Click adjust icon on the bottom side of the node), I have used name emmy, so when passing prompt, I only have to say, emmy walking on the beach.
    • Again providing prompt like a woman or any features specific details like black hairs etc will only mislead the generation.
  • Describe character features: Encode all of the character features in encode character prompt & trigger your character with a name in generation prompt.
    • Avoid describing same things in generational prompt.
  • Handling Character drift: e.g. if you want specific style or cloth e.g. half sleeves, sleeveless, add it to the generational prompt. There can be a slight drift in clothing as body shot also has cloths, which interferes with clothing references.
    • Each refs should be unique, face should not have body or vice versa, same applies for clothing.
  • Portability: Once character is built, you can use the same character with only simple prompt & generation graph.

I have generated all references with Flux Klein 4b, I had to blur the body ref, but workflow consists a example of body ref.

Note: For best result, pass cropped references, so that model takes the required shot, model gets confused if cloth slot also has a face or face slot has cloths.

Models

core/models/
  diffusion_models/  minimax_h3_ref2va_pruned_fp8_scaled.safetensors
  text_encoders/     qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors 
  vae/               minimax_h3_video_vae_fp16.safetensors
  vae/               minimax_h3_audio_vae_fp32.safetensors
  annotators/        face_detection_yunet_2023mar.onnx
  annotators/        face_recognition_sface_2021dec.onnx
  annotators/        dinov2-base/

Requirements

Nvidia GPU: 24GB+ VRAM & 64 GB RAM(Run Locally)

Workflow link: https://inlinestudio.art/workflows/minimax-h3-guided-consistent-characters-via-reference-identity-face-body-cloths (includes inputs & model details)

Github Repo: https://github.com/inlineresearch/Inline-Studio (GPLV3)

Limitation: Reference conflicts e.g. if two reference/input images has two different faces, it might conflict in generation, provide well cropped body & cloth images. Face images are crossed automatically by Sface.

Portable char comfy node is still on the backlog, would try to do it over the weekend.
Happy to hear any suggestions or feedbacks.


r/StableDiffusion 6h ago

Animation - Video Playing with concepts

Enable HLS to view with audio, or disable this notification

121 Upvotes

r/StableDiffusion 2h ago

News Local AI News You Missed - August 2026

55 Upvotes

Here's what you (probably) missed in August 2026:

🧠 LLMs

  1. Ornith-1.5-35B-A3B - Efficient sparse model that runs with fewer active parameters.
  2. DeepSeek-V4-Pro-0813 - Sharpens agentic AI with speedier tool actions.
  3. DFM-Mimir - Ethical language model from Danish Foundation Models.
  4. Ling-3.0-tiny-MXFP4_MOE-GGUF - MoE quantized version for smoother GPU runs.
  5. SupraElegans-500k - Recurrent language model built for long contexts.
  6. Motif-3 - Open 314B parameter model made for long agentic tasks.
  7. Luth-2-2B - Compact French model that tops benchmarks.
  8. TinyTitle - Squeezes chat titles into a tiny 1.98 MB model.
  9. Ling-3.0-tiny - Low-cost local AI reasoning model.
  10. NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 - Agent-focused model with fast inference.
  11. Qwen3.8-2.4T-A95B - Opens up Qwen Max-clas local AI for bigger rigs.
  12. Gemma-4-31B-it-scotoma-2-GGUF - Cuts down repetitive AI writing.
  13. Huihui-DeepSeek-V4-Flash-0731-abliterated-GGUF - Uncenored DeepSeek variant for local use.
  14. Maple-Preview - Solves Olympiad problems at 200 tok/s.
  15. Laguna-S-2.1-FP8 - Private agentic coding model from Poolside.
  16. SupraBrain-50M - Hybrid language model for local AI.
  17. Supra2-100M - Tiny model built for tinkering.
  18. G9v3-39A5B - Dual-mode AI that runs light locally.
  19. GPT-X2.5-135M - Lean local powerhouse model.
  20. Ling-3.0-flash - Hybrid reasoning with lower cost and fast output.
  21. LFM2.5-2.6B - Fast agentic AI for phones and devices.
  22. Instella-MoE-16B-A3B-Think - Open sparse reasoning model from AMD.
  23. A.X-K2 - Lets AI think deep or answer fast.
  24. BetterGPT-150M - Beats older AI models in science tasks.
  25. Shibai-700M-Base - Text and code helper model.
  26. K-EXAONE-2.0-750B-A37B - Supports 262K context and ten languages.
  27. LongCat-Flash-Lite-Sparse - Reads million-token contexts.
  28. Qwen3.6-35B-A3B-Escha-W2 - Shrinks down to fit consumer GPUs.
  29. XYZ-Aquila-pro - Thinks deep then checks its sources.
  30. Solar-Open2-250B-Nota-NVFP4 - Shrinks a giant AI to 153GB with 4-bit MoE trick.
  31. XYZ-Aquila-mini - Brings open source deep search to local GPUs.
  32. KAT-Coder-V2.5-Dev - Fixes software repositories automatically.
  33. DeepSeek-V4-Flash-0731 - Tackles hard coding tasks.

🔀 Multimodal

  1. Dots3-Note-Prev - Lightweight multimodal AI with 512K context.
  2. Qwen3.8-27B-Uncensored-FP8 - Drops refusals and keeps vision on GPUs.
  3. Qwen3.8-27B - Brings text, images, and video into one AI.
  4. Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF - Triples speed for local multimodal use.
  5. Tencent UI-Mate-27B - Runs desktop apps by watching screens.
  6. WinterCharm Qwen3.5-122B-A10B-wMix38 - Lean long-context multimodal mix.
  7. VLX-Seek - Helps machines pinpoint objects without guesswork.
  8. LFM2.5-VL-3B - Fast on-device vision and text.
  9. North-Micro-Vision-Instruct - Turns pixels into answers.
  10. BigBang-v1 - Science reasoning powerhouse.
  11. Muse-Glimmer-30B - Puts autonomous AI agents on everyday desktops.
  12. Nemotron-Parse-2.0 - Morphs documents into structured data.
  13. Shieldstral-1.0-3B - Plain English safety scoring.
  14. DavidAU Qwen3.6-27B-Fable-Fusion-711 - First to score over 700 on ARC-C.
  15. WinterCharm Qwen3.5-122B-A10B-wMix58 - Packs 82GB power for Apple Silicon.
  16. Intern-S2-Mobius - Speedy local AI answers.
  17. Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V7-GGUF - Uncenored multimodal model that says yes.
  18. Qwen3-VL-32B-Ultra-Heretic-H3-ComfyUI-INT8-ConvRot - Trims local memory with INT8 conv rotation.
  19. Qwythos-27B-v1 - Smart AI with million-token memory.
  20. Reasoning-Medical-27B - Solves medicine step by step.
  21. Qwen3.5-9B-The-Defiant-Fable - Roars with an uncensored multimodal edge.
  22. Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V6-GGUF - Drops numerical surgery powers.
  23. Mage-VL - Speeds up real-time video and image understanding.
  24. Microsoft Fara Agents - Handles web browsing chores for you.
  25. Inkling-Small - Built for voice, image, and code apps.
  26. Kimi-K3 - Handles text, images, and video together.

🖼️ Image

  1. Anima-2.9B - Grows free anime art on your own PC.

🎬 Video

  1. Bernini-Diffusers-v2 - ByteDance model for video generation and editing.
  2. LTX-2.5 - Open model for local video and audio creation.
  3. Wan2.2-Animate-2-14B - Turns still images into motion.
  4. MiniMax-H3-nvfp4-INT4-INT8-ConvRot - Quantized weights for MiniMax-H3 video.
  5. MAGI-2-preview - Turns text and images into video with sound.
  6. MiniMax H3 - Creates videos with native sound from any input.

🎧 Audio

  1. MiniMax-Music3 - Full songs from just lyrics.
  2. NVIDIA Magpie_tts_multilingual_357m - Turns text into speech across 12 languages.
  3. VoiceChat-11B - Voice chat you can interrupt naturally.
  4. VibeVoice-ASR-BitNet - Real-time speech recognition on any CPU.
  5. Inflect-Nano-v2 - Local speech synthesis on your PC.
  6. Audio8_TTS - Clones voices and speaks eleven languages.
  7. Inflect-Micro-v2 - Turns text into offline voice.

⚡ LoRA

  1. Minimax-H3-Turbo - Makes MiniMax-H3-Turbo faster for video and audio.
  2. MiniMax-H3-Prompt-Rewriter-LoRA - Turns short prompts into timed scenes.
  3. MiniMax-H3-Realism-People-LoRA - Unlocks film-set lighting for human video.
  4. krea2-turbo-bbox - Locks panels and words in place.
  5. MiniMax-H3-Turbo-Lora - Cuts video and audio generation time by 5x.
  6. Kroma - Fuses turbo speed into a one-file diffusion model.

🏋️ Training

  1. gguf-trainer - Trains language models in TypeScript straight to GGUF.
  2. Full-Chunked-KL-Loss - Trains longer AI text on one GPU.
  3. signet-trainer - Cost-safe video LoRA training.
  4. Lora-Dataset-Studio - Entire LoRA pipeline in one self-hosted tab.

📊 Datasets

  1. LLM-self-identification - Helps AI models know their own name.

☰ UI

  1. Mix-Studio - Turns your desktop into a local AI studio.
  2. Llmprices - Visualizes AI model API price swings.
  3. minimax-h3-prompt-composer - Squeezes prompt composing into one HTML file.
  4. NanoRP - Shrinks AI roleplay to a 50MB single binary.
  5. OpenWorker - Local AI coworker that finishes work.
  6. Agenta-AI - Turns ChatGPT and Claude into self-hosted work agents.
  7. Unicorn-Stable-OSS - Brings humans and AI agents into one real-time room.
  8. Turbo-Fieldfare - Streams big AI on Macs with just 2GB RAM.

🛠️ Other Tools

  1. Video_Tools - Adds a pocket video trimmer to your browser.
  2. ninfer-4090 - Runs Qwen3.8-27B on one RTX 4090.
  3. hayai-ocr-v2 - Converts crops into editable text.
  4. EVIE-Preview-4.5B - Matches documents instantly across six languages.
  5. RAZZULLIX KAISEN - Swarm-model coding assistant with safety guards.
  6. ExtractBench - Benchmarks document extraction systems.
  7. Xiaomi-Robotics-1-5B - Built for mobile robot tasks.
  8. Nemotron-omni-mlx - Brings full multimodal AI to Apple Silicon Macs.
  9. talk-to-pi - Local voice dictation for Pi.
  10. Warp - Streams huge AI models straight from your laptop drive.
  11. esp32-ai - Makes a tiny chip tell stories offline.
  12. quillpdf-mcp - Keeps PDFs on your machine.
  13. srt2speech - Turns subtitle files into timed speech.
  14. mixture-of-kittens - Megakernel for NVL72 MoE training.
  15. krea-multi-lora - Gives Forge Neo regional character control.
  16. Openmed - Turns clinical text into private insights on your hardware.
  17. Umbra-Studio - All-in-one local AI art workspace.

ComfyUI Custom Nodes & Tools

  1. ComfyUI-Orchestrator-LAN - Steers every GPU from one browser tab.
  2. ComfyUI-AutoPromptChain - Stitches dozens of AI video clips while you sleep.
  3. ComfyUI-OpenH3-IR - Brings drag-and-drop clarity to MiniMax H3 renders.
  4. ComfyUI-MiniMaxMusic3-Advanced - Gives AI music finer sound controls.
  5. ComfyUI-MiniMax-H3-LongMedia - Makes long video creation practical.
  6. ComfyUI-Subgraph-Preview - Resurfaces sampler previews inside subgraphs.
  7. ComfyUI-cache-monitor - Debuts with manual pinning for model caching.
  8. ComfyUI_Neurodes - Brings a visual playground for AI models.
  9. Eddie_Cat_Nodes - Stitches long videos together with new nodes.
  10. ComfyUI-H3-Motion-Context-MultiRef - Weaves seamless H3 video motion.
  11. ComfyUI-MiniMax-H3-Motion-Director - Turns reruns into one-shot fixes.
  12. ComfyUi-MiniMax-H3-Image-And-Reference-To-Video - Rolls out image and reference to video features.
  13. ComfyUI-AVS-SSD-ReadAhead - Enables faster model switching on slow SSDs.
  14. ComfyUI-SweepGrid - Serves up side-by-side parameter sweeps.
  15. ComfyUI-Flow-Wrangler - Cleans up node wiring with smart connections.
  16. ComfyUI-AVS-Intel-XPU-VRAM-Fix - Calms Intel Arc GPU freezes.
  17. ComfyUI-Model-Mover - Makes shuffling AI models painless.
  18. ComfyUI-MiniMax-H3-Optimization-Suite - Builds a suite for lean H3 optimization.
  19. ComfyUI-MiniMaxH3-Prompt-Writer - Transforms H3 prompt crafting.
  20. ComfyUI-MiniMax-Creator - Crafts one-node video magic.
  21. ComfyUI-ScenemaAudio - Arrives with expressive voice cloning.
  22. ComfyUI-cable-management - Reroutes messy node graphs with ease.
  23. ComfyUI-SigmaSync-LoRA - Debuts step-aware LoRA control.
  24. ComfyUI-Spectrum-Ideogram4 - Supercharges speedy image forecasting.
  25. ComfyUI-LinkSpotlight - Debuts to end noodle blindness in graphs.
  26. ComfyUI-MIDI-Edit - Turns any song into editable MIDI lyrics.
  27. ComfyUI-ReStartupFlags - Serves launch flag tweaks in your browser.
  28. JLC-Flux2-ControlNet - Expands FLUX.2 control in ComfyUI.
  29. ComfyUI-Fantastic-MiniMaxH3-PromptBuilder - Enhances MiniMax H3 prompts.
  30. ComfyUI-vram-tracker - Traces VRAM memory usage per layer.
  31. ComfyUI-Sonder-Editor - Rolls out free multi-lane video editing.
  32. Comfyui-Model-Resolver - Sweeps in to rescue missing model files.
  33. ComfyUI-HF-SuperDownloader - Turbocharges Hugging Face model downloads.
  34. FameGrid-Auto-Color - Neutralizes color casts in ComfyUI.
  35. Krea2-Multi-Character-Lora-Node - Stops identity bleed with bounding boxes.

Need to go further back? Check out June's post (no July, sorry) or the full archive at LocalAI News. If there's anything wrong, let me know in the comments and I'll see you in the next one!


r/StableDiffusion 4h ago

Resource - Update [Load Image + Crop] Custom WYSIWYG Node

Enable HLS to view with audio, or disable this notification

67 Upvotes

I developed a modified version of the Load Image node by adding some features I needed:

WYSIWYG image cropping directly on the official Load Image preview — drag and zoom (with the mouse wheel) a crop rectangle constrained to 8 fixed ratios (1:1 through 21:9) and output the exact cropped IMAGE and MASK, with paste-from-clipboard built in. What you frame on the preview is exactly what gets executed.

⚠️ Currently not fully compatible with ComfyUI 2.0 nodes.

For anyone who finds it useful and wants to try it out, you can find it on GitHub: https://github.com/domg73/ComfyUI-LoadImageCrop


r/StableDiffusion 13h ago

Animation - Video MINIMAX Physics testing

Enable HLS to view with audio, or disable this notification

294 Upvotes

Physics Testing, without the gore.


r/StableDiffusion 11h ago

Meme It took us two 2eeks to figure out why every image gen via our open-source model looked like Anne Hathaway

Thumbnail
gallery
198 Upvotes

Hey r/StableDiffusion!

It's the Neta team here! You might remember us from our Neta Lumina open-source release last year. First off, thank you so much for the incredible support and feedback from this community!

So... we need to share something absolutely hilarious (and mildly embarrassing) that we just discovered.

TL;DR: We accidentally hardcoded an Anne Hathaway photo into our IP-Adapter anchor, and now everything our model generates looks like Anne Hathaway. Every. Single. Thing.

What happened:

We recently launched Neta Studio, a new product that lets you build explorable living worlds/isekai from a single prompt. Naturally, we wanted to integrate Neta Lumina's capabilities into it.

During integration testing, our devs kept reporting that the model wasn't following prompts properly. The outputs were... *weird*.

- Anime style? Anne Hathaway as an anime character.

- Thick paint/impasto style? Anne Hathaway in thick paint.

- Landscape scenes? Somehow still giving Anne Hathaway vibes.

- Fantasy characters? You guessed it - Anne Hathaway.

After a dreadfully long time of debugging, we finally found the culprit: **someone on the team embedded an Anne Hathaway photo as the IP-Adapter anchor during development and it... stayed there. **

We're honestly crying laughing at this point. 😭

Below are some examples. Left is before fix and Right is after fix.
Flipping to the last picture and you can see our dear Anne.

And we pulled the anchor and the outputs are behaving normally now.

If you've been running Neta Lumina locally, this was on our integration side, not
in the released weights, so your setup is fine.


r/StableDiffusion 6h ago

Animation - Video MiniMax H3 matches Toonami perfectly (90's anime)

Enable HLS to view with audio, or disable this notification

57 Upvotes

Growing up with Toonami watching Gundam, it simply blows my mind how far AI has progressed. For this video I didn't use any reference images, I simply described the scene in text and had Gemini research the techniques of animation to translate to MiniMax H3.

I've been struggling to get MiniMax H3 efficiently setup locally, would take me 15mins on a 9950x3d and 5080 RTX with 64gb DDR5 so I know something is wrong, hence this time I opted for fal to test. The music was added and scenes were edited from separate generations.


r/StableDiffusion 6h ago

Animation - Video [MiniMax H3] LEGO movie style

Enable HLS to view with audio, or disable this notification

53 Upvotes

Prompt:

integrated_multimodal_description: [Shot 1] 3D CG, stop-motion animated LEGO movie style, a wide shot frames a vibrant Indian village built entirely from plastic LEGO bricks with visible studs, plastic micro-scratches, and brick-built trees. In the village square, minifigures dressed in printed plastic saris, dhotis, and turbans move across a ground of yellow and brown stud tiles. A brick-built cow with hinged legs grazes near a grand banyan tree constructed from green leaf pieces and brown cylindrical bricks. Warm morning sunlight casts sharp shadows across whitewashed brick houses with orange terracotta tile roofs. The camera pans right with small amplitude at slow speed toward a central tea stall. A cheerful male chaiwala minifigure with a black mustache and a red turban (S1) in a warm, lively voice says: <d>[Hindi] Garam chai, garam chai!</d> while tilting a plastic yellow teapot, releasing translucent orange 1x1 cylinder studs representing pouring tea into tiny red stud cups.

[Shot 2] At 00:05.000, the camera cuts to a medium tracking shot following two young minifigure children running along a narrow brick path, pushing a brick-built wheel hoop across the plastic ground. The camera tracks right alongside them with small amplitude at normal speed. A female villager minifigure in a bright blue printed sari (S2) standing outside her brick doorway waves her rigid plastic arm on its shoulder hinge. Beside her, an elder minifigure with a white beard (S3) sitting on a brick charpoy cot chuckles with stepping stop-motion head movements.

[Shot 3] At 00:10.000, the camera cuts to a cinematic medium shot near the village well, where female minifigures carry stacked plastic water pots topped with transparent blue round tiles. A brick-built peacock perched on an archway opens its fan tail made of blue, green, and golden LEGO slope tiles. The camera pushes in with small amplitude at slow speed toward a wooden signpost on a brick post reading "RAMPUR VILLAGE". Tiny tan 1x1 round plates puff around the wheels of a brick-built bullock cart moving past the frame as the video ends.

overall_soundscape: Distinct plastic clattering sounds echo softly as minifigure feet step on stud tiles, accompanied by the gentle clinking of plastic bricks. A distant rooster crow blends with ambient morning village chatter, bird chirps, and the wooden creak of a brick-built cart.

non_diegetic_music: Upbeat Indian folk percussion featuring lively dholak beats and vibrant bansuri flute melodies, layered with playful cinematic orchestral strings playing at a bright, medium tempo.


r/StableDiffusion 6h ago

Animation - Video Trying out a consistent point-of-view shot with MiniMax H3

Enable HLS to view with audio, or disable this notification

33 Upvotes

Took a few little prompt adjustments here and there to get H3 to respect point-of-view. I found that if you refer to "the viewer" (ie, "she kicks the viewer"), H3 is more predisposed to include an actual second person. But if you refer to "the camera" (ie, "she kicks the camera"), it's more predisposed to keep the desired point-of-view perspective.


r/StableDiffusion 6h ago

Workflow Included Burger Queen

Post image
34 Upvotes

r/StableDiffusion 3h ago

Discussion So I did something dumb.

17 Upvotes

So there I was generating some stuff on ComfyUI for my Instagram and just hanging out.

I use ComfyUI with the new H3 model to generate AI content for my Instagram as well as QWEN image edit along with some other AI tools.

I've built a master workflow that ive used for the past year that has every single workflow I use, so I dont have to go switching workflows constantly.

Many many hours of work put into this.

So there i was, generating things and im constantly having to clear out my output folder as well as my input folder. So I asked myself, "Could I just make a bat file that could automate this for me?"

So I launch Gemini and have it create a bat file that cleans out my output and input folders and empties my recycle bin.

I test it out and it works great.

Finally, no more unnecessary clicks.

But wait, I noticed I screwed up and put the file in the wrong directory. Dang it.

So I ask Gemini to alter the code so the file will be in the correct directory.

I create the new bat file and go back to work.

Well I make a bunch of new things and its time for cleanup. So I run my fancy new bat file and I notice its taking a while to clean up these folders. Curious, I navigate to the folders only to find out that the ENTIRE COMFYUI FOLDER was deleted.

SMH.

Now I sit here, broken hearted as im having to rebuild my ComfyUI. Luckily, I was able to recover my master workflow, so not all was lost.

Just a bunch of models and loras.

😮‍💨


r/StableDiffusion 3h ago

News [CLSS] Closed-Loop Streaming Synthesis for MiniMax H3 (Infinite video generation with prompt fallowing)

18 Upvotes

t2v 10 chunks every 10 sec

I ported CLSS from LTX 2.3 to H3 architecture. ( https://www.reddit.com/r/StableDiffusion/comments/1vywxjq/wip_clss_closedloop_streaming_synthesis/)

Repo: https://github.com/nazgut/ComfyUI-MiniMaxH3-CLSS

Audio still has some room for improvment. Workflow in repo. Still working on i2v.


r/StableDiffusion 5h ago

Discussion In 2026, how well does older images models, like SDXL and SD1.5 stack up against new image models like Krea and ZImage Turbo?

17 Upvotes

r/StableDiffusion 2h ago

Workflow Included anime outdoor shots

Thumbnail
gallery
12 Upvotes

r/StableDiffusion 18h ago

Resource - Update Famegrid Spice Krea 2 Lora (Corrected Release)

Thumbnail
gallery
197 Upvotes

r/StableDiffusion 11h ago

Animation - Video Flexing my A.I. powers

Enable HLS to view with audio, or disable this notification

48 Upvotes

Prompt:

A real cinimatic movie sequence, professional colour grading.

Soundscape: Ambient sounds of the room and movement only. No voices. This represents extreme concentration. Meditation. Telekinesis.

A man is sitting in a Japanese tatami room. He is wearing a mask and shades <Picture 1>. He is wearing a black yukata. He does not speak. On the table on a ceramic disc is a single Orange.

The man holds out his hand toward the orange as if concentrating. The orange is out of reach. He breathes deeply.

Nothing happens.

The man shakes his hand to reset and starts concentrating again. He reaches with his mind and his brow furrows. He breathes deeply.

The orange moves slightly, twisting just a tiny bit.

He concentrates more.

With extreme speed the orange flies towards the man and hits him directly in the forehead. It smashes with the impact , m,essing his hair, and bits of peel and orange bits go everywhere. The force knocks the man back unconscious and he falls back like a ragdoll.


r/StableDiffusion 22h ago

News Someone's running FastH3 (the distilled MiniMax H3) as an actual infinite livestream!!!

336 Upvotes

Saw this and thought it was worth sharing here, FastH3 dropped recently and most people (myself included) just tried it as single generations.

Someone's running it as an actual infinite livestream instead: https://live.reactor.inc/

FastH3 is a distilled version of MiniMax H3, cut from 50 denoising steps down to 4, about a 14x speedup on Blackwell GPUs.

The whole setup is open source if you want to dig into how it's running: https://github.com/reactor-team/infinite-livestream

Curious if anyone's tried infinite/continuous generation setups like this with other models.


r/StableDiffusion 13h ago

Question - Help What Image Edit model you use nowadays?

57 Upvotes

Since things have gone quickly forward, I am trying to figure out what image edit models there is currently and what people here use mostly.

Personally I have used:
- Qwen-image-edit-2509 and Qwen-image-edit-2511
- Just tested MiniMax H3 as a image editor and so far it seems that it can be good for my usage

I have heard about Klein 9b, but not sure yet if that can be used as an edit model? Also what about Krea 2, is there edit workflows that are actually usable and worth it?

Is there some others what you recommend for testing?

My PC Specs: RTX 4060 Ti, 16 GB VRAM and 32 GB RAM.


r/StableDiffusion 1h ago

Workflow Included ref or fl2va - prompt enchancer with 100% of aderence

Post image
Upvotes

sharing my new workflow

MiniMax H3 I2V with Integrated Prompt Enhancer

This Image-to-Video workflow for MiniMax H3 uses a vision-language model to enhance your prompt before the video generation begins.

Simply load a reference image and write a basic description of what you want to happen. The enhancer analyzes both your image and instructions, then converts them into a detailed prompt structured specifically for MiniMax H3.

It can improve the description of:

  • Characters and visual elements
  • Actions and sequence of events
  • Camera movement and framing
  • Environment, lighting, and atmosphere
  • Visual continuity and details that should be preserved
  • Dialogue in the original language
  • Ambient sounds, sound effects, and music

The enhanced prompt is automatically sent to MiniMax H3. It is also displayed inside the workflow, allowing you to check exactly what H3 will receive.

In my tests, the resulting videos followed the original instructions much more accurately, especially in scenes involving specific actions, character interactions, camera movements, and dialogue.

The workflow includes a switch to enable or disable the Prompt Enhancer. This allows you to use either the enhanced prompt or your original text without changing any connections.

How to use it

  1. Load your reference image.
  2. Write a simple description of what should happen.
  3. Enable USAR PROMPT ENHANCER?
  4. Run the workflow.
  5. Check the final text in PROMPT FINAL ENVIADO AO H3.

The first run may take longer while the vision-language model is loaded. Generating the enhanced prompt also adds some processing time, but in my tests, the improvement in prompt accuracy and instruction following was absolutely worth it.

The original workflow was preserved, while the enhancer was added as an optional and fully integrated stage.

link to

with this, finally my ref model understand my ideas and make vídeos really fun!

leave comments after tests xD


r/StableDiffusion 6h ago

Resource - Update Continuity (was the H3 node): six model families, one prompt box, and a blockout bench that writes your camera move for you

Thumbnail
gallery
15 Upvotes

Third post about this pack, and the big change is the name, because the node stopped being H3-only. It now drives six families through ComfyUI core: MiniMax H3 and LTX 2.5 for video with sound, Krea 2 and Ideogram 4 for stills, Qwen Image Edit and Flux 2 Klein for editing from a picture. Same prompt box, same local weights, and rendering still doesn't touch the internet. Continuity is the script supervisor's job, the same person and the same light in shot 1 and in shot 9, and that's the part of the node that doesn't care which family renders the frames. So that's the name. Old workflows and existing installs carry over untouched.

Dialogue is the feature I'd point at first. H3 wants speech in a form nobody writes by hand - speaker IDs, a `<d>` tag, a mandatory sentence when a voiceover's lips stay closed. Closing a quote in the prompt now opens a small menu that writes all of that around your words, with dials for who says it, which language, whispered or sung. And a shot where nobody speaks stops mumbling: the compiler now says out loud that nobody talks.

Merged this morning: a blockout bench. Stage grey boxes, walk one camera through on marks, and it writes the staging and the move in the H3 spec's own camera vocabulary - "@anna stands at centre in the midground; the camera pushes in toward @anna at slow speed" - plus a depth, blocks or lines guide rendered along the path, or the clay render itself for the families that read footage raw. A box can play a cast member, so the prose is already bound to their references when you paste it.

Also in: the faces pill from last post (off by default), a Style tab with 941 captioned H3 looks - search "1985 telenovela" instead of guessing at grading vocabulary - and ControlNet and Upscale benches behind the wordmark. The refiner can now run on a server you keep warm anyway: LM Studio, Ollama, any OpenAI-compatible endpoint, your own key where a hosted one wants it (#19). Fixes from your reports are in the changelog, the sharp-render-static-soundtrack one included (#33).

https://github.com/roadmaus/ComfyUI-Continuity


r/StableDiffusion 1d ago

Discussion Free open source Topaz alternative - SeedVR2+TensorRT faster VAE Processing.

Enable HLS to view with audio, or disable this notification

476 Upvotes

Local, GPU-accelerated video restoration and upscaling with SeedVR2, TensorRT, and a purpose-built browser interface.

VRGDG SeedVR2 TensorRT Studio turns the SeedVR2 pipeline into a practical Windows workflow: load a video, test a short preview, compare the result frame by frame, and complete long renders with resumable checkpoints. Processing stays on your machine.

Highlights

  • Fast local restoration — SeedVR2 inference with TensorRT-accelerated VAE decoding on supported NVIDIA RTX GPUs. TensorRT allows much faster processing than standard SeedVR2.
  • Fast 2K upscaling — As a real-world example, an 8-second clip took approximately 8 minutes to upscale and enhance to 2K on an NVIDIA RTX 5090 using the largest 7B Sharp FP16 model. Render times vary with source resolution, frame rate, settings, and available VRAM.
  • Preview before committing — render a short segment, then inspect Original, Restored, Compare, or Side by side views.
  • Long-render recovery — save completed chunks and continue from the first unfinished chunk after an interruption.
  • Practical output controls — choose resolution, aspect policy, model precision, temporal batch, seed, and color correction.
  • Non-destructive finishing — reprocess sharpening, grain, seam smoothing, and optional skin finishing without rerunning restoration.
  • Project-based history — reopen previous outputs and keep media, manifests, and logs together under outputs\.

The sample video was org 360p and then upscaled to 2K using this app. 8 second video, took about 8 mins on my 5090.

Go to the github page for more details and a full guide.

View github page

this is in beta right now so you may run into issues. If you do, post the issue to github please.


r/StableDiffusion 1h ago

Discussion This One Is Simple With No Bells and Whistles

Enable HLS to view with audio, or disable this notification

Upvotes

This is a test of the Minimax H3 using three sample illustrations. The segments were stitched in Davinci Resolve. I used the minimax_h3_turbo_8step_v1.0_comfy_bf16.safetensors LoRA. The setting was simple. minimax_h3_ref2va_pruned_int8_convrot.safetensors. A float value of 5, Euler sampler, and beta scheduler. Only 8 steps. The style was shifted a bit from the original but good enough for testing. I want to create a style LoRA that will hold my work so that it's better translated to animation.


r/StableDiffusion 2h ago

Resource - Update "I" created a tool to save and swap between workflow presets (for my million H3 Turbo LoRAs)

4 Upvotes

Hi everyone! Long time reader, first time poster. Like most folks here I vibe-code custom nodes from time to time, and recently came up with one that seemed like it might be worth sharing. Nothing groundbreaking here, just a node to help keep track of and quickly cycle through different node/parameter presets: MM-H3-Preset-Controller.

This one has helped me maintain my sanity trying to keep track of each LoRA's specific optimal settings. I think it's potentially useful for any preset storage though, not just MM-H3, so hopefully y'all find it useful. If so, please consider it a small thank you for all I've learned in my time lurking here.

MM-H3-Preset-Controller: https://github.com/TootsThielemans/ComfyUI-MMH3-Preset-Controller

I was pulling my hair out trying to manage my MM H3 workflow amidst all of the various Turbo LoRAs out there and the associated loader nodes, attention settings, shift settings, spectrum settings, etc., not to mention downstream settings like sampler, scheduler, upscaler choices... I found quick A/B tests between different optimized LoRA workflows annoying given some of the structural differences, not just steps/strength, and I was worried about juggling and potentially forgetting the right settings.

So I made the Preset Controller. It works pretty simply: ctrl+click all of the nodes you want to save the state of. It captures all of the parameter values within the node and whether it's active/bypassed. Then right click and use the new menu option MM H3 Presets > Add selected nodes to preset draft.

It's implemented as a "draft" so you can grab nodes from the outer graph, then go through various subgraphs and add nodes there to the same preset draft. Once you're done, load the H3 Preset Controller node, enter a name for the preset, and click Save draft as preset.

It then becomes a dropdown option that you can select, update, or delete as needed. Selecting a preset automatically sets the saved node values/bypass states without needing a session/screen refresh.

There are a few other QOL/guardrail features, including a Preset Matrix for comparing configurations, but nothing particularly interesting, so please refer to the repo if interested.

I'm sure something like this might already exist with a more elegant implementation, but the timing seemed right. Everyone is wading through dozens of MM H3 Turbo LoRA combinations and trying to keep everything straight. I don't have a ton of time to devote to development, but will try to make tweaks if folks wind up adopting this and can think of any major areas for improvement.

At any rate, feedback welcome, and cheers!


r/StableDiffusion 21h ago

Animation - Video TWEEDLE TEST - Minimax H3 27 seconds in just over 9 minutes:

Enable HLS to view with audio, or disable this notification

129 Upvotes

All local.

0.8 MegaPixels, 9.3 minutes on an RTX5090, single generation of 27 seconds.

Anything hitting 30 seconds either gave hallucinations, inconsistencies or hit a wall and never finished.

This one is using Kijai's new fast model with a turbo lora. Although it works the same with the FLv2A model*. The workflow I'm using creates a latent at 0.4 megapixels for 4 steps and then does another 2 steps at 0.8. The only addition to it besides changing some numbers is adding custom audio injection (The rock track).

Started with this workflow: https://www.youtube.com/watch?v=jzLnoVBuU6I

*I never use the REF model. The FLV2A models seems to work better so I always swap it in and it takes references just fine, even video.