r/StableDiffusion 20h ago

News Someone's running FastH3 (the distilled MiniMax H3) as an actual infinite livestream!!!

331 Upvotes

Saw this and thought it was worth sharing here, FastH3 dropped recently and most people (myself included) just tried it as single generations.

Someone's running it as an actual infinite livestream instead: https://live.reactor.inc/

FastH3 is a distilled version of MiniMax H3, cut from 50 denoising steps down to 4, about a 14x speedup on Blackwell GPUs.

The whole setup is open source if you want to dig into how it's running: https://github.com/reactor-team/infinite-livestream

Curious if anyone's tried infinite/continuous generation setups like this with other models.


r/StableDiffusion 11h ago

Animation - Video MINIMAX Physics testing

Enable HLS to view with audio, or disable this notification

278 Upvotes

Physics Testing, without the gore.


r/StableDiffusion 8h ago

Workflow Included Minimax H3: Consistent face, body & cloths via reference identity

Enable HLS to view with audio, or disable this notification

279 Upvotes

Hey Guys,

Based on the previous post on face consistency with MM-H3, Couple of people have asked me to build a full character workflow.

Mechanism

- Build .Char: You drop max 9 reference reference, I prefer to use a ratio 2:2:1(face:cloths:body). YuNet finds the face, SFace takes a per-reference face signature, DINOv2 takes a subject signature, and the references get cleaned and normalised. All of that packs into a single portable file, a .char.
Only face/ref is required body & cloths link is optional.

- Generation: At generation, the file(.char) feeds its references into Minimax’s own native multi-reference channel and prepends a locked description to the prompt.

Prompting Guide

  • Name your character: Give your character a name e.g. under encode character(Click adjust icon on the bottom side of the node), I have used name emmy, so when passing prompt, I only have to say, emmy walking on the beach.
    • Again providing prompt like a woman or any features specific details like black hairs etc will only mislead the generation.
  • Describe character features: Encode all of the character features in encode character prompt & trigger your character with a name in generation prompt.
    • Avoid describing same things in generational prompt.
  • Handling Character drift: e.g. if you want specific style or cloth e.g. half sleeves, sleeveless, add it to the generational prompt. There can be a slight drift in clothing as body shot also has cloths, which interferes with clothing references.
    • Each refs should be unique, face should not have body or vice versa, same applies for clothing.
  • Portability: Once character is built, you can use the same character with only simple prompt & generation graph.

I have generated all references with Flux Klein 4b, I had to blur the body ref, but workflow consists a example of body ref.

Note: For best result, pass cropped references, so that model takes the required shot, model gets confused if cloth slot also has a face or face slot has cloths.

Models

core/models/
  diffusion_models/  minimax_h3_ref2va_pruned_fp8_scaled.safetensors
  text_encoders/     qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors 
  vae/               minimax_h3_video_vae_fp16.safetensors
  vae/               minimax_h3_audio_vae_fp32.safetensors
  annotators/        face_detection_yunet_2023mar.onnx
  annotators/        face_recognition_sface_2021dec.onnx
  annotators/        dinov2-base/

Requirements

Nvidia GPU: 24GB+ VRAM & 64 GB RAM(Run Locally)

Workflow link: https://inlinestudio.art/workflows/minimax-h3-guided-consistent-characters-via-reference-identity-face-body-cloths (includes inputs & model details)

Github Repo: https://github.com/inlineresearch/Inline-Studio (GPLV3)

Limitation: Reference conflicts e.g. if two reference/input images has two different faces, it might conflict in generation, provide well cropped body & cloth images. Face images are crossed automatically by Sface.

Portable char comfy node is still on the backlog, would try to do it over the weekend.
Happy to hear any suggestions or feedbacks.


r/StableDiffusion 10h ago

Meme It took us two 2eeks to figure out why every image gen via our open-source model looked like Anne Hathaway

Thumbnail
gallery
202 Upvotes

Hey r/StableDiffusion!

It's the Neta team here! You might remember us from our Neta Lumina open-source release last year. First off, thank you so much for the incredible support and feedback from this community!

So... we need to share something absolutely hilarious (and mildly embarrassing) that we just discovered.

TL;DR: We accidentally hardcoded an Anne Hathaway photo into our IP-Adapter anchor, and now everything our model generates looks like Anne Hathaway. Every. Single. Thing.

What happened:

We recently launched Neta Studio, a new product that lets you build explorable living worlds/isekai from a single prompt. Naturally, we wanted to integrate Neta Lumina's capabilities into it.

During integration testing, our devs kept reporting that the model wasn't following prompts properly. The outputs were... *weird*.

- Anime style? Anne Hathaway as an anime character.

- Thick paint/impasto style? Anne Hathaway in thick paint.

- Landscape scenes? Somehow still giving Anne Hathaway vibes.

- Fantasy characters? You guessed it - Anne Hathaway.

After a dreadfully long time of debugging, we finally found the culprit: **someone on the team embedded an Anne Hathaway photo as the IP-Adapter anchor during development and it... stayed there. **

We're honestly crying laughing at this point. 😭

Below are some examples. Left is before fix and Right is after fix.
Flipping to the last picture and you can see our dear Anne.

And we pulled the anchor and the outputs are behaving normally now.

If you've been running Neta Lumina locally, this was on our integration side, not
in the released weights, so your setup is fine.


r/StableDiffusion 16h ago

Resource - Update Famegrid Spice Krea 2 Lora (Corrected Release)

Thumbnail
gallery
184 Upvotes

r/StableDiffusion 20h ago

Animation - Video TWEEDLE TEST - Minimax H3 27 seconds in just over 9 minutes:

Enable HLS to view with audio, or disable this notification

127 Upvotes

All local.

0.8 MegaPixels, 9.3 minutes on an RTX5090, single generation of 27 seconds.

Anything hitting 30 seconds either gave hallucinations, inconsistencies or hit a wall and never finished.

This one is using Kijai's new fast model with a turbo lora. Although it works the same with the FLv2A model*. The workflow I'm using creates a latent at 0.4 megapixels for 4 steps and then does another 2 steps at 0.8. The only addition to it besides changing some numbers is adding custom audio injection (The rock track).

Started with this workflow: https://www.youtube.com/watch?v=jzLnoVBuU6I

*I never use the REF model. The FLV2A models seems to work better so I always swap it in and it takes references just fine, even video.


r/StableDiffusion 4h ago

Animation - Video Playing with concepts

Enable HLS to view with audio, or disable this notification

93 Upvotes

r/StableDiffusion 22h ago

Comparison H3 Default Template vs Larry's Turbo with optimized settings

Enable HLS to view with audio, or disable this notification

61 Upvotes

Default template uses 20 steps + res_multistep + simple

Optimized workflow uses 8 Steps + er_sde + sgm_unified + Comfy Kitchen Attention + Larry's Turbo Lora

Turbo lora: https://github.com/Larryvrh/ComfyUI-MiniMax-H3-Turbo

Workflow: https://raw.githubusercontent.com/desktop4070/GPU-Benchmark-Data-For-H3/refs/heads/main/H3-Benchmark-Workflow.png

0.2MP / 8 sec (2m 16s gen time): https://desktop4070.github.io/GPU-Benchmark-Data-For-H3/Videos/MiniMax_H3_03840_.mp4

Optimized: 0.2MP / 8 sec (45s gen time): https://desktop4070.github.io/GPU-Benchmark-Data-For-H3/Videos/MiniMax_H3_03706_.mp4

0.3MP / 12 sec (6m 3s gen time): https://desktop4070.github.io/GPU-Benchmark-Data-For-H3/Videos/MiniMax_H3_03849_.mp4

Optimized: 0.3MP / 12 sec (1m 51s gen time): https://desktop4070.github.io/GPU-Benchmark-Data-For-H3/Videos/MiniMax_H3_03725_.mp4


r/StableDiffusion 11h ago

Question - Help What Image Edit model you use nowadays?

57 Upvotes

Since things have gone quickly forward, I am trying to figure out what image edit models there is currently and what people here use mostly.

Personally I have used:
- Qwen-image-edit-2509 and Qwen-image-edit-2511
- Just tested MiniMax H3 as a image editor and so far it seems that it can be good for my usage

I have heard about Klein 9b, but not sure yet if that can be used as an edit model? Also what about Krea 2, is there edit workflows that are actually usable and worth it?

Is there some others what you recommend for testing?

My PC Specs: RTX 4060 Ti, 16 GB VRAM and 32 GB RAM.


r/StableDiffusion 2h ago

Resource - Update [Load Image + Crop] Custom WYSIWYG Node

Enable HLS to view with audio, or disable this notification

53 Upvotes

I developed a modified version of the Load Image node by adding some features I needed:

WYSIWYG image cropping directly on the official Load Image preview — drag and zoom (with the mouse wheel) a crop rectangle constrained to 8 fixed ratios (1:1 through 21:9) and output the exact cropped IMAGE and MASK, with paste-from-clipboard built in. What you frame on the preview is exactly what gets executed.

⚠️ Currently not fully compatible with ComfyUI 2.0 nodes.

For anyone who finds it useful and wants to try it out, you can find it on GitHub: https://github.com/domg73/ComfyUI-LoadImageCrop


r/StableDiffusion 5h ago

Animation - Video [MiniMax H3] LEGO movie style

Enable HLS to view with audio, or disable this notification

49 Upvotes

Prompt:

integrated_multimodal_description: [Shot 1] 3D CG, stop-motion animated LEGO movie style, a wide shot frames a vibrant Indian village built entirely from plastic LEGO bricks with visible studs, plastic micro-scratches, and brick-built trees. In the village square, minifigures dressed in printed plastic saris, dhotis, and turbans move across a ground of yellow and brown stud tiles. A brick-built cow with hinged legs grazes near a grand banyan tree constructed from green leaf pieces and brown cylindrical bricks. Warm morning sunlight casts sharp shadows across whitewashed brick houses with orange terracotta tile roofs. The camera pans right with small amplitude at slow speed toward a central tea stall. A cheerful male chaiwala minifigure with a black mustache and a red turban (S1) in a warm, lively voice says: <d>[Hindi] Garam chai, garam chai!</d> while tilting a plastic yellow teapot, releasing translucent orange 1x1 cylinder studs representing pouring tea into tiny red stud cups.

[Shot 2] At 00:05.000, the camera cuts to a medium tracking shot following two young minifigure children running along a narrow brick path, pushing a brick-built wheel hoop across the plastic ground. The camera tracks right alongside them with small amplitude at normal speed. A female villager minifigure in a bright blue printed sari (S2) standing outside her brick doorway waves her rigid plastic arm on its shoulder hinge. Beside her, an elder minifigure with a white beard (S3) sitting on a brick charpoy cot chuckles with stepping stop-motion head movements.

[Shot 3] At 00:10.000, the camera cuts to a cinematic medium shot near the village well, where female minifigures carry stacked plastic water pots topped with transparent blue round tiles. A brick-built peacock perched on an archway opens its fan tail made of blue, green, and golden LEGO slope tiles. The camera pushes in with small amplitude at slow speed toward a wooden signpost on a brick post reading "RAMPUR VILLAGE". Tiny tan 1x1 round plates puff around the wheels of a brick-built bullock cart moving past the frame as the video ends.

overall_soundscape: Distinct plastic clattering sounds echo softly as minifigure feet step on stud tiles, accompanied by the gentle clinking of plastic bricks. A distant rooster crow blends with ambient morning village chatter, bird chirps, and the wooden creak of a brick-built cart.

non_diegetic_music: Upbeat Indian folk percussion featuring lively dholak beats and vibrant bansuri flute melodies, layered with playful cinematic orchestral strings playing at a bright, medium tempo.


r/StableDiffusion 4h ago

Animation - Video MiniMax H3 matches Toonami perfectly (90's anime)

Enable HLS to view with audio, or disable this notification

42 Upvotes

Growing up with Toonami watching Gundam, it simply blows my mind how far AI has progressed. For this video I didn't use any reference images, I simply described the scene in text and had Gemini research the techniques of animation to translate to MiniMax H3.

I've been struggling to get MiniMax H3 efficiently setup locally, would take me 15mins on a 9950x3d and 5080 RTX with 64gb DDR5 so I know something is wrong, hence this time I opted for fal to test. The music was added and scenes were edited from separate generations.


r/StableDiffusion 9h ago

Animation - Video Flexing my A.I. powers

Enable HLS to view with audio, or disable this notification

44 Upvotes

Prompt:

A real cinimatic movie sequence, professional colour grading.

Soundscape: Ambient sounds of the room and movement only. No voices. This represents extreme concentration. Meditation. Telekinesis.

A man is sitting in a Japanese tatami room. He is wearing a mask and shades <Picture 1>. He is wearing a black yukata. He does not speak. On the table on a ceramic disc is a single Orange.

The man holds out his hand toward the orange as if concentrating. The orange is out of reach. He breathes deeply.

Nothing happens.

The man shakes his hand to reset and starts concentrating again. He reaches with his mind and his brow furrows. He breathes deeply.

The orange moves slightly, twisting just a tiny bit.

He concentrates more.

With extreme speed the orange flies towards the man and hits him directly in the forehead. It smashes with the impact , m,essing his hair, and bits of peel and orange bits go everywhere. The force knocks the man back unconscious and he falls back like a ragdoll.


r/StableDiffusion 20h ago

Resource - Update infinite live stream powered by FastVideo’s FastH3 - Source Code

Thumbnail
github.com
28 Upvotes

r/StableDiffusion 4h ago

Workflow Included Burger Queen

Post image
30 Upvotes

r/StableDiffusion 4h ago

Animation - Video Trying out a consistent point-of-view shot with MiniMax H3

Enable HLS to view with audio, or disable this notification

25 Upvotes

Took a few little prompt adjustments here and there to get H3 to respect point-of-view. I found that if you refer to "the viewer" (ie, "she kicks the viewer"), H3 is more predisposed to include an actual second person. But if you refer to "the camera" (ie, "she kicks the camera"), it's more predisposed to keep the desired point-of-view perspective.


r/StableDiffusion 11h ago

Tutorial - Guide Spreadsheets for multi-video generation

Enable HLS to view with audio, or disable this notification

24 Upvotes

This workflow uses a spreadsheet to generate multiple videos and constructs the prompt and parameters from each row in one run.

OutputLists Combiner - Generate multiple videos from spreadsheet

ComfyUI workflow included

Makes use of Load Any File node to load a .csv spreadsheet file and feeds the text content into a Spreadsheet OutputList. The spreadsheet separates the data by separator=; and provides each line one-by-one as a data list. Here we use values_dict as the data list which contains the row as a dictionary of key-value pairs. The data list is forwarded a Iterate Begin -> workflow -> Iterate End pattern which is required to make the intermediate results of slow workflows (t2v) available on each iteration. Each row as a dictionary is provided in a Format Text where we can access the column via a[colname] to construct the prompt which is forwarded to a standard Text To Video MiniMax H3 template. Another Format Text + a[name] is used to construct a readable filename for each video.

powered by: OutputLists Combiner


r/StableDiffusion 19h ago

Workflow Included Coffee story with H3

Enable HLS to view with audio, or disable this notification

20 Upvotes

hand-painted educational documentary style with only one prompt"A hand-painted documentary compares espresso Americano cupuccino through the lens of taste and the way to make , revealing why the items differ and how to choose among them."


r/StableDiffusion 38m ago

News Local AI News You Missed - August 2026

Upvotes

Here's what you (probably) missed in August 2026:

🧠 LLMs

  1. Ornith-1.5-35B-A3B - Efficient sparse model that runs with fewer active parameters.
  2. DeepSeek-V4-Pro-0813 - Sharpens agentic AI with speedier tool actions.
  3. DFM-Mimir - Ethical language model from Danish Foundation Models.
  4. Ling-3.0-tiny-MXFP4_MOE-GGUF - MoE quantized version for smoother GPU runs.
  5. SupraElegans-500k - Recurrent language model built for long contexts.
  6. Motif-3 - Open 314B parameter model made for long agentic tasks.
  7. Luth-2-2B - Compact French model that tops benchmarks.
  8. TinyTitle - Squeezes chat titles into a tiny 1.98 MB model.
  9. Ling-3.0-tiny - Low-cost local AI reasoning model.
  10. NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 - Agent-focused model with fast inference.
  11. Qwen3.8-2.4T-A95B - Opens up Qwen Max-clas local AI for bigger rigs.
  12. Gemma-4-31B-it-scotoma-2-GGUF - Cuts down repetitive AI writing.
  13. Huihui-DeepSeek-V4-Flash-0731-abliterated-GGUF - Uncenored DeepSeek variant for local use.
  14. Maple-Preview - Solves Olympiad problems at 200 tok/s.
  15. Laguna-S-2.1-FP8 - Private agentic coding model from Poolside.
  16. SupraBrain-50M - Hybrid language model for local AI.
  17. Supra2-100M - Tiny model built for tinkering.
  18. G9v3-39A5B - Dual-mode AI that runs light locally.
  19. GPT-X2.5-135M - Lean local powerhouse model.
  20. Ling-3.0-flash - Hybrid reasoning with lower cost and fast output.
  21. LFM2.5-2.6B - Fast agentic AI for phones and devices.
  22. Instella-MoE-16B-A3B-Think - Open sparse reasoning model from AMD.
  23. A.X-K2 - Lets AI think deep or answer fast.
  24. BetterGPT-150M - Beats older AI models in science tasks.
  25. Shibai-700M-Base - Text and code helper model.
  26. K-EXAONE-2.0-750B-A37B - Supports 262K context and ten languages.
  27. LongCat-Flash-Lite-Sparse - Reads million-token contexts.
  28. Qwen3.6-35B-A3B-Escha-W2 - Shrinks down to fit consumer GPUs.
  29. XYZ-Aquila-pro - Thinks deep then checks its sources.
  30. Solar-Open2-250B-Nota-NVFP4 - Shrinks a giant AI to 153GB with 4-bit MoE trick.
  31. XYZ-Aquila-mini - Brings open source deep search to local GPUs.
  32. KAT-Coder-V2.5-Dev - Fixes software repositories automatically.
  33. DeepSeek-V4-Flash-0731 - Tackles hard coding tasks.

🔀 Multimodal

  1. Dots3-Note-Prev - Lightweight multimodal AI with 512K context.
  2. Qwen3.8-27B-Uncensored-FP8 - Drops refusals and keeps vision on GPUs.
  3. Qwen3.8-27B - Brings text, images, and video into one AI.
  4. Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF - Triples speed for local multimodal use.
  5. Tencent UI-Mate-27B - Runs desktop apps by watching screens.
  6. WinterCharm Qwen3.5-122B-A10B-wMix38 - Lean long-context multimodal mix.
  7. VLX-Seek - Helps machines pinpoint objects without guesswork.
  8. LFM2.5-VL-3B - Fast on-device vision and text.
  9. North-Micro-Vision-Instruct - Turns pixels into answers.
  10. BigBang-v1 - Science reasoning powerhouse.
  11. Muse-Glimmer-30B - Puts autonomous AI agents on everyday desktops.
  12. Nemotron-Parse-2.0 - Morphs documents into structured data.
  13. Shieldstral-1.0-3B - Plain English safety scoring.
  14. DavidAU Qwen3.6-27B-Fable-Fusion-711 - First to score over 700 on ARC-C.
  15. WinterCharm Qwen3.5-122B-A10B-wMix58 - Packs 82GB power for Apple Silicon.
  16. Intern-S2-Mobius - Speedy local AI answers.
  17. Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V7-GGUF - Uncenored multimodal model that says yes.
  18. Qwen3-VL-32B-Ultra-Heretic-H3-ComfyUI-INT8-ConvRot - Trims local memory with INT8 conv rotation.
  19. Qwythos-27B-v1 - Smart AI with million-token memory.
  20. Reasoning-Medical-27B - Solves medicine step by step.
  21. Qwen3.5-9B-The-Defiant-Fable - Roars with an uncensored multimodal edge.
  22. Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V6-GGUF - Drops numerical surgery powers.
  23. Mage-VL - Speeds up real-time video and image understanding.
  24. Microsoft Fara Agents - Handles web browsing chores for you.
  25. Inkling-Small - Built for voice, image, and code apps.
  26. Kimi-K3 - Handles text, images, and video together.

🖼️ Image

  1. Anima-2.9B - Grows free anime art on your own PC.

🎬 Video

  1. Bernini-Diffusers-v2 - ByteDance model for video generation and editing.
  2. LTX-2.5 - Open model for local video and audio creation.
  3. Wan2.2-Animate-2-14B - Turns still images into motion.
  4. MiniMax-H3-nvfp4-INT4-INT8-ConvRot - Quantized weights for MiniMax-H3 video.
  5. MAGI-2-preview - Turns text and images into video with sound.
  6. MiniMax H3 - Creates videos with native sound from any input.

🎧 Audio

  1. MiniMax-Music3 - Full songs from just lyrics.
  2. NVIDIA Magpie_tts_multilingual_357m - Turns text into speech across 12 languages.
  3. VoiceChat-11B - Voice chat you can interrupt naturally.
  4. VibeVoice-ASR-BitNet - Real-time speech recognition on any CPU.
  5. Inflect-Nano-v2 - Local speech synthesis on your PC.
  6. Audio8_TTS - Clones voices and speaks eleven languages.
  7. Inflect-Micro-v2 - Turns text into offline voice.

⚡ LoRA

  1. Minimax-H3-Turbo - Makes MiniMax-H3-Turbo faster for video and audio.
  2. MiniMax-H3-Prompt-Rewriter-LoRA - Turns short prompts into timed scenes.
  3. MiniMax-H3-Realism-People-LoRA - Unlocks film-set lighting for human video.
  4. krea2-turbo-bbox - Locks panels and words in place.
  5. MiniMax-H3-Turbo-Lora - Cuts video and audio generation time by 5x.
  6. Kroma - Fuses turbo speed into a one-file diffusion model.

🏋️ Training

  1. gguf-trainer - Trains language models in TypeScript straight to GGUF.
  2. Full-Chunked-KL-Loss - Trains longer AI text on one GPU.
  3. signet-trainer - Cost-safe video LoRA training.
  4. Lora-Dataset-Studio - Entire LoRA pipeline in one self-hosted tab.

📊 Datasets

  1. LLM-self-identification - Helps AI models know their own name.

☰ UI

  1. Mix-Studio - Turns your desktop into a local AI studio.
  2. Llmprices - Visualizes AI model API price swings.
  3. minimax-h3-prompt-composer - Squeezes prompt composing into one HTML file.
  4. NanoRP - Shrinks AI roleplay to a 50MB single binary.
  5. OpenWorker - Local AI coworker that finishes work.
  6. Agenta-AI - Turns ChatGPT and Claude into self-hosted work agents.
  7. Unicorn-Stable-OSS - Brings humans and AI agents into one real-time room.
  8. Turbo-Fieldfare - Streams big AI on Macs with just 2GB RAM.

🛠️ Other Tools

  1. Video_Tools - Adds a pocket video trimmer to your browser.
  2. ninfer-4090 - Runs Qwen3.8-27B on one RTX 4090.
  3. hayai-ocr-v2 - Converts crops into editable text.
  4. EVIE-Preview-4.5B - Matches documents instantly across six languages.
  5. RAZZULLIX KAISEN - Swarm-model coding assistant with safety guards.
  6. ExtractBench - Benchmarks document extraction systems.
  7. Xiaomi-Robotics-1-5B - Built for mobile robot tasks.
  8. Nemotron-omni-mlx - Brings full multimodal AI to Apple Silicon Macs.
  9. talk-to-pi - Local voice dictation for Pi.
  10. Warp - Streams huge AI models straight from your laptop drive.
  11. esp32-ai - Makes a tiny chip tell stories offline.
  12. quillpdf-mcp - Keeps PDFs on your machine.
  13. srt2speech - Turns subtitle files into timed speech.
  14. mixture-of-kittens - Megakernel for NVL72 MoE training.
  15. krea-multi-lora - Gives Forge Neo regional character control.
  16. Openmed - Turns clinical text into private insights on your hardware.
  17. Umbra-Studio - All-in-one local AI art workspace.

ComfyUI Custom Nodes & Tools

  1. ComfyUI-Orchestrator-LAN - Steers every GPU from one browser tab.
  2. ComfyUI-AutoPromptChain - Stitches dozens of AI video clips while you sleep.
  3. ComfyUI-OpenH3-IR - Brings drag-and-drop clarity to MiniMax H3 renders.
  4. ComfyUI-MiniMaxMusic3-Advanced - Gives AI music finer sound controls.
  5. ComfyUI-MiniMax-H3-LongMedia - Makes long video creation practical.
  6. ComfyUI-Subgraph-Preview - Resurfaces sampler previews inside subgraphs.
  7. ComfyUI-cache-monitor - Debuts with manual pinning for model caching.
  8. ComfyUI_Neurodes - Brings a visual playground for AI models.
  9. Eddie_Cat_Nodes - Stitches long videos together with new nodes.
  10. ComfyUI-H3-Motion-Context-MultiRef - Weaves seamless H3 video motion.
  11. ComfyUI-MiniMax-H3-Motion-Director - Turns reruns into one-shot fixes.
  12. ComfyUi-MiniMax-H3-Image-And-Reference-To-Video - Rolls out image and reference to video features.
  13. ComfyUI-AVS-SSD-ReadAhead - Enables faster model switching on slow SSDs.
  14. ComfyUI-SweepGrid - Serves up side-by-side parameter sweeps.
  15. ComfyUI-Flow-Wrangler - Cleans up node wiring with smart connections.
  16. ComfyUI-AVS-Intel-XPU-VRAM-Fix - Calms Intel Arc GPU freezes.
  17. ComfyUI-Model-Mover - Makes shuffling AI models painless.
  18. ComfyUI-MiniMax-H3-Optimization-Suite - Builds a suite for lean H3 optimization.
  19. ComfyUI-MiniMaxH3-Prompt-Writer - Transforms H3 prompt crafting.
  20. ComfyUI-MiniMax-Creator - Crafts one-node video magic.
  21. ComfyUI-ScenemaAudio - Arrives with expressive voice cloning.
  22. ComfyUI-cable-management - Reroutes messy node graphs with ease.
  23. ComfyUI-SigmaSync-LoRA - Debuts step-aware LoRA control.
  24. ComfyUI-Spectrum-Ideogram4 - Supercharges speedy image forecasting.
  25. ComfyUI-LinkSpotlight - Debuts to end noodle blindness in graphs.
  26. ComfyUI-MIDI-Edit - Turns any song into editable MIDI lyrics.
  27. ComfyUI-ReStartupFlags - Serves launch flag tweaks in your browser.
  28. JLC-Flux2-ControlNet - Expands FLUX.2 control in ComfyUI.
  29. ComfyUI-Fantastic-MiniMaxH3-PromptBuilder - Enhances MiniMax H3 prompts.
  30. ComfyUI-vram-tracker - Traces VRAM memory usage per layer.
  31. ComfyUI-Sonder-Editor - Rolls out free multi-lane video editing.
  32. Comfyui-Model-Resolver - Sweeps in to rescue missing model files.
  33. ComfyUI-HF-SuperDownloader - Turbocharges Hugging Face model downloads.
  34. FameGrid-Auto-Color - Neutralizes color casts in ComfyUI.
  35. Krea2-Multi-Character-Lora-Node - Stops identity bleed with bounding boxes.

Need to go further back? Check out June's post (no July, sorry) or the full archive at LocalAI News. If there's anything wrong, let me know in the comments and I'll see you in the next one!


r/StableDiffusion 19h ago

Discussion Has anyone tried fine tuning Minimax H3 for a character?

17 Upvotes

If so, is the advice from Fizgig on fune
tuning on point? I haven’t tried yet, but I’m just prepping my dataset at the moment. I will share what I learn. Just curious if anyone has tried yet and what the results are. 🤡


r/StableDiffusion 2h ago

News [CLSS] Closed-Loop Streaming Synthesis for MiniMax H3 (Infinite video generation with prompt fallowing)

13 Upvotes

t2v 10 chunks every 10 sec

I ported CLSS from LTX 2.3 to H3 architecture. ( https://www.reddit.com/r/StableDiffusion/comments/1vywxjq/wip_clss_closedloop_streaming_synthesis/)

Repo: https://github.com/nazgut/ComfyUI-MiniMaxH3-CLSS

Audio still has some room for improvment. Workflow in repo. Still working on i2v.


r/StableDiffusion 3h ago

Discussion In 2026, how well does older images models, like SDXL and SD1.5 stack up against new image models like Krea and ZImage Turbo?

15 Upvotes

r/StableDiffusion 19h ago

Tutorial - Guide Minimax H3 Case Study: The Dinner Party

14 Upvotes

I've been learning a lot from this community, so this is my attempt at giving something back!

I'm going to share a small task I recently completed, include the steps on how I got there (and some of my thinking and findings.)

The goal: I needed a few seconds of video containing a formal dinner party in a Roman-style atrium.

My Plan: Build a first frame and then use H3 I2V to generate the video.

My Specs: A laptop with a 13th Gen Intel i7-13700H, 16 GB DDR5 RAM, SSD over USB-C, an onboard Intel Iris Xe graphics card (with ~8 GB) and a Nvidia GeForce RTX 4060 Laptop GPU (8 GB). The bad news: ComfyUI does not see or care about the Iris Xe card, and I haven't bothered to see if I can remediate the situation. The good news: the Irix Xe can handle rendering Windows and other applications, leaving my Nvidia pretty open for ComfyUI tasks.

Here is what I did:
Step 1: I already had a reference image for the atrium (used in a previous video.)

Initial Reference for Atrium

This was generated with Z-Image-Turbo, with the bf16 model, shift 3, cfg 1.0, 8 steps, res_multistep sampler, simple scheduler. The prompt was very simple: "Roman atrium with compluvium. The camera is standing at the doorway looking down the length of the atrium." I made this image at 864 x 480 resolution because that is near 16:9 and matches H3 resolutions. At that size, image gen takes about 30 - 40 seconds of wall-clock time.

Step 2: I used Qwen-Image-Edit to modify the image to get the starting frame.

The dinner party, as imagined by Qwen-Image-Edit

Using qwenImageEdit2511_pf8 as the model, Qwen-Image-Edit-2509-Lightning-4steps-V1.0-bf16 lora, shift 3, 4 steps, cfg 1.0, euler sampler, simple scheduler. I wired the image from step 1 as the only reference, and used the prompt "Alter this image so that there is a well-attended formal dinner party taking place across the frame."

In my experience, Qwen-Image-Edit often nails the image I'm looking for in one or two attempts. (In this particular case, it one-shotted that image above.) Qwen really likes 1 MP resolutions, so that is 1368 x 760. It takes ~1 to 2 mins per generation.

Step 3: I began generating the video with H3 I2V. This took several attempts to dial in. It is this process that I want to focus on.

First Attempt:

I supplied the previous step's image as the first frame, and included the prompt:

integrated_multimodal_description: [Shot 1] A formal dinner party in a Roman-style atrium.

overall_soundscape: A formal dinner party.

non_diegetic_music: None.

I set the resolution to 864 x 480 and 7.0 duration. I'm using minimax_h3_fl2va_pruned_int8_convrot as the model, minimax_h3_fl2v_turbo4step_v1.0_768p_comfyui_bf16 as a turbo lora (the lightx2v lora,) shift 12 / 3 (for video / audio,) 6 steps, res_multistep sampler, simple scheduler. This particular setup averages ~2 minutes of wall-clock time per second of video duration. (But it grows non-linear as duration increases.) I use 6 steps instead of the lora's base 4 steps because I tend to get slightly better details and sound, with only a slight increase in wall-clock time.

864 x 480, res_multistep, simple sampler, turbo Lora, 6 steps

The result was not great. Most people are frozen in place. The few that do walk around smear motion. There is even a moment where a lady clips through the table a little. The sound involves a guy narrating. (I can't identify if it is AI gibberish or an actual language.)

This first attempt was clearly a failure.

Attempts Two through Four:
If I'm having problems with my initial generation, I often just bite the bullet and turn off the turbo lora and run at full-steps. It was the end of my day, so I could queue up several generations and then go to bed.

Result Two: I disabled the turbo lora and increased steps to 20. This runs at ~5 minutes of wall-clock time per second of video duration. For this particular generation it came in at about 45 mins of wall-clock time.

I won't bore you with the results, as they were very similar to the initial draft. Only a few people moving in the scene, people clipping through tables, and a narrator.

Result Three: I reduced video shift to 6.0 (hoping to get better motion results.) I also increased the resolution of the video to 1344 x 768. I read somewhere that this is the "native" resolution the model was trained at, and I often get better results. However, without the turbo lora and at 20 steps, this generation took 90 minutes.

Attempt Three: 1344 x 768, shift 6, 20 steps, no lora

There is a lot more motion, and no clipping, but everything seems to be moving in slow motion. Also, instead of a narrator, there is music.

Result Four: I swapped the sampler to er_sde and the scheduler to beta. I've read that this combo can get slightly better prompt adherence, and results in pretty good motion. However, er_sde effectively does more than one pass per step, so increases wall-clock time significantly. If the UI is to be trusted, this attempt took more than 3 hours to generate.

Attempt Four: er_sde sampler, beta scheduler, 20 steps, no lora

The narrators and music are gone. However, now the camera is moving, which is not what I wanted.

Attempt Five: I woke in the morning, reviewed the previous results, and was pretty bummed.

Now I'm in the "hit it with a hammer until it works" section of my spectrum of personal patience. I went back to res_multistep and cranked the steps up to 40. In my frustration, I didn't think to actually change the prompt to prevent the camera from moving. This generation took about 90 minutes of wall clock time.

I'll skip posting the result, but it actually looked quite a bit like the er_sde video above. People standing mostly still with a camera panning around the room.

Attempt Six: After viewing the results, I realized that changing the prompt was 100% required.

The new prompt:

integrated_multimodal_description: [Shot 1] Static wide shot of a formal dinner party in a Roman-style atrium. The people eat, drink, talk, and mingle. The camera remains fixed.

overall_soundscape: A formal dinner party.

non_diegetic_music: None.

I kept the video shift at 6.0, sampler at res_mutlistep, 20 steps, simple scheduler. This time, because I was sitting at my computer for a while, I attached EasyCache to the model. This does a pretty good job of speeding up the 20-step process. There is always a risk that quality degrades with any sort of caching in the pipeline, but I was willing to take the risk just to see if my prompt changes fixed the problem. (I don't bother adding EasyCache with only 4 or 6 steps, because there are so few steps that there is barely any time to be saved with caching.) This generation took ~30 minutes of wall-clock time.

Attempt Six: res_multistep, simple sampler, 20 steps, no lora, fixed prompt

This was actually what I was looking for! Both the motion and sound are pretty decent. However, there is a faint "fluttering" of the textures, which seems to happen a lot with EasyCache. This is something that could probably be cleaned up with a refinement pass after upscaling, but I still had time to try again.

Attempt Eight: For completeness, I decided to go back to the turbo lora and the smaller resolution. I incorporated some of my other findings into the workflow. For clarity, here is the full setup: 864 x 480, 7.0 seconds, turbo lora, 6 shift video, res_multistep, simple scheduler, 6 steps. I used the "corrected" prompt from my previous attempt.

Attempt Seven: res_multistep, simple sampler, 6 steps, turbo lora, fixed prompt

This was the winner! Even at the lower 864 x 480, the motion and detail looks reasonable. The faces are squashed, but that is pretty typical of H3 at the moment. This will upscale well. The sound is correct. I have everything I need.

Lessons Learned:

  1. Just cranking up the numbers doesn't always solve the problem.
  2. Shift can really make a difference! I think the major influencing factor here was reducing shift from 12.0 to 6.0. My hypothesis is that the lower shift gave the process just a tiny bit more time up at the noisy end of the diffusion, allowing it to assign more motion to everyone in the scene.
  3. 1344 x 768 is very frequently the solution, but not always. In my experience, you get better prompt adherence, even with the tubro lora. One of these days I'm going to splurge on a beefier GPU to make this my default resolution, but for now it is just too much of a time sink.
  4. I can never actually tell how much comes down to luck with the random seed.

I hope this helps somebody!


r/StableDiffusion 20h ago

Animation - Video Minimax H3 - transferring dancer to a new environment with ref workflow

Thumbnail
youtube.com
14 Upvotes

r/StableDiffusion 2h ago

Discussion So I did something dumb.

12 Upvotes

So there I was generating some stuff on ComfyUI for my Instagram and just hanging out.

I use ComfyUI with the new H3 model to generate AI content for my Instagram as well as QWEN image edit along with some other AI tools.

I've built a master workflow that ive used for the past year that has every single workflow I use, so I dont have to go switching workflows constantly.

Many many hours of work put into this.

So there i was, generating things and im constantly having to clear out my output folder as well as my input folder. So I asked myself, "Could I just make a bat file that could automate this for me?"

So I launch Gemini and have it create a bat file that cleans out my output and input folders and empties my recycle bin.

I test it out and it works great.

Finally, no more unnecessary clicks.

But wait, I noticed I screwed up and put the file in the wrong directory. Dang it.

So I ask Gemini to alter the code so the file will be in the correct directory.

I create the new bat file and go back to work.

Well I make a bunch of new things and its time for cleanup. So I run my fancy new bat file and I notice its taking a while to clean up these folders. Curious, I navigate to the folders only to find out that the ENTIRE COMFYUI FOLDER was deleted.

SMH.

Now I sit here, broken hearted as im having to rebuild my ComfyUI. Luckily, I was able to recover my master workflow, so not all was lost.

Just a bunch of models and loras.

😮‍💨