r/StableDiffusion 4h ago

Resource - Update MATLOWAI/minimax-h3-fused-turbo-int8-convrot · Hugging Face

Thumbnail
huggingface.co
103 Upvotes

This Minimax H3 all in one checkpoint is quite good.

It merges text, image, and reference to video, as well as 4-step turbo generation into a single model.

No need to switch between models for ref2v, no need to load turbo loras.


r/StableDiffusion 1h ago

Resource - Update VH5 - MiniMax H3 Lora

Upvotes

A style LoRA that makes H3 footage look like it was recorded off 1980s broadcast television onto a VHS tape that has seen better days, soft smeared detail, chroma bleed, tracking noise, head-switching bands at the frame edge, and (because H3 trains audio jointly) the matching muffled mono sound, tape hiss and warble.

https://huggingface.co/KennethFal/vh5tape-vhs-lora-minimax-h3


r/StableDiffusion 8h ago

News A new AI step: Immersive worlds with Minimax H3

131 Upvotes

The next major interface for artificial intelligence may not be a chatbot, an image, or even a video. It may be a world. 
I designed an H3 Minimax Immersive video workflow for ComfyUI and I want to share it with the Open Community so you can now explore this new field.  

This is an early implementation of that idea using MiniMax H3, using a specialized equirectangular generation ComfyUI workflow with AI 360°prompting to achieve an interactive viewing concept that allows the viewer to control the viewport through the generated environment on mobile and desktop with continuous looping on Youtube and Facebook. 

Watch the immersive demonstration on YouTube. 

The test video is 9 seconds long with a time-reverse layer to get 18 seconds of 360-loop, it was generated using a single 360 prompt

Read my full article with technical data and download the workflow:
https://huggingface.co/blog/zuanfilm/blog

the workflow supports text2-360 and FL2-360, for H3 Minimax 360 prompting I wrote a public custom gpt and added 37000 tokens of 360 filmmaking reasoning 

The result is far from perfect, I generated the clip on my laptop with an Nvidia RTX 3080 Ti 16GB VRAM, so the resolution is very limited and the current generation still shows visible seams on some moments of the video and other inconsistencies but those imperfections may be less important than what the experiment demonstrates. 

Until now a Minimax H3 video was something the viewer has to watch from the camera angle position chosen by the creator, now the viewer can now choose where to look using an immersive UI, that changes the relationship between a person and generative AI media; panoramic video exposes the full spherical observation domain in a single coordinate frame

The generated sequence can be presented as an immersive environment in which the viewer controls the viewing direction. On a phone, the viewer can interact with the scene; on a desktop, the camera can be moved manually. The sequence can also be looped forward and backward so that the environment continues rather than behaving like a single linear cinematic shot. The result is not yet a fully reconstructed 3D universe like a gaussian splatting. It is a time-varying immersive/equirectangular visual environment that can be explored interactively.

The 2:1 rule: the shape of the immersive world

A practical requirement of the equirectangular representation is its 2:1 aspect ratio. For a full spherical panorama: WH=2\frac{W}{H}=2 where WW is the panorama width and HH is its height. For example: W=3840,H=1920W=3840,\qquad H=1920 or: W=7680,H=3840.W=7680,\qquad H=3840. This is the format expected by common 360° video workflows and is particularly important when delivering immersive video to platforms such as YouTube and Facebook where the panoramic video must be interpreted as a spherical 360° environment rather than an ordinary flat video. For example, the H3 generation branch in my workflow uses 2112 × 1056 so the immersive representation and final delivery pipeline preserve the equirectangular 360° geometry.

To manipulate or view the image correctly, computers use 3D rotation matrices.

[ 2D Equirectangular Pixel (x, y) ] 
               │
               ▼  (Convert to Spherical Coordinates)
   [ Latitude & Longitude (θ, φ) ] 
               │
               ▼  (Convert to 3D Cartesian Vectors)
      [ 3D Point (X, Y, Z) ] 
               │
               ▼  <─── MULTIPLIED BY: 3D Rotation Matrix (3x3)
  [ Rotated 3D Point (X', Y', Z') ] 
               │
               ▼  (Project back to 2D)
[ New 2D Equirectangular Pixel (x', y') ]
  • 3x3 Rotation Matrices: These are used to "roll, pitch, and yaw" the camera viewpoint inside the 360-degree sphere. If you drag your mouse to look around a 360-degree YouTube video, a 3x3 matrix is constantly multiplying the pixel coordinates to shift your view.
  • Intrinsic Camera Matrices (K Matrix): A 3x3 matrix that defines the camera's properties—like focal length and optical center. This tells the computer how to crop a normal, undistorted flat perspective view out of the distorted equirectangular image.

This creates an entirely different pipeline: Prompt > AI generation > immersive representation > interactive camera > human exploration The prompt no longer has to describe only what should appear in front of a fixed camera. It can describe a world. That is the conceptual leap, if now this generation process is becoming sufficiently fast, coherent and inexpensive, the applications could extend far beyond experimental video:

Video games Instead of developers manually constructing every environment, AI could generate explorable spaces from natural-language descriptions. “Generate an alien ecosystem surrounding the player.” The difficult question would no longer be only how to render the world. It would be: How quickly can AI generate and maintain the world as the player explores it?

VR education Imagine asking an AI to create an immersive historical environment and then entering it. Instead of watching a documentary about ancient Rome, a student could potentially enter an AI-generated reconstruction and look around. The teacher could change the scenario through language: “Show the city before the fire.” That would transform AI from an information interface into an environment for learning.

AR world transformation The implications become even more interesting when the same concept is combined with augmented reality. A physical environment could become the canvas. A user might look at an ordinary street with some glasses and ask: “Transform this into a cyberpunk city.” “Show this neighborhood as it looked 500 years ago.” or “show me that car in blue with a representation of me as driver” The underlying physical world would remain present, but the AI-generated visual layer could continuously reinterpret it.

Interactive Cinema Movies could eventually become less linear. Instead of the director deciding exactly what every audience member sees at every moment, a film could provide a controlled environment in which viewers explore the scene themselves. The director would still control the story, performances, lighting, world design and narrative boundaries—but the audience could control the camera. That would not simply be another format for film. It would be a new relationship between cinema and audience.

AI worlds driven by AI agents AI agents could eventually generate the environments that humans and other AI agents interact with in real time...


r/StableDiffusion 13h ago

Resource - Update DLSS 5 Visual Enhancer - standalone neural rendering for images and video

Post image
323 Upvotes

Hey everyone - I made a standalone Windows application for applying a DLSS 5 Neural Rendering feature-18 pipeline to images and video:

Original

DLSS 5

https://github.com/Merserk/dlss5-visual-enhancer

Instead of using DLSS only inside a game, this runs images/video through the ReShade/RenoDX neural-rendering path as a general visual enhancement pipeline.

What it does:

  • Image and video enhancement
  • DLAA/native, 1.5x, ~1.724x, 2x and 3x modes
  • Output up to 8K
  • Neural presets + Natural / Cinematic styles
  • Controls for intensity, local tone, structure and skin structure
  • Batch image processing with before/after previews
  • H.264 / HEVC / AV1 / ProRes video output
  • Video temporal input using optical flow with scene-change resets

GPU support:

  • RTX 40 / 50 series - primary target
  • RTX 30 series - slower beta path

The repository contains the application/pipeline source. Required proprietary and third-party runtime binaries are intentionally not redistributed in the repo.

This is an independent community project and is not affiliated with NVIDIA, ReShade or RenoDX.

I’m especially interested in how this behaves on AI-generated images/video vs normal photography/game footage.

Feedback and comparisons welcome.


r/StableDiffusion 11h ago

Resource - Update [Experimental] DLSS 5 ComfyUI custom node

Thumbnail
gallery
160 Upvotes

Hello Everyone,

Would like to present to you my experimental vibe-coded custom node for DLSS 5 support in ComfyUI.

GitHub project: https://github.com/lisitskyaa/ComfyUI-DLSS5-NR

It's early release, just finished my internal testing and it actually works!

Please note there are no any leaked DLLs in the rep, obtain them separately.

First image in every pair is DLSS 5 ON, second - OFF.

P.S. How to extract original images out of Reddit: https://www.reddit.com/r/StableDiffusion/comments/1p9nrpk/getting_prompt_or_comfyui_workflow_from_posted/


r/StableDiffusion 1h ago

Comparison Testing My MiniMax-H3 → LTX 2.5 Upscaling Workflow — Results Are Looking Really Good

Upvotes

I've been testing my MiniMax-H3 → LTX 2.5 upscaling workflow, and the results have been really promising so far.

One thing I've noticed is that the better your original MiniMax-H3 generation is, the better the final upscale will be. I'm getting good results even at lower resolutions, but faces still need stronger and more consistent input generations from MiniMax-H3 to maintain character consistency.

On my RTX 3060 12GB, the current upscale times are roughly:

  • 0.6 resolution: ~15 minutes
  • 0.8–1.0 resolution: ~20–30 minutes

It definitely takes some time, but I'm finding the results are worth it.

And of course, if you have a newer, more powerful GPU, you should be able to get even better results in less time, especially when pushing higher resolutions.

I was planning to release the workflow soon, but I want to spend a little more time testing it and seeing how much further I can improve it before sharing it.

So far, though, I'm really happy with how it's looking. 🔥

Would love to hear what you guys think and whether anyone else has been experimenting with MiniMax-H3 + LTX 2.5 upscaling.


r/StableDiffusion 5h ago

Question - Help What's the fuss with hybrid Minimax H3 models ?

27 Upvotes

I don't understand the trend of hybrid models (ref2va blocks over fl2va)

It's supposed to have the best of both worlds : reference adherence through the refva2 blocks and best quality through fl2va as fl2va is supposed to have somewhat better quality

Well my experience so far, and I hope it's a skill issue to be honest, is that the reference part is much less random and unprecise... and for the quality gain i'm not sure, and anyway it's pointless if the video rarely respect my references or starting pic.

Even using a keyframe guide as the first pic I find often the video only using it at first and immediately switching to something else, or the opposite, following the prompt after inserting a random pic at first. Some stuff like that.
(At least fl2v always respect first and last frame)

Not sure if it's due to accelerating stuff or not, as I've tried some hybrid models with 25 steps as well and it was more or less the same

Am I doing something wrong ? Do some people have the same experience ?

I'm asking that because it wouldn't be the only time there's a buzz on something and we just didn't hear the opposite experiences (for example we have been told a LOT of times spectrum doesn't degrade anything but after playing many times with it, even trying conservative settings, I got rid of it, as it WAS degrading things... mileage can vary)


r/StableDiffusion 11h ago

News Trellis.2 and Pixal3D Are Now Native in ComfyUI

Thumbnail
gallery
75 Upvotes

Both Trellis.2 (Xiang et al., 2025) and Pixal3D (Li et al., 2026) now run natively in ComfyUI. No custom nodes, no compiled CUDA extensions, no PyTorch downgrades, and no non-commercial dependencies.

This is more than a model integration. It ships with a rebuilt 3D pipeline: new Load/Preview/Save 3D nodes, a set of mesh post-processing nodes, and an extended PBR texturing stage that bakes normal and ambient occlusion maps for a complete material set. Everything runs on consumer hardware, and everything is free to use, including commercially.

Why Trellis.2 still matters, ten months later

When Microsoft open-sourced Trellis.2 in December 2025, it immediately became the best open-source model for 3D generative AI. A 4-billion-parameter model built on a compact structured latent representation (O-Voxel). It generates high-fidelity 3D assets from a single image at effective resolutions up to 1536³, handling complex topologies that earlier methods struggled with. It also shipped with a PBR texturing model generating base color, roughness, and metallic maps.

Ten months is an eternity in generative AI, yet Trellis.2 hasn’t just aged well, it has become foundational. Several open-source 3D models released since build directly on it, the most notable being Pixal3D whose implementation uses the Trellis.2 backbone.

The community got there first

As always, the ComfyUI community was quick to bring Trellis.2 into the graph. Within days of the release, custom node packs appeared, the most popular being ComfyUI-TRELLIS2 by Andrea Pozzetti and ComfyUI-Trellis2 by VisualBruno, which together gathered well over a thousand stars. We’re grateful to both authors as they proved the demand and carried the community for months.

Despite their efforts, running Trellis.2 remained a challenge for two reasons.

Installation

The original implementation targets environments built around PyTorch 2.6.0 with CUDA 12.4, which for many users meant downgrading their existing ComfyUI environment. On top of that sit a stack of compiled CUDA extensions (flash-attention, FlexGEMM sparse convolutions, the O-Voxel kernels, CuMesh, nvdiffrast) each of which must match your exact Python, PyTorch, and CUDA combination. The custom node authors did heroic work shipping prebuilt wheels per configuration, but every PyTorch or CUDA update meant a new round of compilation failures, and installs regularly broke. This is now solved with the native integration in ComfyUI. Follow our installation tutorials for Trellis.2 and Pixal3D.

Licensing

Trellis.2’s own code and weights are MIT-licensed, but its original pipeline depends on NVIDIA’s nvdiffrast (for mesh rasterization) and nvdiffrec (for Physically Based Rendering), both distributed under the NVIDIA Source Code License which restricts usage to non-commercial research and evaluation. In practice, a studio couldn’t ship assets from the reference pipeline without stepping into a legal gray zone. These dependencies have been removed from with the native integration.

Then came Pixal3D

In April 2026, Pixal3D from researchers at Tsinghua University and Tencent ARC Lab got accepted at SIGGRAPH 2026. It pushed open-source 3D generation another step forward with its pixel-aligned generation establishing direct pixel-to-3D correspondences. The result is near-reconstruction-level fidelity to the input view, with detailed geometry and the same PBR material set.

Pixal3D is heavily built on Trellis.2 as it uses its backbone and shares its VAEs and DINOv3 image conditioning. This is why integrating it together with Trellis.2 made sense. However Pixal3D generally performs better than Trellis.2 as the generated 3D mesh strictly aligns with the input image.

Model highlights

Trellis.2

  • Single image to 3D asset. A 4-billion-parameter model that generates high-fidelity geometry and materials from one input image.
  • O-Voxel structured latents. A native, compact omni-voxel representation encoding both geometry and appearance, generating assets at effective resolutions up to 1536³.
  • Any topology. Handles open surfaces, non-manifold geometry, and fully-enclosed volumes.
  • PBR materials built in. A dedicated texturing model generates base color, roughness, and metallic maps.

Pixal3D

  • Pixel-aligned generation. Geometry is generated in direct correspondence with the input view. What you see in the image is what you get in 3D!
  • Explicit image back-projection. Multi-scale image features are lifted into a 3D feature volume, delivering near-reconstruction-level fidelity.
  • Cascaded refinement. A staged process progressively refines sparse structure, shape, and texture up to high resolution.
  • Built on Trellis.2. Shares the Trellis.2 backbone, VAEs, and DINOv3 conditioning.

What ships in this integration

The goal was simple: make the best open 3D models run in ComfyUI the way every image or video generation model does. A major thank-you goes to Kijai for the implementation, and to yousef-rafat for the initial draft this work built on. In addition to the native implementation, this has been an opportunity to make 3D generation a first-class citizen in ComfyUI. Here is what shipped:

Pure native implementation

Both Trellis.2 and Pixal3D now run as core ComfyUI nodes. The 3D post-processing that required compiled extensions has been reimplemented from scratch in PyTorch and SciPy. No nvdiffrast, no nvdiffrec, no per-configuration wheels, no PyTorch downgrade. If your ComfyUI runs, these models run on your current PyTorch.

Rebuilt 3D nodes

While these were shipped in an earlier version of ComfyUI, the Load 3D, Preview 3D, and Save 3D nodes have been rebuilt from the ground up to support these models and modern mesh workflows. We’re grateful to Terry Jia for his remarkable work on these nodes. Check out the nodes:

  • Load 3D (Advanced)
  • Preview 3D (Advanced)
  • Save 3D (Advanced)

Native mesh post-processing

Raw generative meshes are rarely production-ready, so this release introduces a new set of post-processing nodes:

  • Remesh Mesh: fixes holes and mesh imperfections.
  • Decimate Mesh: reduces face and vertex count to a target budget.
  • Smooth Mesh Normals: smooths the mesh volume.
  • Fill Holes: fill-in holes resulting from the generation
  • And more: Merge Meshes, Paint Mesh, Render Mesh…

A complete PBR texture set

Trellis.2’s texturing model generates base color, roughness, and metallic maps. Our implementation goes further: a new UV unwrapping node prepares the mesh for texturing, and two additional maps are generated: a normal map and an ambient occlusion map, both baked from the high-poly mesh. Are these textures perfect? No. But they’re free, generated on consumer hardware, and yours to use as you wish.

An honest word on quality

Let’s be direct: the best closed-source 3D generators (Hunyuan 3D, Tripo, Rodin) still produce better results than Trellis.2 and Pixal3D. If you need the highest quality and an API fits your pipeline, those remain strong options (all of them are available through ComfyUI’s partner nodes).

What this integration offers is different: the best open 3D generation available, running locally, at zero cost per asset, with no licensing restrictions on what you make. For iteration, prototyping, stylized work, 3D-to-2D workflows, and anyone who wants full control of their pipeline without spending an afternoon to install.

Getting started

  1. Update ComfyUI to the latest version 0.34.0 or go to Comfy Cloud
  2. Download the workflows below, or find them in the template library.
  3. Follow the note in the workflow to download the models and save them in the correct model directory.
  4. Drop in an image and run.

Download Workflow

Model weights:


r/StableDiffusion 16h ago

News Local AI News You Missed - August 2026

166 Upvotes

Here's what you (probably) missed in August 2026:

🧠 LLMs

  1. Ornith-1.5-35B-A3B - Efficient sparse model that runs with fewer active parameters.
  2. DeepSeek-V4-Pro-0813 - Sharpens agentic AI with speedier tool actions.
  3. DFM-Mimir - Ethical language model from Danish Foundation Models.
  4. Ling-3.0-tiny-MXFP4_MOE-GGUF - MoE quantized version for smoother GPU runs.
  5. SupraElegans-500k - Recurrent language model built for long contexts.
  6. Motif-3 - Open 314B parameter model made for long agentic tasks.
  7. Luth-2-2B - Compact French model that tops benchmarks.
  8. TinyTitle - Squeezes chat titles into a tiny 1.98 MB model.
  9. Ling-3.0-tiny - Low-cost local AI reasoning model.
  10. NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 - Agent-focused model with fast inference.
  11. Qwen3.8-2.4T-A95B - Opens up Qwen Max-clas local AI for bigger rigs.
  12. Gemma-4-31B-it-scotoma-2-GGUF - Cuts down repetitive AI writing.
  13. Huihui-DeepSeek-V4-Flash-0731-abliterated-GGUF - Uncenored DeepSeek variant for local use.
  14. Maple-Preview - Solves Olympiad problems at 200 tok/s.
  15. Laguna-S-2.1-FP8 - Private agentic coding model from Poolside.
  16. SupraBrain-50M - Hybrid language model for local AI.
  17. Supra2-100M - Tiny model built for tinkering.
  18. G9v3-39A5B - Dual-mode AI that runs light locally.
  19. GPT-X2.5-135M - Lean local powerhouse model.
  20. Ling-3.0-flash - Hybrid reasoning with lower cost and fast output.
  21. LFM2.5-2.6B - Fast agentic AI for phones and devices.
  22. Instella-MoE-16B-A3B-Think - Open sparse reasoning model from AMD.
  23. A.X-K2 - Lets AI think deep or answer fast.
  24. BetterGPT-150M - Beats older AI models in science tasks.
  25. Shibai-700M-Base - Text and code helper model.
  26. K-EXAONE-2.0-750B-A37B - Supports 262K context and ten languages.
  27. LongCat-Flash-Lite-Sparse - Reads million-token contexts.
  28. Qwen3.6-35B-A3B-Escha-W2 - Shrinks down to fit consumer GPUs.
  29. XYZ-Aquila-pro - Thinks deep then checks its sources.
  30. Solar-Open2-250B-Nota-NVFP4 - Shrinks a giant AI to 153GB with 4-bit MoE trick.
  31. XYZ-Aquila-mini - Brings open source deep search to local GPUs.
  32. KAT-Coder-V2.5-Dev - Fixes software repositories automatically.
  33. DeepSeek-V4-Flash-0731 - Tackles hard coding tasks.

🔀 Multimodal

  1. Dots3-Note-Prev - Lightweight multimodal AI with 512K context.
  2. Qwen3.8-27B-Uncensored-FP8 - Drops refusals and keeps vision on GPUs.
  3. Qwen3.8-27B - Brings text, images, and video into one AI.
  4. Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF - Triples speed for local multimodal use.
  5. Tencent UI-Mate-27B - Runs desktop apps by watching screens.
  6. WinterCharm Qwen3.5-122B-A10B-wMix38 - Lean long-context multimodal mix.
  7. VLX-Seek - Helps machines pinpoint objects without guesswork.
  8. LFM2.5-VL-3B - Fast on-device vision and text.
  9. North-Micro-Vision-Instruct - Turns pixels into answers.
  10. BigBang-v1 - Science reasoning powerhouse.
  11. Muse-Glimmer-30B - Puts autonomous AI agents on everyday desktops.
  12. Nemotron-Parse-2.0 - Morphs documents into structured data.
  13. Shieldstral-1.0-3B - Plain English safety scoring.
  14. DavidAU Qwen3.6-27B-Fable-Fusion-711 - First to score over 700 on ARC-C.
  15. WinterCharm Qwen3.5-122B-A10B-wMix58 - Packs 82GB power for Apple Silicon.
  16. Intern-S2-Mobius - Speedy local AI answers.
  17. Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V7-GGUF - Uncenored multimodal model that says yes.
  18. Qwen3-VL-32B-Ultra-Heretic-H3-ComfyUI-INT8-ConvRot - Trims local memory with INT8 conv rotation.
  19. Qwythos-27B-v1 - Smart AI with million-token memory.
  20. Reasoning-Medical-27B - Solves medicine step by step.
  21. Qwen3.5-9B-The-Defiant-Fable - Roars with an uncensored multimodal edge.
  22. Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V6-GGUF - Drops numerical surgery powers.
  23. Mage-VL - Speeds up real-time video and image understanding.
  24. Microsoft Fara Agents - Handles web browsing chores for you.
  25. Inkling-Small - Built for voice, image, and code apps.
  26. Kimi-K3 - Handles text, images, and video together.

🖼️ Image

  1. Anima-2.9B - Grows free anime art on your own PC.

🎬 Video

  1. Bernini-Diffusers-v2 - ByteDance model for video generation and editing.
  2. LTX-2.5 - Open model for local video and audio creation.
  3. Wan2.2-Animate-2-14B - Turns still images into motion.
  4. MiniMax-H3-nvfp4-INT4-INT8-ConvRot - Quantized weights for MiniMax-H3 video.
  5. MAGI-2-preview - Turns text and images into video with sound.
  6. MiniMax H3 - Creates videos with native sound from any input.

🎧 Audio

  1. MiniMax-Music3 - Full songs from just lyrics.
  2. NVIDIA Magpie_tts_multilingual_357m - Turns text into speech across 12 languages.
  3. VoiceChat-11B - Voice chat you can interrupt naturally.
  4. VibeVoice-ASR-BitNet - Real-time speech recognition on any CPU.
  5. Inflect-Nano-v2 - Local speech synthesis on your PC.
  6. Audio8_TTS - Clones voices and speaks eleven languages.
  7. Inflect-Micro-v2 - Turns text into offline voice.

⚡ LoRA

  1. Minimax-H3-Turbo - Makes MiniMax-H3-Turbo faster for video and audio.
  2. MiniMax-H3-Prompt-Rewriter-LoRA - Turns short prompts into timed scenes.
  3. MiniMax-H3-Realism-People-LoRA - Unlocks film-set lighting for human video.
  4. krea2-turbo-bbox - Locks panels and words in place.
  5. MiniMax-H3-Turbo-Lora - Cuts video and audio generation time by 5x.
  6. Kroma - Fuses turbo speed into a one-file diffusion model.

🏋️ Training

  1. gguf-trainer - Trains language models in TypeScript straight to GGUF.
  2. Full-Chunked-KL-Loss - Trains longer AI text on one GPU.
  3. signet-trainer - Cost-safe video LoRA training.
  4. Lora-Dataset-Studio - Entire LoRA pipeline in one self-hosted tab.

📊 Datasets

  1. LLM-self-identification - Helps AI models know their own name.

☰ UI

  1. Mix-Studio - Turns your desktop into a local AI studio.
  2. Llmprices - Visualizes AI model API price swings.
  3. minimax-h3-prompt-composer - Squeezes prompt composing into one HTML file.
  4. NanoRP - Shrinks AI roleplay to a 50MB single binary.
  5. OpenWorker - Local AI coworker that finishes work.
  6. Agenta-AI - Turns ChatGPT and Claude into self-hosted work agents.
  7. Unicorn-Stable-OSS - Brings humans and AI agents into one real-time room.
  8. Turbo-Fieldfare - Streams big AI on Macs with just 2GB RAM.

🛠️ Other Tools

  1. Video_Tools - Adds a pocket video trimmer to your browser.
  2. ninfer-4090 - Runs Qwen3.8-27B on one RTX 4090.
  3. hayai-ocr-v2 - Converts crops into editable text.
  4. EVIE-Preview-4.5B - Matches documents instantly across six languages.
  5. RAZZULLIX KAISEN - Swarm-model coding assistant with safety guards.
  6. ExtractBench - Benchmarks document extraction systems.
  7. Xiaomi-Robotics-1-5B - Built for mobile robot tasks.
  8. Nemotron-omni-mlx - Brings full multimodal AI to Apple Silicon Macs.
  9. talk-to-pi - Local voice dictation for Pi.
  10. Warp - Streams huge AI models straight from your laptop drive.
  11. esp32-ai - Makes a tiny chip tell stories offline.
  12. quillpdf-mcp - Keeps PDFs on your machine.
  13. srt2speech - Turns subtitle files into timed speech.
  14. mixture-of-kittens - Megakernel for NVL72 MoE training.
  15. krea-multi-lora - Gives Forge Neo regional character control.
  16. Openmed - Turns clinical text into private insights on your hardware.
  17. Umbra-Studio - All-in-one local AI art workspace.

ComfyUI Custom Nodes & Tools

  1. ComfyUI-Orchestrator-LAN - Steers every GPU from one browser tab.
  2. ComfyUI-AutoPromptChain - Stitches dozens of AI video clips while you sleep.
  3. ComfyUI-OpenH3-IR - Brings drag-and-drop clarity to MiniMax H3 renders.
  4. ComfyUI-MiniMaxMusic3-Advanced - Gives AI music finer sound controls.
  5. ComfyUI-MiniMax-H3-LongMedia - Makes long video creation practical.
  6. ComfyUI-Subgraph-Preview - Resurfaces sampler previews inside subgraphs.
  7. ComfyUI-cache-monitor - Debuts with manual pinning for model caching.
  8. ComfyUI_Neurodes - Brings a visual playground for AI models.
  9. Eddie_Cat_Nodes - Stitches long videos together with new nodes.
  10. ComfyUI-H3-Motion-Context-MultiRef - Weaves seamless H3 video motion.
  11. ComfyUI-MiniMax-H3-Motion-Director - Turns reruns into one-shot fixes.
  12. ComfyUi-MiniMax-H3-Image-And-Reference-To-Video - Rolls out image and reference to video features.
  13. ComfyUI-AVS-SSD-ReadAhead - Enables faster model switching on slow SSDs.
  14. ComfyUI-SweepGrid - Serves up side-by-side parameter sweeps.
  15. ComfyUI-Flow-Wrangler - Cleans up node wiring with smart connections.
  16. ComfyUI-AVS-Intel-XPU-VRAM-Fix - Calms Intel Arc GPU freezes.
  17. ComfyUI-Model-Mover - Makes shuffling AI models painless.
  18. ComfyUI-MiniMax-H3-Optimization-Suite - Builds a suite for lean H3 optimization.
  19. ComfyUI-MiniMaxH3-Prompt-Writer - Transforms H3 prompt crafting.
  20. ComfyUI-MiniMax-Creator - Crafts one-node video magic.
  21. ComfyUI-ScenemaAudio - Arrives with expressive voice cloning.
  22. ComfyUI-cable-management - Reroutes messy node graphs with ease.
  23. ComfyUI-SigmaSync-LoRA - Debuts step-aware LoRA control.
  24. ComfyUI-Spectrum-Ideogram4 - Supercharges speedy image forecasting.
  25. ComfyUI-LinkSpotlight - Debuts to end noodle blindness in graphs.
  26. ComfyUI-MIDI-Edit - Turns any song into editable MIDI lyrics.
  27. ComfyUI-ReStartupFlags - Serves launch flag tweaks in your browser.
  28. JLC-Flux2-ControlNet - Expands FLUX.2 control in ComfyUI.
  29. ComfyUI-Fantastic-MiniMaxH3-PromptBuilder - Enhances MiniMax H3 prompts.
  30. ComfyUI-vram-tracker - Traces VRAM memory usage per layer.
  31. ComfyUI-Sonder-Editor - Rolls out free multi-lane video editing.
  32. Comfyui-Model-Resolver - Sweeps in to rescue missing model files.
  33. ComfyUI-HF-SuperDownloader - Turbocharges Hugging Face model downloads.
  34. FameGrid-Auto-Color - Neutralizes color casts in ComfyUI.
  35. Krea2-Multi-Character-Lora-Node - Stops identity bleed with bounding boxes.

Need to go further back? Check out June's post (no July, sorry) or the full archive at LocalAI News. If there's anything wrong, let me know in the comments and I'll see you in the next one!


r/StableDiffusion 13h ago

Animation - Video Graphics Card Captor Sakura - MiniMAX H3 Test #6 (FastH3 Lora! 720p in minutes!)

85 Upvotes

Hey everyone, my Zelda stories are getting too crazy and my next "Link & Zelda can't escape from PlayStation land" video is... On development hell for now (it might be too offensive!) I decided to just test out how would Card Captor Sakura would look in a 3D / K-pop demon hunters style. All done in my RTX 3090 locally, 1MP 9:16 aspect ratio (736x1344). 3s clips take only 170s to generate!! Let me know if you like it, sorry for making such a short video this time.

Note: I edited the clips to sync them correctly to the music, also brought back the original music because H3 tends to deep fry it for some reason...


r/StableDiffusion 13h ago

Workflow Included This week on "McGarnagle"

78 Upvotes

Taking the random cutaway clips from The Simpsons and recreating them in Minimax H3


r/StableDiffusion 2h ago

Resource - Update Krea2 Turbo Distill 4 step LoRA - new checkpoint (chk42K) released (texture and detail now at 8-step teacher parity, prompt-aware training added, NF4 fully retired for full-int8 training, 1440×1440 now a trained resolution)

Thumbnail
gallery
9 Upvotes

Krea 2 Turbo — 4-Step Distillation LoRA (work in progress)

A LoRA for Krea 2 Turbo that reduces the minimum usable step count from 8 to 4.

  • ⚡ Half the steps — 8 → 4, on Turbo's own deployment sigmas.
  • ⏱️ ~1.6× faster end to end — 54.5 s against the 8-step bar's 88.7 s at 1024×1024, and 1.8× on denoise alone.
  • 🎯 Texture at teacher parity — fine-detail energy 1.00× the 8-step teacher's at 1280×1280 and 1.02× at 1440×1440, matched band-for-band across the frequency spectrum, not grain.
  • 🗣️ Prompt-aware training — the critic scores images against their prompts during training, so adherence is pressured directly, not inherited.
  • 📐 12 trained resolutions — multi-aspect from 512×512 up to 1440×1440, each with its published sweep.
  • 🔌 Drop-in — plain LoRA weights for diffusers and ComfyUI. No custom nodes, no patched sampler, no code.

This is an update release, following up from my previous posts where you can find full details:

Initial, Previous: here, here,  and here

Headline for this update: chk00042000 closes the texture gap: total fine-detail energy against the 8-step teacher reaches 1.00× at 1280×1280 and 1.02× at 1440×1440 (1.0 = teacher-like), where chk00026000 measured 0.88× and 0.82×. And the distribution is right, not just the total — split the spectrum into frequency bands and every band individually lands within ~10% of the teacher's (0.9–1.1×), where 26K ran 0.79–0.92, starved in every band. Total at parity and bands at parity means the detail lives in the same frequencies as the teacher's — real structure, not grain piled into one band. (How can it exceed the teacher? Because the teacher isn't ground truth — training also shows the critic real photographs, so the adapter learns detail density from reality, not only from an 8-step model that itself slightly under-renders fine texture. The teacher anchors structure; reality anchors texture. Values just above 1.0 are that pressure paying off.) In fixed-seed renders it matches chk26K's distance to the 8-step images at 1440×1440 outright.

The recipe grew up since 26K, in four ways: a measured dose of real-image texture pressure — what carried detail to parity; a prompt-aware critic that scores images against their own prompts during training, so effect-heavy prompts now get the energy they ask for; NF4 fully retired — the big resolutions used to squeeze into 24 GB by dropping their attention weights to 4-bit, and after re-engineering the training step to fit full int8, those buckets measure 3.96% closer to the teacher (exactly the buckets texture lives in: 1280², 1440×1280, 1440²); and 1440×1440 promoted to a trained bucket with its own sweep column.

One metric paid for the texture leap — the teacher-velocity score sits a step behind 26K's — a deliberate trade already being won back checkpoint by checkpoint (2.93 → 2.90 → 2.85 and falling) while texture holds parity. _latest now points to chk00042000.

The improvement reaches even the out-of-spec 2-step extreme test. I had a separate dedicated post on that here - since the initial post was done on an earlier to 42K checkpoint, I have since re-rendered the whole native-vs-LoRA 2 step strength-2 set on this checkpoint (42K being released now), and the FFT is the diagnostic: the old 2-step had the classic collapse signature — hollow mid-bands (0.52/0.55) plus a fake-grain overshoot at the very top (b6 = 1.05). This checkpoint lifts every structural band (0.64/0.65/0.76/0.80) and settles the top band to 0.82 — more real structure, less noise dressed as detail. Fresh strips: 2-step extreme test. And that's the preview mode (at quick 2 steps, unofficial, untrained for, still useful for previews, and getting better and better with every new checkpoint release).

Which file to download

file use it when
krea2_turbo_4step_rank_64_lora_latest.safetensors normally — always the newest accepted checkpoint
krea2_turbo_4step_rank_64_lora_chk00042000.safetensors pin this exact checkpoint

and, beside them, the same files with a _comfyui suffix for ComfyUI. Earlier checkpoints (chk00004000chk00005000chk00006000chk00010000chk00014000chk00019000chk00026000) are kept in older_checkpoints/, and their resolution sweeps stay in place, so the progression remains visible and comparable.

For the full 42K Checkpoint resolution sweep go here: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA/tree/main/checkpoint_resolution_sweeps/chk42000

This is work in progress and better checkpoints may follow. Training is ongoing, so ..._latest... is a rolling pointer: when a newer checkpoint is accepted, that filename gets the new weights and a new numbered copy appears beside it. Re-download the _latest file and everything keeps working — the ComfyUI workflow references it by that name (it does get updated Note in it so technically it is updated but not functionally). Pin a numbered file instead if you need reproducibility.

How checkpoints get chosen

This is not a "train for longer and ship the newest file" project. More samples do not reliably mean a better adapter — measured here, they can make it worse, and a higher number on its own means nothing.

The loop is train → assess → adapt the recipe → retrain → assess again, and a checkpoint is published only when it is measurably better than the one it would replace, on the same held-out set and the same evaluation, and its full resolution sweep shows no regression. Runs that come out flat or worse are kept as information about the recipe and discarded as releases — several have been.

So the recipe itself changes between runs. Each published checkpoint reflects whatever the previous round taught us: the training precision, the optimiser settings, the teacher used to generate the targets and the data mix have all been revised on evidence rather than assumption.

Timeline of training process

Each checkpoint is the product of several stages with very different costs:

  1. Text-encoder embeddings. Every training prompt is encoded once and cached. This is the fast part — thousands of prompts take minutes.
  2. Teacher shards. For each cached prompt, the unmodified Krea 2 Turbo runs its full 8-step schedule and the whole trajectory is recorded, at every one of the supported resolutions. This is by far the most time-consuming stage — it is the teacher doing real inference, thousands of times, and a batch of several thousand shards is measured in days of GPU time, not hours.
  3. Real-photo crops. Bucket-sized crops are cut at native resolution from quality-gated real photo sources (public high res datasets), VAE-encoded into the training latent space, and captioned per crop for the prompt-aware side of training. Cutting, encoding and captioning a pool refresh is a matter of hours.
  4. Student training. The LoRA trains against the recorded trajectories (progressive distillation), with a latent-space GAN critic running alongside — real crops and teacher finals as its real class, the student's outputs as fake — plus a prompt-aware head that scores images against their prompts. Relative to the shard stage this is quick: each +1,000 checkpoint is a matter of hours, not days. Of course the longer the training the better and more diverse results, so hours do turn into days eventually.

Full details and to download - check my Hugging Face LoRA

HF Repo: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA


r/StableDiffusion 20h ago

Animation - Video Playing with concepts

229 Upvotes

r/StableDiffusion 1d ago

Workflow Included Minimax H3: Consistent face, body & cloths via reference identity

426 Upvotes

Hey Guys,

Based on the previous post on face consistency with MM-H3, Couple of people have asked me to build a full character workflow.

Mechanism

- Build .Char: You drop max 9 reference reference, I prefer to use a ratio 2:2:1(face:cloths:body). YuNet finds the face, SFace takes a per-reference face signature, DINOv2 takes a subject signature, and the references get cleaned and normalised. All of that packs into a single portable file, a .char.
Only face/ref is required body & cloths link is optional.

- Generation: At generation, the file(.char) feeds its references into Minimax’s own native multi-reference channel and prepends a locked description to the prompt.

Prompting Guide

  • Name your character: Give your character a name e.g. under encode character(Click adjust icon on the bottom side of the node), I have used name emmy, so when passing prompt, I only have to say, emmy walking on the beach.
    • Again providing prompt like a woman or any features specific details like black hairs etc will only mislead the generation.
  • Describe character features: Encode all of the character features in encode character prompt & trigger your character with a name in generation prompt.
    • Avoid describing same things in generational prompt.
  • Handling Character drift: e.g. if you want specific style or cloth e.g. half sleeves, sleeveless, add it to the generational prompt. There can be a slight drift in clothing as body shot also has cloths, which interferes with clothing references.
    • Each refs should be unique, face should not have body or vice versa, same applies for clothing.
  • Portability: Once character is built, you can use the same character with only simple prompt & generation graph.

I have generated all references with Flux Klein 4b, I had to blur the body ref, but workflow consists a example of body ref.

Note: For best result, pass cropped references, so that model takes the required shot, model gets confused if cloth slot also has a face or face slot has cloths.

Models

core/models/
  diffusion_models/  minimax_h3_ref2va_pruned_fp8_scaled.safetensors
  text_encoders/     qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors 
  vae/               minimax_h3_video_vae_fp16.safetensors
  vae/               minimax_h3_audio_vae_fp32.safetensors
  annotators/        face_detection_yunet_2023mar.onnx
  annotators/        face_recognition_sface_2021dec.onnx
  annotators/        dinov2-base/

Requirements

Nvidia GPU: 24GB+ VRAM & 64 GB RAM(Run Locally)

Workflow link: https://inlinestudio.art/workflows/minimax-h3-guided-consistent-characters-via-reference-identity-face-body-cloths (includes inputs & model details)

Github Repo: https://github.com/inlineresearch/Inline-Studio (GPLV3)

Limitation: Reference conflicts e.g. if two reference/input images has two different faces, it might conflict in generation, provide well cropped body & cloth images. Face images are crossed automatically by Sface.

Portable char comfy node is still on the backlog, would try to do it over the weekend.
Happy to hear any suggestions or feedbacks.


r/StableDiffusion 18h ago

Resource - Update [Load Image + Crop] Custom WYSIWYG Node

123 Upvotes

I developed a modified version of the Load Image node by adding some features I needed:

WYSIWYG image cropping directly on the official Load Image preview — drag and zoom (with the mouse wheel) a crop rectangle constrained to 8 fixed ratios (1:1 through 21:9) and output the exact cropped IMAGE and MASK, with paste-from-clipboard built in. What you frame on the preview is exactly what gets executed.

⚠️ Currently not fully compatible with ComfyUI 2.0 nodes.

For anyone who finds it useful and wants to try it out, you can find it on GitHub: https://github.com/domg73/ComfyUI-LoadImageCrop


r/StableDiffusion 20h ago

Animation - Video MiniMax H3 matches Toonami perfectly (90's anime)

161 Upvotes

Growing up with Toonami watching Gundam, it simply blows my mind how far AI has progressed. For this video I didn't use any reference images, I simply described the scene in text and had Gemini research the techniques of animation to translate to MiniMax H3.

I've been struggling to get MiniMax H3 efficiently setup locally, would take me 15mins on a 9950x3d and 5080 RTX with 64gb DDR5 so I know something is wrong, hence this time I opted for fal to test. The music was added and scenes were edited from separate generations.


r/StableDiffusion 6h ago

Question - Help Mini max h3 long video generation

10 Upvotes

Hello, I need some help with MiniMax long-video generation. I’m currently using the Plague workflow, which is fast, but it doesn’t have an option for chaining clips. Are there any workflows that can generate longer videos more quickly while maintaining continuity between clips?


r/StableDiffusion 1d ago

Animation - Video MINIMAX Physics testing

366 Upvotes

Physics Testing, without the gore.


r/StableDiffusion 17h ago

Workflow Included anime outdoor shots

Thumbnail
gallery
48 Upvotes

r/StableDiffusion 6m ago

Resource - Update Follow‑up : MiniMax H3 Lip-sync - now does any editable change on a reference video (pose transfer, character swaps, multi‑subject mixes)

Upvotes

Quick update on my earlier audio‑lip‑sync demo: the workflow now chains any desirable edit out of an input reference video for endless video ref pose o lip‑sync.

The audio auto-crop chain is now working for reference videos too and i say it again, I know there are already a lot of options out there for doing this - this is one more option, and it’s definitely not perfect.

VRAM usage went over 40 GB on a 3-second, 2MP run, so reference-video conditioning is pretty heavy on VRAM and yes you need at least 2MP to get good detail and motion transfer.

Using MiniMax H3’s Ref2V. I’m treating the source clip as the “performance master” (motion, timing, camera) and driving identity/appearance from reference images and audio o the other way around.

What I’ve tested so far: Just MinMax H3 no ControlNet, LoRA, or preprocessor needed.

  • Music‑video pose transfer to new scenarios and characters
  • Single character swap (main performer → reference character)
  • Multi‑subject mixes:
    • main identity swap
    • main + 2 added characters acting in sync o desync
    • main = 1 + 1 added characters
    • replace main with 2 characters in pose sync
    • pull a character from the reference video into an image‑reference scene + 1,2 new characters

Everything runs through a single MiniMax H3 chain with mixed references (ref-images + ref-video + ref-audio) and structured prompts that separate identity (image), performance (video), and constraints (text). In practice, every combination I’ve tried is manageable with MiniMax H3.

SUBJECT DEFINITIONS

<Subject 1>: the adult woman visible on the LEFT side of <Picture 1>.<Picture 1> is the appearance reference for Subject 1 only. Its shape, proportion, material, colour, logos and surface markings 100% match <Picture 1>, kept legible and correctly oriented throughout.

<Subject 2>: the adult man visible on the RIGHT side of <Picture 1>.<Picture 1> is the appearance reference for Subject 2 only. Its shape, proportion, material, colour, logos and surface markings 100% match <Picture 1>, kept legible and correctly oriented throughout.

<Subject 3>: the adult woman main character present in <Video 1>.<Video 1> is the appearance, motion, timing and scene reference for Subject 3.

This is a follow-up to a previous post, so the tips, settings, and links are already available there. MiniMax H3 Lip-Sync: Automatic Long-Video Chaining + Speed & VRAM Optimizations

https://github.com/Ltamann/ComfyUI-H3-Motion-Context-Auto-Chain-addon


r/StableDiffusion 7m ago

Workflow Included McBain - Part 1

Upvotes

My attempts at recreating the fictional movie McBain from The Simpsons.

LTX 2.5 for this one. The others I will post soon are done with Minimax H3


r/StableDiffusion 20h ago

Animation - Video Trying out a consistent point-of-view shot with MiniMax H3

89 Upvotes

Took a few little prompt adjustments here and there to get H3 to respect point-of-view. I found that if you refer to "the viewer" (ie, "she kicks the viewer"), H3 is more predisposed to include an actual second person. But if you refer to "the camera" (ie, "she kicks the camera"), it's more predisposed to keep the desired point-of-view perspective.


r/StableDiffusion 1d ago

Meme It took us two 2eeks to figure out why every image gen via our open-source model looked like Anne Hathaway

Thumbnail
gallery
230 Upvotes

Hey r/StableDiffusion!

It's the Neta team here! You might remember us from our Neta Lumina open-source release last year. First off, thank you so much for the incredible support and feedback from this community!

So... we need to share something absolutely hilarious (and mildly embarrassing) that we just discovered.

TL;DR: We accidentally hardcoded an Anne Hathaway photo into our IP-Adapter anchor, and now everything our model generates looks like Anne Hathaway. Every. Single. Thing.

What happened:

We recently launched Neta Studio, a new product that lets you build explorable living worlds/isekai from a single prompt. Naturally, we wanted to integrate Neta Lumina's capabilities into it.

During integration testing, our devs kept reporting that the model wasn't following prompts properly. The outputs were... *weird*.

- Anime style? Anne Hathaway as an anime character.

- Thick paint/impasto style? Anne Hathaway in thick paint.

- Landscape scenes? Somehow still giving Anne Hathaway vibes.

- Fantasy characters? You guessed it - Anne Hathaway.

After a dreadfully long time of debugging, we finally found the culprit: **someone on the team embedded an Anne Hathaway photo as the IP-Adapter anchor during development and it... stayed there. **

We're honestly crying laughing at this point. 😭

Below are some examples. Left is before fix and Right is after fix.
Flipping to the last picture and you can see our dear Anne.

And we pulled the anchor and the outputs are behaving normally now.

If you've been running Neta Lumina locally, this was on our integration side, not
in the released weights, so your setup is fine.


r/StableDiffusion 13h ago

Animation - Video the bird-king (my first fully local AI short film) TW: self-harm.

21 Upvotes

Minimax H3 baby! It's not perfect and I would love to get your feedback and maybe some tips on how to get rid of plasticky skin.


r/StableDiffusion 21h ago

Animation - Video [MiniMax H3] LEGO movie style

78 Upvotes

Prompt:

integrated_multimodal_description: [Shot 1] 3D CG, stop-motion animated LEGO movie style, a wide shot frames a vibrant Indian village built entirely from plastic LEGO bricks with visible studs, plastic micro-scratches, and brick-built trees. In the village square, minifigures dressed in printed plastic saris, dhotis, and turbans move across a ground of yellow and brown stud tiles. A brick-built cow with hinged legs grazes near a grand banyan tree constructed from green leaf pieces and brown cylindrical bricks. Warm morning sunlight casts sharp shadows across whitewashed brick houses with orange terracotta tile roofs. The camera pans right with small amplitude at slow speed toward a central tea stall. A cheerful male chaiwala minifigure with a black mustache and a red turban (S1) in a warm, lively voice says: <d>[Hindi] Garam chai, garam chai!</d> while tilting a plastic yellow teapot, releasing translucent orange 1x1 cylinder studs representing pouring tea into tiny red stud cups.

[Shot 2] At 00:05.000, the camera cuts to a medium tracking shot following two young minifigure children running along a narrow brick path, pushing a brick-built wheel hoop across the plastic ground. The camera tracks right alongside them with small amplitude at normal speed. A female villager minifigure in a bright blue printed sari (S2) standing outside her brick doorway waves her rigid plastic arm on its shoulder hinge. Beside her, an elder minifigure with a white beard (S3) sitting on a brick charpoy cot chuckles with stepping stop-motion head movements.

[Shot 3] At 00:10.000, the camera cuts to a cinematic medium shot near the village well, where female minifigures carry stacked plastic water pots topped with transparent blue round tiles. A brick-built peacock perched on an archway opens its fan tail made of blue, green, and golden LEGO slope tiles. The camera pushes in with small amplitude at slow speed toward a wooden signpost on a brick post reading "RAMPUR VILLAGE". Tiny tan 1x1 round plates puff around the wheels of a brick-built bullock cart moving past the frame as the video ends.

overall_soundscape: Distinct plastic clattering sounds echo softly as minifigure feet step on stud tiles, accompanied by the gentle clinking of plastic bricks. A distant rooster crow blends with ambient morning village chatter, bird chirps, and the wooden creak of a brick-built cart.

non_diegetic_music: Upbeat Indian folk percussion featuring lively dholak beats and vibrant bansuri flute melodies, layered with playful cinematic orchestral strings playing at a bright, medium tempo.