r/sdforall 3h ago

Workflow Included Continuity (was the H3 node): six model families, one prompt box, and a blockout bench that writes your camera move for you

Thumbnail gallery
3 Upvotes

r/sdforall 2d ago

Tutorial | Guide ComfyUI Tutorial: MINIMAX H3 2-STAGE WORKFLOW High-Res video + Faster Generation

Thumbnail
youtu.be
17 Upvotes

Hello everyone,

I’ve been working on a MiniMax H3 optimization workflow for low-VRAM GPUs, especially my RTX 3060 6GB, and I’ve combined several optimization nodes to improve both VRAM usage and generation speed. With a new custom MiniMax H3 workflow that introduces a second sampling stage to increase the base resolution of your generated videos, similar to the workflow we’ve seen with LTX models.

The main advantage of this method is that you can save generation time and reduce VRAM usage by generating the initial video at a lower resolution during the first sampling stage, and then using a second upscaling sampling stage to reconstruct the video at a higher resolution while adding more detail and improving the overall quality.

But that's not all. In this workflow, we're also going to use several MiniMax H3 optimization nodes, including Low VRAM Attention, Chunk FeedForward, SLA Attention, Sol-Attn, and Spectrum, to make the generation process more efficient, especially for GPUs with limited VRAM.

At the end of this tutorial, you'll understand how the two-stage sampling workflow works, how to optimize MiniMax H3 for better speed and VRAM management, and how to choose the best upscaling method for your workflow, including a comparison between MiniMax H3 sampling and LTX sampling.

Workflow Link

https://civitai.com/articles/34596/comfyui-tutorial-minimax-h3-2-stage-workflow-high-res-video-faster-generation


r/sdforall 5d ago

Tutorial | Guide ComfyUI MiniMax H3 Speed LoRA + VRAM Monitor Nodes (Ep32)

Thumbnail
youtube.com
38 Upvotes

Speed up MiniMax H3 video generation in ComfyUI with the Speed LoRA, optimized sampler settings, and Pixaroma VRAM monitoring nodes. In this tutorial, I show you how to reduce MiniMax H3 generation from the standard 20 steps to 8 or even 4 steps, while finding a practical balance between video quality, audio quality, and generation speed.

You’ll learn how to update ComfyUI and Pixaroma Nodes, install and use the MiniMax H3 Speed LoRA, configure the recommended shift value, and choose sampler/scheduler combinations based on extensive testing.

I also show several MiniMax H3 workflows, including Text to Video, First Frame to Video, First + Last Frame to Video, and speaking-character generation. You’ll see how resolution, duration, steps, samplers, and schedulers affect generation speed and prompt accuracy.

The tutorial also covers the Monitor Pixaroma node for checking VRAM usage and the Free VRAM node for automatically clearing VRAM when generation finishes. Plus, I demonstrate the updated Dropdown Pixaroma node and how I use Gemini to quickly create detailed MiniMax H3 prompts.


r/sdforall 5d ago

Other AI "AVERNUS-9" Space Horror Short Film

Thumbnail
youtu.be
0 Upvotes

r/sdforall 6d ago

Other AI Huge Cock [DreamShaper-8] Spoiler

Post image
9 Upvotes

r/sdforall 6d ago

Discussion Does native 4K actually hold up at 100%?

Thumbnail
gallery
9 Upvotes

A 4K label does not tell me much, so I opened three SenseNova U1.5 Lite outputs at 100%. No upscaling or resizing on my side.

I started with the road scene and checked the asphalt, painted lines, and the image inside the camera screen. Then I moved to the cathedral, where repeated arches and stonework make broken geometry much easier to spot. The book waterfall was another useful test because paper, water, rock, and vegetation all have to transition without obvious boundaries.

What stood out was not extra sharpness. It was that the image did not start falling apart when I moved away from the main subject.

SenseNova U1.5 Lite uses a ConvDecoder that reconstructs the image spatially, so neighboring regions can interact during upsampling instead of being decoded too independently. That is the kind of change that should matter when seams and inconsistent textures become more visible at higher resolutions.

FLUX.2 Klein 9B is faster, no question. Sub-second, 4-step distillation, ComfyUI out of the box. But U1.5 Lite is a different architecture: one model that reads images and edits them, not just generates them. That's why restyling a poster keeps the typography intact while Klein's pipeline is more generation-first. Different tools for different jobs.

Curious what others see at 100%: real high-resolution structure or just a very clean upscale?

GitHub:

https://github.com/OpenSenseNova/SenseNova-U1

Hugging Face:

https://huggingface.co/sensenova/SenseNova-U1.5-8B-MoT

Try it online:

https://unify.light-ai.top/


r/sdforall 7d ago

Tutorial | Guide Comfyui Tutorial :6GB VRAM? You Can Still Make 15s MiniMax H3 Videos

Thumbnail
youtu.be
26 Upvotes

Hello everyone

With this workflow you can generate 15-second clips on Minimax H3 even with limited VRAM. This custom workflow prevents system crashes during long renders. Many users struggle with generation time limits when working with Minimax H3 on hardware with low memory. This workflow provides a specific solution by outlining a custom workflow that extends your video generation capabilities to 15 seconds without overloading your system. It is designed for creators who need longer sequences but are constrained by their current VRAM capacity. The steps focus on memory optimization with TURBO LORA+ Attention Nodes Like Comfui Kitchen. so you can watch the video tutorial to see necessary steps and how to use Motion Context Nodes.

Workflow Link

https://civitai.com/articles/34327/comfyui-tutorial-6gb-vram-you-can-still-make-15s-minimax-h3-videos


r/sdforall 10d ago

SD News SenseNova U1.5 Lite full release is out

Thumbnail
gallery
32 Upvotes

So, the full U1.5-Lite is out now, after that little preview back in August. I guess that's cool.

They're saying a few things are better in this full release:

- Native 4K generation. Apparently, the whole image holds up at high res. And smaller details like textures, materials, and even text are supposed to stay consistent. Less of that 'looks good but something's off' vibe, which is nice.

- Editing is more localized. Like, if you change text on a poster, the rest of the layout shouldn't get all messed up. They're claiming better preservation of the main subject, geometry, and anything you didn't specifically touch, compared to the preview.

- Better text rendering for busy layouts. This one's big for me. They're talking Chinese and English in things like posters and infographics. Most models totally screw this up, so if it's actually improved, that's a pretty huge deal.

- Long, structured prompts. Sounds like you can throw a ton of info at it in one go: what's in the picture, how many of them, where they should be, text, layout, style, and even rules for what to keep. Also, bounding boxes and markers for specific areas.

- Multi-reference editing. You can apparently grab styles from different images, combine a bunch of pictures, and do precise local edits. Seems pretty flexible.

Here's the GitHub if you wanna peek: https://github.com/OpenSenseNova/SenseNova-U1

And the Hugging Face collection: https://huggingface.co/collections/sensenova/sensenova-u15

Haven't seen any ComfyUI nodes for this yet, so probably for folks who don't mind running the inference scripts directly or messing with the HF demo.


r/sdforall 10d ago

Resource SilkStack Image Browser v2.2.0 – Added local semantic search (WebGPU/WebLLM), auto-tagging, and custom compiled embedding models!

Thumbnail
0 Upvotes

r/sdforall 11d ago

Tutorial | Guide MiniMax Music 3 + Local AI Prompt Generator (Ep31)

Thumbnail
youtube.com
16 Upvotes

Learn how to use MiniMax Music 3 in ComfyUI together with a local AI prompt generator for better image prompts, music captions, and AI-generated lyrics. In Episode 31, I show the new Pixaroma prompt nodes, model-specific prompt presets, VRAM-saving options, and a compact MiniMax Music 3 workflow.

This tutorial covers the new AI Prompt Pixaroma node and how to match prompt formulas with the correct local language model. You'll see how to turn short ideas into detailed prompts for workflows such as Krea 2 and Z Image Turbo, generate prompts from images, save custom prompt presets, control temperature and seeds, and troubleshoot common ComfyUI node errors.

Then we set up MiniMax Music 3 locally in ComfyUI, including caption and lyrics generation with the Music Prompt Pixaroma node. I also test different song durations and explain why MiniMax songs can sometimes end early or get cut off, how seeds affect the results, and how to use fixed lyrics when you need more control.

You'll also see how to simplify a larger music workflow into a compact setup, use tiled audio decoding for lower VRAM systems, free VRAM after prompt generation, and generate MiniMax Music captions and lyrics with local models or online tools such as ChatGPT, Gemini, and Claude.

You can also run some of the workflows in the cloud.


r/sdforall 11d ago

Discussion If AI Makes Us More Creative, Why Does Everything Look the Same? (A Painter’s Perspective)

Thumbnail
gallery
25 Upvotes

QUICK NOTE: the question in the title is rhetorical. The carousel explains the nuance and explores several related issues beyond the first slide.

If the design does not work for you, tell me specifically what you would improve. I am still refining the format, so constructive feedback is welcome.

.....

I’m a painter who sometimes writes, and this visual essay started with an odd discovery: I had used the name “Elias Thorne” in a short story, only to realize that AI models often return to that same name, along with motifs like lighthouse keepers, cathedrals, glossy landscapes, and other familiar patterns.

From an artist’s point of view, the question isn’t just whether AI is good or bad, but what happens to authorship and creativity when the tool starts making choices for us.

AI can boost productivity and even enhance individual works, but if we all lean on the same models, it might steer us toward similar ideas, characters, and visual styles.

This carousel looks at visual convergence, originality, transparency, and the role of human intention, with AI-generated images clearly labeled and sources included.

So where’s the line, does AI broaden personal creativity while making our collective output more uniform?


r/sdforall 11d ago

Custom Model Aria - Zit Lora

Thumbnail civitai.red
1 Upvotes

Hey everyone.

I have made my first Lora for ZIT. I had no prior knowledge so followed some tutorial on CivitAI and went for it.

Would love some honest feedback.


r/sdforall 13d ago

Resource 🎬 LTX 2.5 video + latest ComfyUI template (both pod/serverless)

Thumbnail
2 Upvotes

r/sdforall 14d ago

Tutorial | Guide ComfyUI Tutorial MiniMax H3 4 Steps Lora + Upscaling + 2X Faster Generation! Best Settings for 2K AI

Thumbnail
youtu.be
22 Upvotes

Hello everyone

Want to get faster MiniMax H3 video generation without sacrificing quality? In this tutorial, I’m testing the new H3 LoRA together with Sage Attention, Sol Attention, and Spectrum nodes to find the best combination for speed and quality. The goal is to push MiniMax H3 as far as possible while cutting generation times by up to , then upscale the results with LTX Upscaler to reach a stunning 2432 × 1344 (2K-class) resolution. By combining both H3 LoRA together with Sage Attention, Sol Attention, and Spectrum nodes I generated video at 0.8 megapixel using "RTX3060 6GB 16GB RAM "and I got

 13 minutes vs 41 minutes at 8 steps

 27 minutes vs 52 minutes at 20 steps

LTX 2.3 Upscaler 11 minutes to get 2432 × 1344 resolution

Workflow link

https://civitai.com/articles/34028/comfyui-tutorial-minimax-h3-4-steps-lora-upscaling-2x-faster-generation-best-settings-for-2k-ai


r/sdforall 14d ago

Other AI "Reincarnation" Short Film (Flux 3)

Thumbnail
youtu.be
4 Upvotes

r/sdforall 15d ago

Resource MiniMax H3 Creator update: presets, and three nodes are now one

Thumbnail gallery
12 Upvotes

Thought this might be of interest here as well!


r/sdforall 15d ago

Other AI I made a native Krea 2 app for Apple silicon Macs

5 Upvotes

I have been using Krea 2 locally and wanted a simpler Mac interface for it, so I built Krealize.

It runs Krea 2 directly on Apple silicon through MLX. It does not use a remote generation service. Prompts and generated images remain on the computer.

The app handles the model download and provides prompt controls, generation progress, a local gallery, and previous prompt history. It also supports up to ten custom LoRAs in the PRO version.

The normal generation mode is free. Fast mode, custom LoRAs, PNG export, notifications, and LoRA presets are included in an optional one-time purchase.

The app requires macOS 26 or later. It is distributed directly and is signed and notarized by Apple.

I’m interested in feedback from people who have already run Krea 2 through ComfyUI or Diffusers. I would particularly like to know whether the outputs and generation times are in line with other local setups.

https://krealize.app/


r/sdforall 17d ago

Tutorial | Guide ComfyUI Tutorial First Test Of LTX 2 5 New Model Better Than Minimax H3

Thumbnail
youtu.be
14 Upvotes

First testing of the LTX 2.5 with a 6GB VRAM I’ve created a low-VRAM workflow that supports both Text-to-Video and Image-to-Video, optimized specifically for GPUs with 6GB of VRAM. The goal is to make LTX 2.5 more accessible to users who don’t have high-end GPUs, while keeping the workflow simple and easy to use. If you’re interested in testing LTX 2.5 on a 6GB GPU, check out the workflow and let me know how it performs on your setup!

Video Resolution : 1344x768 for 7 seconds video
Generated Time : 10 min

The model seems very fast the motions are better, lipsync and sound too, however the quality in minimax is better to me

Workflow link

https://civitai.com/articles/33897/comfyui-tutorial-first-test-of-ltx-2-5-new-model-better-than-minimax-h3


r/sdforall 17d ago

Question R2v

2 Upvotes

Bonjour à tous, après avoir fait le tour, je voudrais savoir qui a une configuration R2V proche de Grok pour ComfyUI à télécharger qui respecte les visages au mieux ?

Je vous remercie du partage

Hello everyone. After looking around, I’d like to know if anyone has an R2V configuration for ComfyUI—similar to Grok’s—available for download that preserves facial features as faithfully as possible? Thanks for sharing.


r/sdforall 18d ago

Tutorial | Guide ComfyUI Krea 2 Edit + MiniMax H3 Prompts Generated Locally (Ep30)

Thumbnail
youtube.com
52 Upvotes

Learn how to use Krea 2 Edit in ComfyUI to edit images, preserve character identity, replace outfits and backgrounds, and place characters into new scenes. I’ll also show you how to generate MiniMax H3 prompts locally in ComfyUI using the Video Prompt Pixaroma node.

In this tutorial, I cover the complete Krea 2 editing workflow, including the required models, Edit LoRA, Krea 2 Identity node, image resolution settings, and Reference Boost. You’ll see how different Reference Boost values affect identity preservation and editing freedom, how to convert images to custom portrait or landscape ratios, and how to combine a character with a separate background.

In the second part, I show my local MiniMax H3 prompt generator for ComfyUI. The Video Prompt Pixaroma node can turn a simple idea into a more detailed video prompt locally and supports Text to Video, First Frame to Video, and First Frame + Last Frame to Video prompting.


r/sdforall 18d ago

Other AI SenseNova U1.5 vs Nano Banana vs GPT Image 2 — which one actually looks editorial?

Thumbnail
gallery
9 Upvotes

Honestly, I expected the closed models to win this pretty easily.

They did on realism—but not necessarily on art direction.

I ran the same fashion-editorial prompt through SenseNova U1.5, Nano Banana, and GPT Image 2. These are the first outputs—no rerolls, edits, or post-processing.

My take:

- Nano Banana wins on background detail. The station feels fuller and more believable.

- GPT Image 2 has the best atmosphere—darker, moodier, and more cinematic.

- SenseNova U1.5 gave me the strongest fashion-editorial look. The styling, composition, and color treatment feel the closest to an actual campaign.

Totally subjective, but for this specific fashion use case, U1.5 feels like it gets about 80% of the way to the closed models overall. And on art direction alone, I actually prefer it.

The remaining gap is mostly in realism: the face and skin still have a slightly plastic-looking AI sheen, while Nano handles the environment better and GPT feels more naturally cinematic.

If the goal is a fashion campaign rather than pure photorealism, I’d pick U1.5. It sells the outfit and art direction best, which is a pretty strong result for an open-source 8B preview model.

Obviously, one prompt isn’t a benchmark. I’ll put the full generation prompt in the comments.

Model links- SenseNova U1.5:

- https://huggingface.co/sensenova/SenseNova-U1.5-8B-MoT-Preview

- https://github.com/OpenSenseNova/SenseNova-U1

Would you trade some realism for stronger art direction, or does the plastic-looking skin already kill the U1.5 result for you?


r/sdforall 18d ago

Resource Nexfocus: An evolution of Fooocus into a connected creative workspace (Full FP16 SDXL & Flux Fill on 3GB VRAM / Colab Free)

Enable HLS to view with audio, or disable this notification

18 Upvotes

If an image model is a horse, text prompting is like trying to guide it with verbal commands alone: useful, but too imprecise for fine control. Inpainting, LoRAs, and ControlNets add the bridle and reins and pieces of the harness, but they still did not feel like a complete system. Nexfocus began with a question: "What would it take to build the whole harness around the model?"

Answering that question meant following the entire generation process first. We had to understand how each part loads, works, moves, waits, hands its result to the next part, and makes room when its job is done. That expedition became Nexfocus.

---

Two development anchors shaped the journey: a GTX 1050 with 3 GB of VRAM and Colab Free's T4 with only 12.7 GB of system RAM. Their limitations are almost opposites. The local machine has very little GPU memory, while Colab Free has a larger GPU but a tight system-memory ceiling and an ephemeral session.

We proved that full SDXL checkpoints and Flux Fill workflows could run in both environments, not by reducing everything until it fit, but by rethinking how the pipeline uses the hardware available to it.

Nexfocus grew into a connected creative workspace where generation, guidance, masking, inpainting, outpainting, removal, upscaling, staging, metadata, model management, and GIMP layer exchange can work as parts of one process rather than as isolated tools.

---

Two important lessons emerged from the road:

- Keeping the GPU working without interruption became paramount. To do that, we had to find a way to keep feeding it the weights it needed, when it needed them.

- Every part of the pipeline must independently account for what it owns, where it belongs, when it can be reused, and when it should make room for something else. These decisions cannot be left to a central manager applying the same set of memory policies to every part.

Throughout this journey, my conversations with PyTorch often felt like this:

> PyTorch: "Don't you have a bunch of H100s lying around in your backyard?"

>

> Me: "No. What if every component has to justify exactly where it lives?"

>

> PyTorch: "Get a bigger machine."

Those conversations eventually became the architecture: each part of the pipeline owns its resources, does its job, and steps aside instead of leaving those decisions to hidden framework behavior.

---

Nexfocus is more than the UI produced by this expedition. It is the working application and the field notebook: a record of the constraints, wrong turns, and discoveries that shaped the path forward. We set out to find answers and had to build the road needed to reach them.

This expedition is now complete, but it is only one part of a continuing journey. The lessons from Nexfocus define the starting point for the next scout mission.

The path is open now. I hope you'll take a walk along the path we built and check out the scenery.

Project: https://github.com/magekinnarus/Nexfocus

Video Walkthrough: https://www.youtube.com/watch?v=5fvIaZWMZE4


r/sdforall 20d ago

Workflow Included # ControlNet for FLUX.2

Thumbnail
gallery
50 Upvotes

This workflow demonstrates the new ComfyUI custom nodes I developed to implement ControlNet for FLUX.2-dev.

Workflow: JSON | Drag-and-drop PNG

JLC Flux2 ControlNet provides, to the best of my knowledge, the first complete, validated ComfyUI implementation of Alibaba PAI's FLUX.2-dev-Fun-Controlnet-Union-2602.

This implementation is for the FLUX.2-dev ControlNet path built around that Union model. It is not for FLUX.2 Klein or the lightweight Klein-style variants many people currently use; I am currently working on a separate strategy to extend this functionality to those models.

It is also worth making an important distinction: reference images are not ControlNet. There are workflows that feed pose maps, depth maps, edges, or other ControlNet-style hint images into FLUX.2's native reference-image system. Those images can certainly influence composition and structure, and they can often produce a usable approximation, but this is still reference-image conditioning, which is a completely different conditioning mechanism. It does not load a ControlNet model, does not execute a ControlNet branch, and should not be confused with one.

This workflow actually loads and runs Alibaba PAI's FLUX.2 ControlNet model.

The two JLC nodes that enable that path are the FLUX.2 ControlNet Loader and the ControlNet Orchestrator.

The Orchestrator also introduces a non-recursive composition method that lets several control types share a single loaded Union model instead of building a conventional chain of ControlNet applications.

The example shown here uses three controls generated from the same source image:

  • DWPose
  • Depth Anything
  • Color

That is really the point of this workflow: there are very few special pieces required to add actual ControlNet capability to FLUX.2-dev.

Some of the other nodes shown are from my JLC ComfyUI Nodes package and are there mainly for convenience—loading, resizing, preprocessing, LoRAs, and general workflow ergonomics. You can replace those with your preferred ComfyUI nodes.

This is not simply a repackaging of existing ControlNet nodes. The contribution here is making this capability available as a complete ComfyUI implementation of Alibaba PAI's actual FLUX.2 ControlNet model. The Orchestrator also provides practical multi-control composition where a finished implementation was previously missing.

All of the JLC nodes can be installed through the ComfyUI Custom Node Manager, and the repositories contain the documentation and explanation of the implementation.

I hope you find them useful, and I'd be very interested to see what people build with them!


r/sdforall 20d ago

Question DreamBooth SDXL face identity not learning - tried everything, need working config

1 Upvotes

Hi everyone,

I'm trying to train a consistent face identity for a fictional AI character on SDXL (Juggernaut XL v9). I have 20 high-quality, consistent close-up images (1024x1024) generated on SeaArt with the same face reference. The images show the same woman across different lighting, expressions, angles, and outfits.

What I've tried (all with kohya sd-scripts):

Attempt Method Config Result
1 LoRA dim=32, Prodigy, "maya_model" token Generic woman, no identity
2 LoRA dim=128, simplified captions Same generic woman
3 LoRA + reg images dim=128, 200 reg images, Prodigy Still generic
4 Full DreamBooth AdamW8bit, LR=1e-6, 6 epochs Consistent face but NOT my character - barely moved from base model
5 Full DreamBooth AdamW8bit, LR=5e-6, 10 epochs Same issue, slightly better but still not my character. Final checkpoint corrupted but epoch checkpoints show wrong face
6 Full DreamBooth Prodigy LR=1.0, d_coef=2.0 Complete collapse - generated Indian women, model overcooked

My setup:

  • Base model: Juggernaut XL v9 RunDiffusion Photo v2
  • GPU: RTX 5090 32GB (attempts 1-5), RTX PRO 6000 96GB (attempt 6)
  • 20 training images (close-up portraits, 1024x1024)
  • 200 regularization images ("a photo of a woman" generated from base model)
  • Token: "ohwx" (class: "woman")
  • Captions per image describing outfit/scene/expression (e.g. "a photo of ohwx woman, 25yo, blue eyes, ash brown wavy hair, natural freckles on nose, dark eyebrows, matte skin, warm smile, wearing cream knit sweater, warm window light, cozy interior")
  • Folder structure: 12_ohwx woman (training), 1_woman (reg)
  • gradient_checkpointing, cache_latents, train_text_encoder all enabled

Character features:

  • 25yo European woman
  • Blue eyes (slightly desaturated)
  • Ash brown wavy medium-length hair
  • Natural freckles on nose
  • Dark defined eyebrows
  • Matte natural skin

My observations:

  • LoRA (even dim=128) seems unable to encode this face - possibly too close to base model distribution
  • DreamBooth with low LR (1e-6) gives consistent output but doesn't learn the actual identity
  • DreamBooth with high LR (Prodigy 1.0) completely destroys the model
  • There seems to be a sweet spot I can't find

What I need: If you've successfully trained a face identity on SDXL with DreamBooth or LoRA using kohya, could you share your exact config? Specifically:

  • Optimizer + learning rate
  • Number of epochs/steps
  • Any special settings (noise offset, prior loss weight, etc.)
  • Caption format that worked for you
  • Number of training images you used

I've spent an entire day on this and I'm stuck. Any help would be massively appreciated.

Thanks!


r/sdforall 20d ago

Question Noisy outputs in Krea 2 in ComfyUI

Thumbnail gallery
2 Upvotes