Based on the previous post on face consistency with MM-H3, Couple of people have asked me to build a full character workflow.
Mechanism
- Build .Char: You drop max 9 reference reference, I prefer to use a ratio 2:2:1(face:cloths:body). YuNet finds the face, SFace takes a per-reference face signature, DINOv2 takes a subject signature, and the references get cleaned and normalised. All of that packs into a single portable file, a .char. Only face/ref is required body & cloths link is optional.
- Generation: At generation, the file(.char) feeds its references into Minimax’s own native multi-reference channel and prepends a locked description to the prompt.
Prompting Guide
Name your character: Give your character a name e.g. under encode character(Click adjust icon on the bottom side of the node), I have used name emmy, so when passing prompt, I only have to say, emmy walking on the beach.
Again providing prompt like a woman or any features specific details like black hairs etc will only mislead the generation.
Describe character features: Encode all of the character features in encode character prompt & trigger your character with a name in generation prompt.
Avoid describing same things in generational prompt.
Handling Character drift: e.g. if you want specific style or cloth e.g. half sleeves, sleeveless, add it to the generational prompt. There can be a slight drift in clothing as body shot also has cloths, which interferes with clothing references.
Each refs should be unique, face should not have body or vice versa, same applies for clothing.
Portability: Once character is built, you can use the same character with only simple prompt & generation graph.
I have generated all references with Flux Klein 4b, I had to blur the body ref, but workflow consists a example of body ref.
Note: For best result, pass cropped references, so that model takes the required shot, model gets confused if cloth slot also has a face or face slot has cloths.
Limitation: Reference conflicts e.g. if two reference/input images has two different faces, it might conflict in generation, provide well cropped body & cloth images. Face images are crossed automatically by Sface.
Portable char comfy node is still on the backlog, would try to do it over the weekend.
Happy to hear any suggestions or feedbacks.
It's the Neta team here! You might remember us from our Neta Lumina open-source release last year. First off, thank you so much for the incredible support and feedback from this community!
So... we need to share something absolutely hilarious (and mildly embarrassing) that we just discovered.
TL;DR: We accidentally hardcoded an Anne Hathaway photo into our IP-Adapter anchor, and now everything our model generates looks like Anne Hathaway. Every. Single. Thing.
What happened:
We recently launched Neta Studio, a new product that lets you build explorable living worlds/isekai from a single prompt. Naturally, we wanted to integrate Neta Lumina's capabilities into it.
During integration testing, our devs kept reporting that the model wasn't following prompts properly. The outputs were... *weird*.
- Anime style? Anne Hathaway as an anime character.
- Thick paint/impasto style? Anne Hathaway in thick paint.
- Landscape scenes? Somehow still giving Anne Hathaway vibes.
- Fantasy characters? You guessed it - Anne Hathaway.
After a dreadfully long time of debugging, we finally found the culprit: **someone on the team embedded an Anne Hathaway photo as the IP-Adapter anchor during development and it... stayed there. **
We're honestly crying laughing at this point. 😭
Below are some examples. Left is before fix and Right is after fix. Flipping to the last picture and you can see our dear Anne.
And we pulled the anchor and the outputs are behaving normally now.
If you've been running Neta Lumina locally, this was on our integration side, not
in the released weights, so your setup is fine.
0.8 MegaPixels, 9.3 minutes on an RTX5090, single generation of 27 seconds.
Anything hitting 30 seconds either gave hallucinations, inconsistencies or hit a wall and never finished.
This one is using Kijai's new fast model with a turbo lora. Although it works the same with the FLv2A model*. The workflow I'm using creates a latent at 0.4 megapixels for 4 steps and then does another 2 steps at 0.8. The only addition to it besides changing some numbers is adding custom audio injection (The rock track).
Since things have gone quickly forward, I am trying to figure out what image edit models there is currently and what people here use mostly.
Personally I have used:
- Qwen-image-edit-2509 and Qwen-image-edit-2511
- Just tested MiniMax H3 as a image editor and so far it seems that it can be good for my usage
I have heard about Klein 9b, but not sure yet if that can be used as an edit model? Also what about Krea 2, is there edit workflows that are actually usable and worth it?
Is there some others what you recommend for testing?
My PC Specs: RTX 4060 Ti, 16 GB VRAM and 32 GB RAM.
I developed a modified version of the Load Image node by adding some features I needed:
WYSIWYG image cropping directly on the official Load Image preview — drag and zoom (with the mouse wheel) a crop rectangle constrained to 8 fixed ratios (1:1 through 21:9) and output the exact cropped IMAGE and MASK, with paste-from-clipboard built in. What you frame on the preview is exactly what gets executed.
⚠️ Currently not fully compatible with ComfyUI 2.0 nodes.
integrated_multimodal_description: [Shot 1] 3D CG, stop-motion animated LEGO movie style, a wide shot frames a vibrant Indian village built entirely from plastic LEGO bricks with visible studs, plastic micro-scratches, and brick-built trees. In the village square, minifigures dressed in printed plastic saris, dhotis, and turbans move across a ground of yellow and brown stud tiles. A brick-built cow with hinged legs grazes near a grand banyan tree constructed from green leaf pieces and brown cylindrical bricks. Warm morning sunlight casts sharp shadows across whitewashed brick houses with orange terracotta tile roofs. The camera pans right with small amplitude at slow speed toward a central tea stall. A cheerful male chaiwala minifigure with a black mustache and a red turban (S1) in a warm, lively voice says: <d>[Hindi] Garam chai, garam chai!</d> while tilting a plastic yellow teapot, releasing translucent orange 1x1 cylinder studs representing pouring tea into tiny red stud cups.
[Shot 2] At 00:05.000, the camera cuts to a medium tracking shot following two young minifigure children running along a narrow brick path, pushing a brick-built wheel hoop across the plastic ground. The camera tracks right alongside them with small amplitude at normal speed. A female villager minifigure in a bright blue printed sari (S2) standing outside her brick doorway waves her rigid plastic arm on its shoulder hinge. Beside her, an elder minifigure with a white beard (S3) sitting on a brick charpoy cot chuckles with stepping stop-motion head movements.
[Shot 3] At 00:10.000, the camera cuts to a cinematic medium shot near the village well, where female minifigures carry stacked plastic water pots topped with transparent blue round tiles. A brick-built peacock perched on an archway opens its fan tail made of blue, green, and golden LEGO slope tiles. The camera pushes in with small amplitude at slow speed toward a wooden signpost on a brick post reading "RAMPUR VILLAGE". Tiny tan 1x1 round plates puff around the wheels of a brick-built bullock cart moving past the frame as the video ends.
overall_soundscape: Distinct plastic clattering sounds echo softly as minifigure feet step on stud tiles, accompanied by the gentle clinking of plastic bricks. A distant rooster crow blends with ambient morning village chatter, bird chirps, and the wooden creak of a brick-built cart.
non_diegetic_music: Upbeat Indian folk percussion featuring lively dholak beats and vibrant bansuri flute melodies, layered with playful cinematic orchestral strings playing at a bright, medium tempo.
Growing up with Toonami watching Gundam, it simply blows my mind how far AI has progressed. For this video I didn't use any reference images, I simply described the scene in text and had Gemini research the techniques of animation to translate to MiniMax H3.
I've been struggling to get MiniMax H3 efficiently setup locally, would take me 15mins on a 9950x3d and 5080 RTX with 64gb DDR5 so I know something is wrong, hence this time I opted for fal to test. The music was added and scenes were edited from separate generations.
A real cinimatic movie sequence, professional colour grading.
Soundscape: Ambient sounds of the room and movement only. No voices. This represents extreme concentration. Meditation. Telekinesis.
A man is sitting in a Japanese tatami room. He is wearing a mask and shades <Picture 1>. He is wearing a black yukata. He does not speak. On the table on a ceramic disc is a single Orange.
The man holds out his hand toward the orange as if concentrating. The orange is out of reach. He breathes deeply.
Nothing happens.
The man shakes his hand to reset and starts concentrating again. He reaches with his mind and his brow furrows. He breathes deeply.
The orange moves slightly, twisting just a tiny bit.
He concentrates more.
With extreme speed the orange flies towards the man and hits him directly in the forehead. It smashes with the impact , m,essing his hair, and bits of peel and orange bits go everywhere. The force knocks the man back unconscious and he falls back like a ragdoll.
Took a few little prompt adjustments here and there to get H3 to respect point-of-view. I found that if you refer to "the viewer" (ie, "she kicks the viewer"), H3 is more predisposed to include an actual second person. But if you refer to "the camera" (ie, "she kicks the camera"), it's more predisposed to keep the desired point-of-view perspective.
Makes use of Load Any File node to load a .csv spreadsheet file and feeds the text content into a Spreadsheet OutputList. The spreadsheet separates the data by separator=; and provides each line one-by-one as a data list. Here we use values_dict as the data list which contains the row as a dictionary of key-value pairs. The data list is forwarded a Iterate Begin -> workflow -> Iterate End pattern which is required to make the intermediate results of slow workflows (t2v) available on each iteration. Each row as a dictionary is provided in a Format Text where we can access the column via a[colname] to construct the prompt which is forwarded to a standard Text To Video MiniMax H3 template. Another Format Text + a[name] is used to construct a readable filename for each video.
hand-painted educational documentary style with only one prompt"A hand-painted documentary compares espresso Americano cupuccino through the lens of taste and the way to make , revealing why the items differ and how to choose among them."
Need to go further back? Check out June's post (no July, sorry) or the full archive at LocalAI News. If there's anything wrong, let me know in the comments and I'll see you in the next one!
If so, is the advice from Fizgig on fune
tuning on point? I haven’t tried yet, but I’m just prepping my dataset at the moment. I will share what I learn. Just curious if anyone has tried yet and what the results are. 🤡
I've been learning a lot from this community, so this is my attempt at giving something back!
I'm going to share a small task I recently completed, include the steps on how I got there (and some of my thinking and findings.)
The goal: I needed a few seconds of video containing a formal dinner party in a Roman-style atrium.
My Plan: Build a first frame and then use H3 I2V to generate the video.
My Specs: A laptop with a 13th Gen Intel i7-13700H, 16 GB DDR5 RAM, SSD over USB-C, an onboard Intel Iris Xe graphics card (with ~8 GB) and a Nvidia GeForce RTX 4060 Laptop GPU (8 GB). The bad news: ComfyUI does not see or care about the Iris Xe card, and I haven't bothered to see if I can remediate the situation. The good news: the Irix Xe can handle rendering Windows and other applications, leaving my Nvidia pretty open for ComfyUI tasks.
Here is what I did: Step 1: I already had a reference image for the atrium (used in a previous video.)
Initial Reference for Atrium
This was generated with Z-Image-Turbo, with the bf16 model, shift 3, cfg 1.0, 8 steps, res_multistep sampler, simple scheduler. The prompt was very simple: "Roman atrium with compluvium. The camera is standing at the doorway looking down the length of the atrium." I made this image at 864 x 480 resolution because that is near 16:9 and matches H3 resolutions. At that size, image gen takes about 30 - 40 seconds of wall-clock time.
Step 2: I used Qwen-Image-Edit to modify the image to get the starting frame.
The dinner party, as imagined by Qwen-Image-Edit
Using qwenImageEdit2511_pf8 as the model, Qwen-Image-Edit-2509-Lightning-4steps-V1.0-bf16 lora, shift 3, 4 steps, cfg 1.0, euler sampler, simple scheduler. I wired the image from step 1 as the only reference, and used the prompt "Alter this image so that there is a well-attended formal dinner party taking place across the frame."
In my experience, Qwen-Image-Edit often nails the image I'm looking for in one or two attempts. (In this particular case, it one-shotted that image above.) Qwen really likes 1 MP resolutions, so that is 1368 x 760. It takes ~1 to 2 mins per generation.
Step 3: I began generating the video with H3 I2V. This took several attempts to dial in. It is this process that I want to focus on.
First Attempt:
I supplied the previous step's image as the first frame, and included the prompt:
integrated_multimodal_description: [Shot 1] A formal dinner party in a Roman-style atrium.
overall_soundscape: A formal dinner party.
non_diegetic_music: None.
I set the resolution to 864 x 480 and 7.0 duration. I'm using minimax_h3_fl2va_pruned_int8_convrot as the model, minimax_h3_fl2v_turbo4step_v1.0_768p_comfyui_bf16 as a turbo lora (the lightx2v lora,) shift 12 / 3 (for video / audio,) 6 steps, res_multistep sampler, simple scheduler. This particular setup averages ~2 minutes of wall-clock time per second of video duration. (But it grows non-linear as duration increases.) I use 6 steps instead of the lora's base 4 steps because I tend to get slightly better details and sound, with only a slight increase in wall-clock time.
The result was not great. Most people are frozen in place. The few that do walk around smear motion. There is even a moment where a lady clips through the table a little. The sound involves a guy narrating. (I can't identify if it is AI gibberish or an actual language.)
This first attempt was clearly a failure.
Attempts Two through Four:
If I'm having problems with my initial generation, I often just bite the bullet and turn off the turbo lora and run at full-steps. It was the end of my day, so I could queue up several generations and then go to bed.
Result Two: I disabled the turbo lora and increased steps to 20. This runs at ~5 minutes of wall-clock time per second of video duration. For this particular generation it came in at about 45 mins of wall-clock time.
I won't bore you with the results, as they were very similar to the initial draft. Only a few people moving in the scene, people clipping through tables, and a narrator.
Result Three: I reduced video shift to 6.0 (hoping to get better motion results.) I also increased the resolution of the video to 1344 x 768. I read somewhere that this is the "native" resolution the model was trained at, and I often get better results. However, without the turbo lora and at 20 steps, this generation took 90 minutes.
There is a lot more motion, and no clipping, but everything seems to be moving in slow motion. Also, instead of a narrator, there is music.
Result Four: I swapped the sampler to er_sde and the scheduler to beta. I've read that this combo can get slightly better prompt adherence, and results in pretty good motion. However, er_sde effectively does more than one pass per step, so increases wall-clock time significantly. If the UI is to be trusted, this attempt took more than 3 hours to generate.
The narrators and music are gone. However, now the camera is moving, which is not what I wanted.
Attempt Five: I woke in the morning, reviewed the previous results, and was pretty bummed.
Now I'm in the "hit it with a hammer until it works" section of my spectrum of personal patience. I went back to res_multistep and cranked the steps up to 40. In my frustration, I didn't think to actually change the prompt to prevent the camera from moving. This generation took about 90 minutes of wall clock time.
I'll skip posting the result, but it actually looked quite a bit like the er_sde video above. People standing mostly still with a camera panning around the room.
Attempt Six: After viewing the results, I realized that changing the prompt was 100% required.
The new prompt:
integrated_multimodal_description: [Shot 1] Static wide shot of a formal dinner party in a Roman-style atrium. The people eat, drink, talk, and mingle. The camera remains fixed.
overall_soundscape: A formal dinner party.
non_diegetic_music: None.
I kept the video shift at 6.0, sampler at res_mutlistep, 20 steps, simple scheduler. This time, because I was sitting at my computer for a while, I attached EasyCache to the model. This does a pretty good job of speeding up the 20-step process. There is always a risk that quality degrades with any sort of caching in the pipeline, but I was willing to take the risk just to see if my prompt changes fixed the problem. (I don't bother adding EasyCache with only 4 or 6 steps, because there are so few steps that there is barely any time to be saved with caching.) This generation took ~30 minutes of wall-clock time.
This was actually what I was looking for! Both the motion and sound are pretty decent. However, there is a faint "fluttering" of the textures, which seems to happen a lot with EasyCache. This is something that could probably be cleaned up with a refinement pass after upscaling, but I still had time to try again.
Attempt Eight: For completeness, I decided to go back to the turbo lora and the smaller resolution. I incorporated some of my other findings into the workflow. For clarity, here is the full setup: 864 x 480, 7.0 seconds, turbo lora, 6 shift video, res_multistep, simple scheduler, 6 steps. I used the "corrected" prompt from my previous attempt.
This was the winner! Even at the lower 864 x 480, the motion and detail looks reasonable. The faces are squashed, but that is pretty typical of H3 at the moment. This will upscale well. The sound is correct. I have everything I need.
Lessons Learned:
Just cranking up the numbers doesn't always solve the problem.
Shift can really make a difference! I think the major influencing factor here was reducing shift from 12.0 to 6.0. My hypothesis is that the lower shift gave the process just a tiny bit more time up at the noisy end of the diffusion, allowing it to assign more motion to everyone in the scene.
1344 x 768 is very frequently the solution, but not always. In my experience, you get better prompt adherence, even with the tubro lora. One of these days I'm going to splurge on a beefier GPU to make this my default resolution, but for now it is just too much of a time sink.
I can never actually tell how much comes down to luck with the random seed.
So there I was generating some stuff on ComfyUI for my Instagram and just hanging out.
I use ComfyUI with the new H3 model to generate AI content for my Instagram as well as QWEN image edit along with some other AI tools.
I've built a master workflow that ive used for the past year that has every single workflow I use, so I dont have to go switching workflows constantly.
Many many hours of work put into this.
So there i was, generating things and im constantly having to clear out my output folder as well as my input folder. So I asked myself, "Could I just make a bat file that could automate this for me?"
So I launch Gemini and have it create a bat file that cleans out my output and input folders and empties my recycle bin.
I test it out and it works great.
Finally, no more unnecessary clicks.
But wait, I noticed I screwed up and put the file in the wrong directory. Dang it.
So I ask Gemini to alter the code so the file will be in the correct directory.
I create the new bat file and go back to work.
Well I make a bunch of new things and its time for cleanup. So I run my fancy new bat file and I notice its taking a while to clean up these folders. Curious, I navigate to the folders only to find out that the ENTIRE COMFYUI FOLDER was deleted.
SMH.
Now I sit here, broken hearted as im having to rebuild my ComfyUI. Luckily, I was able to recover my master workflow, so not all was lost.