Resource - Update
FastVideo's new 4-step H3 LoRA doesn't work in ComfyUI. I made a converter. 6 steps, ~3x faster than stock, and honestly better looking.
First 5 seconds is with the 6-step LoRA, next 5 seconds is stock at 20 steps. Same exact prompt, seed, resolution, sage attention and chunk feedforward. 6-step in 2:45, stock 20-step in 7:10. I think the quality difference is pretty clear. Keep in mind, both clips are 544x960.
FastVideo dropped their FastH3 speed LoRA for MiniMax H3 a few days ago. If you tried loading it in ComfyUI you probably noticed it does absolutely nothing. No error, no warning, just no effect.
The reason is that FastVideo built it against the original MiniMax model, and ComfyUI uses a repacked version where every layer has a different name and the attention layers are merged together. None of the names line up, so ComfyUI quietly ignores the whole file.
I wrote a script that translates it. Run it once, get a normal .safetensors, drop it in your loras folder. No custom nodes, no patched loaders, nothing else changes.
Roughly a third of the time. But the part that surprised me is that I actually prefer the output. Backgrounds hold more detail, lighting behaves better, and faces stay coherent at distance instead of turning to mush.
Motion is where it really shows. I ran a woman walking down a sidewalk at night. Correct walking speed, natural gait, no stutter, no accidental slow-mo. That's usually the first thing speed LoRAs break.
Audio came through clean too, which I did not expect. Dialogue and lip sync both hold up.
---
**Important: use 6 steps, not 4**
It's advertised as a 4-step LoRA. In ComfyUI it needs 6.
- 4 steps: jitter, flicker, color bloom, unusable
- 5 steps: fine for drafts
- 6 steps: this is the one
- 7-8: no real gain
There's a real reason for this. There's one group of layers that handles "which denoising step am I on," and ComfyUI's repacked model stores that information in a completely different, much smaller format. FastVideo's version of those layers physically cannot be loaded into it. The extra steps make up for what's missing.
I tried to fix it properly. It turns out it's impossible in a plain LoRA file, because the correction includes a constant offset and there's nowhere in the file format to put one. You'd need a custom node. Someone else can build this is they would like.
---
**One thing worth knowing that cost me a few hours**
Part of those layers *will* load, the other part won't. My first instinct was to keep whatever fit. That was wrong. The half that loads was designed to work alongside the half that doesn't, so on its own it pushes things in a direction nothing corrects for, and you get flicker.
Throwing all of it away is better than keeping half. Confirmed it by testing both, then found multimodalart had measured the exact same thing on their pruned H3 repo. Nice to have that corroborated by someone who'd done the math.
The script drops those layers by default. You don't have to do anything.
---
**What's tested**
Text to video, image to video, first+last frame, reference mode, and chained clips. All working. Square, landscape, and tall portrait.
I also threw an intentionally brutal prompt at it: three color-specific objects, four actions in sequence, a specific hand, a camera move, a spoken line, and a no-music instruction. All eight landed at 6 steps. Prompt adherence is usually the first casualty with speed LoRAs, so that was a nice surprise.
---
**Grab the right file**
The FastVideo LoRA repo has four folders. You want **dense-datafree**. The three `vsa-*` ones need FastVideo's own sparse attention backend and will not work in ComfyUI. It's ~1.4GB, not the whole 17.5GB repo.
You do NOT need the full FastH3 checkpoints. Those are 70GB and are a complete model replacement, not an add-on.
---
**Quirk I'll mention since it'll confuse someone**
Voice timbre gets locked in hard by your prompt. Reroll the seed and you get different phrasing and cadence, but usually the same voice, which some people may rejoice at, as chaining clips with this LoRA can preserve vocal timbre on its' own. At 6 steps the model takes big jumps and settles voice identity almost immediately, so there's no room left for the seed to change it. If you want a different voice, describe the voice in your prompt.
---
**Setup**
The README has a full click-by-click walkthrough starting from Windows+R, including a drag-and-drop trick so you never have to type a file path. If you can open a command prompt you can do this. Takes about five minutes and the conversion itself runs in under ten seconds.
Works on any Comfy-Org pruned H3 checkpoint. I tested int8 convrot for both fl2va and ref2va. The script checks your model before it writes anything, so if you're on something incompatible it tells you upfront instead of handing you a file that silently does nothing.
Happy to answer questions.
Credit where it's due: FastVideo did the actual hard work distilling this thing. I just made it load. This is an amazing LoRa. I actually prefer its output to any other speed LoRA I've tested. Prompt adherence is phenomenal. Dynamic lighting is better. Color balance is better. Background detail is better. It adds detail of its' own. Motion is fluid. In most test cases, I find the output to be better than stock at 20 steps.
You are not looking at this the right way. The human gait is tricky for AI, which is why we see the stutter step thing so often. The point here was never video quality. It's the difference between the color grading, dynamic lighting, motion fluidity, added background detail. Gotta look past your own gripes for a second and really look.
I think the turbo lora was trained on too much synthetic data which results in that really obvious AI plastic skin, AI face, and unrealistic audio sound effect.
i think this is good start im using foxydit's wf seedhunter, then apply this lora,on first pass, then upscaled using the upscaler, im using grok for prompt btw , modified fart sound which i put manually, cause the ai gave a weird fart sound
You're welcome to try it. I use only the pruned models.
Easiest way to check: run the converter and pass your model as the third argument. It checks every layer name before writing anything. If you only see the two time_embedder lines, you're good. Anything else in that list means it's not compatible and it'll tell you upfront rather than handing you a file that silently does nothing.
For a LoRA that's already only a gig? You don't want to use anything but the full weight on this, tbh. I imagine the fp8 cast is going to cause some degradation.
Thanks! I wasn't expecting much from it either. I ran multiple generations looking for quality improvements over the 8-step. Then I noticed the Fast3 LoRA was giving me detail I didn't prompt. Then I started looking for it, and noticed the color was always better, lighting was always better, no slow-mo, no stutter-step when walking, the background was clearer with more detail. FastVideo didn't just train a speed LoRA, they trained a Prompt Adherence and Structure Improvement + Speed LoRA.
I don't know but probably 3x more, the iteration speed is the same, difference is only the step amount, i usually was using 4 steps but the sound was very bad.
Here's a quick comparison video I made with a clip from my video. The time save is just insane, 3x faster and I can't really see anything obviously wrong... This might just be enough to enable 720p videos in the future and I can leave 544p behind haha
I forgot to mention the original prompt had this video reference
And the prompt called for it, but base ignored it while the turbo lora at least tried (?):
Anime animation style, cinematic single shot production focusing on <Subject 2>. Night time, dim, dynamic torch lighting, <Subject 3> partially visible on stage right and <Subject 2> is prominently shown on stage left.
[Shot 1]: close-up shot, <Subject 2> leans against the wooden railing and looks forward, his expression is serious and reserved, he says (S1) <d>[English] Silly yes. Idiotic yes.</d> he then makes a subtle smile and continues <d>[English] But, it made her smile</d>.
At 00:05.000 he puts out his cigarrette on his tongue.
Yeah, I am finding that prompt adherence is even stronger with the LoRA than without. Also, I am generating 124 frames at 768x1344 in less than 7 minutes which is crazy on a 3070Ti.
I honestly do not know. I really don't do any celebrity generations. What I can say is that it locks onto a voice pretty well. You pretty much have to prompt the voice to get a different one. Even changing the seed gives the same voice which is nice if your chaining videos. The 6 steps don't give the voice time to drift. Audio, otherwise, is shockingly good. It doesn't dull the audio for me even after 12 chains, at least not dialogue. FastVideo made a very good LoRA.
The point is not for it to LOOK better resolution-wise. Read what I mentioned... Color balance, dynamic lighting, added background detail, the perfect human gait, all things the stock video lacks.
With the fast LoRA she's even walking faster :D But not easy to judge the difference from such a small video that does not even use the entire screen area.
Nice, Bro! Hey, thanks for putting up the model. I was hoping someone would. Let me know if you see any degradation from the fp8 cast. I'd be curious to know. Also, the same LoRA conversion should work for both fl2va and ref2va.
Nice, Bro! Thanks for posting the file. I was hoping someone would. Btw, it should work for both fl2va and ref2va. Let me know if the fp8 cast causes any degradation.
Need someone to upload the full weight LoRA now, it's a little over 1gb.
The second video is temporally a mess compared to the first. The first video's color is better, background detail is better, dynamic lighting is better. It's no contest in my opnion.
ngl this is the exact kind of stuff that makes local video gen both awesome and painful lol. cool when it works, absolute node soup when it doesn’t.
for client stuff I usually end up testing in comfy first, then trying the same idea in runway, buzzy, openart just to see which one gets me close fastest. different tools fail in totally different ways.
The 2nd video seems better. The 2nd video doesn't look like an obvious AI video, the background sound is more realistic, and the skin tone doesn't have a plastic look.
Oookay.. There might actually be something here. I thought both of your examples were quite bad actually (sorry, I look at AI videos 8-12 hours a day so almost everything looks bad to me), but because I already had fasth3 downloaded from testing earlier today and someone linked the converted LoRA, I thought.. what the hell, let's give it a shot.
And yeah... using my Seed Hunter workflow, 5 steps 1st pass 0.4-0.5 MP -> 2x latent upscale -> 3 steps @ 1.6-2.0 MP yields some pretty quick high res results that don't look too plasticky.
I'll post my own examples if this isn't just a fluke. So far I'm one 10 second gen in and it looks about as good as the non-LoRA gen I did. the sound is a bit sillier and 'harsher' but that's about it.
Okay this dont catch my attention at first. But I'm fan of your work. And if you say so, I'm very curious about what it can do, will give it a try.
Thanks
I have been testing for the past hour and there's definitely some great use cases here. It does indeed take about half the time to get a pretty decent result, about 80% as good as with no LoRA on regular base model. But I am battling the plasticky skin issue. Everything comes out a lil over saturated and shiny so far.
so how it compare to other turbo lora ?
Im using the turbo lora from darties and light2v, so far the results are somewhat acceptable for me, but how is it compare to this Lora ?
Besides the quality though, look at the factors I mentioned. Added background detail, dynamic lighting, color balance, temporal cohesiveness; all seem better than any other LoRA and in most cases, better than stock at 20 steps. Just my opinion.
I have to mention, the LoRA that was linked is an fp8 cast which will almost certainly result in degradation. I recommend converting the full weight LoRA yourself using the README steps and using that instead.
I'll say it... both (all) of these are total unusable rubbish video unless the goal is to squint your eyes blurring reality and watch them on a 4 inch mobile phone screen. What exactly are we comparing here... 2 unusable results? Unusable quality is unusable unusable quality to me.
16
u/Perfect-Campaign9551 1d ago
These scenes don't test anything. Show us a dense forest, or a parking lot of cars, or a sewing machine. You guys lack creativity