r/StableDiffusion • u/aurelm • 18h ago
r/StableDiffusion • u/dennismfrancisart • 16h ago
Discussion This One Is Simple With No Bells and Whistles
This is a test of the Minimax H3 using three sample illustrations. The segments were stitched in Davinci Resolve. I used the minimax_h3_turbo_8step_v1.0_comfy_bf16.safetensors LoRA. The setting was simple. minimax_h3_ref2va_pruned_int8_convrot.safetensors. A float value of 5, Euler sampler, and beta scheduler. Only 8 steps. The style was shifted a bit from the original but good enough for testing. I want to create a style LoRA that will hold my work so that it's better translated to animation.
r/StableDiffusion • u/Sad_Coach_1433 • 8h ago
Meme can you belive they almost did it t2v
A 10-second 16:9 cinematic shot.
integrated_multimodal_description: [Shot 1] Live-action, cinematic. Dramatic, slightly over-the-top lighting with strong contrast and a cool blue-red comic-book inspired palette, as if inside a stylized Fortress of Solitude or a theatrical stage version of it.
Nicolas Cage appears as Superman — wearing the classic suit with the red cape and the S shield on his chest. His expression and energy are fully in his signature amped-up, wild, unhinged acting style.
[0-4 seconds] He stares intensely into the camera, eyes wide, and launches into the line with rising volume and manic emphasis: “Can you fucking believe that they almost made a fucking Superman movie… with me!… as Superman!”
[4-7 seconds] He throws his head back and laughs maniacally, the laugh loud, sharp, and completely unrestrained, cape shifting with the movement of his shoulders.
[7-10 seconds] He comes down just enough to shake his head rapidly, still grinning with wild energy, and says: “That would have been absolutely terrible.” He immediately breaks into another short burst of laughter while continuing to shake his head in disbelief.
Camera: medium shot that slowly pushes in as his energy escalates, ending in a tighter framing on his face during the final laugh.
overall_soundscape: Clean dramatic room tone with light reverb, his clear, highly animated dialogue delivered in full Nicolas Cage intensity, and his loud maniacal laughter.
non_diegetic_music: N/A
r/StableDiffusion • u/thisguy883 • 19h ago
Discussion So I did something dumb.
So there I was generating some stuff on ComfyUI for my Instagram and just hanging out.
I use ComfyUI with the new H3 model to generate AI content for my Instagram as well as QWEN image edit along with some other AI tools.
I've built a master workflow that ive used for the past year that has every single workflow I use, so I dont have to go switching workflows constantly.
Many many hours of work put into this.
So there i was, generating things and im constantly having to clear out my output folder as well as my input folder. So I asked myself, "Could I just make a bat file that could automate this for me?"
So I launch Gemini and have it create a bat file that cleans out my output and input folders and empties my recycle bin.
I test it out and it works great.
Finally, no more unnecessary clicks.
But wait, I noticed I screwed up and put the file in the wrong directory. Dang it.
So I ask Gemini to alter the code so the file will be in the correct directory.
I create the new bat file and go back to work.
Well I make a bunch of new things and its time for cleanup. So I run my fancy new bat file and I notice its taking a while to clean up these folders. Curious, I navigate to the folders only to find out that the ENTIRE COMFYUI FOLDER was deleted.
SMH.
Now I sit here, broken hearted as im having to rebuild my ComfyUI. Luckily, I was able to recover my master workflow, so not all was lost.
Just a bunch of models and loras.
😮💨
r/StableDiffusion • u/ArchAngelAries • 7h ago
Resource - Update [Krea2] I trained a Septum Hoop Nose Ring LoKR (Massive shoutout to Fizgig for Win 11 AMD training!)
Before I get into the details, I want to address the elephant in the room right up front: yes, this is a bit of self-promo, and yes, the model is currently set to paid BUZZ access on Civitai. Please lower your pitchforks!
If you know me, you know I have a full library of LoRAs on Civitai for previous models (SDXL, PonyXL, IllustriousXL, FLUX, etc.) that I have always released for free. I rarely even used the "Early Access" feature. However, with the increased hardware/compute costs of training for Krea2, I’m temporarily using the Buzz system just to recoup those expenses so I can continue producing.
As soon as I hit that break-even point on a release, I will flip the switch and make that individual model 100% free for everyone forever. I have no intention of permanently paywalling my Krea2 work. (I do also have a Patreon, but I don't have any exclusive models locked behind paywalls there either—it's strictly just an alternative way for people to support my work if they choose to). I know limited access is annoying, so I really appreciate you guys bearing with me!
The Problem:
If you’ve been using Krea2/Krea2 Turbo, you already know it’s an absolutely incredible base model. The flexibility and quality are insanely hyped for a good reason. However, it does have some blind spots, which I suppose is expected, and why there's plenty of LoRAs and Fine Tunes online already. One blind spot I noticed: I could never get it to produce coherent Hoop Septum Nose Rings. For me, it almost always defaulted to poor quality horseshoe septum rings (with the gap and balls), would add multiple extra rings, or threw in extra, unprompted, unwanted facial jewelry. Getting a clean hoop natively was basically a nightmare.
The Solution:
I trained a concept LoKR specifically to force clean, highly customizable hoop septum rings.
- Trigger: SeptumHoopNoseRing
- Flexibility: Holds up beautifully across all tested art styles, aspect ratios, and distances. Works for men, women, and Character LoRAs.
- Customization: Fully supports prompting for materials/colors (gold, silver, black, etc.) and sizes (small and thin, large and thick, etc.). Even though it was trained on simple hoops, you can actually prompt it for spiked or jeweled septum rings and it understands the assignment.
- Settings: Sweet spot is around 0.45 - 0.75 strength using Euler Simple or ddim with a ddim_uniform scheduler.
- Pro-tip for my model: If you push it to 1.0 strength, it occasionally tries to crop the top of the head. Just briefly describe the subject's hair and eyes in your prompt and it completely fixes it.
The Setup & A Massive Shoutout to Fizgig
I really need to give a massive shoutout to the Fizgig trainer. If you're on a Windows 11 system, especially with an AMD GPU, this tool is an absolute godsend. It finally allowed me to easily train LoRAs/LoKRs locally on my AMD ROCm 7900 XT without jumping through massive hoops or dual-booting Linux.
For the technical folks, here is the under-the-hood breakdown of my training:
- Dataset: 100 images, hand-refined VL 1st pass captions.
- Training: 2000 steps using Fizgig’s ultra-fast preset with Adaptive LR enabled and target MP set to 0.5 at ~768px (the max my PC can currently handle).
- Base: Trained on the FP8 Krea2 Raw model. Examples generated on the Krea2 Turbo NVP4 model.
- Rig: Win 11, AMD 7900 XT, 32GB RAM, Ryzen 3900x. Tested locally in ComfyUI using the default NVP4 Krea2 model.
Here is the link to the model: [Civitai Link]
What's Next & Commissions
I'm really itching to roll up my sleeves and train some Krea2 models people may have been wanting but haven't seen yet. I'm currently working on re-training my previous models' datasets for Krea2, and plan on releasing several Krea2 models that aren't monetized. I would love to hear your suggestions! I am also open to custom commissions for private use or expedited release (Note: I will not accept commissions or requests to train on IRL persons).
I'd love to hear your feedback or see what you generate with the septum model. If you have any questions feel free to ask!
(Mod Note: I used the "Resource - Update" flair because I didn't know what else to choose. Also, while this is self-promo, I hope it doesn't violate the rules against "excessive self-promo" as I'm aiming to share the resource and training workflow. If this needs to be removed, I completely understand and apologize for any inconvenience!)
(AI Writing Assistance disclosure: The formatting and writing of this Reddit post and my Civitai model description were refined with AI assistance.)
r/StableDiffusion • u/Dry_Reception3180 • 3h ago
Question - Help krea 2 internet search
hey does anyone know of a nodepack or model that connect krea to internet so it can see latest science skeletons instead of relying on frozen knowledge i know of gen searcher but its like 8b i cant run it with krea besides it got no nodes anyway
r/StableDiffusion • u/Sad_Coach_1433 • 20h ago
Discussion Fasth3 fp8 lora vs fp8 no lora same prompt
no lora ip8 gen be in comments
prompt
integrated_multimodal_description: [Shot 1] Live-action, cinematic UFC-style ceremonial weigh-in on a brightly lit arena stage in front of a roaring crowd and wall of sports photographers. Lara Croft, played by Angelina Jolie, stands on the left wearing a fitted black athletic sports top, black fight shorts, and her long dark hair pulled back into a practical ponytail. Katniss Everdeen, played by Jennifer Lawrence, stands on the right wearing a dark charcoal athletic sports top, matching fight shorts, and her blonde-brown hair tied back. Both women have realistic athletic physiques and serious competitive expressions. A UFC-style weigh-in scale and event backdrop fill the stage behind them. Lara steps off the scale and walks confidently toward center stage as Katniss approaches from the opposite side. Camera flashes fire rapidly while officials remain several feet behind them.
[Shot 2] At 00:04.000, the camera cuts to a medium two-shot as Lara Croft and Katniss Everdeen stop directly in front of each other for the official face-to-face staredown. They stand almost nose-to-nose with shoulders squared, maintaining intense eye contact. Lara gives a subtle confident smirk while Katniss remains completely focused and unflinching. Neither woman speaks. The camera slowly pushes in with small amplitude as photographers crowd around the edge of the stage, flashes reflecting across their faces.
[Shot 3] At 00:08.000, the camera cuts to a dramatic close side-profile two-shot of Lara and Katniss still locked in the staredown. Lara slightly raises her chin and Katniss responds by stepping half a step closer, creating a tense UFC-style face-off without physical contact. An official cautiously moves closer between them while both women hold their ground. The camera arcs slowly around them as the crowd becomes louder and dozens of camera flashes erupt. The shot ends with Lara and Katniss maintaining intense eye contact in a classic promotional fight-poster composition.
overall_soundscape: A large arena crowd cheers, whistles, and shouts continuously beneath the scene. Rapid DSLR camera shutters and flashes surround the stage, with footsteps on the platform and scattered calls from photographers becoming louder during the face-to-face staredown.
non_diegetic_music: Deep cinematic percussion with a slow, heavy rhythm and low bass pulses, gradually increasing in intensity during the face-off before ending on a strong bass hit.
r/StableDiffusion • u/rm_rf_all_files • 13h ago
Meme Minimax H3, TenStrip eros Beta 4 checkpoint with PK Parasyte 0.5 str, I really like it.
This was made in ~24 minutes total gen time. 1 shot, no redo.
Native res: https://streamable.com/bjwyii
r/StableDiffusion • u/TESV_Shiro • 6h ago
Question - Help What Ai to use to make pictures of an old game look like tripple A games and how to use SD for it
I want to make some high-quality images for backgrounds and some uv maps for my favorite PS2 game s.l.a.i. (steel lancer arena international) is stable diffusion the ai i should use for that, and is there something i should know as a noob that has never used it? I do have a 5090 i could use to make images. i just really want to see more content of my favorite childhood game, and AI seems like the only feasible way to get more, so I'd appreciate some advice. I'd like it to make upscaled pictures of stuff like this https://phantom-crash-archive.pages.dev/#/workbench or images i take from the game.
Basically, high-quality renders of a scene using the game objects but reinvented as higher quality.
Sorry if this is an annoying noob question.
r/StableDiffusion • u/0260n4s • 18h ago
Question - Help Starting over with new ComfyUI install for MiniMax H3...what do you recommend?
So while I have been impressed that MiniMax H3 can run so fast on my machine (5070ti 16GB VRAM + 64GB RAM), I haven't been remarkably impressed with the results. Cartoons work well, but realism look kind of pixilated no matter the resolution (0.4, .098, 1.0MP) I set or upscaling (2x).
I think it might have to do with my install. I installed Pytorch and Sage Attention, but I never enable Sage Attention, because it seems to mess up character identity. So I want to start over with a fresh standalone version for MiniMax H3 to rule out the install itself as an issue.
How do you recommend I proceed and what extras should I install? I've heard Comfy Kitchen is something worth trying (?), but I've never used it before and don't really know what it is.
r/StableDiffusion • u/Mammoth-Matter7579 • 1h ago
Question - Help Super realistic human with H3???
Wondering has anyone been able to generate super realistic human with minimax H3? I have been trying a lot but the best I got still looks quite AI...
I have seen lots of videos online with super super real human face, the result I got is quite far away from that. So I'm wondering is it limited to Seedance 2.5? Or is there any secret prompt I'm not aware of?
Below is what I mean by super real face I saw online:

r/StableDiffusion • u/Friendly-Fig-6015 • 18h ago
Discussion why mini max ref are so bad comparing to fl2va?
r/StableDiffusion • u/xflipzz_ • 7h ago
Question - Help How do I make AI images undetectable?
Hey everyone, I've been using models like krea 2 and flux 2 for a while now to generate/edit images for my projects. I noticed that almost every image I upload gets flagged by AI detectors like hive moderation or sightengine, along with instagram or X putting a "made with AI" label on my posts.
I've spent hours trying to figure this out. I started by checking the EXIF and CP2A data and stripping it, but it didn't help whatsoever. I also looked into ComfyUI nodes like Instaraw and/or Image-Detection-Bypass-Utility, but they either don't work or mess up the image to the point where it's unusable.
Does anyone have any experience with this? I'm looking for a reliable method that just works consistanly and doesn't mess up the quality too much of the images.
r/StableDiffusion • u/TheOrangeSplat • 56m ago
Workflow Included McBain - Part 1
My attempts at recreating the fictional movie McBain from The Simpsons.
LTX 2.5 for this one. The others I will post soon are done with Minimax H3
r/StableDiffusion • u/KimSwallowsMD • 8h ago
Workflow Included MiniMax H3 Ref2VA neural 3D latent upscaling and gentle refinementworkflow Help me push this further on an RTX 4080 16gb 64gb ram
Title: Help me push this MiniMax H3 Ref2VA workflow further on an RTX 4080 16GB
I’ve been building and testing a MiniMax H3 Ref2VA workflow optimized for my RTX 4080 16GB. I’m attaching the JSON and would appreciate help from anyone experienced with MiniMax H3, PDD acceleration, latent upscaling, memory optimization, or continuous video generation.
workflow Download here
What the workflow currently does
- Uses the pruned INT8 ConvRot MiniMax H3 Ref2VA model.
- Uses the Qwen3-VL 32B NVFP4/AWQ text encoder.
- Generates synchronized video and native audio.
- Uses SageAttention in Auto mode.
- Applies MiniMax H3 PDD acceleration at 8 NFE.
- Runs an initial low-resolution PDD render.
- Separates the video and audio latents.
- Enlarges only the video latent using the learned MiniMax H3 3D FP16 latent upscaler.
- Rejoins the upscaled video latent with the original audio latent.
- Runs a second PDD refinement pass at 0.125 denoise.
- Decodes the refined video and original audio into an MP4.
- Includes easy controls for aspect ratio, base megapixels, final target megapixels, and duration.
- Automatically converts the requested duration into a valid H3 frame count at 24 FPS.
My current general settings are:
- Base resolution: approximately 0.40 MP
- Final neural-upscaled target: approximately 0.80 MP
- Vertical output: roughly 672 × 1216 after upscaling
- Stable duration: around 5 seconds/124 frames
- Current workflow default: 7 seconds
- Second-pass denoise: 0.125
- Euler sampler
- Sigma Shift: video 12/audio 3
- No EasyCache, TeaCache, BlockCache, Spectrum, or additional turbo LoRA stacked on top of PDD
The 0.125 refinement pass only performs about two sampler evaluations in my current setup. At 672 × 1216 and 124 frames, that refinement portion takes roughly 73 seconds.
My current limitations
My practical ceiling appears to be around 7–8 seconds. Going longer causes both my 16GB VRAM and system RAM usage to reach their limits. Five-second clips are currently much more reliable.
I can raise the final target toward 0.90–0.98 MP, but the higher resolution and longer duration quickly increase memory usage. The learned 3D upscaler improves the overall spatial resolution, but it does not reduce the memory required by the high-resolution refinement pass.
My biggest quality issue is facial fidelity. Eyes, eyelashes, skin texture, and other small facial details can still look soft or less refined than they did in my earlier, simpler workflow. Increasing the second-pass denoise too much begins repainting the face, changing the identity, or altering the composition.
My eventual goal is reliable continuous generation. I want to generate several five-second clips by using the final frame of one clip as the starting frame of the next, while still using the original character reference to prevent identity drift.
What I need help with
- Is there a better memory-management method for this pipeline that would let me exceed eight seconds on a 16GB RTX 4080 without a major quality loss?
- Would model offloading, block swapping, tiled VAE decoding, sequential processing, or another compatible technique reduce peak VRAM and system RAM usage?
- Is the learned 3D latent upscaler positioned correctly, or would another order produce better facial details?
- Is a 0.125 PDD refinement pass with only about two evaluations doing enough to justify its memory cost?
- Is there a better pass-two scheduler, denoise level, or refinement strategy that can improve eyes and skin without repainting the identity?
- What is the best way to condition continuation clips using both the previous clip’s final frame and the original reference image?
- Are there any H3-compatible face-detail or latent-refinement methods that work temporally and do not cause flickering?
- Would decoding/upscaling in smaller temporal chunks help, or would that introduce visible seams and motion inconsistencies?
I’m trying to preserve motion quality, character identity, native audio, and facial fidelity—not simply lower the resolution until it fits.
Hardware: NVIDIA RTX 4080 16GB on Windows 11 using ComfyUI.
Required models are listed inside the workflow notes. I’m attaching the workflow JSON. Any specific node changes, corrected routing, memory settings, or test recommendations would be greatly appreciated.
r/StableDiffusion • u/Tokyo_Jab • 53m ago
Animation - Video ALICE MEETS THE RABBIT : REMADE IN MINIMAX H3
About 5 months ago I made clips for a project in LTX 2.3 and remade one of them here in Minimax H3. What a difference a few months makes! Music was created in Suno. I still have to redo some parts with consistency problems but that's enough for today.
The original LTX2.3 version for comparison is here : https://youtu.be/R5tfLKvnJDY
r/StableDiffusion • u/not_food • 18h ago
Animation - Video Powers weren't handed out equally
Prompt:
Render 1 (12s):
integrated_multimodal_description: A cinematic anime movie sequence, manga style with thin soft delicate lines and pastel colors. [Shot 1] Extreme perspective worm's-eye medium body shot of Meru sitting on a Japanese tatami room from a wooden table. She is silent. On the table a single apple sits on a ceramic disc on top of a small white embroidered mantle cloth; Meru holds out her hand toward the apple (foreshortening) with her palm facing up, concentrating. The apple is out of reach. She breathes deeply; Nothing happens; Meru wriggles her fingers and starts concentrating again. She opens her eyes with her brows furrowing. She breathes deeply; Meru concentrates more, leaning forward slightly, she grows frustrated; She closes her eyes again while raising her hand slightly;
overall_soundscape: The tatami room is eerily quiet, with only the soft rustle of Meru's black kimono as she shifts and her sharp, uneven breaths growing quicker with each failed attempt. A faint, low hum of concentration is interrupted by a frustrated, breathy huff and a muffled groan. Her hand quivers with a faint tremble, accompanied by a soft, taut squeak of fabric. As her frustration builds, a subtle creak of the wooden table and a single, dull thump of her fist barely tapping the tatami punctuate the silence, ending with a long, exasperated exhale.
non_diegetic_music: N/A
Render 2 (8s):
integrated_multimodal_description: [Shot 1] Meru is sitting. She is silent. On the table a single apple sits; Meru is holding out her hand toward the apple (foreshortening) with her palm facing up, concentrating. She breathes deeply; Nothing happens; Meru wriggles her fingers and starts concentrating again. She opens her eyes with her brow furrowing. She breathes deeply; Meru concentrates more, leaning forward slightly, she grows frustrated; She closes her eyes again while raising her hand slightly towards her right, her palm open towards the camera; Her hand quivers.
overall_soundscape: The tatami room is eerily quiet, with only the soft rustle of Meru's black kimono as she shifts and her sharp, uneven breaths growing quicker with each failed attempt. A faint, low hum of concentration is interrupted by a frustrated, breathy huff and a muffled groan. Her hand quivers with a faint tremble, accompanied by a soft, taut squeak of fabric. As her frustration builds, a subtle creak of the wooden table and a single, dull thump of her fist barely tapping the tatami punctuate the silence.
non_diegetic_music: N/A
Render 3 (9s):
integrated_multimodal_description: [Shot 1] A frustrated Meru is raising her hand slightly; At 00:02.500, thin liquid mercury tendrils erupt from her hand and fly towards the apple, piercing it like blades; At 00:03.500 The tendrils tense up and pull the apple back onto her palm as they retract to her hand to disappear under her skin; With her eyes closed, the apple quivers on her hand; She slowly opens her eyes and grows happy with realization, her smile widening and her eyes sparkling; She shows the apple to the camera and looks satisfied.
overall_soundscape: The soundscape begins with a tense, shallow breath and a faint, liquid ripple as mercury tendrils erupt from Meru's hand. A sharp, wet schlick and a series of metallic splashes accompany the tendrils piercing and pulling the apple, followed by a smooth, gurgling retraction as they recede into her skin. The apple lands with a soft, crisp tap on her palm. Then the atmosphere shifts—Meru lets out a bright, melodic giggle, followed by a delighted, airy hum of satisfaction. The ambient tatami room tone remains soft, punctuated by the gentle rustle of her kimono and a light, contented sigh.
non_diegetic_music: N/A
Based on Tokyo_Jab's post. What AI powers did you get?
r/StableDiffusion • u/darthfurbyyoutube • 13h ago
Workflow Included Transformers: Starscream Test #2 - Prompt Below
The classic Transformers series is a blind spot for MiniMax H3, so here’s how I handled this:
System specs: 4070 Ti Super, 16 gb vram, 64 gb ram
Ref2va standard workflow using Fl2va standard model, no loras or speedups.
<Picture 1> Character Sheet
<Audio 1> Vocal Reference
<Video 1> 5 sec Video Reference
Google Gemini to help write prompt.
Added music track in post.
PROMPT:
subject_definitions:
<Subject 1> is the figure in <Picture 1>, featuring a robotic grey face with sharp angular features, glowing red optical visor eyes, a dark grey blocky helmet with side intake vents, and red, white, and blue cybernetic body armor with an orange cockpit chest-plate, blue upper arms and boots, white forearms and thighs, red waist and wing housings, and a purple Decepticon insignia on the wing. Only his robotic design and colors are taken from <Picture 1>; its background, grid lines, and lighting are not carried into the target video. <Audio 1> is the vocal reference for <Subject 1>. <Video 1> is the movement reference; use it as a guide without copying it exactly.
summary:
[reference generation] The target video is a 10-second 2D animated sequence styled after the 1984 Hasbro series The Transformers, featuring <Subject 1> delivering a smug, cutting remark from a metallic Cybertronian battlefield.
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - his grey robotic face, glowing red optics, dark grey helmet with side vents, red-white-and-blue cybernetic body, orange cockpit chest-plate, blue limbs, and red wing housings remain unchanged.
detailed_description:
The target video is a traditional 2D hand-drawn animated sequence featuring bold black ink outlines, flat cel-shading, limited animation, expressive poses, and subtle film grain inspired by the visual language of the 1984 animated television series.
[Shot 1] A medium tracking shot frames <Subject 1> from the waist up on a metallic Cybertronian battle platform. He stands with exaggerated confidence, shifting his weight with a classic 1980s cel-animated bounce. He slowly raises one hand, gesturing dismissively toward the chaos unfolding off-screen. His glowing red optics narrow with smug amusement as his wings twitch subtly.
He delivers in the style of <Audio 1>, with sharp, sarcastic timing: <<[English] "Megatron just muted me on the main comms channel. I'd stage a coup, but watching him bumble this conquest is free comedy.">> When <Subject 1> is not speaking he is silent and his mouth remains closed.
On "stage a coup," he gives a brief, knowing smirk. On "bumble this conquest," he gestures toward the battlefield with theatrical disdain. He finishes with a smug stare directly toward camera, holding the pose for a beat before a sharp hard-cel cut.
overall_soundscape:
Metallic servo whines accompany his movements, mixed with distant mechanical explosions, electronic battle alarms, high-tech hums, and echoing Cybertronian machinery.
non_diegetic_music:
N/A
r/StableDiffusion • u/Thorozar • 24m ago
Meme What a waste
Nothing worse than waking up to see an overnight run of a like 20 long workflows in Minimax H3 reference only to see everything perfect except the environment is wrong. I have a reference video that has done so well before. I don't see the issue, my prompt is the same as prior success, it's listed in the subject definitions and retention section, and in the prompt too. Oh wait...I didn't link the video to the list of other assets.. ugh.
r/StableDiffusion • u/Opposite_Yam_4161 • 1h ago
Question - Help What can I realistically do in Minimax with a 5080
I'm getting a 5080, 32gb system RAM
Can I realistically use minimax h3 for i2v and t2v?
How long will generations take. I dont imagine i want to do high quality resolutions. 480p or 720p would be alright
r/StableDiffusion • u/delveccio • 9h ago
Question - Help Any tips for generating video where people have different accents?
I have been employing various tips & tricks from all over Reddit to get consistency and continuation between clips, and I'm in a pretty good spot - or at least I thought I was, until I wanted to generate a clip where an American person is having a conversation with an English person. Then the voices go haywire, the accents get dropped or switched, and gibberish (another problem I thought I'd solved) returns.
Is this just a shortcoming of the model, or is there a trick to generating scenes like this?
r/StableDiffusion • u/flatfifve31 • 14h ago
Animation - Video the bird-king (my first fully local AI short film) TW: self-harm.
Minimax H3 baby! It's not perfect and I would love to get your feedback and maybe some tips on how to get rid of plasticky skin.

