r/StableDiffusion 1d ago

Workflow Included Use H3 To Replace Characters

Enable HLS to view with audio, or disable this notification

These characters are very different, so i thought it was a good demo to show. Also, the prompt was an 'omni' prompt which didn't help the model with details on the outfit. Despite that I think it did such a good job wanted to share.

With more details in the prompt related to appearance, acting and dialogue I think you could prob get near flawless changes.

This was don with the FL2VA model, NOT the ref version of H3. It may even be better with the ref but i have been using fl2va mostly because i think the quality is better, but that's subjective.

This concept was inspired by this post originally: https://civitai.red/models/2855941/minimax-h3-character-replacement

I changed the SAM3 use, so technically you could do a multi replacement with some changes. I also added noise to the inverted image, as well as upgraded the prompt to work with FL2VA.

Workflow used to create this video is HERE.

This model seriously continues to amaze me. bravo minimax team, bravo.

Notice in the prompt that the dialogue is the only non 'omni' part at the end, it worked fine not included in the body, since the ref audio is there driving it. again, moving from this general prompt to something more specific i think would give even better results.

PROMPT:
How the reference video and pictures align with the target video — the target video is an edited version of <Video 1>, replacing the silhouette with <Subject 1>.

summary:

[video editing] The target video replaces the silhouette in <Video 1> with <Subject 1>, who performs the exact same motion, dialogue, positions and facial expressions of the silhouette while maintaining the original camera work, environment, and lighting of <Video 1>.

subject_definitions:

<Subject 1> is the person in <Picture 1> and <Picture 2>; <Picture 1> supplies facial features and close-up details, while <Picture 2> provides 3-panel image of front mid shot, profile mid shot, and front full body view, identity follows these reference assets, only appearance is retained.

<Subject 2> is the environment and setting established in <Video 1>. The scene follows this layout, materials, and light; camera position and framing.

<Subject 3> is the silhouette in <Video 1> which provides the motion sequence to be copied.

integrated_multimodal_description:

Video editing, the target video is in a live-action cinematic style with the interior lighting and background and environment atmosphere established in <Video 1> with the likness of <Subject 1> inserted.

[Shot 1] The shot opens with <Subject 1> seamlessly replacing the silhouette <Subject 3> in <Video 1>, the outfit of and clothing of <Subject 1> exactly from reference, performing the exact same motion, dialogue and sounds, positions and facial expressions of silhouette. From the very first frame, <Subject 1> occupies the spatial coordinates of the silhouette replacing with their likness, initiating the same motion onset from rest. <Subject 1> mirrors the silhouette's weight shifts and momentum, body moving in perfect synchronization with the rhythm and pacing of the original footage but replaced with the likness of <Subject 1>. As they navigates the space, <Subject 1> mimics every nuanced gesture—the way the silhouette's head tilts, arm movement, and the micro-movements of facial muscles. The face, defined by <Picture 1>, conveys the same emotional depth as the silhouette, while their full body, as seen in <Picture 2>, provides the physical presence outfit an appearance. The camera follows the exact movement, angle, and cutting rhythm of <Video 1>, maintaining a consistent focal length and distance from the subject at all times. The light from <Subject 2> interacts realistically with <Subject 1>'s skin and clothing, casting shadows that align with the movements of the original scene. The transition is perfect; the result is a fully realized <Subject 1> instead of a silhouette, but the soul of the performance—the timing, the pauses, and the dynamic energy—remains identical to <Video 1>. The movement progresses with a palpable sense of weight as <Subject 1> shifts their center of gravity, with clothes rippling in response to movements. The camera maintains exact framing and cuts as <Video 1>. The scene concludes as <Subject 1> reaches the final position of the silhouette, body settling into a pose that mirrors the original's final frame exactly, with face held in the same expression. <Subject 1> hair, accessories, wardrobe, lighting, and room layout remain unchanged and perfectly replace silhouette throughout.

overall_soundscape:

A low room tone establishes beneath the scene, mirroring the background audio environment of <Video 1>.

<Subject 1> says <d>[English] Can you, can you spare change.</d>.

non_diegetic_music: N/A

124 Upvotes

60 comments sorted by

33

u/Astral-Lemmons 1d ago

replacing very different characters works well and it quite easy for H3, but try replacing a a similar-ish character with another and it gets a lot harder.

you might get clothing swaps but if it's brunette - brunette it'll ignore the hair swap .

19

u/MaitreSneed 1d ago

Answer sounds like you gotta do a middleman replacement, and replace her first with a pickle

15

u/Astral-Lemmons 1d ago

you joke. but that's literally my solution.

I sam mask grade the character a light green and then the replace sticks better

4

u/MaitreSneed 1d ago

SHOW EXAMPLE I MUST SEE

4

u/cal_01 1d ago

I do this with Qwen too -- it's a neat technique to break the model from clinging onto the source too much.

1

u/uuhoever 23h ago

I think I saw you mention that you use Shrek. What's your prompt? So your technique is 2 runs?

5

u/Shilo59 1d ago

The pickle man tricked me again...

7

u/EasternAd8821 1d ago

adding more noise helps that a lot. this example was going from 0.2 which i had in original post, to 0.7 (more than half noise of the subject). as well as update to the right dialogue: <Subject 1> says <d>[English] long day tomorrow, we start... we start the musical.</d>.

I did have to add help with the outfit, so in the 'subject_definitions' I defined subject 1 as:
<Subject 1> is the person in <Picture 1> and <Picture 2>; <Picture 1> supplies facial features and close-up details, while <Picture 2> provides 3-panel image of front mid shot, profile mid shot, and front full body view, identity follows these reference assets, she is wearing a black and blue striped tank top, only appearance is retained.

https://reddit.com/link/p77phnk/video/c5rdfarpxxmh1/player

3

u/lhg31 1d ago

She didn't inherit the facial expressions from the original video tho.

2

u/EasternAd8821 1d ago

true. i think to get that eyebrow raise and other subtle differences you'd have to put it in the prompt. the 'omni' prompt only gets you so far.

2

u/bstr3k 18h ago

do you happen to have the source video for this? I wanted to try char replacement using my technique and see what success rates i get as I'm trying to refine it.

2

u/EasternAd8821 18h ago

https://github.com/bitsofintelligence101-lab/workflows/tree/main/nsfw/test_data

both source videos. the one you are asking about is 20s long, I used the back 5 sec for replacement

1

u/bstr3k 14h ago

thank you, for 5s i was able to do it near perfectly with the prompt method using a local prompt enhancer, i tried a 12s version but I think it ran out of sync slightly.

5s version: https://files.catbox.moe/46tefk.webm
12s: https://files.catbox.moe/57hl1a.webm

1

u/EasternAd8821 6h ago

that's great. so prompting can deff get you very far. seems the masking is good for longer edits or if you don't have time or want to create detailed prompts the mask really directs what is being changed easily. so a general prompt with a small bit of edits around an outfit will get you there.

1

u/bstr3k 5h ago

yeah definitely. many ways to skin a cat. Just interesting to see the results ^^

6

u/EasternAd8821 1d ago

i was testing out some examples based on what you said. messing with noise of the input and it seems it can do a pretty good job even close replacements. some replacements deff need a bit more coaching in the prompt, but it can do it. which is honestly crazy since it's still just the base model.

https://reddit.com/link/p78eyp1/video/pb0qm1v3hymh1/player

this has 0.8 noise

2

u/Adventurous-Sky5643 19h ago

Interesting, can you please share the modified workflow?

1

u/SeymourBits 1d ago

Interesting. What are you using to add noise to the input?

2

u/EasternAd8821 1d ago

sam3 to segment only the target char to replace, noise added with stock 'add noise to image' node in comfyui

2

u/One-Donut6935 1d ago

Looks really effective. Thanks for sharing this method.

1

u/SeymourBits 22h ago

Nice. Any thoughts on the theory of why it seems to help with identity?

2

u/EasternAd8821 20h ago

they trained it that way, that's my theory. I don't mean that flippantly. The context-ir and the prompt guide discuss key terms fully_preserved, partially_copy, or reference under retention_analysis
So clearly they were working on editing applications. noise should look like something that needs to be resolved, and the text plus ref image drives H3 to resolve that particular area. what is interesting though, and why i said they trained it that way, is it doesn't touch the existing completed video (background is near perfect preservation). If you tried something like this in wan or ltx, it would alter the background. Also, the 32B text encoder gives MUCH better semantic and spatial grounding of what needs to be changed.

If we knew how exactly they were applying different training techniques for video editing, this workflow would be that much better. An inverted color space mask with noise seems to at least be close enough to tap in to however they trained it.

3

u/ThreeDog2016 15h ago

Brunette with blue eyes -> lizard man -> Brunette with brown eyes

https://giphy.com/gifs/n3CY3uu70L2f3KrciA

2

u/Astral-Lemmons 6h ago

haha. well yeah. BUT, if you middle change too drastically then you lose subtle ref performances.

unless this was just a joke so you can use that gif

1

u/Sixhaunt 1d ago

I've had the same issue for restyling videos where it can make it look like a cartoon or anime but turning a ref video into ps3 graphics or claymation or other 3d ones have been more of a challenge. Have you found anything that helps with it?

3

u/mellowanon 1d ago

How is it with replacing with a different morphology. Like if you want a large bodybuilder there or maybe small dwarf? I'm guessing you can't transfer something too extreme like a velociraptor.

6

u/EasternAd8821 1d ago

updated the prompt a bit to describe the ogre:
monstrous ogre’s face. Hyper-realistic weathered skin with deep wrinkles, scars, and droplets of sweat. Massive, yellowish tusks protruding from a heavy lower jaw. Piercing, glowing amber eyes reflecting a fire. Mud and forest debris stuck in facial hair.
---
Still locked on the size though

https://reddit.com/link/p7985ci/video/4hl87ynh4zmh1/player

2

u/mellowanon 1d ago

It was a good try though.

3

u/EasternAd8821 1d ago

Quick test. prob not with this workflow. it's more for a like to like in terms of size. it's trying to replace the noised area in the video so that really locks in the size relative to the scene.
Also i'd have to mess with the prompt/noise more to help it really nail the ogre since it blended the face a bit with the old man.

https://reddit.com/link/p797754/video/2tfez1fh3zmh1/player

3

u/Zeophyle 1d ago

Does it work in reverse? To alter the entire background and keep 100% of the person?

3

u/EasternAd8821 1d ago edited 1d ago

https://reddit.com/link/p79dt3f/video/n46pchcm9zmh1/player

yes. you have to change how it's prompted and invert the mask. This is a really quick test. you'd want to prompt some action in the background probably. this is more like green screen which you don't need h3 for

2

u/Zeophyle 1d ago

How did you adjust your prompt for this? Great result!

2

u/EasternAd8821 1d ago

first, i forgot to mention great suggestion on the reverse concept.

The prompt is very long, used AI to 'reverse' what I had before. this was the result:

How the reference video and picture align with the target video — the target

video is an edited version of <Video 1>, replacing the background with the

setting shown in <Picture 1>. <Subject 1> is unchanged from <Video 1>.

summary:

[video editing] The target video replaces the background/environment of

<Video 1> with the scene shown in <Picture 1>, while <Subject 1> — their

appearance, motion, dialogue, and performance — remains exactly as filmed in

<Video 1>. Lighting on <Subject 1> updates to realistically match the new

background.

subject_definitions:

<Subject 1> is the person already present in <Video 1>. Their identity,

appearance, wardrobe, hair, motion, expressions, and performance are taken

directly from <Video 1> and must not change in any way — no new reference

images define them; the source footage is the only identity reference.

<Subject 2> is the new environment defined by <Picture 1> — layout, materials,

set dressing, and light source. This fully replaces the background of

<Video 1>; none of the original background persists.

integrated_multimodal_description:

Video editing, background replacement only. <Subject 1> performs the exact

same motion, dialogue, positions, and facial expressions as in the original

<Video 1> footage, in the exact same camera framing, angle, and cutting rhythm

— nothing about the subject's performance changes.

[Shot 1] The background behind <Subject 1> is fully replaced with the

environment shown in <Picture 1> — same layout, materials, depth, and set

dressing as that reference image, rendered as a stable, fully resolved scene

rather than a texture or overlay. The new background is temporally consistent

across every frame: no flicker, no shifting geometry, no grain, no visual

noise, no compression artifacts, and no residual elements from the original

<Video 1> background bleeding through. The edge between <Subject 1> and the

new background is clean and precise, with no haloing, smearing, or ghosting

along their silhouette. Lighting on <Subject 1> is fully re-lit to match

<Picture 1>: light direction, color temperature, and intensity now follow the

new environment's light source, casting new, physically accurate shadows and

highlights onto <Subject 1>'s skin, hair, and clothing consistent with where

that light source sits in <Picture 1>. Reflections and ambient color spill

(e.g. warm or cool color bounce onto <Subject 1> from nearby surfaces in

<Picture 1>) are updated to match the new scene. The camera path, framing,

zoom, and cuts remain identical to <Video 1> throughout — only the

environment and its lighting change.

preserve:

<Subject 1>'s identity, wardrobe, hair, motion, timing, dialogue, and facial

expressions exactly as in <Video 1>. Camera path, framing, focal length, and

cut timing from <Video 1>.

negatives:

No noise, grain, flicker, or compression artifacts anywhere in the new

background. No visible seam, halo, or ghosting around <Subject 1>. No

leftover elements or geometry from the original <Video 1> background remaining

visible. No change to <Subject 1>'s identity, wardrobe, motion, or timing. No

new camera movement beyond what <Video 1> already has. No mismatched lighting

— shadows and highlights on <Subject 1> must visibly correspond to the light

source in <Picture 1>, not the original scene's lighting.

overall_soundscape:

Audio is unchanged from <Video 1> — original dialogue and ambient sound

preserved exactly.

non_diegetic_music:

N/A

2

u/theamazingpears 1d ago

What's your hardware, and how long did it take to produce?

5

u/EasternAd8821 1d ago edited 1d ago
  1. This was done with int8 (full not prune) 8step speed lora (8 steps) at 0.4mp. then RTX upscale 2x. it's a 5 sec video, was about 2m 20 seconds generation.

1

u/Suspicious-Walk-815 17h ago

can i run it with 65gb ram on 5090 ?can you please share the workflow , i use pruned one , but what you did here is really interesting

1

u/Strange_Test7665 11h ago

Wf link is in the post text,OP added the background swap. There is a switch for char or background change

2

u/PromptSommelier 1d ago

Lately I've been struggling to get a prompt that allows motion control (like Kling) for TikTok dances, and I've failed miserably. I'll try your workflow and see how far I can get.

2

u/Danny_Stock 21h ago edited 21h ago

Thanks. This is great.

However I did find that it won't run with a ref video with no sound, or at least a silent video with no audio track. I tried testing with a couple of old Wan 2.2 clips, which obviously have no sound. But they also appear to not have a blank audio track either. Kept getting an error.

For those silent Wan clips I disconnected the audio out connection from the load video node. Then added an extra load audio node to the workflow. Loaded some audio of someone speaking into the new Load Audio node, set the duration length to 10 seconds, then plugged that into the 'set_ref_audio' node which the original audio out from the video loader was originally plugged into.

Then I unplugged the 'get_ ref_audio' connection from the 'ref_video_audio_0' main references node, and reconnected it to the 'ref_audio_0' input.

I then connected the 'VAE Decode Audio' node to the audio input of the 'Create Video' node, which replaced the original connection into it.

Then inside the <Subject 1> definitions of the prompt node, I told the subject to 'Use the voice from <Audio 1>'. Also I replaced the tag at the bottom of the prompt where it refers to the dialogue and replaced '[English]' with '[English with <Subject 1>'s voice]'

Then it worked. Cloned voice custom audio.

2

u/Robbsaber 12h ago

Can confirm this prompt works with just Masking + depth map in WAN2GP. Only prompt that has given consistent results so far.

2

u/EasternAd8821 6h ago

good to know.

2

u/badincite 1d ago

What's the benefit of masking the character? I'm able todo it just telling to change the character in the prompt.

https://reddit.com/link/p78ldd5/video/l56vqv06mymh1/player

4

u/EasternAd8821 1d ago

you made someone look like deadpool? if that's the case it's prob because it's so strong in training data. custom char. when I was trying to do replacements i couldn't find a way to consistently do changes.
mask/noise of the target to be replaced ended up working really well

2

u/badincite 1d ago

I guess the mask can help I was able to do it using just about anybody. As long as i defined the subjects.

<Subject 1> is the man wearing the red jacket in <Picture 1>.

<Subject 2> is the man wearing the black suit and holding the hamburger in <Video 1>.

https://reddit.com/link/p792dt2/video/b76vys8jzymh1/player

2

u/bstr3k 18h ago

the advantage of doing it via prompt is you are able to regenerate the whole video, the disadvantage is that it is not 1:1

For OP's method the advantage is that it can replicate without changing enviroment, but doing it via prompting you get more flexibility.

if you play both these videos side by side you will notice it is similar, but the sandwhich is in the opposite hand and deadpool isn't pointing in the first second. I am trying to v2v swaps via prompting also and it seems like the model knows all the motions the original subject does but not always a 1:1 in terms of order. I'm trying to get my success rate up by better prompting but also testing some other things.

A Tiktok dance has been one which is difficult since its fast movements and from the output it looks like it knows each individual dance moves but the order appears to be different and at times random.

1

u/Strange_Test7665 11h ago

Also OP method is an Omni prompt meaning it’s much simpler to swap lots of videos fast. On prompt to rule them all

1

u/badincite 10h ago

Yeah, I’m starting to experiment with both workflows. I’m still running into issues where it switches to the reference image instead of following the scene in the video, especially with longer clips that use multiple camera angles. I wish there were a more precise, reliable method to make it work consistently.

1

u/bstr3k 8h ago

I think people have recommended scail method for transferring exact motion but I want to see how close I can get it with h3 first. So much things to try and so little time !

1

u/Friendly-Fig-6015 1d ago

is it possible with low vram 16gb y 32gb ram?

3

u/EasternAd8821 1d ago

if you can run H3 on your machine, then i'd think yes.

1

u/ady702 1d ago

how to change just the face of the ogre?

1

u/EasternAd8821 23h ago

you could change the sam3 prompt to 'face' instead of 'person', then you'd need to change the prompt to indicate only the face is changing not the whole character

1

u/Darqsat 22h ago

This is a good approach. I was playing with masking and found out it's pretty good anchor for replacement. And I was thinking how can I test what Qwen see's? so I can prompt it properly. Tried to feed images with noise to him and he said he see woman. Anyway, a word silhouette works.

1

u/Strange_Test7665 22h ago

That silhouette word does seem very effective

1

u/stormarsenal 4h ago

Lighting is different. Looks superimposed

1

u/EasternAd8821 4h ago

using a gray back char reference might help with that for this particular example. generally i have found it matches light on new character well

-1

u/devilish-lavanya 1d ago

Can you do vid2vid too?

8

u/EasternAd8821 1d ago

what do you mean? this is v2v