You're looking at a specific version of this model. Jump to the model overview.
Input schema
The fields you can use to run this model with an API. If you don’t give a value for a field its default value will be used.
| Field | Type | Default value | Description |
|---|---|---|---|
| image |
string
|
Reference character image. Transparent PNGs are composited on white and their alpha becomes the reference mask.
|
|
| video |
string
|
Driving video (mp4/mov/webm/mkv). Its motion is transferred to the character; its aspect ratio sets the output size unless width/height are given.
|
|
| prompt |
string
|
|
Describe the character and the motion, e.g. 'A cartoon robot walking in place, side view'. Describes the final video; not instructions.
|
| negative_prompt |
string
|
|
What to avoid, e.g. 'distorted limbs, camera movement, blurry'. Only matters with guidance_scale > 1 (the quality preset). Wan's stock Chinese negative prompt is deliberately not applied: its 'painting / artwork / style' terms push stylized characters toward a 3D-CG look.
|
| mode |
None
|
animation
|
animation: the reference character (and its background) performs the driving motion. replacement: the character is placed into the driving video, keeping its background and lighting.
|
| image_mask |
string
|
Optional mask for the reference: a grayscale/black-and-white matte (white = character) or a SCAIL-2 palette mask (blue = identity 0). If omitted and auto_mask is on, one is derived from the image's alpha channel or BiRefNet.
|
|
| video_mask |
string
|
Optional per-frame mask video for the driving video, same conventions as image_mask (grayscale matte or SCAIL-2 colours). If omitted and auto_mask is on, BiRefNet masks every frame.
|
|
| auto_mask |
boolean
|
True
|
Derive missing masks automatically (alpha channel or BiRefNet single-subject matting). Off = run without masks.
|
| resolution |
None
|
512p
|
Short side of the output; the driving video's aspect ratio is kept (e.g. 896x512 for 16:9, 512x512 for square). SCAIL-2 was trained at both.
|
| width |
integer
|
0
Max: 1536 |
Explicit output width (multiple of 32). Set together with height to override resolution; the driving video is centre-cropped to this aspect.
|
| height |
integer
|
0
Max: 1536 |
Explicit output height (multiple of 32).
|
| num_frames |
integer
|
0
|
None
|
| fps |
integer
|
0
Max: 60 |
Resample the driving video to this frame rate before animating (frames are held, never interpolated); the output uses the same rate. 0 = keep the driving video's rate.
|
| preset |
None
|
fast
|
fast (default): the official ComfyUI recipe — lightx2v step/CFG-distill LoRA, Euler, 6 steps, CFG 1, shift 5; about 12x cheaper. quality: the paper's sampler (UniPC, 40 steps, CFG 5, shift 3) — 81 frames at 512p takes ~10 min on an H100.
|
| steps |
integer
|
0
Max: 100 |
Sampling steps; 0 = preset default.
|
| guidance_scale |
number
|
0
Max: 20 |
Classifier-free guidance; 0 = preset default (5 quality / 1 fast).
|
| shift |
number
|
0
Max: 20 |
Flow-matching schedule shift; 0 = preset default (3 quality / 5 fast).
|
| dpo_lora |
number
|
1
Max: 2 |
Strength of the official Bias-Aware DPO LoRA (0 = off). 1.0 is what the official ComfyUI template ships with; the released base checkpoint is pre-DPO.
|
| relight_lora |
number
|
0
Max: 2 |
Strength of the official relighting LoRA for replacement mode (0 = off). Improves lighting consistency with the driving scene.
|
| pose_strength |
number
|
1
Max: 2 |
Weight of the driving-motion conditioning.
|
| seed |
integer
|
Random seed; leave empty for random.
|
|
| return_masks |
boolean
|
False
|
Also return the reference/driving masks that were used (for debugging or re-use as image_mask / video_mask).
|
Output schema
The shape of the response you’ll get when you run this model with an API.
Schema
{'properties': {'driving_mask': {'format': 'uri',
'nullable': True,
'title': 'Driving Mask',
'type': 'string'},
'reference_mask': {'format': 'uri',
'nullable': True,
'title': 'Reference Mask',
'type': 'string'},
'seed': {'title': 'Seed', 'type': 'integer'},
'video': {'format': 'uri', 'title': 'Video', 'type': 'string'}},
'required': ['video', 'seed'],
'title': 'Output',
'type': 'object'}