You're looking at a specific version of this model. Jump to the model overview.
Input schema
The fields you can use to run this model with an API. If you don’t give a value for a field its default value will be used.
| Field | Type | Default value | Description |
|---|---|---|---|
| mode |
None
|
animation
|
animation: the reference character (and its background) performs the driving motion. replacement: the character is placed into the driving video, keeping its background and lighting.
|
| resolution |
None
|
512p
|
Output resolution, measured by the shorter side. The driving video's aspect ratio is preserved unless width and height are specified.
|
| preset |
None
|
fast
|
fast (default): the official ComfyUI recipe — lightx2v step/CFG-distill LoRA, Euler, 6 steps, CFG 1, shift 5; balanced: UniPC, 8 steps, CFG 1, shift 5, DPO before LightX2V. The model uses FP8 scaled weights and defaults to BF16 VAE. quality: the paper's sampler (UniPC, 40 steps, CFG 5, shift 3); substantially slower.
|
| vae_precision |
None
|
bf16
|
VAE precision. BF16 matches the ComfyUI template.
|
| sampler_name |
None
|
preset
|
Override preset sampler.
|
| scheduler |
None
|
simple
|
ComfyUI sampling schedule.
|
| output_format |
None
|
mp4
|
Output file format: MP4 (H.264), WebM (VP9), or looping animated WebP. Does not change inference.
|
| fps |
integer
|
0
Max: 60.0 |
Resample the driving video to this frame rate before animating (frames are held, never interpolated); the output uses the same rate. 0 = keep the driving video's rate.
|
| seed |
integer
|
Random seed; leave empty for random.
|
|
| image |
string
|
Reference character image. Transparent PNGs are composited on white and their alpha becomes the reference mask.
|
|
| shift |
number
|
0
Max: 20.0 |
Flow-matching schedule shift; 0 = preset default (3 quality / 5 fast).
|
| steps |
integer
|
0
Max: 100.0 |
Sampling steps; 0 = preset default.
|
| video |
string
|
Driving video or animated WebP. Its motion is transferred to the character. WebP transparency is used as the driving mask unless video_mask is supplied.
|
|
| width |
integer
|
0
Max: 1536.0 |
Explicit output width (multiple of 32). Set together with height to override resolution; the driving video is centre-cropped to this aspect.
|
| height |
integer
|
0
Max: 1536.0 |
Explicit output height (multiple of 32).
|
| prompt |
string
|
|
Describe the character and the motion, e.g. 'A cartoon robot walking in place, side view'. Describes the final video; not instructions.
|
| denoise |
number
|
1
Min: 0.001 Max: 1.0 |
KSampler denoise strength.
|
| dpo_lora |
number
|
1.0
Max: 2.0 |
Strength of the official Bias-Aware DPO LoRA (0 = off). 1.0 is what the official ComfyUI template ships with; the released base checkpoint is pre-DPO.
|
| pose_end |
number
|
1
Max: 1.0 |
End fraction of pose conditioning.
|
| auto_mask |
boolean
|
True
|
Derive missing masks automatically (alpha channel or BiRefNet single-subject matting). Off disables automatic matting; supplied masks and driving WebP transparency still apply.
|
| image_mask |
string
|
Optional mask for the reference: a grayscale/black-and-white matte (white = character) or a SCAIL-2 palette mask (blue = identity 0). If omitted and auto_mask is on, one is derived from the image's alpha channel or BiRefNet.
|
|
| num_frames |
integer
|
0
Max: 161.0 |
Maximum output frames, taken from the start of the driving video after applying fps. 0 uses the available video, up to 161 frames. 81 frames at 24 fps is about 3.4 seconds. Does not extend a shorter video. Counts are trimmed to a supported length (5, 9, 13, ...).
|
| pose_start |
number
|
0
Max: 1.0 |
Start fraction of pose conditioning.
|
| video_mask |
string
|
Optional per-frame mask video for the driving video, same conventions as image_mask (grayscale matte or SCAIL-2 colours). If omitted, animated WebP transparency is used when available; otherwise auto_mask uses BiRefNet.
|
|
| relight_lora |
number
|
0
Max: 2.0 |
Strength of the official relighting LoRA for replacement mode (0 = off). Improves lighting consistency with the driving scene.
|
| return_masks |
boolean
|
False
|
Also return the reference/driving masks that were used (for debugging or re-use as image_mask / video_mask).
|
| lightx2v_lora |
number
|
-1
Min: -1.0 Max: 2.0 |
LightX2V strength: -1 uses preset, 0 disables. The fast and balanced presets use 0.8.
|
| pose_strength |
number
|
1.0
Max: 10.0 |
Weight of the driving-motion conditioning.
|
| return_frames |
boolean
|
False
|
Return original decoded PNG frames as ZIP for lossless local Toonout and sprite assembly. Encoded video/WebP is compressed.
|
| guidance_scale |
number
|
0
Max: 20.0 |
Classifier-free guidance; 0 = preset default (5 quality / 1 fast).
|
| output_quality |
integer
|
80
Min: 1.0 Max: 100.0 |
Compression quality: higher means better fidelity and usually larger files. Codec-relative, not a model quality score; 100 does not guarantee lossless output. Use PNG frames for lossless originals.
|
| negative_prompt |
string
|
|
What to avoid, e.g. 'distorted limbs, camera movement, blurry'. Only matters with guidance_scale > 1 (the quality preset). Wan's stock Chinese negative prompt is deliberately not applied: its 'painting / artwork / style' terms push stylized characters toward a 3D-CG look.
|
| prepared_inputs |
boolean
|
False
|
Already-composited RGB grids and exact RGB SCAIL palette masks. No resize, crop, recoloring or automatic masks. Inputs must have matching dimensions; video masks must match frame count and FPS. Use lossless WebP or FFV1 MKV.
|
| previous_frames |
string
|
Previous output as an image or lossless video. Its tail anchors the beginning of this request. The driving video must include that overlapping interval, and returned frames include the anchor interval.
|
|
| additional_images |
array
|
Up to 7 additional reference views. Supply a matching additional_image_masks list and image_mask; palette colors bind views to the same identities. CLIP vision uses the primary image.
|
|
| previous_frame_count |
integer
|
5
Min: 1.0 Max: 77.0 |
Tail frames used for anchoring and chunk overlap; must be 4n+1. Use 1 for a single-image anchor. SCAIL-2 was trained with 5.
|
| additional_image_masks |
array
|
Palette mask for each additional reference, in the same order.
|
Output schema
The shape of the response you’ll get when you run this model with an API.
{'properties': {'driving_mask': {'format': 'uri',
'nullable': True,
'title': 'Driving Mask',
'type': 'string'},
'frames': {'format': 'uri',
'nullable': True,
'title': 'Frames',
'type': 'string'},
'metadata': {'format': 'uri',
'nullable': True,
'title': 'Metadata',
'type': 'string'},
'reference_mask': {'format': 'uri',
'nullable': True,
'title': 'Reference Mask',
'type': 'string'},
'seed': {'title': 'Seed', 'type': 'integer'},
'video': {'format': 'uri', 'title': 'Video', 'type': 'string'}},
'required': ['video', 'seed'],
'title': 'Output',
'type': 'object'}