You're looking at a specific version of this model. Jump to the model overview.
Input schema
The fields you can use to run this model with an API. If you don’t give a value for a field its default value will be used.
| Field | Type | Default value | Description |
|---|---|---|---|
| image |
string
|
Reference character image. Transparent PNGs are composited on white and their alpha becomes the reference mask.
|
|
| video |
string
|
Driving video (mp4/mov/webm/mkv or constant-rate animated WebP). Its motion is transferred to the character; its aspect ratio sets the output size unless width/height are given.
|
|
| prompt |
string
|
|
Describe the character and the motion, e.g. 'A cartoon robot walking in place, side view'. Describes the final video; not instructions.
|
| negative_prompt |
string
|
|
What to avoid, e.g. 'distorted limbs, camera movement, blurry'. Only matters with guidance_scale > 1 (the quality preset). Wan's stock Chinese negative prompt is deliberately not applied: its 'painting / artwork / style' terms push stylized characters toward a 3D-CG look.
|
| mode |
None
|
animation
|
animation: the reference character (and its background) performs the driving motion. replacement: the character is placed into the driving video, keeping its background and lighting.
|
| image_mask |
string
|
Optional mask for the reference: a grayscale/black-and-white matte (white = character) or a SCAIL-2 palette mask (blue = identity 0). If omitted and auto_mask is on, one is derived from the image's alpha channel or BiRefNet.
|
|
| video_mask |
string
|
Optional per-frame mask video for the driving video, same conventions as image_mask (grayscale matte or SCAIL-2 colours). If omitted and auto_mask is on, BiRefNet masks every frame.
|
|
| auto_mask |
boolean
|
True
|
Derive missing masks automatically (alpha channel or BiRefNet single-subject matting). Off = run without masks.
|
| resolution |
None
|
512p
|
Short side of the output; the driving video's aspect ratio is kept (e.g. 896x512 for 16:9, 512x512 for square). SCAIL-2 was trained at both.
|
| width |
integer
|
0
Max: 1536 |
Explicit output width (multiple of 32). Set together with height to override resolution; the driving video is centre-cropped to this aspect.
|
| height |
integer
|
0
Max: 1536 |
Explicit output height (multiple of 32).
|
| num_frames |
integer
|
0
|
None
|
| fps |
integer
|
0
Max: 60 |
Resample the driving video to this frame rate before animating (frames are held, never interpolated); the output uses the same rate. 0 = keep the driving video's rate.
|
| preset |
None
|
fast
|
fast (default): the official ComfyUI recipe — lightx2v step/CFG-distill LoRA, Euler, 6 steps, CFG 1, shift 5; sprute: UniPC, 8 steps, CFG 1, shift 5, DPO before LightX2V. Model/VAE precision are selected separately. quality: the paper's sampler (UniPC, 40 steps, CFG 5, shift 3) — 81 frames at 512p takes ~10 min on an H100.
|
| prepared_inputs |
boolean
|
False
|
Already-composited RGB grids and exact RGB SCAIL palette masks from local Sprute. No resize, crop, recoloring or automatic masks. Inputs must have matching dimensions; video masks must match frame count and FPS. Use lossless WebP or FFV1 MKV.
|
| additional_images |
array
|
Up to 7 additional reference views. Supply a matching additional_image_masks list and image_mask; palette colors bind views to the same identities. CLIP vision uses the primary image.
|
|
| additional_image_masks |
array
|
Palette mask for each additional reference, in the same order.
|
|
| previous_frames |
string
|
Previous output as an image or lossless video. Its tail anchors the beginning of this request. The driving video must include that overlapping interval, and returned frames include the anchor interval.
|
|
| previous_frame_count |
integer
|
5
Min: 1 Max: 77 |
Tail frames used for anchoring and chunk overlap; must be 4n+1. Use 1 for a single-image anchor. SCAIL-2 was trained with 5.
|
| vae_precision |
None
|
default
|
Select VAE checkpoint; Sprute uses bf16.
|
| sampler_name |
None
|
preset
|
Override preset sampler.
|
| scheduler |
None
|
simple
|
ComfyUI sampling schedule.
|
| lightx2v_lora |
number
|
-1
Min: -1 Max: 2 |
LightX2V strength: -1 uses preset, 0 disables. Sprute uses 0.8.
|
| pose_start |
number
|
0
Max: 1 |
Start fraction of pose conditioning.
|
| pose_end |
number
|
1
Max: 1 |
End fraction of pose conditioning.
|
| denoise |
number
|
1
Min: 0.001 Max: 1 |
KSampler denoise strength.
|
| return_frames |
boolean
|
False
|
Return original decoded PNG frames as ZIP for lossless local Toonout and sprite assembly. MP4 is only a preview.
|
| steps |
integer
|
0
Max: 100 |
Sampling steps; 0 = preset default.
|
| guidance_scale |
number
|
0
Max: 20 |
Classifier-free guidance; 0 = preset default (5 quality / 1 fast).
|
| shift |
number
|
0
Max: 20 |
Flow-matching schedule shift; 0 = preset default (3 quality / 5 fast).
|
| dpo_lora |
number
|
1
Max: 2 |
Strength of the official Bias-Aware DPO LoRA (0 = off). 1.0 is what the official ComfyUI template ships with; the released base checkpoint is pre-DPO.
|
| relight_lora |
number
|
0
Max: 2 |
Strength of the official relighting LoRA for replacement mode (0 = off). Improves lighting consistency with the driving scene.
|
| pose_strength |
number
|
1
Max: 10 |
Weight of the driving-motion conditioning.
|
| seed |
integer
|
Random seed; leave empty for random.
|
|
| return_masks |
boolean
|
False
|
Also return the reference/driving masks that were used (for debugging or re-use as image_mask / video_mask).
|
Output schema
The shape of the response you’ll get when you run this model with an API.
{'properties': {'driving_mask': {'format': 'uri',
'nullable': True,
'title': 'Driving Mask',
'type': 'string'},
'frames': {'format': 'uri',
'nullable': True,
'title': 'Frames',
'type': 'string'},
'metadata': {'format': 'uri',
'nullable': True,
'title': 'Metadata',
'type': 'string'},
'reference_mask': {'format': 'uri',
'nullable': True,
'title': 'Reference Mask',
'type': 'string'},
'seed': {'title': 'Seed', 'type': 'integer'},
'video': {'format': 'uri', 'title': 'Video', 'type': 'string'}},
'required': ['video', 'seed'],
'title': 'Output',
'type': 'object'}