You're looking at a specific version of this model. Jump to the model overview.

sprited /scail-2:e2c47272

Input schema

The fields you can use to run this model with an API. If you don’t give a value for a field its default value will be used.

Field Type Default value Description
image
string
Reference character image. Transparent PNGs are composited on white and their alpha becomes the reference mask.
video
string
Driving video or animated WebP. Its motion is transferred to the character. WebP transparency is used as the driving mask unless video_mask is supplied.
prompt
string
Describe the character and the motion, e.g. 'A cartoon robot walking in place, side view'. Describes the final video; not instructions.
negative_prompt
string
What to avoid, e.g. 'distorted limbs, camera movement, blurry'. Only matters with guidance_scale > 1 (the quality preset). Wan's stock Chinese negative prompt is deliberately not applied: its 'painting / artwork / style' terms push stylized characters toward a 3D-CG look.
mode
None
animation
animation: the reference character (and its background) performs the driving motion. replacement: the character is placed into the driving video, keeping its background and lighting.
image_mask
string
Optional mask for the reference: a grayscale/black-and-white matte (white = character) or a SCAIL-2 palette mask (blue = identity 0). If omitted and auto_mask is on, one is derived from the image's alpha channel or BiRefNet.
video_mask
string
Optional per-frame mask video for the driving video, same conventions as image_mask (grayscale matte or SCAIL-2 colours). If omitted, animated WebP transparency is used when available; otherwise auto_mask uses BiRefNet.
auto_mask
boolean
True
Derive missing masks automatically (alpha channel or BiRefNet single-subject matting). Off disables automatic matting; supplied masks and driving WebP transparency still apply.
resolution
None
512p
Short side of the output; the driving video's aspect ratio is kept (e.g. 896x512 for 16:9, 512x512 for square). SCAIL-2 was trained at both.
width
integer
0

Max: 1536

Explicit output width (multiple of 32). Set together with height to override resolution; the driving video is centre-cropped to this aspect.
height
integer
0

Max: 1536

Explicit output height (multiple of 32).
num_frames
integer
0
None
fps
integer
0

Max: 60

Resample the driving video to this frame rate before animating (frames are held, never interpolated); the output uses the same rate. 0 = keep the driving video's rate.
preset
None
fast
fast (default): the official ComfyUI recipe — lightx2v step/CFG-distill LoRA, Euler, 6 steps, CFG 1, shift 5; sprute: UniPC, 8 steps, CFG 1, shift 5, DPO before LightX2V. Model/VAE precision are selected separately. quality: the paper's sampler (UniPC, 40 steps, CFG 5, shift 3) — 81 frames at 512p takes ~10 min on an H100.
prepared_inputs
boolean
False
Already-composited RGB grids and exact RGB SCAIL palette masks from local Sprute. No resize, crop, recoloring or automatic masks. Inputs must have matching dimensions; video masks must match frame count and FPS. Use lossless WebP or FFV1 MKV.
additional_images
array
Up to 7 additional reference views. Supply a matching additional_image_masks list and image_mask; palette colors bind views to the same identities. CLIP vision uses the primary image.
additional_image_masks
array
Palette mask for each additional reference, in the same order.
previous_frames
string
Previous output as an image or lossless video. Its tail anchors the beginning of this request. The driving video must include that overlapping interval, and returned frames include the anchor interval.
previous_frame_count
integer
5

Min: 1

Max: 77

Tail frames used for anchoring and chunk overlap; must be 4n+1. Use 1 for a single-image anchor. SCAIL-2 was trained with 5.
vae_precision
None
default
Select VAE checkpoint; Sprute uses bf16.
sampler_name
None
preset
Override preset sampler.
scheduler
None
simple
ComfyUI sampling schedule.
lightx2v_lora
number
-1

Min: -1

Max: 2

LightX2V strength: -1 uses preset, 0 disables. Sprute uses 0.8.
pose_start
number
0

Max: 1

Start fraction of pose conditioning.
pose_end
number
1

Max: 1

End fraction of pose conditioning.
denoise
number
1

Min: 0.001

Max: 1

KSampler denoise strength.
return_frames
boolean
False
Return original decoded PNG frames as ZIP for lossless local Toonout and sprite assembly. MP4 is only a preview.
steps
integer
0

Max: 100

Sampling steps; 0 = preset default.
guidance_scale
number
0

Max: 20

Classifier-free guidance; 0 = preset default (5 quality / 1 fast).
shift
number
0

Max: 20

Flow-matching schedule shift; 0 = preset default (3 quality / 5 fast).
dpo_lora
number
1

Max: 2

Strength of the official Bias-Aware DPO LoRA (0 = off). 1.0 is what the official ComfyUI template ships with; the released base checkpoint is pre-DPO.
relight_lora
number
0

Max: 2

Strength of the official relighting LoRA for replacement mode (0 = off). Improves lighting consistency with the driving scene.
pose_strength
number
1

Max: 10

Weight of the driving-motion conditioning.
seed
integer
Random seed; leave empty for random.
return_masks
boolean
False
Also return the reference/driving masks that were used (for debugging or re-use as image_mask / video_mask).

Output schema

The shape of the response you’ll get when you run this model with an API.

Schema
{'properties': {'driving_mask': {'format': 'uri',
                                 'nullable': True,
                                 'title': 'Driving Mask',
                                 'type': 'string'},
                'frames': {'format': 'uri',
                           'nullable': True,
                           'title': 'Frames',
                           'type': 'string'},
                'metadata': {'format': 'uri',
                             'nullable': True,
                             'title': 'Metadata',
                             'type': 'string'},
                'reference_mask': {'format': 'uri',
                                   'nullable': True,
                                   'title': 'Reference Mask',
                                   'type': 'string'},
                'seed': {'title': 'Seed', 'type': 'integer'},
                'video': {'format': 'uri', 'title': 'Video', 'type': 'string'}},
 'required': ['video', 'seed'],
 'title': 'Output',
 'type': 'object'}