You're looking at a specific version of this model. Jump to the model overview.
nicolascoutureau /ac:cb5cf186
Input schema
The fields you can use to run this model with an API. If you don’t give a value for a field its default value will be used.
| Field | Type | Default value | Description |
|---|---|---|---|
| video |
string
|
Input horizontal video to convert to vertical format
|
|
| aspect_ratio |
None
|
9:16
|
Output aspect ratio
|
| speed_preset |
None
|
balanced
|
Processing speed preset. Fast uses smaller YOLO model and lower resolution analysis.
|
| detect_speaker |
boolean
|
True
|
Detect and focus on the active speaker using TalkNet ASD (audio-visual neural network). When multiple people are detected, follows whoever is talking (switching with hysteresis) and shows the two main speakers in split screen during a conversation.
|
| tracking_mode |
None
|
smooth
|
Camera tracking mode. Smooth = cinematic spring-damped movement. Fast = more responsive. Static = fixed per scene.
|
| debug_overlay |
boolean
|
False
|
Draw debug info on video (scene, strategy, speaker, ASD scores, crop).
|
| start |
number
|
Start of the range to process, in seconds (optional). Audio is trimmed too.
|
|
| end |
number
|
End of the range to process, in seconds (optional, default: end of video).
|
|
| focus |
None
|
auto
|
What to frame. auto = active speaker when detected, else the main person, motion when nobody is on screen. speaker = always follow the active speaker (no split screen). action = follow motion, snapping to the nearest person. center = fixed centre crop.
|
| allow_letterbox |
boolean
|
False
|
For scenes without people, show the whole frame over a blurred background instead of cropping on motion / saliency.
|
| min_output_height |
integer
|
0
Max: 4096 |
Minimum output height in pixels. 0 = standard size (1080x1920 for 9:16, 1080x1350 for 4:5, 1080x1080 for 1:1, 1920x1080 for 16:9).
|
| detect_interval |
number
|
0.5
Min: 0.1 Max: 5 |
Seconds between person detections used by the tracking camera.
|
| return_track |
boolean
|
False
|
Also return the crop track as JSON (per keyframe: time, box, strategy, speaker, confidence). When true the output is an object {video, track} instead of a single video URL.
|
Output schema
The shape of the response you’ll get when you run this model with an API.
Schema
{'title': 'Output'}