You're looking at a specific version of this model. Jump to the model overview.

usamaehsan /chatterbox-mtl-batched:27e7bfca

Input schema

The fields you can use to run this model with an API. If you don’t give a value for a field its default value will be used.

Field Type Default value Description
text
string
Text to narrate. No length cap; it is split into sentences and synthesised in one batched pass.
language
None
es
Language of the text.
reference_audio
string
Optional 7-20s clip of the voice to clone. Omitted uses the built-in voice.
exaggeration
number
0.5

Min: 0.25

Max: 2

Expressiveness. 0.5 is neutral.
cfg_weight
number
0.5

Max: 1

Pace/guidance. 0 disables CFG and halves the cost.
temperature
number
0.8

Min: 0.05

Max: 5

None
flow_steps
integer
10

Min: 1

Max: 10

Token-to-mel flow-matching steps. Fewer is cheaper.
chunk_chars
integer
40

Min: 30

Max: 300

Target characters per chunk. Lower means more rows decoded in parallel and a shorter critical path, at some risk of audible seams.
fast
boolean
True
Use the hand-written decoder. Off falls back to transformers' eager loop, for comparison.
graph
boolean
False
Replay CUDA graphs instead of launching kernels. Experimental and off: capture succeeds but the replayed decode produces non-finite logits, and no graph is captured at boot unless GRAPH=1.
batch_sentences
boolean
True
Decode every sentence in one pass. Off synthesises serially, for cost comparison.
silence
number
0.25

Max: 2

Seconds of silence between sentences.
watermark
boolean
True
Apply Resemble's Perth watermark to the output.
output_format
None
mp3
None
mp3_bitrate
None
96k
None
seed
integer
0
0 for random.

Output schema

The shape of the response you’ll get when you run this model with an API.

Schema
{'format': 'uri', 'title': 'Output', 'type': 'string'}