usamaehsan/chatterbox-mtl-batched
Chatterbox Multilingual V3 (MIT) with batched sentence synthesis: a whole story is decoded in one pass instead of chunk by chunk. Voice cloning from a short reference, 23 languages.
Run usamaehsan/chatterbox-mtl-batched with an API
Use one of our client libraries to get started quickly. Clicking on a library will take you to the Playground tab where you can tweak different inputs, see the results, and copy the corresponding code to use in your own project.
Input schema
The fields you can use to run this model with an API. If you don't give a value for a field its default value will be used.
| Field | Type | Default value | Description |
|---|---|---|---|
| text |
string
|
Text to narrate. No length cap; it is split into sentences and synthesised in one batched pass.
|
|
| language |
None
|
es
|
Language of the text.
|
| reference_audio |
string
|
Optional 7-20s clip of the voice to clone. Omitted uses the built-in voice.
|
|
| exaggeration |
number
|
0.5
Min: 0.25 Max: 2 |
Expressiveness. 0.5 is neutral.
|
| cfg_weight |
number
|
0.5
Max: 1 |
Pace/guidance. 0 disables CFG and halves the cost.
|
| temperature |
number
|
0.8
Min: 0.05 Max: 5 |
None
|
| flow_steps |
integer
|
10
Min: 1 Max: 10 |
Token-to-mel flow-matching steps. Fewer is cheaper.
|
| chunk_chars |
integer
|
40
Min: 30 Max: 300 |
Target characters per chunk. Lower means more rows decoded in parallel and a shorter critical path, at some risk of audible seams.
|
| fast |
boolean
|
True
|
Use the hand-written decoder. Off falls back to transformers' eager loop, for comparison.
|
| graph |
boolean
|
False
|
Replay CUDA graphs instead of launching kernels. Experimental and off: capture succeeds but the replayed decode produces non-finite logits, and no graph is captured at boot unless GRAPH=1.
|
| batch_sentences |
boolean
|
True
|
Decode every sentence in one pass. Off synthesises serially, for cost comparison.
|
| silence |
number
|
0.25
Max: 2 |
Seconds of silence between sentences.
|
| watermark |
boolean
|
True
|
Apply Resemble's Perth watermark to the output.
|
| output_format |
None
|
mp3
|
None
|
| mp3_bitrate |
None
|
96k
|
None
|
| seed |
integer
|
0
|
0 for random.
|
{
"type": "object",
"title": "Input",
"required": [
"text"
],
"properties": {
"fast": {
"type": "boolean",
"title": "Fast",
"default": true,
"x-order": 8,
"description": "Use the hand-written decoder. Off falls back to transformers' eager loop, for comparison."
},
"seed": {
"type": "integer",
"title": "Seed",
"default": 0,
"x-order": 15,
"description": "0 for random."
},
"text": {
"type": "string",
"title": "Text",
"x-order": 0,
"description": "Text to narrate. No length cap; it is split into sentences and synthesised in one batched pass."
},
"graph": {
"type": "boolean",
"title": "Graph",
"default": false,
"x-order": 9,
"description": "Replay CUDA graphs instead of launching kernels. Experimental and off: capture succeeds but the replayed decode produces non-finite logits, and no graph is captured at boot unless GRAPH=1."
},
"silence": {
"type": "number",
"title": "Silence",
"default": 0.25,
"maximum": 2,
"minimum": 0,
"x-order": 11,
"description": "Seconds of silence between sentences."
},
"language": {
"enum": [
"ar",
"da",
"de",
"el",
"en",
"es",
"fi",
"fr",
"he",
"hi",
"it",
"ja",
"ko",
"ms",
"nl",
"no",
"pl",
"pt",
"ru",
"sv",
"sw",
"tr",
"zh"
],
"type": "string",
"title": "language",
"description": "Language of the text.",
"default": "es",
"x-order": 1
},
"watermark": {
"type": "boolean",
"title": "Watermark",
"default": true,
"x-order": 12,
"description": "Apply Resemble's Perth watermark to the output."
},
"cfg_weight": {
"type": "number",
"title": "Cfg Weight",
"default": 0.5,
"maximum": 1,
"minimum": 0,
"x-order": 4,
"description": "Pace/guidance. 0 disables CFG and halves the cost."
},
"flow_steps": {
"type": "integer",
"title": "Flow Steps",
"default": 10,
"maximum": 10,
"minimum": 1,
"x-order": 6,
"description": "Token-to-mel flow-matching steps. Fewer is cheaper."
},
"chunk_chars": {
"type": "integer",
"title": "Chunk Chars",
"default": 40,
"maximum": 300,
"minimum": 30,
"x-order": 7,
"description": "Target characters per chunk. Lower means more rows decoded in parallel and a shorter critical path, at some risk of audible seams."
},
"mp3_bitrate": {
"enum": [
"64k",
"96k",
"128k",
"192k"
],
"type": "string",
"title": "mp3_bitrate",
"description": "An enumeration.",
"default": "96k",
"x-order": 14
},
"temperature": {
"type": "number",
"title": "Temperature",
"default": 0.8,
"maximum": 5,
"minimum": 0.05,
"x-order": 5
},
"exaggeration": {
"type": "number",
"title": "Exaggeration",
"default": 0.5,
"maximum": 2,
"minimum": 0.25,
"x-order": 3,
"description": "Expressiveness. 0.5 is neutral."
},
"output_format": {
"enum": [
"mp3",
"wav"
],
"type": "string",
"title": "output_format",
"description": "An enumeration.",
"default": "mp3",
"x-order": 13
},
"batch_sentences": {
"type": "boolean",
"title": "Batch Sentences",
"default": true,
"x-order": 10,
"description": "Decode every sentence in one pass. Off synthesises serially, for cost comparison."
},
"reference_audio": {
"type": "string",
"title": "Reference Audio",
"format": "uri",
"x-order": 2,
"description": "Optional 7-20s clip of the voice to clone. Omitted uses the built-in voice."
}
}
}
Output schema
The shape of the response you’ll get when you run this model with an API.
{
"type": "string",
"title": "Output",
"format": "uri"
}