platform-kit / mars5-tts

A novel speech model for insane prosody.

  • Public
  • 438 runs
  • A100 (80GB)
  • GitHub
  • License

Input

Video Player is loading.
Current Time 00:00:000
Duration 00:00:000
Loaded: 0%
Stream Type LIVE
Remaining Time 00:00:000
 
1x
integer

Default: 95

number
(minimum: 0, maximum: 5)

Default: 0.5

integer

Default: 3

integer
(minimum: 0, maximum: 100)

Default: 95

string
Shift + Return to add a new line

Text to synthesize

Default: "Hi there, I'm your new voice clone, powered by Mars5."

file

Reference audio file to clone from <= 10 seconds

Default: "https://replicate.delivery/pbxt/L9a6SelzU0B2DIWeNpkNR0CKForWSbkswoUP69L0NLjLswVV/voice_sample.wav"

string
Shift + Return to add a new line

Text in the reference audio file

Default: "Hi there. I'm your new voice clone. Try your best to upload quality audio."

Output

Video Player is loading.
Current Time 00:00:000
Duration 00:00:000
Loaded: 0%
Stream Type LIVE
Remaining Time 00:00:000
 
1x
Generated in

Run time and cost

This model costs approximately $0.11 to run on Replicate, or 9 runs per $1, but this varies depending on your inputs. It is also open source and you can run it on your own computer with Docker.

This model runs on Nvidia A100 (80GB) GPU hardware. Predictions typically complete within 76 seconds. The predict time for this model varies significantly based on the inputs.

Readme

This is a demo for the MARS5 English speech model (TTS) from CAMB.AI.

The model follows a two-stage AR-NAR pipeline with a distinctively novel NAR component (see more info in the Architecture).

With just 5 seconds of audio and a snippet of text, MARS5 can generate speech even for prosodically hard and diverse scenarios like sports commentary, anime and more.