









Miso One is an open-weights, 8B-parameter text-to-speech (TTS) system developed by Miso Labs. It is designed specifically for producing highly realistic, expressive, and emotionally varied English conversational speech, making it ideal for voice-agent research and developer workflows. Built on a Sesame-style conversational speech model (CSM) architecture with Mimi audio codes, it features a highly optimized inference capability boasting a published low latency of 110 ms. In addition to text-to-speech generation, the model supports voice continuation and one-shot voice cloning from audio context with clear consent boundaries.
Users can evaluate Miso One by reading its official model card on the repository or Hugging Face page, trying out the hosted web demo to check voice quality, or downloading the public 8B weights and inference code to run local benchmarks within their own CUDA environment. For hosted creator workflows, users can sign up and choose a subscription plan based on their required annual or monthly character capacity.
Here is the Miso One support email for customer service: [email protected] . More Contact, visit the contact us page()
Miso One Company name: Miso One .
Miso One Company address: .
More about Miso One, Please visit the about us page().
Miso One Github Link: https://github.com/MisoLabsAI/MisoTTS

Free tier
$0
Includes free credits for initial testing with a maximum of 120 characters per conversion.
Basic (Annual Plan)
$4.95 per month
Billed annually ($9.90/month if monthly). Includes 960,000 TTS characters per year, 9,600 voice credits, up to 480 instant voice clones, private voice model creation, and email support.
Pro (Annual Plan)
$14.95 per month
Billed annually ($29.90/month if monthly). Includes 4,200,000 TTS characters per year, 42,000 voice credits, up to 2,100 instant voice clones, and priority support for voice workflows.
Enterprise (Annual Plan)
$24.95 per month
Billed annually ($49.90/month if monthly). Includes 9,600,000 TTS characters per year, 96,000 voice credits, up to 4,800 instant voice clones, and dedicated team priority support.

Social Listening