fishaudio / s1-mini

huggingface.co
Total runs: 2.0K
24-hour runs: 0
7-day runs: -185
30-day runs: -367
Model's Last Updated: February 06 2026
text-to-speech

Introduction of s1-mini

Model Details of s1-mini

FishAudio S1

FishAudio S1 is a leading text-to-speech (TTS) model trained on more than 2 million hours of audio data in multiple languages.

Supported languages:

  • English (en)
  • Chinese (zh)
  • Japanese (ja)
  • German (de)
  • French (fr)
  • Spanish (es)
  • Korean (ko)
  • Arabic (ar)
  • Russian (ru)
  • Dutch (nl)
  • Italian (it)
  • Polish (pl)
  • Portuguese (pt)

Please refer to Fish Speech Github for more info. Demo available at Fish Audio Playground . Visit the Fish Audio website for blog & tech report.

Emotion and Tone Support

FishAudio S1 supports a variety of emotional, tone, and special markers to enhance speech synthesis:

1. Emotional markers: (angry) (sad) (disdainful) (excited) (surprised) (satisfied) (unhappy) (anxious) (hysterical) (delighted) (scared) (worried) (indifferent) (upset) (impatient) (nervous) (guilty) (scornful) (frustrated) (depressed) (panicked) (furious) (empathetic) (embarrassed) (reluctant) (disgusted) (keen) (moved) (proud) (relaxed) (grateful) (confident) (interested) (curious) (confused) (joyful) (disapproving) (negative) (denying) (astonished) (serious) (sarcastic) (conciliative) (comforting) (sincere) (sneering) (hesitating) (yielding) (painful) (awkward) (amused)

2. Tone markers: (in a hurry tone) (shouting) (screaming) (whispering) (soft tone)

3. Special markers: (laughing) (chuckling) (sobbing) (crying loudly) (sighing) (panting) (groaning) (crowd laughing) (background laughter) (audience laughing)

Special markers with corresponding onomatopoeia:

  • Laughing: Ha,ha,ha
  • Chuckling: Hmm,hmm
Model Variants and Performance

FishAudio S1 includes the following models:

  • S1 (4B, proprietary): The full-sized model.
  • S1-mini (0.5B): A distilled version of S1.

Both S1 and S1-mini incorporate online Reinforcement Learning from Human Feedback (RLHF).

Seed TTS Eval Metrics (English, auto eval, based on OpenAI gpt-4o-transcribe, speaker distance using Revai/pyannote-wespeaker-voxceleb-resnet34-LM):

  • S1:
    • WER (Word Error Rate): 0.008
    • CER (Character Error Rate): 0.004
    • Distance: 0.332
  • S1-mini:
    • WER (Word Error Rate): 0.011
    • CER (Character Error Rate): 0.005
    • Distance: 0.380
License

This model is permissively licensed under the CC-BY-NC-SA-4.0 license.

Runs of fishaudio s1-mini on huggingface.co

2.0K
Total runs
0
24-hour runs
0
3-day runs
-185
7-day runs
-367
30-day runs

More Information About s1-mini huggingface.co Model

s1-mini huggingface.co

s1-mini huggingface.co is an AI model on huggingface.co that provides s1-mini's model effect (), which can be used instantly with this fishaudio s1-mini model. huggingface.co supports a free trial of the s1-mini model, and also provides paid use of the s1-mini. Support call s1-mini model through api, including Node.js, Python, http.

fishaudio s1-mini online free

s1-mini huggingface.co is an online trial and call api platform, which integrates s1-mini's modeling effects, including api services, and provides a free online trial of s1-mini, you can try s1-mini online for free by clicking the link below.

fishaudio s1-mini online free url in huggingface.co:

https://huggingface.co/fishaudio/s1-mini

s1-mini install

s1-mini is an open source model from GitHub that offers a free installation service, and any user can find s1-mini on GitHub to install. At the same time, huggingface.co provides the effect of s1-mini install, users can directly use s1-mini installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

s1-mini install url in huggingface.co:

https://huggingface.co/fishaudio/s1-mini

Url of s1-mini

Provider of s1-mini huggingface.co

fishaudio
ORGANIZATIONS

Other API from fishaudio

huggingface.co

Total runs: 63.2K
Run Growth: -337.0K
Growth Rate: -533.02%
Updated:March 11 2026