Fish Audio S2 Pro
is a leading text-to-speech (TTS) model with fine-grained inline control of prosody and emotion. Trained on over 10M+ hours of audio data across 80+ languages, the system combines reinforcement learning alignment with a dual-autoregressive architecture. The release includes model weights, fine-tuning code, and an SGLang-based streaming inference engine.
Architecture
S2 Pro builds on a decoder-only transformer combined with an RVQ-based audio codec (10 codebooks, ~21 Hz frame rate) using a
Dual-Autoregressive (Dual-AR)
architecture:
Slow AR
(4B parameters): Operates along the time axis and predicts the primary semantic codebook.
Fast AR
(400M parameters): Generates the remaining 9 residual codebooks at each time step, reconstructing fine-grained acoustic detail.
This asymmetric design keeps inference efficient while preserving audio fidelity. Because the Dual-AR architecture is structurally isomorphic to standard autoregressive LLMs, it inherits all LLM-native serving optimizations from SGLang — including continuous batching, paged KV cache, CUDA graph replay, and RadixAttention-based prefix caching.
Fine-Grained Inline Control
S2 Pro enables localized control over speech generation by embedding natural-language instructions directly within the text using
[tag]
syntax. Rather than relying on a fixed set of predefined tags, S2 Pro accepts
free-form textual descriptions
— such as
[whisper in small voice]
,
[professional broadcast tone]
, or
[pitch up]
— allowing open-ended expression control at the word level.
Tier 2:
Korean (ko), Spanish (es), Portuguese (pt), Arabic (ar), Russian (ru), French (fr), German (de)
Other supported languages:
sv, it, tr, no, nl, cy, eu, ca, da, gl, ta, hu, fi, pl, et, hi, la, ur, th, vi, jw, bn, yo, sl, cs, sw, nn, he, ms, uk, id, kk, bg, lv, my, tl, sk, ne, fa, af, el, bo, hr, ro, sn, mi, yi, am, be, km, is, az, sd, br, sq, ps, mn, ht, ml, sr, sa, te, ka, bs, pa, lt, kn, si, hy, mr, as, gu, fo, and more.
Production Streaming Performance
On a single NVIDIA H200 GPU:
Real-Time Factor (RTF):
0.195
Time-to-first-audio:
~100 ms
Throughput:
3,000+ acoustic tokens/s while maintaining RTF below 0.5
This model is licensed under the
Fish Audio Research License
. Research and non-commercial use is permitted free of charge. Commercial use requires a separate license from Fish Audio — contact
[email protected]
.
Runs of fishaudio s2-pro on huggingface.co
541.8K
Total runs
0
24-hour runs
101.5K
3-day runs
127.5K
7-day runs
241.6K
30-day runs
More Information About s2-pro huggingface.co Model
s2-pro huggingface.co is an AI model on huggingface.co that provides s2-pro's model effect (), which can be used instantly with this fishaudio s2-pro model. huggingface.co supports a free trial of the s2-pro model, and also provides paid use of the s2-pro. Support call s2-pro model through api, including Node.js, Python, http.
s2-pro huggingface.co is an online trial and call api platform, which integrates s2-pro's modeling effects, including api services, and provides a free online trial of s2-pro, you can try s2-pro online for free by clicking the link below.
fishaudio s2-pro online free url in huggingface.co:
s2-pro is an open source model from GitHub that offers a free installation service, and any user can find s2-pro on GitHub to install. At the same time, huggingface.co provides the effect of s2-pro install, users can directly use s2-pro installed effect in huggingface.co for debugging and trial. It also supports api for free installation.