MOSS‑TTS Family is an open‑source
speech and sound generation model family
from
MOSI.AI
and the
OpenMOSS team
. It is designed for
high‑fidelity
,
high‑expressiveness
, and
complex real‑world scenarios
, covering stable long‑form speech, multi‑speaker dialogue, voice/character design, environmental sound effects, and real‑time streaming TTS.
Introduction
When a single piece of audio needs to
sound like a real person
,
pronounce every word accurately
,
switch speaking styles across content
,
remain stable over tens of minutes
, and
support dialogue, role‑play, and real‑time interaction
, a single TTS model is often not enough. The
MOSS‑TTS Family
breaks the workflow into five production‑ready models that can be used independently or composed into a complete pipeline.
MOSS‑TTS
: MOSS-TTS is the flagship production TTS foundation model, centered on high-fidelity zero-shot voice cloning with controllable long-form synthesis, pronunciation, and multilingual/code-switched speech. It serves as the core engine for scalable narration, dubbing, and voice-driven products.
MOSS‑TTSD
: MOSS-TTSD is a production long-form dialogue model for expressive multi-speaker conversational audio at scale. It supports long-duration continuity, turn-taking control, and zero-shot voice cloning from short references for podcasts, audiobooks, commentary, dubbing, and entertainment dialogue.
MOSS‑VoiceGenerator
: MOSS-VoiceGenerator is an open-source voice design model that creates speaker timbres directly from free-form text, without reference audio. It unifies timbre design, style control, and content synthesis, and can be used standalone or as a voice-design layer for downstream TTS.
MOSS‑SoundEffect
: MOSS-SoundEffect is a high-fidelity text-to-sound model with broad category coverage and controllable duration for real content production. It generates stable audio from prompts across ambience, urban scenes, creatures, human actions, and music-like clips for film, games, interactive media, and data synthesis.
MOSS‑TTS‑Realtime
: MOSS-TTS-Realtime is a context-aware, multi-turn streaming TTS model for real-time voice agents. By conditioning on dialogue history across both text and prior user acoustics, it delivers low-latency synthesis with coherent, consistent voice responses across turns.
MOSS-SoundEffect
is the
environment sound & sound effect generation model
in the
MOSS‑TTS Family
. It generates ambient soundscapes and concrete sound effects directly from text descriptions, and is designed to complement speech content with immersive context in production workflows.
1. Overview
1.1 TTS Family Positioning
MOSS-SoundEffect is designed as an audio generation backbone for creating high-fidelity environmental and action sounds from text, serving both scalable content pipelines and a strong research baseline for controllable audio generation.
Design goals
Coverage & richness
: broad sound taxonomy with layered ambience and realistic texture
Composability
: easy integration into creative pipelines (games/film/tools) and synthetic data generation setups
1.2 Key Capabilities
MOSS‑SoundEffect focuses on
contextual audio completion
beyond speech, enabling creators and systems to enrich scenes with believable acoustic environments and action‑level cues.
What it can generate
Natural environments
: e.g., “fresh snow crunching under footsteps.”
Urban environments
: e.g., “a sports car roaring past on the highway.”
Animals & creatures
: e.g., “early morning park with birds chirping in a quiet atmosphere.”
Human actions
: e.g., “clear footsteps echoing on concrete at a steady rhythm.”
Why it matters
Completes
scene immersion
for narrative content, film/TV, documentaries, games, and podcasts.
Supports
voice agents
and interactive systems that need ambient context, not just speech.
Acts as the
sound‑design layer
of the MOSS‑TTS Family’s end‑to‑end workflow.
1.3 Model Architecture
MOSS-SoundEffect
employs the
MossTTSDelay
architecture (see
moss_tts_delay/README.md
), reusing the same discrete token generation backbone for audio synthesis. A text prompt (optionally with simple control tags such as
duration
) is tokenized and fed into the Delay-pattern autoregressive model to predict
RVQ audio tokens
over time. The generated tokens are then decoded by the audio tokenizer/vocoder to produce high-fidelity sound effects, enabling consistent quality and controllable length across diverse SFX categories.
1.4 Released Models
Recommended decoding hyperparameters
Model
audio_temperature
audio_top_p
audio_top_k
audio_repetition_penalty
MOSS-SoundEffect
1.5
0.6
50
1.2
2. Quick Start
Environment Setup
We recommend a clean, isolated Python environment with
Transformers 5.0.0
to avoid dependency conflicts.
MOSS-SoundEffect huggingface.co is an AI model on huggingface.co that provides MOSS-SoundEffect's model effect (), which can be used instantly with this OpenMOSS-Team MOSS-SoundEffect model. huggingface.co supports a free trial of the MOSS-SoundEffect model, and also provides paid use of the MOSS-SoundEffect. Support call MOSS-SoundEffect model through api, including Node.js, Python, http.
MOSS-SoundEffect huggingface.co is an online trial and call api platform, which integrates MOSS-SoundEffect's modeling effects, including api services, and provides a free online trial of MOSS-SoundEffect, you can try MOSS-SoundEffect online for free by clicking the link below.
OpenMOSS-Team MOSS-SoundEffect online free url in huggingface.co:
MOSS-SoundEffect is an open source model from GitHub that offers a free installation service, and any user can find MOSS-SoundEffect on GitHub to install. At the same time, huggingface.co provides the effect of MOSS-SoundEffect install, users can directly use MOSS-SoundEffect installed effect in huggingface.co for debugging and trial. It also supports api for free installation.