LFM2.5-Audio-1.5B-JP is
Liquid AI
's first Japanese capable audio model and the first Japanese speech-to-speech model at this lightweight scale built from the foundations of
LFM2.5-Audio-1.5B.
LFM2.5-Audio-1.5B-JP is an end-to-end multimodal speech and text language model, and as such does not require separate ASR and TTS components.
Designed with low latency and real time conversation in mind, at only 1.5 billion parameters LFM2.5-Audio-JP enables seamless Japanese conversational interaction, achieving capabilities on par with much larger models.
Our model consists of a pretrained LFM2.5 model as its multimodal backbone, along with a FastConformer based audio encoder to handle continuous audio inputs, and an RQ-transformer generating discrete tokens coupled with a lightweight audio detokenizer for audio output.
LFM2.5-Audio-JP supports two distinct generation routines, each suitable for a set of tasks.
Interleaved generation enables real-time speech-to-speech conversational chatbot capabilities, where audio generation latency is key.
Sequential generation is suited for non-conversational tasks such as ASR or TTS, and allows the model to switch generated modality on the fly.
LFM2.5-Audio-1.5B-JPは、
Liquid AI
初となる日本語対応の音声モデルであり、
LFM2.5-Audio-1.5B
の基盤から構築された、この軽量な規模においては初となる日本語のSpeech-to-Speechモデルです。
We use GPT-4o as an LLM-as-a-judge and report the highest score assigned to each sample. Elyza and spoken_Elyza are scored on a 0–5 scale, while M-ifeval is scored on a 0–1 scale. Average performance was computed after rescaling all benchmark scores to a 0–5 scale.
pip install liquid-audio
pip install "liquid-audio [demo]"# optional, to install demo dependencies
pip install flash-attn --no-build-isolation # optional, to use flash attention 2. Will fallback to torch SDPA if not installed
Multi-turn, multi-modal chat
The
liquid-audio
library provides a lower lever interface to the model and generation routines, ideal for custom usecases.
We demonstrate this with a simple multi-turn chat, where the first turn is given as audio, and the second turn is given as text.
For multi-turn chat with text and audio output, we use interleaved generation. The system prompt should be set to
Respond with interleaved text and audio.
. Here we use audio as the first user turn, and text as the second one.
import torch
import soundfile as sf
from liquid_audio import LFM2AudioModel, LFM2AudioProcessor, ChatState, LFMModality
# Load models
HF_REPO = "LiquidAI/LFM2.5-Audio-1.5B-JP"
processor = LFM2AudioProcessor.from_pretrained(HF_REPO).eval()
model = LFM2AudioModel.from_pretrained(HF_REPO).eval()
# Set up inputs for the model
chat = ChatState(processor)
chat.new_turn("system")
chat.add_text("Respond with interleaved text and audio.")
chat.end_turn()
chat.new_turn("user")
wav, sampling_rate = sf.read("assets/question_jp.wav", dtype="float32")
wav = torch.from_numpy(wav).unsqueeze(0)
chat.add_audio(wav, sampling_rate)
chat.end_turn()
chat.new_turn("assistant")
# Generate text and audio tokens.
text_out: list[torch.Tensor] = []
audio_out: list[torch.Tensor] = []
modality_out: list[LFMModality] = []
for t in model.generate_interleaved(**chat, max_new_tokens=512, audio_temperature=1.0, audio_top_k=4):
if t.numel() == 1:
print(processor.text.decode(t), end="", flush=True)
text_out.append(t)
modality_out.append(LFMModality.TEXT)
else:
audio_out.append(t)
modality_out.append(LFMModality.AUDIO_OUT)
# output: こんにちは。私はリキッドリリーと申します。質問に答えたり、アドバイスを提供したりするためのAIボイスアシスタントです。リアルタイムでさまざまな言語タスクをお手伝いするよう設計されています。# Detokenize audio, removing the last "end-of-audio" codes# Mimi returns audio at 24kHz
audio_codes = torch.stack(audio_out[:-1], 1).unsqueeze(0)
waveform = processor.decode(audio_codes)
sf.write("answer_jp1.wav", waveform.cpu()[0], 24_000)
# Append newly generated tokens to chat history
chat.append(
text = torch.stack(text_out, 1),
audio_out = torch.stack(audio_out, 1),
modality_flag = torch.tensor(modality_out),
)
chat.end_turn()
# Start new turn
chat.new_turn("user")
chat.add_text("富士山の高さは何メートルですか。")
chat.end_turn()
chat.new_turn("assistant")
# Generate second turn text and audio tokens.
audio_out: list[torch.Tensor] = []
for t in model.generate_interleaved(**chat, max_new_tokens=512, audio_temperature=1.0, audio_top_k=4):
if t.numel() == 1:
print(processor.text.decode(t), end="", flush=True)
else:
audio_out.append(t)
# output: 富士山の高さは約3,776メートルです。# Detokenize second turn audio, removing the last "end-of-audio" codes
audio_codes = torch.stack(audio_out[:-1], 1).unsqueeze(0)
waveform = processor.decode(audio_codes)
sf.write("answer_jp2.wav", waveform.cpu()[0], 24_000)
ASR
For ASR, we use sequential generation, with the fixed system prompt
Perform ASR in japanese.
.
import torch
import soundfile as sf
from liquid_audio import LFM2AudioModel, LFM2AudioProcessor, ChatState, LFMModality
# Load models
HF_REPO = "LiquidAI/LFM2.5-Audio-1.5B-JP"
processor = LFM2AudioProcessor.from_pretrained(HF_REPO).eval()
model = LFM2AudioModel.from_pretrained(HF_REPO).eval()
# Set up inputs for the model
chat = ChatState(processor)
chat.new_turn("system")
chat.add_text("Perform ASR in japanese.")
chat.end_turn()
chat.new_turn("user")
wav, sampling_rate = sf.read("assets/asr_jp.wav", dtype="float32")
wav = torch.from_numpy(wav).unsqueeze(0)
chat.add_audio(wav, sampling_rate)
chat.end_turn()
chat.new_turn("assistant")
# Generate textfor t in model.generate_sequential(**chat, max_new_tokens=512):
if t.numel() == 1:
print(processor.text.decode(t), end="", flush=True)
# Output: この度は弊社の確認不足により多大なご迷惑をおかけしましたことを深くお詫び申し上げます。今後はこのようなことが二度と起こらないよう社内のチェック体制を徹底してまいります。
TTS
For TTS, we also use sequential generation, with the fixed system prompt
Perform TTS in japanese.
.
import torch
import soundfile as sf
from liquid_audio import LFM2AudioModel, LFM2AudioProcessor, ChatState, LFMModality
# Load models
HF_REPO = "LiquidAI/LFM2.5-Audio-1.5B-JP"
processor = LFM2AudioProcessor.from_pretrained(HF_REPO).eval()
model = LFM2AudioModel.from_pretrained(HF_REPO).eval()
# Set up inputs for the model
chat = ChatState(processor)
chat.new_turn("system")
chat.add_text("Perform TTS in japanese.")
chat.end_turn()
chat.new_turn("user")
chat.add_text("先週ご相談いただいた新しいプロジェクトの件ですが、社内で検討した結果、ぜひ前向きに進めさせていただきたいと考えております。つきましては、具体的なスケジュールについて一度お打ち合わせの機会をいただけますでしょうか。")
chat.end_turn()
chat.new_turn("assistant")
# Generate text
audio_out: list[torch.Tensor] = []
for t in model.generate_sequential(**chat, max_new_tokens=512, audio_temperature = 0.8, audio_top_k=64):
if t.numel() > 1:
audio_out.append(t)
# Detokenize audio
audio_codes = torch.stack(audio_out[:-1], 1).unsqueeze(0)
waveform = processor.decode(audio_codes)
sf.write("tts_jp.wav", waveform.cpu()[0], 24_000)
train a model from the preprocessed dataset with
LFM2DataLoader
To finetune our Japanese model on your own data in interleaved generation mode, instantiate the
LFM2AudioChatMapper
class with
interleaved_text_tokens=6
and
interleaved_audio_tokens=9
. These values reflect the predefined Japanese interleaving ratio of 6 text tokens to 9 audio tokens, based on tokenization statistics.
LFM2.5-Audio-1.5B-JP huggingface.co is an AI model on huggingface.co that provides LFM2.5-Audio-1.5B-JP's model effect (), which can be used instantly with this LiquidAI LFM2.5-Audio-1.5B-JP model. huggingface.co supports a free trial of the LFM2.5-Audio-1.5B-JP model, and also provides paid use of the LFM2.5-Audio-1.5B-JP. Support call LFM2.5-Audio-1.5B-JP model through api, including Node.js, Python, http.
LFM2.5-Audio-1.5B-JP huggingface.co is an online trial and call api platform, which integrates LFM2.5-Audio-1.5B-JP's modeling effects, including api services, and provides a free online trial of LFM2.5-Audio-1.5B-JP, you can try LFM2.5-Audio-1.5B-JP online for free by clicking the link below.
LiquidAI LFM2.5-Audio-1.5B-JP online free url in huggingface.co:
LFM2.5-Audio-1.5B-JP is an open source model from GitHub that offers a free installation service, and any user can find LFM2.5-Audio-1.5B-JP on GitHub to install. At the same time, huggingface.co provides the effect of LFM2.5-Audio-1.5B-JP install, users can directly use LFM2.5-Audio-1.5B-JP installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
LFM2.5-Audio-1.5B-JP install url in huggingface.co: