SeamlessM4T is a collection of models designed to provide high quality translation, allowing people from different
linguistic communities to communicate effortlessly through speech and text.
This repository hosts 🤗 Hugging Face's
implementation
of SeamlessM4T. You can find the original weights, as well as a guide on how to run them in the original hub repositories (
large
and
medium
checkpoints).
🌟 SeamlessM4T v2, an improved version of this version with a novel architecture, has been released
here
.
This new model improves over SeamlessM4T v1 in quality as well as inference speed in speech generation tasks.
This is the "medium" variant of the unified model, which enables multiple tasks without relying on multiple separate models:
Speech-to-speech translation (S2ST)
Speech-to-text translation (S2TT)
Text-to-speech translation (T2ST)
Text-to-text translation (T2TT)
Automatic speech recognition (ASR)
You can perform all the above tasks from one single model,
SeamlessM4TModel
, but each task also has its own dedicated sub-model.
🤗 Usage
First, load the processor and a checkpoint of the model:
>>> from transformers import AutoProcessor, SeamlessM4TModel
>>> processor = AutoProcessor.from_pretrained("facebook/hf-seamless-m4t-medium")
>>> model = SeamlessM4TModel.from_pretrained("facebook/hf-seamless-m4t-medium")
You can seamlessly use this model on text or on audio, to generated either translated text or translated audio.
Here is how to use the processor to process text and audio:
>>> # let's load an audio sample from an Arabic speech corpus>>> from datasets import load_dataset
>>> dataset = load_dataset("arabic_speech_corpus", split="test", streaming=True)
>>> audio_sample = next(iter(dataset))["audio"]
>>> # now, process it>>> audio_inputs = processor(audios=audio_sample["array"], return_tensors="pt")
>>> # now, process some English test as well>>> text_inputs = processor(text = "Hello, my dog is cute", src_lang="eng", return_tensors="pt")
Speech
SeamlessM4TModel
can
seamlessly
generate text or speech with few or no changes. Let's target Russian voice translation:
With basically the same code, I've translated English text and Arabic speech to Russian speech samples.
Text
Similarly, you can generate translated text from audio files or from text with the same model. You only have to pass
generate_speech=False
to
SeamlessM4TModel.generate
.
This time, let's translate to French.
SeamlessM4TModel
is transformers top level model to generate speech and text, but you can also use dedicated models that perform the task without additional components, thus reducing the memory footprint.
For example, you can replace the audio-to-audio generation snippet with the model dedicated to the S2ST task, the rest is exactly the same code:
>>> from transformers import SeamlessM4TForSpeechToSpeech
>>> model = SeamlessM4TForSpeechToSpeech.from_pretrained("facebook/hf-seamless-m4t-medium")
Or you can replace the text-to-text generation snippet with the model dedicated to the T2TT task, you only have to remove
generate_speech=False
.
>>> from transformers import SeamlessM4TForTextToText
>>> model = SeamlessM4TForTextToText.from_pretrained("facebook/hf-seamless-m4t-medium")
You have the possibility to change the speaker used for speech synthesis with the
spkr_id
argument. Some
spkr_id
works better than other for some languages!
3. Change the generation strategy
You can use different
generation strategies
for speech and text generation, e.g
.generate(input_ids=input_ids, text_num_beams=4, speech_do_sample=True)
which will successively perform beam-search decoding on the text model, and multinomial sampling on the speech model.
4. Generate speech and text at the same time
Use
return_intermediate_token_ids=True
with
SeamlessM4TModel
to return both speech and text !
Runs of facebook hf-seamless-m4t-medium on huggingface.co
88.1K
Total runs
0
24-hour runs
-10.7K
3-day runs
-26.2K
7-day runs
-34.6K
30-day runs
More Information About hf-seamless-m4t-medium huggingface.co Model
hf-seamless-m4t-medium huggingface.co is an AI model on huggingface.co that provides hf-seamless-m4t-medium's model effect (), which can be used instantly with this facebook hf-seamless-m4t-medium model. huggingface.co supports a free trial of the hf-seamless-m4t-medium model, and also provides paid use of the hf-seamless-m4t-medium. Support call hf-seamless-m4t-medium model through api, including Node.js, Python, http.
hf-seamless-m4t-medium huggingface.co is an online trial and call api platform, which integrates hf-seamless-m4t-medium's modeling effects, including api services, and provides a free online trial of hf-seamless-m4t-medium, you can try hf-seamless-m4t-medium online for free by clicking the link below.
facebook hf-seamless-m4t-medium online free url in huggingface.co:
hf-seamless-m4t-medium is an open source model from GitHub that offers a free installation service, and any user can find hf-seamless-m4t-medium on GitHub to install. At the same time, huggingface.co provides the effect of hf-seamless-m4t-medium install, users can directly use hf-seamless-m4t-medium installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
hf-seamless-m4t-medium install url in huggingface.co: