laion / voiceclap-small

huggingface.co
Total runs: 131
24-hour runs: 0
7-day runs: 65
30-day runs: 103
Model's Last Updated: May 11 2026
feature-extraction

Introduction of voiceclap-small

Model Details of voiceclap-small

VoiceCLAP-Small

Voice-text contrastive (CLAP-style) embedding model trained on dense vocal-style captions for the VoiceNet suite.

VoiceCLAP-Small is the smaller of the two voice-text contrastive anchors released with VoiceNet. It is a dual-tower model: a BUD-E-Whisper_V1.1 audio encoder paired with sentence-transformers/all-MiniLM-L6-v2 on the text side, joined by an MLP projection on each side and trained with the SigLIP sigmoid contrastive loss.

Architecture dual-tower CLAP (BUD-E-Whisper-Small + MiniLM-L6-v2)
Audio encoder Whisper-style: 12 layers × 768 dim × 12 heads, 80-mel input @ 16 kHz
Text encoder BERT/MiniLM, 6 layers × 384 dim, mean-pooled
Joint embedding 768-d, L2-normalised
Loss SigLIP (sigmoid contrastive)
Total parameters ~110 M
Epochs 1
Training data

Trained for 1 epoch on the open mixture (9 datasets) used in the VoiceNet paper:

  • emolia-balanced-5M-subset (annotated subset of Emilia )
  • laions_got_talent_clean_with_captions
  • majestrino-data
  • synthetic_vocal_bursts
  • improved_synthetic_vocal_bursts
  • ears
  • expresso
  • voxceleb1
  • voxceleb2

All clips are captioned with MOSS-Audio-8B-Thinking -derived dense vocal-style captions covering emotions, talking-style attributes, and demographics.

Standalone load example

Only transformers and torchaudio are required (both on PyPI).

import torch, torchaudio
from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained("VoiceNet/voiceclap-small", trust_remote_code=True).eval()
tok   = AutoTokenizer.from_pretrained("VoiceNet/voiceclap-small")

# Audio: any-length 16 kHz waveform, mono
wav, sr = torchaudio.load("clip.wav")
if sr != 16000:
    wav = torchaudio.functional.resample(wav, sr, 16000)
wav = wav.mean(0)                                         # (T,)
audio_emb = model.encode_waveform(wav)                    # (1, 768), L2-normed

# Text: short caption(s)
enc      = tok(["a calm and steady voice"], padding=True, return_tensors="pt")
text_emb = model.encode_text(enc.input_ids, enc.attention_mask)

# Cosine similarity (embeddings already L2-normalised)
print((audio_emb @ text_emb.T).item())

encode_waveform accepts clips up to 30 s; longer clips should be chunked or truncated before being passed in. Embeddings are 768-d and unit-norm, so a @ t.T is the cosine similarity used in zero-shot retrieval.

Citation

If you use this model, please cite the VoiceNet paper.

Runs of laion voiceclap-small on huggingface.co

131
Total runs
0
24-hour runs
0
3-day runs
65
7-day runs
103
30-day runs

More Information About voiceclap-small huggingface.co Model

More voiceclap-small license Visit here:

https://choosealicense.com/licenses/cc-by-4.0

voiceclap-small huggingface.co

voiceclap-small huggingface.co is an AI model on huggingface.co that provides voiceclap-small's model effect (), which can be used instantly with this laion voiceclap-small model. huggingface.co supports a free trial of the voiceclap-small model, and also provides paid use of the voiceclap-small. Support call voiceclap-small model through api, including Node.js, Python, http.

voiceclap-small huggingface.co Url

https://huggingface.co/laion/voiceclap-small

laion voiceclap-small online free

voiceclap-small huggingface.co is an online trial and call api platform, which integrates voiceclap-small's modeling effects, including api services, and provides a free online trial of voiceclap-small, you can try voiceclap-small online for free by clicking the link below.

laion voiceclap-small online free url in huggingface.co:

https://huggingface.co/laion/voiceclap-small

voiceclap-small install

voiceclap-small is an open source model from GitHub that offers a free installation service, and any user can find voiceclap-small on GitHub to install. At the same time, huggingface.co provides the effect of voiceclap-small install, users can directly use voiceclap-small installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

voiceclap-small install url in huggingface.co:

https://huggingface.co/laion/voiceclap-small

Url of voiceclap-small

voiceclap-small huggingface.co Url

Provider of voiceclap-small huggingface.co

laion
ORGANIZATIONS

Other API from laion

huggingface.co

Total runs: 5.7K
Run Growth: 4.7K
Growth Rate: 82.98%
Updated:June 20 2025
huggingface.co

Total runs: 89
Run Growth: -56
Growth Rate: -62.92%
Updated:May 06 2026