NexaAI / OmniAudio-2.6B

huggingface.co
Total runs: 742
24-hour runs: 0
7-day runs: 54
30-day runs: 54
Model's Last Updated: December 14 2024
audio-text-to-text

Introduction of OmniAudio-2.6B

Model Details of OmniAudio-2.6B

OmniAudio-2.6B

Example

OmniAudio is the world's fastest and most efficient audio-language model for on-device deployment - a 2.6B-parameter multimodal model that processes both text and audio inputs. It integrates three components: Gemma-2-2b, Whisper turbo, and a custom projector module, enabling secure, responsive audio-text processing directly on edge devices.

Unlike traditional approaches that chain ASR and LLM models together, OmniAudio-2.6B unifies both capabilities in a single efficient architecture for minimal latency and resource overhead.

Quick Links
  1. Interactive Demo in our HuggingFace Space
  2. Quickstart for local setup
  3. Learn more in our Blogs

Feedback: Send questions or suggestions about the model in our Discord

Demo

Performance Benchmarks on Consumer Hardware

On a 2024 Mac Mini M4 Pro, Qwen2-Audio-7B-Instruct running on 🤗 Transformers achieves an average decoding speed of 6.38 tokens/second, while Omni-Audio-2.6B through Nexa SDK reaches 35.23 tokens/second in FP16 GGUF version and 66 tokens/second in Q4_K_M quantized GGUF version - delivering 5.5x to 10.3x faster performance on consumer hardware.

Use Cases
  • Voice QA without Internet : Process offline voice queries like "I am at camping, how do I start a fire without fire starter?" OmniAudio provides practical guidance even without network connectivity.
  • Voice-in Conversation : Have conversations about personal experiences. When you say "I am having a rough day at work," OmniAudio engages in supportive talk and active listening.
  • Creative Content Generation : Transform voice prompts into creative pieces. Ask "Write a haiku about autumn leaves" and receive poetic responses inspired by your voice input.
  • Recording Summary : Simply ask "Can you summarize this meeting note?" to convert lengthy recordings into concise, actionable summaries.
  • Voice Tone Modification : Transform casual voice memos into professional communications. When you request "Can you make this voice memo more professional?" OmniAudio adjusts the tone while preserving the core message.
How to Use On Device

Step 1: Install Nexa-SDK (local on-device inference framework)

🚀 Install Nexa-SDK

Nexa-SDK is a open-sourced, local on-device inference framework, supporting text generation, image generation, vision-language models (VLM), audio-language models, speech-to-text (ASR), and text-to-speech (TTS) capabilities. Installable via Python Package or Executable Installer.

Step 2: Then run the following code in your terminal

nexa run omniaudio -st

💻 OmniAudio-2.6B q4_K_M version requires 1.30GB RAM and 1.60GB storage space.

Training

We developed OmniAudio through a three-stage training pipeline:

  • Pretraining: The initial stage focuses on core audio-text alignment using MLS English 10k transcription dataset. We introduced a special <|transcribe|> token to enable the model to distinguish between transcription and completion tasks, ensuring consistent performance across use cases.
  • Supervised Fine-tuning (SFT): We enhance the model's conversation capabilities using synthetic datasets derived from MLS English 10k transcription. This stage leverages a proprietary model to generate contextually appropriate responses, creating rich audio-text pairs for effective dialogue understanding.
  • Direct Preference Optimization (DPO): The final stage refines model quality using GPT-4o API as a reference. The process identifies and corrects inaccurate responses while maintaining semantic alignment. We additionally leverage Gemma2's text responses as a gold standard to ensure consistent quality across both audio and text inputs.
What's Next for OmniAudio?

OmniAudio is in active development and we are working to advance its capabilities:

  • Building direct audio generation for two-way voice communication
  • Implementing function calling support via Octopus_v2 integration

In the long term, we aim to establish OmniAudio as a comprehensive solution for edge-based audio-language processing.

Join Community

Discord | X(Twitter)

Runs of NexaAI OmniAudio-2.6B on huggingface.co

742
Total runs
0
24-hour runs
10
3-day runs
54
7-day runs
54
30-day runs

More Information About OmniAudio-2.6B huggingface.co Model

More OmniAudio-2.6B license Visit here:

https://choosealicense.com/licenses/apache-2.0

OmniAudio-2.6B huggingface.co

OmniAudio-2.6B huggingface.co is an AI model on huggingface.co that provides OmniAudio-2.6B's model effect (), which can be used instantly with this NexaAI OmniAudio-2.6B model. huggingface.co supports a free trial of the OmniAudio-2.6B model, and also provides paid use of the OmniAudio-2.6B. Support call OmniAudio-2.6B model through api, including Node.js, Python, http.

OmniAudio-2.6B huggingface.co Url

https://huggingface.co/NexaAI/OmniAudio-2.6B

NexaAI OmniAudio-2.6B online free

OmniAudio-2.6B huggingface.co is an online trial and call api platform, which integrates OmniAudio-2.6B's modeling effects, including api services, and provides a free online trial of OmniAudio-2.6B, you can try OmniAudio-2.6B online for free by clicking the link below.

NexaAI OmniAudio-2.6B online free url in huggingface.co:

https://huggingface.co/NexaAI/OmniAudio-2.6B

OmniAudio-2.6B install

OmniAudio-2.6B is an open source model from GitHub that offers a free installation service, and any user can find OmniAudio-2.6B on GitHub to install. At the same time, huggingface.co provides the effect of OmniAudio-2.6B install, users can directly use OmniAudio-2.6B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

OmniAudio-2.6B install url in huggingface.co:

https://huggingface.co/NexaAI/OmniAudio-2.6B

Url of OmniAudio-2.6B

OmniAudio-2.6B huggingface.co Url

Provider of OmniAudio-2.6B huggingface.co

NexaAI
ORGANIZATIONS

Other API from NexaAI

huggingface.co

Total runs: 1.3K
Run Growth: -484
Growth Rate: -38.05%
Updated:August 09 2025
huggingface.co

Total runs: 793
Run Growth: -109
Growth Rate: -13.75%
Updated:August 21 2025
huggingface.co

Total runs: 634
Run Growth: -567
Growth Rate: -89.43%
Updated:November 14 2025
huggingface.co

Total runs: 556
Run Growth: -1
Growth Rate: -0.18%
Updated:May 21 2024
huggingface.co

Total runs: 486
Run Growth: 0
Growth Rate: 0.00%
Updated:July 22 2025
huggingface.co

Total runs: 258
Run Growth: -17
Growth Rate: -6.59%
Updated:September 28 2025
huggingface.co

Total runs: 155
Run Growth: 17
Growth Rate: 10.97%
Updated:November 07 2025
huggingface.co

Total runs: 74
Run Growth: 4
Growth Rate: 5.41%
Updated:July 22 2025
huggingface.co

Total runs: 58
Run Growth: 40
Growth Rate: 68.97%
Updated:December 03 2025
huggingface.co

Total runs: 54
Run Growth: -1
Growth Rate: -1.85%
Updated:May 05 2024
huggingface.co

Total runs: 52
Run Growth: 4
Growth Rate: 7.69%
Updated:January 10 2026
huggingface.co

Total runs: 37
Run Growth: 25
Growth Rate: 67.57%
Updated:January 13 2026
huggingface.co

Total runs: 23
Run Growth: 19
Growth Rate: 86.36%
Updated:November 07 2025
huggingface.co

Total runs: 16
Run Growth: -2
Growth Rate: -12.50%
Updated:September 04 2024
huggingface.co

Total runs: 14
Run Growth: 10
Growth Rate: 71.43%
Updated:October 28 2025
huggingface.co

Total runs: 13
Run Growth: 6
Growth Rate: 50.00%
Updated:December 17 2025
huggingface.co

Total runs: 10
Run Growth: 2
Growth Rate: 20.00%
Updated:January 13 2026
huggingface.co

Total runs: 5
Run Growth: 0
Growth Rate: 0.00%
Updated:December 16 2025