1. SaladCloud: GPU Cloud Computing, AI Transcription API, Container Engine, Gateway Service, Virtual Kubelets, Distributed File Storage (Coming Soon), Object Storage (Coming Soon), Managed Databases (Coming Soon).
2. Pulse — Speech-to-Text by Smallest AI: Real-time speech-to-text with latency as low as 64 milliseconds, Transcription across 38+ languages, accents, and dialects, Speaker diarization and automated speaker labeling, Real-time sentiment, emotion recognition, and language identification, Speech-to-text models for both live and recorded audio, Voice agent orchestration with telephony and campaign support, Node.js and Python SDKs, Knowledge bases, post-call analytics, and web widgets, Enterprise security and compliance, including SOC 2, GDPR, HIPAA, and ISO 27001.
3. Gradium: Ultra-low-latency streaming text-to-speech with expressive voices, Accurate speech-to-text transcription with code-switching and custom vocabulary, Real-time speech-to-speech and live voice translation, Instant voice cloning from approximately 10 seconds of audio, High-fidelity Pro Voice Clones, Prompt-based voice design, Offline on-device text-to-speech running on CPU, Bidirectional WebSocket real-time API, Timestamp-accurate speech and pronunciation dictionaries, Semantic voice activity detection for conversational turn-taking, Scalable concurrency with stable latency under load, Cloud, dedicated, self-hosted, on-premises, and cloud marketplace deployment, Python and Rust SDKs, Integrations with LiveKit, Pipecat, Gradbot, and major agent frameworks, Telephony audio formats, Enterprise SLAs and zero data retention.
4. Speko: Multilingual benchmarking of speech and language models across 10 languages, Managed routing for speech-to-text, LLM, and text-to-speech models, Unified provider-neutral API for voice AI applications, Measured model selection and pre-response failover, Provider-direct streaming through the Speko Gateway, LiveKit and Pipecat integrations, BYOK credential support and optional Speko-managed routes, Public OpenAPI, AsyncAPI, and MCP access, Performance and cost comparison for voice models, Enterprise deployment reviews and dedicated support options.
5. Humalike: Turn-taking API for knowing when to speak, interrupt, or stay silent, Theory of Mind for tracking what people think, want, and feel, Social memory for remembering people across conversations, Norms and persona components for adapting to group tone and community behavior, Social signals and social observability for reading engagement and mood, One-shot Hermes integration.
6. KugelAudio: Ultra-low latency with a 39ms inference time to first audio for the turbo model, Grammar-aware normalization for natural reading of phone numbers, addresses, and edge cases, 100% European infrastructure with full GDPR compliance and data sovereignty, Support for voice cloning and word-level timestamps with IPA support, Seamless 2-line code integration with LiveKit, Pipecat, and Vapi adapters, On-premise deployment options for enterprise customers.