









Doubleword is an AI inference platform for high-volume workloads, offering OpenAI-compatible access to open-weight language, vision, OCR, and embedding models. It supports real-time, async, and approximately 24-hour batch inference, allowing users to trade latency for substantially lower token costs. Doubleword is designed for agentic reasoning, evaluations, synthetic data generation, annotation, extraction, coding workflows, and large-scale data processing. Its inference stack includes optimization technologies such as FlashOffload, Speculative KV Coding, and Cloudburst to improve throughput, cache efficiency, and startup times.
Create a Doubleword API key, select a supported open-weight model, and use the OpenAI-compatible API by changing your existing API base URL to api.doubleword.ai. Choose real-time, async, or batch delivery based on your latency and cost requirements. Submit prompts, agent tasks, evaluation jobs, extraction workflows, or batch files, then monitor and retrieve results through the API or Doubleword's interface. Teams can also request migration assistance, prompt optimization, queue tuning, custom pricing, or dedicated deployments.

Real-time inference
Usage-based; full-price real-time rates
Lowest latency for interactive applications, chat, and live sessions.
Async inference
Usage-based; up to 50% off real-time rates
Results typically delivered in minutes to hours for background agents and high-throughput workloads.
Batch inference
Usage-based; up to 80% off real-time rates
Approximately 24-hour delivery for large jobs where minimizing cost is the priority.
Lowest listed batch rate
$20 per 1 billion input tokens
Qwen3-Embedding-8B batch pricing; output tokens are not charged for this embedding model.
Lowest listed generative batch rate
$90 per 1 billion input and 1 billion output tokens
GPT-OSS-20B batch pricing.
Custom enterprise pricing
Contact for pricing
Available for bulk discounts, large workloads, and dedicated deployments.




Social Listening