









oMLX is a native macOS inference server for running local AI models on Apple Silicon using Apple's MLX framework. It provides paged SSD KV caching to reduce time-to-first-token for coding agents, continuous batching for higher throughput, and OpenAI- and Anthropic-compatible APIs. oMLX supports LLM, vision-language, embedding, and reranker models, along with tool calling, MCP integration, multi-model serving, and a native menu bar application with a web dashboard.
Download the signed macOS DMG and drag oMLX to the Applications folder, then configure the model directory, start the server, and download or select an MLX-format model. Alternatively, install from source with Python 3.10+ using the provided Git commands and run omlx serve with a model directory. Connect Claude Code, OpenClaw, Cursor, or other compatible clients through the OpenAI or Anthropic API endpoints.



7.95%

Social Listening