High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.

anthropic anthropic-api apple-silicon claude-code continuous-batching inference-server llm local-llm macos mcp mlx multimodal-ai openai openai-api openai-compatible speech-to-text text-to-speech tool-calling vision-language-model vllm
1 Open Issue Need Help Last updated: Sep 2, 2026

Open Issues Need Help

View All on GitHub

High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.

Python
#anthropic#anthropic-api#apple-silicon#claude-code#continuous-batching#inference-server#llm#local-llm#macos#mcp#mlx#multimodal-ai#openai#openai-api#openai-compatible#speech-to-text#text-to-speech#tool-calling#vision-language-model#vllm