OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous batching, MCP tool calling, and multimodal support. Native MLX backend, 400+ tok/s. Works with Claude Code.
Read this in other languages: English · Español · Français · 中文 Continuous batching + OpenAI + Anthropic APIs in one server. Native Apple Silicon inference. A vLLM-style inference server for Apple Silicon Macs. Unlike or used directly, it ships continuous batching, paged KV cache, prefix caching, and SSD-tiered cache, and exposes both OpenAI and Anthropic from a single process. Run LLMs, vision…
La verificación confirma la identidad del publicador (la propiedad del repo), no la seguridad del código. El análisis de seguridad cubre los CVE conocidos y los scripts de instalación sospechosos.
Forge no tiene ningún análisis registrado de esta entrada, así que no tiene ninguna observación de su superficie de herramientas. Eso es ausencia de pruebas, no prueba de que no exponga ninguna herramienta.
Read this in other languages: English · Español · Français · 中文 Continuous batching + OpenAI + Anthropic APIs in one server. Native Apple Silicon inference. A vLLM-style inference server for Apple Silicon Macs. Unlike or used directly, it ships continuous batching, paged KV cache, prefix caching, and SSD-tiered cache, and exposes both OpenAI and Anthropic from a single process. Run LLMs, vision models, audio, and embeddings on Metal with unified memory, no conversion step. Anthropic SDK /…