OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous batching, MCP tool calling, and multimodal support. Native MLX backend, 400+ tok/s. Works with Claude Code.
Read this in other languages: English · Español · Français · 中文 Continuous batching + OpenAI + Anthropic APIs in one server. Native Apple Silicon inference. A vLLM-style inference server for Apple Silicon Macs. Unlike or used directly, it ships continuous batching, paged KV cache, prefix caching, and SSD-tiered cache, and exposes both OpenAI and Anthropic from a single process. Run LLMs, vision…
Verification confirms publisher identity (repo ownership), not code safety. The security scan covers known CVEs and suspicious install scripts.
The repository archive could not be read at scan time (The operation was aborted due to timeout), so no tool declarations could be extracted from it. Nothing here is a statement about what this exposes.
Read this in other languages: English · Español · Français · 中文 Continuous batching + OpenAI + Anthropic APIs in one server. Native Apple Silicon inference. A vLLM-style inference server for Apple Silicon Macs. Unlike or used directly, it ships continuous batching, paged KV cache, prefix caching, and SSD-tiered cache, and exposes both OpenAI and Anthropic from a single process. Run LLMs, vision models, audio, and embeddings on Metal with unified memory, no conversion step. Anthropic SDK /…
This entry publishes no npm package, so Forge has no dependency tree for it. That is a gap in coverage — not a statement that it has no dependencies.