OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous batching, MCP tool calling, and multimodal support. Native MLX backend, 400+ tok/s. Works with Claude Code.
Read this in other languages: English · Español · Français · 中文 Continuous batching + OpenAI + Anthropic APIs in one server. Native Apple Silicon inference. A vLLM-style inference server for Apple Silicon Macs. Unlike or used directly, it ships continuous batching, paged KV cache, prefix caching, and SSD-tiered cache, and exposes both OpenAI and Anthropic from a single process. Run LLMs, vision…
La vérification confirme l’identité de l’éditeur (la propriété du dépôt), pas la sûreté du code. L’analyse de sécurité couvre les CVE connues et les scripts d’installation suspects.
l’archive du dépôt n’a pas pu être lue au moment de l’analyse (The operation was aborted due to timeout), aucune déclaration d’outil n’a donc pu en être extraite. Rien ici n’est une affirmation sur ce que cette entrée expose.
Read this in other languages: English · Español · Français · 中文 Continuous batching + OpenAI + Anthropic APIs in one server. Native Apple Silicon inference. A vLLM-style inference server for Apple Silicon Macs. Unlike or used directly, it ships continuous batching, paged KV cache, prefix caching, and SSD-tiered cache, and exposes both OpenAI and Anthropic from a single process. Run LLMs, vision models, audio, and embeddings on Metal with unified memory, no conversion step. Anthropic SDK /…
Cette entrée ne publie aucun paquet npm : Forge n'a donc pas d'arbre de dépendances pour elle. C'est une lacune de couverture — pas une affirmation qu'elle n'a aucune dépendance.