waybarrios/vllm-mlx

MCPcommunity
waybarriosApache-2.0Updated 2mo agoGitHub

OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous batching, MCP tool calling, and multimodal support. Native MLX backend, 400+ tok/s. Works with Claude Code.

Read this in other languages: English · Español · Français · 中文 Continuous batching + OpenAI + Anthropic APIs in one server. Native Apple Silicon inference. A vLLM-style inference server for Apple Silicon Macs. Unlike or used directly, it ships continuous batching, paged KV cache, prefix caching, and SSD-tiered cache, and exposes both OpenAI and Anthropic from a single process. Run LLMs, vision…

Automatically indexed from public sources. Not yet verified by the developer on Forge.Claim this listing →
1kGitHub stars
188Forks
2mo agoLast update
Package
Authorwaybarrios
LicenseApache-2.0
Sourcegithub
Trust Status
B
60/100Good
Listed in Forge index+10/10
Publisher identity verified+0/30
Publisher: run `forge publish` from the repo to claim ownership
Domain verification+0/10
Not currently available for this listing type — the domain-verification check only runs for npm-backed packages today, so this row cannot be earned here yet regardless of what's hosted at the domain.
Prompt-injection scan · clean+30/30
Obfuscation / exfil scan · clean+20/20
Paste into Claude Code, Cursor, or any AI assistant to fix all gaps
StatusCommunity-indexed
PublisherUnverified
SignatureUnsigned
Domain
Provenance
DependenciesNot audited
Tool surface
Security scan✓ CleanvHEAD · 2mo ago
EvalsNone
IndexedMay 24, 2026

Verification confirms publisher identity (repo ownership), not code safety. The security scan covers known CVEs and suspicious install scripts — it cannot prove the absence of malicious code.

About

Read this in other languages: English · Español · Français · 中文 Continuous batching + OpenAI + Anthropic APIs in one server. Native Apple Silicon inference. A vLLM-style inference server for Apple Silicon Macs. Unlike or used directly, it ships continuous batching, paged KV cache, prefix caching, and SSD-tiered cache, and exposes both OpenAI and Anthropic from a single process. Run LLMs, vision models, audio, and embeddings on Metal with unified memory, no conversion step. Anthropic SDK /…

Keywords
anthropicapple-siliconaudio-processingclaude-codecomputer-visionimage-understandinginferencellmmachine-learningmacosmllmmlxmultimodal-aispeech-to-textstttext-to-speechttsvideo-understandingvision-language-modelvllm