waybarrios/vllm-mlx

MCPcommunity
waybarriosApache-2.0Updated 20d agoGitHub

OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous batching, MCP tool calling, and multimodal support. Native MLX backend, 400+ tok/s. Works with Claude Code.

Read this in other languages: English · Español · Français · 中文 Continuous batching + OpenAI + Anthropic APIs in one server. Native Apple Silicon inference. A vLLM-style inference server for Apple Silicon Macs. Unlike or used directly, it ships continuous batching, paged KV cache, prefix caching, and SSD-tiered cache, and exposes both OpenAI and Anthropic from a single process. Run LLMs, vision…

Automatically indexed from public sources. Not yet verified by the developer on Forge.Claim this listing →
1kGitHub stars
188Forks
20d agoLast update
Package
Authorwaybarrios
LicenseApache-2.0
Sourcegithub
Trust Status
B
60/100Good
Listed in Forge index+10/10
Publisher identity verified+0/25
Publisher: run `forge publish` from the package repo to claim ownership
Ed25519 publish signature+0/10
Included automatically when the publisher runs `forge publish`
Domain verification+0/5
Publisher: host /.well-known/forge.json on the package homepage with { "publisher": "<github-login>" }
CVE scan · not run+0/30
Not yet scanned — package must be on npm
Static analysis · clean+20/20
npm provenance (Sigstore)+0/5
Publish from GitHub Actions with the --provenance flag
Paste into Claude Code, Cursor, or any AI assistant to fix all gaps
StatusCommunity-indexed
PublisherUnverified
SignatureUnsigned
Domain
Provenance
DependenciesNot audited
Tool surface
Security scan✓ CleanvHEAD · 19d ago
EvalsNone
IndexedMay 24, 2026

Verification confirms publisher identity (repo ownership), not code safety. The security scan covers known CVEs and suspicious install scripts — it cannot prove the absence of malicious code.

About

Read this in other languages: English · Español · Français · 中文 Continuous batching + OpenAI + Anthropic APIs in one server. Native Apple Silicon inference. A vLLM-style inference server for Apple Silicon Macs. Unlike or used directly, it ships continuous batching, paged KV cache, prefix caching, and SSD-tiered cache, and exposes both OpenAI and Anthropic from a single process. Run LLMs, vision models, audio, and embeddings on Metal with unified memory, no conversion step. Anthropic SDK /…

Keywords
anthropicapple-siliconaudio-processingclaude-codecomputer-visionimage-understandinginferencellmmachine-learningmacosmllmmlxmultimodal-aispeech-to-textstttext-to-speechttsvideo-understandingvision-language-modelvllm