model-serving

SKILLFlujo de trabajocomunidad
v0.0.0ancolemanMITActualizado hace 8 mFuente →

LLM and ML model deployment for inference. Use when serving models in production, building AI APIs, or optimizing inference. Covers vLLM (LLM serving), TensorRT-LLM (GPU optimization), Ollama (local), BentoML (ML deployment), Triton (multi-model), LangChain (orchestration), LlamaIndex (RAG), and str

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
508Estrellas del repo
1Clientes
1Formatos
hace 8 mÚltima actualización
Skill
Autorancoleman
Versión0.0.0
LicenciaMIT
CategoríaFlujo de trabajo
Formatosskill.md
PromptAbrir (ver la pestaña Prompt)
Compatibilidad
Claude✓ Compatible
Cursor
Copilot
ChatGPT
Gemini
Acerca de

LLM and ML model deployment for inference. Use when serving models in production, building AI APIs, or optimizing inference. Covers vLLM (LLM serving), TensorRT-LLM (GPU optimization), Ollama (local), BentoML (ML deployment), Triton (multi-model), LangChain (orchestration), LlamaIndex (RAG), and streaming patterns.

Palabras clave
skillclaude