llm-inference-scaling

SKILLFlujo de trabajocomunidad
v0.0.0BagelHoleMITActualizado hace 3 mFuente →

Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling. Handle traffic spikes, implement queue-based scaling, and optimize cost with spot instances for AI workloads.

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
47Estrellas del repo
1Clientes
1Formatos
hace 3 mÚltima actualización
Skill
AutorBagelHole
Versión0.0.0
LicenciaMIT
CategoríaFlujo de trabajo
Formatosskill.md
PromptNo publicado
Compatibilidad
Claude✓ Compatible
Cursor
Copilot
ChatGPT
Gemini
Acerca de

Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling. Handle traffic spikes, implement queue-based scaling, and optimize cost with spot instances for AI workloads.

Palabras clave
skillclaude