llm-inference-scaling

SKILLFlusso di lavorocommunity
v0.0.0BagelHoleMITAggiornato 3 mesi faFonte →

Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling. Handle traffic spikes, implement queue-based scaling, and optimize cost with spot instances for AI workloads.

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
47Stelle del repo
1Client
1Formati
3 mesi faUltimo aggiornamento
Skill
AutoreBagelHole
Versione0.0.0
LicenzaMIT
CategoriaFlusso di lavoro
Formatiskill.md
PromptNon pubblicato
Compatibilità
Claude✓ Supportato
Cursor
Copilot
ChatGPT
Gemini
Descrizione

Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling. Handle traffic spikes, implement queue-based scaling, and optimize cost with spot instances for AI workloads.

Parole chiave
skillclaude