llm-inference-scaling

SKILLWorkflowCommunity
v0.0.0BagelHoleMITAktualisiert vor 3 Mon.Quelle →

Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling. Handle traffic spikes, implement queue-based scaling, and optimize cost with spot instances for AI workloads.

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
47Repo-Sterne
1Clients
1Formate
vor 3 Mon.Letzte Aktualisierung
Skill
AutorBagelHole
Version0.0.0
LizenzMIT
KategorieWorkflow
Formateskill.md
PromptNicht veröffentlicht
Kompatibilität
Claude✓ Unterstützt
Cursor
Copilot
ChatGPT
Gemini
Über

Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling. Handle traffic spikes, implement queue-based scaling, and optimize cost with spot instances for AI workloads.

Schlagwörter
skillclaude