llm-inference-scaling

SKILLWorkflowcommunauté
v0.0.0BagelHoleMITMis à jour il y a 3 moisSource →

Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling. Handle traffic spikes, implement queue-based scaling, and optimize cost with spot instances for AI workloads.

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
47Étoiles du dépôt
1Clients
1Formats
il y a 3 moisDernière mise à jour
Skill
AuteurBagelHole
Version0.0.0
LicenceMIT
CatégorieWorkflow
Formatsskill.md
PromptNon publié
Compatibilité
Claude✓ Pris en charge
Cursor
Copilot
ChatGPT
Gemini
À propos

Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling. Handle traffic spikes, implement queue-based scaling, and optimize cost with spot instances for AI workloads.

Mots-clés
skillclaude