inference-performance

SKILLFlujo de trabajocomunidad
v0.0.0jpoindexterUnknownActualizado hace 19 dFuente →

Use when serving or optimizing LLM inference in production — diagnosing or improving TTFT/TPOT/throughput, choosing batching strategy, sizing GPUs, picking vLLM/TensorRT-LLM, or debugging low GPU utilization, TTFT spikes, and OOM. Covers prefill vs decode, the roofline, continuous batching, PagedAtt

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
1Estrellas del repo
1Clientes
1Formatos
hace 19 dÚltima actualización
Skill
Autorjpoindexter
Versión0.0.0
LicenciaUnknown
CategoríaFlujo de trabajo
Formatosskill.md
PromptNo publicado
Compatibilidad
Claude✓ Compatible
Cursor
Copilot
ChatGPT
Gemini
Acerca de

Use when serving or optimizing LLM inference in production — diagnosing or improving TTFT/TPOT/throughput, choosing batching strategy, sizing GPUs, picking vLLM/TensorRT-LLM, or debugging low GPU utilization, TTFT spikes, and OOM. Covers prefill vs decode, the roofline, continuous batching, PagedAttention, chunked prefill, disaggregation, and FlashAttention.

Palabras clave
skillclaude