inference-caching-and-kv

SKILLFlusso di lavorocommunity
v0.0.0jpoindexterUnknownAggiornato 18 g faFonte →

Reference-grade guide to caching in LLM inference — provider prompt caching (Anthropic cache_control breakpoints, OpenAI automatic prefix caching, Gemini implicit/explicit), semantic caching, and KV-cache internals & management (PagedAttention/vLLM, RadixAttention/SGLang, eviction, quantized KV, mem

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
1Stelle del repo
1Client
1Formati
18 g faUltimo aggiornamento
Skill
Autorejpoindexter
Versione0.0.0
LicenzaUnknown
CategoriaFlusso di lavoro
Formatiskill.md
PromptNon pubblicato
Compatibilità
Claude✓ Supportato
Cursor
Copilot
ChatGPT
Gemini
Descrizione

Reference-grade guide to caching in LLM inference — provider prompt caching (Anthropic cache_control breakpoints, OpenAI automatic prefix caching, Gemini implicit/explicit), semantic caching, and KV-cache internals & management (PagedAttention/vLLM, RadixAttention/SGLang, eviction, quantized KV, memory pressure, multi-tenant safety). Concrete numbers, formulas, failure modes.

Parole chiave
skillclaude