inference-caching-and-kv

SKILLWorkflowcommunauté
v0.0.0jpoindexterUnknownMis à jour il y a 19 jSource →

Reference-grade guide to caching in LLM inference — provider prompt caching (Anthropic cache_control breakpoints, OpenAI automatic prefix caching, Gemini implicit/explicit), semantic caching, and KV-cache internals & management (PagedAttention/vLLM, RadixAttention/SGLang, eviction, quantized KV, mem

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
1Étoiles du dépôt
1Clients
1Formats
il y a 19 jDernière mise à jour
Skill
Auteurjpoindexter
Version0.0.0
LicenceUnknown
CatégorieWorkflow
Formatsskill.md
PromptNon publié
Compatibilité
Claude✓ Pris en charge
Cursor
Copilot
ChatGPT
Gemini
À propos

Reference-grade guide to caching in LLM inference — provider prompt caching (Anthropic cache_control breakpoints, OpenAI automatic prefix caching, Gemini implicit/explicit), semantic caching, and KV-cache internals & management (PagedAttention/vLLM, RadixAttention/SGLang, eviction, quantized KV, memory pressure, multi-tenant safety). Concrete numbers, formulas, failure modes.

Mots-clés
skillclaude