optimizing-attention-flash

SKILLFlusso di lavorocommunity
v0.0.0Orchestra-ResearchMITAggiornato 3 mesi faFonte →

Optimizes transformer attention with Flash Attention for 2-4x speedup and 10-20x memory reduction. Use when training/running transformers with long sequences (>512 tokens), encountering GPU memory issues with attention, or need faster inference. Supports PyTorch native SDPA, flash-attn library, H100

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
12kStelle del repo
1Client
1Formati
3 mesi faUltimo aggiornamento
Skill
AutoreOrchestra-Research
Versione0.0.0
LicenzaMIT
CategoriaFlusso di lavoro
Formatiskill.md
PromptApri (vedi la scheda Prompt)
Compatibilità
Claude✓ Supportato
Cursor—
Copilot—
ChatGPT—
Gemini—
Descrizione

Optimizes transformer attention with Flash Attention for 2-4x speedup and 10-20x memory reduction. Use when training/running transformers with long sequences (>512 tokens), encountering GPU memory issues with attention, or need faster inference. Supports PyTorch native SDPA, flash-attn library, H100 FP8, and sliding window attention.

Parole chiave
skillclaude

Nessuna copertura delle dipendenze

Questa voce non pubblica alcun pacchetto npm, quindi Forge non ha un albero delle dipendenze per essa. È una lacuna di copertura, non l'affermazione che non abbia dipendenze.