Skip to main content

attention-variants-from-papers

Implement paper-defined attention variants by extracting mechanism invariants, mapping tensor shapes, preserving module interfaces, validating numerical behavior, and integrating the custom attention into an existing transformer stack.

설치로 이동

소스 정보

저장소
benchflow-ai/skillsbench
최근 소스 활동
2026년 5월 30일 04:55
감지된 SKILL.md 언어
영어
스타
1,816
포크
368

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.

파일 탐색기
6 개 파일

SKILL.md 표시 중

SKILL.md
소스 지침 · 읽기 전용 미리보기
name
attention-variants-from-papers
description
Implement paper-defined attention variants by extracting mechanism invariants, mapping tensor shapes, preserving module interfaces, validating numerical behavior, and integrating the custom attention into an existing transformer stack.
# Attention Variants from Papers Use this skill when a paper changes how attention scores, branches, normalization, or head sharing work, but the module still needs to behave like a drop-in transformer attention block. ## Workflow 1. Read the paper for invariants, not names. Use [paper-to-implementation.md](references/paper-to-implementation.md) to extract the external contract, the changed computation, and the training-time constraints. 2. Build a shape ledger before coding. Use [shape-ledger.md](references/shape-ledger.md) to track projections, head grouping, branch count, and output width. 3. Choose the mechanism pattern. Use [mechanism-patterns.md](references/mechanism-patterns.md) for subtractive attention, branch mixing, learned gates, and extra normalization. 4. Preserve the module boundary. Keep the same input and output shape, mask semantics, positional encoding flow, and cache behavior unless the task explicitly changes them. 5. Validate in layers. Start with random-tensor smoke tests, then compare against a baseline attention path. Use [stability-and-validation.md](references/stability-and-validation.md). 6. Integrate into the stack last. Swap the new module into one transformer block, verify the residual path, then roll it through the full model. Use [transformer-integration.md](references/transformer-integration.md). ## Checklist - extract the paper's invariants before writing code - account for every reshape, branch, and repeat in a shape ledger - preserve output width at concatenation or output projection - apply masks and positional terms at the intended stage - confirm random smoke tests stay finite - compare unchanged behaviors against a baseline attention implementation ## Reference Map - [paper-to-implementation.md](references/paper-to-implementation.md): turn paper text into module invariants and coding decisions - [shape-ledger.md](references/shape-ledger.md): keep dimensions consistent while branch structure changes - [mechanism-patterns.md](references/mechanism-patterns.md): reusable patterns for nonstandard score composition and mixing - [stability-and-validation.md](references/stability-and-validation.md): numerical checks, smoke tests, and baseline comparisons - [transformer-integration.md](references/transformer-integration.md): wire the custom module into an existing transformer block
GitHub에서 보기