Principal/Senior-level Spark playbook for batch and streaming architecture, data layout, shuffle control, cost-efficient compute, reliability, and operating large-scale data processing platforms.
Use when: designing Spark workloads, tuning jobs, optimizing cluster economics, or operating Spark platforms in production.
インストール
Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。
Principal/Senior-level Spark playbook for batch and streaming architecture, data layout, shuffle control, cost-efficient compute, reliability, and operating large-scale data processing platforms.
Use when: designing Spark workloads, tuning jobs, optimizing cluster economics, or operating Spark platforms in production.
Spark Mastery (Senior → Principal)
Operate
Start from data volume, compute economics, shuffle behavior, and correctness requirements.
Treat Spark as a distributed execution system with real storage, network, and scheduling tradeoffs.
Prefer explicit workload design over vague “big data” assumptions.
Optimize for predictable cost, reliability, and debuggable pipelines.
Default Standards
Data layout and partitioning must match workload reality.
Shuffle-heavy patterns require scrutiny.
Memory and executor tuning should follow evidence.
Streaming and batch semantics must be separated clearly.
Platform cost and job performance should be evaluated together.