Principal/Senior-level Spark playbook for batch and streaming architecture, data layout, shuffle control, cost-efficient compute, reliability, and operating large-scale data processing platforms.
Use when: designing Spark workloads, tuning jobs, optimizing cluster economics, or operating Spark platforms in production.
Instalação
Instalar com Codex ou Claude Copie este prompt, cole no Codex, Claude ou outro assistente e deixe que ele revise a página da skill e instale para você.
Principal/Senior-level Spark playbook for batch and streaming architecture, data layout, shuffle control, cost-efficient compute, reliability, and operating large-scale data processing platforms.
Use when: designing Spark workloads, tuning jobs, optimizing cluster economics, or operating Spark platforms in production.
Spark Mastery (Senior → Principal)
Operate
Start from data volume, compute economics, shuffle behavior, and correctness requirements.
Treat Spark as a distributed execution system with real storage, network, and scheduling tradeoffs.
Prefer explicit workload design over vague “big data” assumptions.
Optimize for predictable cost, reliability, and debuggable pipelines.
Default Standards
Data layout and partitioning must match workload reality.
Shuffle-heavy patterns require scrutiny.
Memory and executor tuning should follow evidence.
Streaming and batch semantics must be separated clearly.
Platform cost and job performance should be evaluated together.