Skip to main content

aws-lakehouse

Stars11
Forks4
UpdatedJune 28, 2026 at 15:44

S3-based lakehouse on Apache Iceberg โ€” Amazon S3 Tables (managed Iceberg table buckets with auto compaction/snapshot maintenance), self-managed Iceberg on plain S3 (Glue/REST catalogs), querying with Athena (engine v3, MERGE/UPDATE/DELETE), and processing with Spark on EMR/Glue/EMR Serverless. Covers catalog choice, Iceberg V3 (deletion vectors, row lineage) and its Athena incompatibility, Spark engine/language speed (Scala vs PySpark vs Kotlin), native accelerators (Comet/Gluten/Photon), and S3 Tables cost pitfalls. Use when building or querying a data lake on S3, choosing between S3 Tables and self-managed Iceberg, picking a query engine, or tuning Spark performance. Triggers: S3 Tables, table bucket, s3tables, Apache Iceberg, Iceberg REST catalog, Glue Data Catalog, Athena Iceberg, MERGE INTO, time travel, deletion vectors, Iceberg V3, EMR Spark, Glue ETL, PySpark slow, Spark accelerator, Comet, Gluten, lakehouse, partition projection.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

SKILL.md
readonly