Skip to main content

bigquery-optimization

Provides workflows to optimize BigQuery environments (capacity planning, editions), storage assets (partitioning, clustering, storage lifecycles, billing models), and SQL queries. Use when optimizing cost, modeling Edition migrations, rightsizing reservations, evaluating logical vs. physical storage, designing table partitioning/clustering, generating table DDL, migrating unpartitioned tables, managing partition expiration, or optimizing individual SQL queries. Do not use for raw usage reporting (use bigquery-observability), query execution plan analysis, error troubleshooting, or diagnosing why a specific job was slow (use bigquery-troubleshooting).

来源信息

仓库
google/skills
最近来源活动
2026年9月29日 04:23
检测到的 SKILL.md 语言
英语
星标
20,469
分支
1,692

安装方式

默认使用会先检查来源的 Prompt;你也可以切换为直接命令,或下载本地副本。

检查来源文件

决定是否安装前,请先阅读 SKILL.md,以及 SkillsMP 当前展示的配套文件。

文件资源管理器
6 个文件

正在显示 SKILL.md

SKILL.md
来源说明 · 只读预览
name
bigquery-optimization
metadata
{"version":"1.0.0","category":"BigDataAndAnalytics"}
description
Provides workflows to optimize BigQuery environments (capacity planning, editions), storage assets (partitioning, clustering, storage lifecycles, billing models), and SQL queries. Use when optimizing cost, modeling Edition migrations, rightsizing reservations, evaluating logical vs. physical storage, designing table partitioning/clustering, generating table DDL, migrating unpartitioned tables, managing partition expiration, or optimizing individual SQL queries. Do not use for raw usage reporting (use bigquery-observability), query execution plan analysis, error troubleshooting, or diagnosing why a specific job was slow (use bigquery-troubleshooting).
# BigQuery Optimization Workflow ## Prerequisites & Environment Setup Before executing optimization analyses, evaluating editions, or applying DDL modifications: 1. **Google Cloud SDK**: Ensure the [Google Cloud SDK](https://cloud.google.com/sdk/docs/install) is installed and configured. 2. **Project Selection**: Set the active Google Cloud project: ```bash gcloud config set project {project_id} ``` 3. **API Enablement**: Ensure BigQuery and BigQuery Reservation APIs are enabled: ```bash gcloud services enable \ bigquery.googleapis.com bigqueryreservation.googleapis.com ``` 4. **Authentication**: Authenticate the environment: * CLI tools and `bq` commands: `gcloud auth login` * SDKs and automation: `gcloud auth application-default login` * Service accounts: Set `GOOGLE_APPLICATION_CREDENTIALS="/path/to/key.json"` 5. **Billing & IAM Roles**: * Verify an active Google Cloud Billing account is attached to `{project_id}`. * Ensure appropriate IAM roles: * `roles/bigquery.admin` or `roles/bigquery.resourceAdmin`: Reservation and capacity commitment management. * `roles/bigquery.dataEditor` or `roles/bigquery.admin`: Modifying table schemas, partitioning, clustering, and storage billing models. * `roles/bigquery.jobUser`: Running evaluation queries. 6. **Companion Skills Installation**: This skill is part of a 3-pillar operations suite (`bigquery-observability`, `bigquery-optimization`, `bigquery-troubleshooting`). If any companion skill is not yet installed in your environment, install the full suite: ```bash npx skills add google/skills --skill bigquery-observability --skill bigquery-optimization --skill bigquery-troubleshooting ``` *(If `bigquery-observability` is not installed, use the self-contained baseline formulas and query templates provided directly in the reference sections below).* ## Workflows Determine the optimization focus of the user's request and follow the relevant workflow: - **Telemetry & Observability Baseline:** For direct raw usage telemetry, `INFORMATION_SCHEMA` queries, and baseline metric calculations, consult [bigquery-observability](../bigquery-observability) (`bigquery_observability`). If the `bigquery-observability` companion skill is not available in the active environment, all optimization guidelines, DDL templates, and decision models across this skill and its reference guides are fully self-contained. - **Capacity & Editions Modeling:** Evaluate the cost-efficiency of migrating workloads from On-Demand to Editions, as well as rightsizing active Edition reservations, baseline commitments, and autoscaling caps. * *Instructions:* Read `references/capacity_planning_editions.md` to provide deep links to BigQuery's built-in recommendation UIs (e.g., Slot Estimator) and guide the user through UI navigation: 1. navigate to the Slot Estimator tab, 2. select 'On-Demand' as the source to analyze historical query volume, and 3. review the Cost-Optimized Recommendations and Slot Usage Chart. - **Table & Storage Optimization:** Optimize storage costs from a billing model, physical layout, and lifecycle perspective. * *Billing Architecture:* Read `references/storage_billing_models.md` for guidance on evaluating aggregate compression ratios (e.g. >2:1 threshold in US) to recommend Physical vs. Logical billing, noting that the break-even ratio depends on specific regional rates and custom enterprise contracts. When providing `TABLE_STORAGE` queries, always scope with `WHERE table_schema = '{dataset_id}'`, use the regional dataset view, and warn that 0 rows indicates a region mismatch or lack of native tables rather than zero billable usage. * *Partitioning & Clustering Strategy:* Read `references/table_partitioning_clustering.md` to generate production DDL templates (CREATE TABLE, CTAS migrations for unpartitioned tables, and modifying clustering specifications), enforce pruning with `require_partition_filter = true`, and manage partition limits (up to 10,000 partitions/table). * *Lifecycle Management:* Read `references/storage_lifecycle_management.md` to pinpoint inactive data and define precise Time-to-Live (TTL) partition expirations, dataset expirations, and Time Travel window reductions. - **SQL Optimization:** Optimize individual SQL queries to reduce slot-time and the amount of data read. * *Instructions:* Follow the instructions in `references/sql_optimization.md` to provide recommendations to the user on how to rewrite their SQL query to reduce slot-time and the amount of data read. ## Execution Guardrails - **Terminology & Cost Framing:** Never promise or guarantee "cost-reduction" or "reducing expenditure." Always frame recommendations using the terminology **"optimizing your bill"** or **"improving cost-efficiency."** - **Explicit Scope Framing & Region Resolution:** Always state the target `project_id` and `region` at the very top of your response so the user immediately knows the exact scope being evaluated. Follow this 3-tier resolution hierarchy: 1. *Explicit Region:* Use the region specified in the user's prompt (e.g., `europe-west1`). 2. *Contextual Region:* Resolve the region from the specific dataset or resource mentioned in the context. 3. *Unspecified Fallback:* Default to `us` / `region-us`, explicitly state that `us` was assumed as the default, and instruct the user to substitute their region if their resources reside elsewhere. *Region Formatting:* In Cloud Console deep links, use the region identifier directly (e.g., `region=us`, `region=europe-west1`). In SQL queries against `INFORMATION_SCHEMA`, use the regional dataset qualifier (e.g., `region-us`, `region-europe-west1`). - **Zero-Row Result Guard:** If querying `TABLE_STORAGE` with `WHERE table_schema = '{dataset_id}'` returns 0 rows, do not proceed with an empty or zero-usage evaluation. Treat this as an indicator that the dataset may reside in a different region or have no native tables; stop and prompt the user to confirm the dataset's regional location. - **Populate Concrete Parameters:** When generating URLs and SQL queries, always substitute known `project_id` and `region` values directly into the code and links. Never leave literal `{project_id}` or `{location}` placeholders for the user to manually edit. - **No Autonomous Purchasing or Financial Mutations:** Never provide the user with executable scripts (e.g., `gcloud` or `bq` shell commands like `bq update --storage_billing_model=...`) designed to autonomously purchase annual commitments, alter edition tier bindings, or mutate storage billing models. Always guide the user to execute commitment purchases, reservation changes, and storage billing model updates manually via the Cloud Console UI.
在 GitHub 查看