This skill provides a workflow to design and implement a governed, secure
pipeline for agentic analytics solution across structured and unstructured data
that's distributed across Google Cloud, on-premises systems, and other cloud
providers.
-
Request the user to describe the functional requirements (business
processes, activities, and use cases) of their workload. Ask the user the
following questions, one question at a time:
- What are your primary inventory data sources? Are they unstructured
(e.g., PDF flavor recipes, invoices) or structured (e.g., historical
sales in Iceberg)?
- Where are these sources hosted? Are they split across AWS S3, Azure
Blob, Google Cloud Storage, or databases like AlloyDB?
- How do you manage and federate metadata across your data
sources within Google Cloud and in external locations (such as other
cloud providers)?
- What are your analytical and computational requirements to join, clean,
and run forecast models over large-scale distributed data?
- What types of natural language prompts do your data scientists or
operational agents expect to execute in their agentic IDE (VS Code or
Antigravity IDE)?
-
Request the user to describe the non-functional requirements of their
workload.
The following are examples of questions you can ask to gather non-functional
requirements:
- Security, privacy, and compliance: What data privacy rules,
regulatory compliance (e.g., GDPR, HIPAA), or data governance
requirements must the system adhere to?
- Reliability: What are your uptime, high-availability,
fault-tolerance, and disaster recovery objectives (RTO/RPO)?
- Performance: What target query latencies and SLA expectations does
your workload require?
- Operations: What operational monitoring metrics do your data
scientists and engineers need?
- Cost & Sustainability: Do you have specific budget constraints and
data egress/transfer cost requirements?
-
Ask the user whether the workload currently runs on other cloud providers or
on-premises.
- If the user answers "yes", then ask the user to describe the
architecture of the current deployment.
- If the user answer "no", then proceed to the next step.
-
Request the user to describe dependencies, if any, on other workloads,
products, or tools. The following are examples of questions that you can
ask to get information about the dependencies:
- Do you have any upstream or downstream dependencies on external systems
(e.g., identity providers, data curation platforms, CI/CD pipelines, or
active data catalogs)?
- Are there any requirements for your general data-engineering software
delivery lifecycle (e.g., version control, testing, data quality
assurance)? Provide the path to a directory or examples of these
artifacts.
-
Review the input that the user has provided so far, and check whether there
are any ambiguities or contradictions.
If you identify any ambiguities or contradictions in the requirements that
the user has provided (e.g., zero-copy vs copying data to a repository), then
do the following for each ambiguity or contradiction that you identify:
- Describe the ambiguity or contradiction (e.g., explain why copying data
contradicts the zero-copy requirement and also incurs data-transfer
costs).
- Ask the user how they wish to resolve the ambiguity or contradiction.
- If the user delegates the choice to you (e.g., the user replies with
"do what you think is best" or "you decide"), then provide a clear
suggestion to resolve the ambiguity or contradiction (e.g., suggest
prioritizing zero-copy remote queries), explain your reasoning
(e.g., to eliminate multi-cloud fees and data duplication), and ask
the user to approve your suggestion.
Critical: Until all the ambiguities and contradictions that you identify
are resolved according to the preceding guidance, you must NOT recommend or
generate any architecture design, technical decomposition, or Google Cloud
product recommendations.
-
Important: DON'T start this step if there are unresolved contradictions
or ambiguities from Step 5.
Generate a technical decomposition of the components of the workload.
- The technical decomposition must break down the solution into logical
components.
- The decomposition MUST address role-based security and credentials
within the relevant layers.
- The decomposition MUST be organized under the following four layers,
which represent a standard architectural pattern for agentic analytics
solutions, flowing from user interaction through data context and
governance to core data processing:
- User-interaction layer (IDE): e.g., agentic development
environment.
- Grounding and trusted data: e.g., foundation model, MCP servers,
and data warehouse in the cloud.
- Metadata curation: e.g., metadata scanning.
- Data processing and analytics: e.g., analytics workflows, Spark
data processing, and external data stores.
-
Request the user to approve the generated technical decomposition.
-
If the user requests changes, then generate an updated technical
decomposition.
-
Repeat steps 5 through 8 until the user approves the generated technical
decomposition.
-
After the user approves the technical decomposition, proceed to Phase 2.
Important: Don't proceed to the next phase until the user approves the
generated technical decomposition of the workload.
For each task in this phase, to ensure that the generated content aligns with
the latest and official Google Cloud guidance, ground the generated content by
using the following resources: