| name | gemini-batch |
| version | 1 |
| description | This skill should be used when the user asks to "use Gemini Batch API", "process documents at scale", "submit a batch job", "upload files to Gemini", or needs large-scale LLM processing. |
| user-invocable | false |
Gemini Batch API Skill
Large-scale asynchronous document processing using Google's Gemini models.
When to Use
- Process thousands of documents with the same prompt
- Cost-effective bulk extraction (50% cheaper than synchronous API)
- Jobs that can tolerate 24-hour completion windows
IRON LAW: Use Examples First, Never Guess API
READ EXAMPLES BEFORE WRITING ANY CODE. NO EXCEPTIONS.
The Rule
User asks for batch API work
โ
MANDATORY: Read examples/batch_processor.py or examples/icon_batch_vision.py
โ
Copy the pattern exactly
โ
DO NOT guess parameter names
DO NOT try wrapper types
DO NOT improvise API calls
Why This Matters
The Batch API has non-obvious requirements that will fail silently:
- Metadata must be flat primitives - Nested objects cause cryptic errors
dest is a config field, not a kwarg - Pass via config={"dest": "gs://..."}. Older SDKs accepted dest= directly; newer ones raise TypeError.
- Config is plain dict - Not a wrapper type
- Examples are authoritative - Working code beats assumptions
Rationale: Previous agents wasted hours debugging API errors that the examples would have prevented. The patterns in examples/ are battle-tested production code.
Red Flags
- About to pass
dest= as a kwarg โ STOP. That works on older SDKs only; the current SDK puts dest inside config={}. Read the examples.
- About to instantiate a
CreateBatchJobConfig object โ STOP. The config is a plain dict, not a wrapper type.
- About to nest metadata like a normal API โ STOP. Nested objects trigger BigQuery type errors; flatten the data.
- About to assume this works like other Google APIs โ STOP. This API is different; the examples are authoritative.
- About to improvise the JSONL format โ STOP. Copy the structure from the examples instead.
MANDATORY Checklist Before ANY Batch API Code
Enforcement: Writing batch API code without reading examples first violates this IRON LAW and will result in preventable errors.
Prerequisites
Install gcloud SDK
curl https://sdk.cloud.google.com | bash
Authentication Setup
gcloud auth login
gcloud auth application-default login
gcloud services enable aiplatform.googleapis.com
Why both auth methods?
gcloud auth login: For gsutil and gcloud CLI commands
gcloud auth application-default login: For google-generativeai Python library
- CRITICAL: Vertex AI requires ADC (step 2), not just API key
Create GCS Bucket
gsutil mb -l us-central1 gs://your-batch-bucket
gsutil ls -L -b gs://your-batch-bucket | grep "Location"
See references/gcs-setup.md for complete setup guide.
Quick Start
Standard Gemini API (API Key)
Uses the Gemini File API for input. Results returned via batch_job.dest.file_name.
from google import genai
client = genai.Client()
uploaded = client.files.upload(
file="requests.jsonl",
config={"mime_type": "application/jsonl"}
)
job = client.batches.create(
model="gemini-2.5-flash-lite",
src=uploaded.name,
config={"display_name": "my-batch-job"}
)
Vertex AI (Recommended for GCS workflows)
Uses GCS URIs directly. dest is a field of the config dict in the
current SDK (older SDKs accepted dest= as a kwarg โ that now raises
TypeError: Batches.create() got an unexpected keyword argument 'dest').
from google import genai
client = genai.Client(
vertexai=True,
project="your-project-id",
location="us-central1"
)
job = client.batches.create(
model="gemini-2.5-flash-lite",
src="gs://bucket/requests.jsonl",
config={
"display_name": "my-job",
"dest": "gs://bucket/outputs/",
},
)
Verify your SDK before changing: inspect.signature(client.batches.create).
If dest is in the kwargs, the kwarg form works; otherwise use config.
Key difference: Standard API uses File API (files/...), Vertex AI uses GCS (gs://...) with dest (now a config field).
Core Workflow
Standard API:
- Create JSONL request file with prompts
- Upload JSONL to File API via
client.files.upload()
- Submit batch job via
client.batches.create(src=uploaded.name)
- Monitor for completion โ use Monitor tool (jobs expire after 24 hours)
- Download results from
job.dest.file_name
Vertex AI:
- Upload files to GCS bucket (us-central1 region required)
- Create JSONL request file with document URIs and prompts
- Submit batch job via
client.batches.create(src=..., config={"dest": ...})
- Monitor for completion โ use Monitor tool (jobs expire after 24 hours)
- Download and parse results from GCS output URI
- Handle failures gracefully (partial failures are common)
Monitoring Batch Jobs with Monitor Tool
After submitting a batch job, use Monitor instead of sleep-polling in Python:
Monitor(
description="Gemini batch job progress",
persistent=true,
timeout_ms=3600000,
command="while true; do uv run python3 -c \"import google.genai as genai; j=genai.batches.get(name='$JOB_NAME'); print(f'{j.state} | {j.name}'); exit(0 if j.state in ('JOB_STATE_SUCCEEDED','JOB_STATE_FAILED','JOB_STATE_CANCELLED') else 1)\" && break; sleep 60; done"
)
This frees the conversation to continue working while the batch runs. You get notified when the job completes or fails โ no polling loop blocking your context.
Key Gotchas (API Structure)
Metadata must be flat primitives (no nested objects โ BigQuery-backed storage). dest is a config field, not a top-level kwarg in the current SDK (Vertex AI only). Config is a plain dict (not a wrapper type).
See the Red Flags in the first Iron Law section above โ the same gotchas apply here. The Key Gotchas table below summarizes all critical issues.
Key Gotchas
| Issue | Solution |
|---|
| Nested metadata fails | Use flat primitives or json.dumps() for complex data |
TypeError: unexpected keyword dest | Move dest inside config={} (Vertex AI; current SDK) |
| Mixing API patterns | Standard API: File API + no dest. Vertex AI: GCS + dest |
| Auth errors with Vertex AI | Run gcloud auth application-default login |
| vertexai=True requires ADC | API key is ignored with vertexai=True |
| Missing aiplatform API | Run gcloud services enable aiplatform.googleapis.com |
| Region mismatch (Vertex) | Use us-central1 bucket only |
| Wrong URI format (Vertex) | Use gs:// not https:// |
| Invalid JSONL | Use scripts/validate_jsonl.py |
| Image batch: inline data | Use fileData.fileUri for batch, not inline |
| Duplicate IDs | Hash file content + prompt for unique IDs |
| Large PDFs fail | Split at 50 pages / 50MB max |
| JSON parsing fails | Use robust extraction (see gotchas.md) |
| Output not found (Vertex) | Output URI is prefix, not file path |
uploadToFileSearchStore 503 for files >10KB | Use two-step: files.upload() then fileSearchStores.importFile() |
| File stuck in PROCESSING state | Poll files.get() until state is ACTIVE before importing |
| SDK Pager stops after first page | Use pager.hasNextPage() + pager.nextPage(), NOT for await |
|
Top 3 mistakes (bolded above):
- Using nested objects in metadata instead of flat primitives
- Mixing Standard API and Vertex AI patterns
- Passing
dest= as a kwarg instead of inside config={} (Vertex AI; current SDK)
See references/gotchas.md for detailed solutions (now with Gotchas 10-17).
Rate Limits
| Limit | Value |
|---|
| Max requests per JSONL | 10,000 |
| Max concurrent jobs | 10 |
| Max job size | 100MB |
| Job expiration | 24 hours |
Recommended Models
ALWAYS verify model IDs and pricing against the live docs
Never recall a model ID or a price from training data โ it is always stale. Fetch the .md.txt variants (LLM-optimized, far easier to parse than the HTML):
- Models:
https://ai.google.dev/gemini-api/docs/models.md.txt
- Pricing:
https://ai.google.dev/gemini-api/docs/pricing.md.txt
Real failures this prevents (encountered 2026-08-03):
- A plan specified
gemini-3-pro โ that ID does not exist.
- A config carried
gemini-3.1-flash-lite priced at {input 0.125, output 0.75}; the current lineup has gemini-3.5-flash-lite at {input 0.30, output 2.50} standard, {0.15, 1.25} batch.
- Version numbers do not stay in parity across lines. As of 2026-08-03 there is a
gemini-3.6-flash but no gemini-3.6-flash-lite; the newest Flash-Lite is gemini-3.5-flash-lite.
Batch API pricing is 50% of standard across models.
Model selection: default to Flash / Flash-Lite for extraction
For structured information extraction โ schema-constrained JSON pulled out of documents โ default to Flash or Flash-Lite. Reserve Pro for tasks needing genuine reasoning. Do not reach for Pro by default just because the task feels important.
Measured 2026-08-03 on the realpage project (SEC IPO prospectus extraction; ~16,700 input tokens/doc, ~350-600 output; identical prompt, identical 100 documents):
| model | finds the target provision | quote-verification | judge | cost/doc | 1,926-doc run |
|---|
gemini-3.1-pro-preview | 47% | 97.9% | 1.00 | $0.0187 | ~$36 |
gemini-3.6-flash | 62% | 95.2% | 0.85 | $0.0151 | ~$29 |
gemini-3.5-flash | 70% | 94.3% | 1.00 | $0.0149 | ~$29 |
Pro was the most conservative extractor, not the best one. It found the target provision in 47% of documents where Flash found 62-70% of the same documents. On an extraction task Pro's extra reasoning showed up as under-extraction โ the failure mode that silently biases a research dataset. Scored against held-out human hand-coding (20 rows the research team coded before the pipeline existed, never having seen a machine output), all four models were identical โ 85.0% exact agreement, 90% recall on real entitlements, 80% exact on those โ and they failed on the same three rows. So Pro's extra reasoning bought nothing measurable, while its conservatism cost 15-23 points of detection.
Two honest caveats. The human sample was small (20 rows, 11 companies), and identical failures on identical rows says the residual errors were structural โ a provision filed in an exhibit rather than the prospectus, a right held through a GP entity โ not model quality. And the detection gap itself stayed unresolved: on the documents where models disagreed there was no ground truth, so which model is right on that 23-point spread was still open. Do not read this table as "Flash is more accurate"; read it as "Pro was not measurably better, and was measurably quieter."
Cost savings from Pro โ Flash are smaller than people expect when the task is input-dominated. Here it was only ~20%, because Flash input is $0.75/1M against Pro's $1.00, while output โ where Flash is much cheaper โ was a rounding error at ~350 tokens. Flash-Lite is the only tier that cuts input price materially ($0.15/1M, ~85% saving). Work out whether the job is input- or output-dominated before assuming a Flash switch saves real money: compute mean_input_tokens * input_price vs mean_output_tokens * output_price from a Stage 2 sample (see references/scale-up-testing.md).
| Model | Use Case | Cost | Location | Thinking default |
|---|
gemini-2.5-flash-lite | Most batch jobs | Lowest | us-central1 | OFF |
gemini-2.5-flash | Complex extraction | Medium | us-central1 | OFF |
gemini-2.5-pro | Highest accuracy | Highest | us-central1 | ON (cannot disable) |
gemini-3-flash-preview | New gen, larger context | 5ร flash-lite | global | HIGH (set MINIMAL!) |
gemini-3.1-flash-lite-preview | Cheapest gen-3 | ~2ร 2.5 flash-lite | global | HIGH (set MINIMAL!) |
gemini-embedding-001 | Default for text-only (short titles, classification, retrieval over text) | Low | Standard API | n/a |
gemini-embedding-2 | Multimodal (text+image) inputs | Low | Standard API | n/a |
text-embedding-005 | Need Vertex Batch console visibility (legacy) | Low | us-central1 | n/a |
Critical for Gemini 3.x: Always pin thinkingConfig: {thinkingLevel: ...} in generationConfig or batch responses will silently fail with MAX_TOKENS and empty content. The level is not the same across tiers: Flash and Flash-Lite accept MINIMAL, but Pro rejects it ("Thinking level MINIMAL is not supported for this model", verified 2026-08-03) and needs LOW. Use a helper that picks the level per model โ a single hardcoded constant breaks when you switch tiers. See references/gotchas.md Gotcha 17.
Critical for embedding batches: Embedding work has its own rules and failure modes โ use file-based JSONL with per-row key on the Standard API; never inlined_requests (scrambles order at scale). Default to gemini-embedding-001 for text-only tasks. See references/embeddings.md and examples/embeddings_batch.py.
Additional Resources
References
references/embeddings.md - NEW: Dedicated reference for embedding batches (model choice, file-based + keyed pattern, sentinel verification)
references/gcs-setup.md - Complete GCS and Vertex AI setup guide
references/gotchas.md - 17 critical production gotchas (Gemini 3.x thinking_level per tier, location='global'; embedding gotcha now lives in embeddings.md)
references/best-practices.md - Idempotent IDs, state tracking, validation
references/scale-up-testing.md - Incremental scale-up testing (LangExtract prototyping, LLM-as-judge, Vertex AI batch, gate design, input- vs output-dominated cost)
references/troubleshooting.md - Common errors and debugging
references/vertex-ai.md - Enterprise alternative with comparison
references/cli-reference.md - gsutil and gcloud commands
references/files-api.md - Files API: upload, poll-until-ACTIVE, 48h expiry, size limits
references/file-search.md - File Search (managed RAG): store creation, metadata filtering, grounding metadata
references/structured-output.md - responseJsonSchema / responseSchema: the supported schema subset, enums
Examples
examples/icon_batch_vision.py - NEW: Batch vision analysis with Vertex AI
examples/batch_processor.py - Complete GeminiBatchProcessor class
examples/embeddings_batch.py - NEW: gemini-embedding-2 via client.batches.create_embeddings() (the only supported production path; Vertex Batch rejects this model)
examples/pipeline_template.py - Customizable pipeline template
Scripts
scripts/validate_jsonl.py - Validate JSONL before submission
scripts/test_single.py - Test single request before batch
External Documentation
Date Awareness
Gemini API evolves rapidly. For API features or model names with uncertainty, verify against current documentation.