Integrate Gemini API with correct current SDK (@google/genai v1.27+, NOT deprecated @google/generative-ai).
Supports text generation, multimodal (images/video/audio/PDFs), function calling, and thinking mode. 1M input tokens.
Use when: integrating Gemini API, implementing multimodal AI, using thinking mode for reasoning, function calling
with parallel execution, streaming responses, deploying to Cloudflare Workers, building chat, or troubleshooting
SDK deprecation, context window, model not found, function calling, or multimodal format errors.
Keywords: gemini api, @google/genai, gemini-2.5-pro, gemini-2.5-flash, gemini-2.5-flash-lite,
gemini-3-pro-preview, multimodal gemini, thinking mode, google ai, genai sdk, function calling gemini,
streaming gemini, gemini vision, gemini video, gemini audio, gemini pdf, system instructions,
multi-turn chat, DEPRECATED @google/generative-ai, gemini context window, gemini models 2025,
gemini 1m tokens, gemini tool use, parallel function calling, compositional function ca
Instalar com Codex ou Claude Copie este prompt, cole no Codex, Claude ou outro assistente e deixe que ele revise a página da skill e instale para você.
Um comando direto ignora o prompt de revisão. Verifique a origem antes de executá-lo.
Instruções da origem · Visualização somente leitura
name
google-gemini-api
description
Integrate Gemini API with correct current SDK (@google/genai v1.27+, NOT deprecated @google/generative-ai).
Supports text generation, multimodal (images/video/audio/PDFs), function calling, and thinking mode. 1M input tokens.
Use when: integrating Gemini API, implementing multimodal AI, using thinking mode for reasoning, function calling
with parallel execution, streaming responses, deploying to Cloudflare Workers, building chat, or troubleshooting
SDK deprecation, context window, model not found, function calling, or multimodal format errors.
Keywords: gemini api, @google/genai, gemini-2.5-pro, gemini-2.5-flash, gemini-2.5-flash-lite,
gemini-3-pro-preview, multimodal gemini, thinking mode, google ai, genai sdk, function calling gemini,
streaming gemini, gemini vision, gemini video, gemini audio, gemini pdf, system instructions,
multi-turn chat, DEPRECATED @google/generative-ai, gemini context window, gemini models 2025,
gemini 1m tokens, gemini tool use, parallel function calling, compositional function calling, gemini 3
license
MIT
Google Gemini API - Complete Guide
Version: Phase 2 Complete + Gemini 3 ✅
Package: @google/genai@1.27.0 (⚠️ NOT @google/generative-ai)
Last Updated: 2025-11-19 (Gemini 3 preview release)
⚠️ CRITICAL SDK MIGRATION WARNING
DEPRECATED SDK: @google/generative-ai (sunset November 30, 2025)
CURRENT SDK: @google/genai v1.27+
If you see code using @google/generative-ai, it's outdated!
This skill uses the correct current SDK and provides a complete migration guide.
const response = await ai.models.generateContentStream({
model: 'gemini-2.5-flash',
contents: 'Write a 200-word story about time travel'
});
for await (const chunk of response) {
process.stdout.write(chunk.text);
}
Streaming with Fetch (SSE Parsing)
const response = await fetch(
`https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash:streamGenerateContent`,
{
method: 'POST',
headers: {
'Content-Type': 'application/json',
'x-goog-api-key': env.GEMINI_API_KEY,
},
body: JSON.stringify({
contents: [{ parts: [{ text: 'Write a 200-word story about time travel' }] }]
}),
}
);
const reader = response.body.getReader();
const decoder = new TextDecoder();
let buffer = '';
while (true) {
const { done, value } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
const lines = buffer.split('\n');
buffer = lines.pop() || '';
for (const line of lines) {
if (line.trim() === '' || line.startsWith('data: [DONE]')) continue;
if (!line.startsWith('data: ')) continue;
try {
const data = JSON.parse(line.slice(6));
const text = data.candidates[0]?.content?.parts[0]?.text;
if (text) {
process.stdout.write(text);
}
} catch (e) {
// Skip invalid JSON
}
}
}
Key Points:
Use streamGenerateContent endpoint (not generateContent)
Gemini supports function calling (tool use) to connect models with external APIs and systems.
Basic Function Calling (SDK)
import { GoogleGenAI, FunctionCallingConfigMode } from '@google/genai';
const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY });
// Define function declarations
const getCurrentWeather = {
name: 'get_current_weather',
description: 'Get the current weather for a location',
parametersJsonSchema: {
type: 'object',
properties: {
location: {
type: 'string',
description: 'City name, e.g. San Francisco'
},
unit: {
type: 'string',
enum: ['celsius', 'fahrenheit']
}
},
required: ['location']
}
};
// Make request with tools
const response = await ai.models.generateContent({
model: 'gemini-2.5-flash',
contents: 'What\'s the weather in Tokyo?',
config: {
tools: [
{ functionDeclarations: [getCurrentWeather] }
]
}
});
// Check if model wants to call a function
const functionCall = response.candidates[0].content.parts[0].functionCall;
if (functionCall) {
console.log('Function to call:', functionCall.name);
console.log('Arguments:', functionCall.args);
// Execute the function (your implementation)
const weatherData = await fetchWeather(functionCall.args.location);
// Send function result back to model
const finalResponse = await ai.models.generateContent({
model: 'gemini-2.5-flash',
contents: [
'What\'s the weather in Tokyo?',
response.candidates[0].content, // Original assistant response with function call
{
parts: [
{
functionResponse: {
name: functionCall.name,
response: weatherData
}
}
]
}
],
config: {
tools: [
{ functionDeclarations: [getCurrentWeather] }
]
}
});
console.log(finalResponse.text);
}
Function Calling (Fetch)
const response = await fetch(
`https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash:generateContent`,
{
method: 'POST',
headers: {
'Content-Type': 'application/json',
'x-goog-api-key': env.GEMINI_API_KEY,
},
body: JSON.stringify({
contents: [
{ parts: [{ text: 'What\'s the weather in Tokyo?' }] }
],
tools: [
{
functionDeclarations: [
{
name: 'get_current_weather',
description: 'Get the current weather for a location',
parameters: {
type: 'object',
properties: {
location: {
type: 'string',
description: 'City name'
}
},
required: ['location']
}
}
]
}
]
}),
}
);
const data = await response.json();
const functionCall = data.candidates[0]?.content?.parts[0]?.functionCall;
if (functionCall) {
// Execute function and send result back (same flow as SDK)
}
Parallel Function Calling
Gemini can call multiple independent functions simultaneously:
const tools = [
{
functionDeclarations: [
{
name: 'get_weather',
description: 'Get weather for a location',
parametersJsonSchema: {
type: 'object',
properties: {
location: { type: 'string' }
},
required: ['location']
}
},
{
name: 'get_population',
description: 'Get population of a city',
parametersJsonSchema: {
type: 'object',
properties: {
city: { type: 'string' }
},
required: ['city']
}
}
]
}
];
const response = await ai.models.generateContent({
model: 'gemini-2.5-flash',
contents: 'What is the weather and population of Tokyo?',
config: { tools }
});
// Model may return MULTIPLE function calls in parallel
const functionCalls = response.candidates[0].content.parts.filter(
part => part.functionCall
);
console.log(`Model wants to call ${functionCalls.length} functions in parallel`);
Function Calling Modes
import { FunctionCallingConfigMode } from '@google/genai';
const response = await ai.models.generateContent({
model: 'gemini-2.5-flash',
contents: 'What\'s the weather?',
config: {
tools: [{ functionDeclarations: [getCurrentWeather] }],
toolConfig: {
functionCallingConfig: {
mode: FunctionCallingConfigMode.ANY, // Force function call
// mode: FunctionCallingConfigMode.AUTO, // Model decides (default)
// mode: FunctionCallingConfigMode.NONE, // Never call functions
allowedFunctionNames: ['get_current_weather'] // Optional: restrict to specific functions
}
}
}
});
Modes:
AUTO (default): Model decides whether to call functions
ANY: Force model to call at least one function
NONE: Disable function calling for this request
System Instructions
System instructions guide the model's behavior and set context. They are separate from the conversation messages.
SDK Approach
const response = await ai.models.generateContent({
model: 'gemini-2.5-flash',
systemInstruction: 'You are a helpful AI assistant that always responds in the style of a pirate. Use nautical terminology and end sentences with "arrr".',
contents: 'Explain what a database is'
});
console.log(response.text);
// Output: "Ahoy there! A database be like a treasure chest..."
Fetch Approach
const response = await fetch(
`https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash:generateContent`,
{
method: 'POST',
headers: {
'Content-Type': 'application/json',
'x-goog-api-key': env.GEMINI_API_KEY,
},
body: JSON.stringify({
systemInstruction: {
parts: [
{ text: 'You are a helpful AI assistant that always responds in the style of a pirate.' }
]
},
contents: [
{ parts: [{ text: 'Explain what a database is' }] }
]
}),
}
);
Key Points:
System instructions are NOT part of contents array
They are set once at the top level of the request
They persist for the entire conversation (when using multi-turn chat)
They don't count as user or model messages
Multi-turn Chat
For conversations with history, use the SDK's chat helpers or manually manage conversation state.
SDK Chat Helpers (Recommended)
const chat = await ai.models.createChat({
model: 'gemini-2.5-flash',
systemInstruction: 'You are a helpful coding assistant.',
history: [] // Start empty or with previous messages
});
// Send first message
const response1 = await chat.sendMessage('What is TypeScript?');
console.log('Assistant:', response1.text);
// Send follow-up (context is automatically maintained)
const response2 = await chat.sendMessage('How do I install it?');
console.log('Assistant:', response2.text);
// Get full chat history
const history = chat.getHistory();
console.log('Full conversation:', history);
For creative tasks: Use high temperature (0.7-1.5)
topP and topK both control randomness; use one or the other (not both)
Always set maxOutputTokens to prevent excessive generation
Context Caching
Context caching allows you to cache frequently used content (like system instructions, large documents, or video files) to reduce costs by up to 90% and improve latency.
How It Works
Create a cache with your repeated content
Reference the cache in subsequent requests
Save tokens - cached tokens cost significantly less
TTL management - caches expire after specified time
Benefits
Cost savings: Up to 90% reduction on cached tokens
Reduced latency: Faster responses by reusing processed content
Consistent context: Same large context across multiple requests
Cache Creation (SDK)
import { GoogleGenAI } from '@google/genai';
import fs from 'fs';
const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY });
// Create a cache for a large document
const documentText = fs.readFileSync('./large-document.txt', 'utf-8');
const cache = await ai.caches.create({
model: 'gemini-2.5-flash',
config: {
displayName: 'large-doc-cache', // Identifier for the cache
systemInstruction: 'You are an expert at analyzing legal documents.',
contents: documentText,
ttl: '3600s', // Cache for 1 hour
}
});
console.log('Cache created:', cache.name);
console.log('Expires at:', cache.expireTime);
// Generate content using the cache
const response = await ai.models.generateContent({
model: cache.name, // Use cache name as model
contents: 'Summarize the key points in the document'
});
console.log(response.text);
// Set specific expiration time (must be timezone-aware)
const in10Minutes = new Date(Date.now() + 10 * 60 * 1000);
await ai.caches.update({
name: cache.name,
config: {
expireTime: in10Minutes
}
});
List and Delete Caches (SDK)
// List all caches
const caches = await ai.caches.list();
for (const cache of caches) {
console.log(cache.name, cache.displayName);
}
// Delete a specific cache
await ai.caches.delete({ name: cache.name });
Caching with Video Files
import { GoogleGenAI } from '@google/genai';
import fs from 'fs';
const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY });
// Upload video file
const videoFile = await ai.files.upload({
file: fs.createReadStream('./video.mp4')
});
// Wait for processing
while (videoFile.state.name === 'PROCESSING') {
await new Promise(resolve => setTimeout(resolve, 2000));
videoFile = await ai.files.get({ name: videoFile.name });
}
// Create cache with video
const cache = await ai.caches.create({
model: 'gemini-2.5-flash',
config: {
displayName: 'video-analysis-cache',
systemInstruction: 'You are an expert video analyzer.',
contents: [videoFile],
ttl: '300s' // 5 minutes
}
});
// Use cache for multiple queries
const response1 = await ai.models.generateContent({
model: cache.name,
contents: 'What happens in the first minute?'
});
const response2 = await ai.models.generateContent({
model: cache.name,
contents: 'Describe the main characters'
});
Key Points
When to Use Caching:
Large system instructions used repeatedly
Long documents analyzed multiple times
Video/audio files queried with different prompts
Consistent context across conversation sessions
TTL Guidelines:
Short sessions: 300s (5 min) to 3600s (1 hour)
Long sessions: 3600s (1 hour) to 86400s (24 hours)
Maximum: 7 days
Cost Savings:
Cached input tokens: ~90% cheaper than regular tokens
Output tokens: Same price (not cached)
Important:
You must use explicit model version suffixes (e.g., gemini-2.5-flash-001, NOT just gemini-2.5-flash)
Caches are automatically deleted after TTL expires
Update TTL before expiration to extend cache lifetime
Code Execution
Gemini models can generate and execute Python code to solve problems requiring computation, data analysis, or visualization.
How It Works
Model generates executable Python code
Code runs in secure sandbox
Results are returned to the model
Model incorporates results into response
Supported Operations
Mathematical calculations
Data analysis and statistics
File processing (CSV, JSON, etc.)
Chart and graph generation
Algorithm implementation
Data transformations
Available Python Packages
Standard Library:
math, statistics, random, datetime, json, csv, re
collections, itertools, functools
Data Science:
numpy, pandas, scipy
Visualization:
matplotlib, seaborn
Note: Limited package availability compared to full Python environment
Basic Code Execution (SDK)
import { GoogleGenAI, Tool, ToolCodeExecution } from '@google/genai';
const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY });
const response = await ai.models.generateContent({
model: 'gemini-2.5-flash',
contents: 'What is the sum of the first 50 prime numbers? Generate and run code for the calculation.',
config: {
tools: [{ codeExecution: {} }]
}
});
// Parse response parts
for (const part of response.candidates[0].content.parts) {
if (part.text) {
console.log('Text:', part.text);
}
if (part.executableCode) {
console.log('Generated Code:', part.executableCode.code);
}
if (part.codeExecutionResult) {
console.log('Execution Output:', part.codeExecutionResult.output);
}
}
Basic Code Execution (Fetch)
const response = await fetch(
`https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash:generateContent`,
{
method: 'POST',
headers: {
'Content-Type': 'application/json',
'x-goog-api-key': env.GEMINI_API_KEY,
},
body: JSON.stringify({
tools: [{ code_execution: {} }],
contents: [
{
parts: [
{ text: 'What is the sum of the first 50 prime numbers? Generate and run code.' }
]
}
]
}),
}
);
const data = await response.json();
for (const part of data.candidates[0].content.parts) {
if (part.text) {
console.log('Text:', part.text);
}
if (part.executableCode) {
console.log('Code:', part.executableCode.code);
}
if (part.codeExecutionResult) {
console.log('Result:', part.codeExecutionResult.output);
}
}
Chat with Code Execution (SDK)
const chat = await ai.chats.create({
model: 'gemini-2.5-flash',
config: {
tools: [{ codeExecution: {} }]
}
});
let response = await chat.sendMessage('I have a math question for you.');
console.log(response.text);
response = await chat.sendMessage(
'Calculate the Fibonacci sequence up to the 20th number and sum them.'
);
// Model will generate and execute code, then provide answer
for (const part of response.candidates[0].content.parts) {
if (part.text) console.log(part.text);
if (part.executableCode) console.log('Code:', part.executableCode.code);
if (part.codeExecutionResult) console.log('Output:', part.codeExecutionResult.output);
}
Data Analysis Example
const response = await ai.models.generateContent({
model: 'gemini-2.5-flash',
contents: `
Analyze this sales data and calculate:
1. Total revenue
2. Average sale price
3. Best-selling month
Data (CSV format):
month,sales,revenue
Jan,150,45000
Feb,200,62000
Mar,175,53000
Apr,220,68000
`,
config: {
tools: [{ codeExecution: {} }]
}
});
// Model will generate pandas/numpy code to analyze data
for (const part of response.candidates[0].content.parts) {
if (part.text) console.log(part.text);
if (part.executableCode) console.log('Analysis Code:', part.executableCode.code);
if (part.codeExecutionResult) console.log('Results:', part.codeExecutionResult.output);
}
Visualization Example
const response = await ai.models.generateContent({
model: 'gemini-2.5-flash',
contents: 'Create a bar chart showing the distribution of prime numbers under 100 by their last digit. Generate the chart and describe the pattern.',
config: {
tools: [{ codeExecution: {} }]
}
});
// Model generates matplotlib code, executes it, and describes results
for (const part of response.candidates[0].content.parts) {
if (part.text) console.log(part.text);
if (part.executableCode) console.log('Chart Code:', part.executableCode.code);
if (part.codeExecutionResult) {
// Note: Chart image data would be in output
console.log('Execution completed');
}
}
Response Structure
{
candidates: [
{
content: {
parts: [
{ text: "I'll calculate that for you." },
{
executableCode: {
language: "PYTHON",
code: "def is_prime(n):\n if n <= 1:\n return False\n ..."
}
},
{
codeExecutionResult: {
outcome: "OUTCOME_OK", // or "OUTCOME_FAILED"
output: "5117\n"
}
},
{ text: "The sum of the first 50 prime numbers is 5117." }
]
}
}
]
}
Error Handling
for (const part of response.candidates[0].content.parts) {
if (part.codeExecutionResult) {
if (part.codeExecutionResult.outcome === 'OUTCOME_FAILED') {
console.error('Code execution failed:', part.codeExecutionResult.output);
} else {
console.log('Success:', part.codeExecutionResult.output);
}
}
}
Key Points
When to Use Code Execution:
Complex mathematical calculations
Data analysis and statistics
Algorithm implementations
File parsing and processing
Chart generation
Computational problems
Limitations:
Sandbox environment (limited file system access)
Limited Python package availability
Execution timeout limits
No network access from code
No persistent state between executions
Best Practices:
Specify what calculation or analysis you need clearly
Request code generation explicitly ("Generate and run code...")
Check outcome field for errors
Use for deterministic computations, not for general programming
Important:
Available on all Gemini 2.5 models (Pro, Flash, Flash-Lite)
Code runs in isolated sandbox for security
Supports Python with standard library and common data science packages
Grounding with Google Search
Grounding connects the model to real-time web information, reducing hallucinations and providing up-to-date, fact-checked responses with citations.
How It Works
Model determines if it needs current information
Automatically performs Google Search
Processes search results
Incorporates findings into response
Provides citations and source URLs
Benefits
Real-time information: Access to current events and data
Reduced hallucinations: Answers grounded in web sources
Verifiable: Citations allow fact-checking
Up-to-date: Not limited to model's training cutoff
Two Grounding APIs
1. Google Search (googleSearch) - Recommended for Gemini 2.5
const groundingTool = {
googleSearch: {}
};
Features:
Simple configuration
Automatic search when needed
Available on all Gemini 2.5 models
2. Google Search Retrieval (googleSearchRetrieval) - Legacy (Gemini 1.5)
import { GoogleGenAI } from '@google/genai';
const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY });
const response = await ai.models.generateContent({
model: 'gemini-2.5-flash',
contents: 'Who won the euro 2024?',
config: {
tools: [{ googleSearch: {} }]
}
});
console.log(response.text);
// Check if grounding was used
if (response.candidates[0].groundingMetadata) {
console.log('Search was performed!');
console.log('Sources:', response.candidates[0].groundingMetadata);
}
import { GoogleGenAI, DynamicRetrievalConfigMode } from '@google/genai';
const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY });
const response = await ai.models.generateContent({
model: 'gemini-1.5-flash',
contents: 'Who won the euro 2024?',
config: {
tools: [
{
googleSearchRetrieval: {
dynamicRetrievalConfig: {
mode: DynamicRetrievalConfigMode.MODE_DYNAMIC,
dynamicThreshold: 0.7 // Search only if confidence < 70%
}
}
}
]
}
});
console.log(response.text);
if (!response.candidates[0].groundingMetadata) {
console.log('Model answered from its own knowledge (high confidence)');
}
Grounding Metadata Structure
{
groundingMetadata: {
searchQueries: [
{ text: "euro 2024 winner" }
],
webPages: [
{
url: "https://example.com/euro-2024-results",
title: "UEFA Euro 2024 Final Results",
snippet: "Spain won UEFA Euro 2024..."
}
],
citations: [
{
startIndex: 42,
endIndex: 47,
uri: "https://example.com/euro-2024-results"
}
],
retrievalQueries: [
{
query: "who won euro 2024 final"
}
]
}
}
Chat with Grounding (SDK)
const chat = await ai.chats.create({
model: 'gemini-2.5-flash',
config: {
tools: [{ googleSearch: {} }]
}
});
let response = await chat.sendMessage('What are the latest developments in quantum computing?');
console.log(response.text);
// Check grounding sources
if (response.candidates[0].groundingMetadata) {
const sources = response.candidates[0].groundingMetadata.webPages || [];
console.log(`Sources used: ${sources.length}`);
sources.forEach(source => {
console.log(`- ${source.title}: ${source.url}`);
});
}
// Follow-up still has grounding enabled
response = await chat.sendMessage('Which company made the biggest breakthrough?');
console.log(response.text);
Combining Grounding with Function Calling
const weatherFunction = {
name: 'get_current_weather',
description: 'Get current weather for a location',
parametersJsonSchema: {
type: 'object',
properties: {
location: { type: 'string', description: 'City name' }
},
required: ['location']
}
};
const response = await ai.models.generateContent({
model: 'gemini-2.5-flash',
contents: 'What is the weather like in the city that won Euro 2024?',
config: {
tools: [
{ googleSearch: {} },
{ functionDeclarations: [weatherFunction] }
]
}
});
// Model will:
// 1. Use Google Search to find Euro 2024 winner
// 2. Call get_current_weather function with the city
// 3. Combine both results in response
Checking if Grounding was Used
const response = await ai.models.generateContent({
model: 'gemini-2.5-flash',
contents: 'What is 2+2?', // Model knows this without search
config: {
tools: [{ googleSearch: {} }]
}
});
if (!response.candidates[0].groundingMetadata) {
console.log('Model answered from its own knowledge (no search needed)');
} else {
console.log('Search was performed');
}
Key Points
When to Use Grounding:
Current events and news
Real-time data (stock prices, sports scores, weather)
Fact-checking and verification
Questions about recent developments
Information beyond model's training cutoff
When NOT to Use:
General knowledge questions
Mathematical calculations
Code generation
Creative writing
Tasks requiring internal reasoning only
Cost Considerations:
Grounding adds latency (search takes time)
Additional token costs for retrieved content
Use dynamicThreshold to control when searches happen (Gemini 1.5)
Important Notes:
Grounding requires Google Cloud project (not just API key)
Search results quality depends on query phrasing
Citations may not cover all facts in response
Search is performed automatically based on confidence
Gemini 2.5 vs 1.5:
Gemini 2.5: Use googleSearch (simple, recommended)
Gemini 1.5: Use googleSearchRetrieval with dynamicThreshold
Best Practices:
Always check groundingMetadata to see if search was used
Display citations to users for transparency
Use specific, well-phrased questions for better search results
Combine with function calling for hybrid workflows
Error Handling
Common Errors
1. Invalid API Key (401)
{
error: {
code: 401,
message: 'API key not valid. Please pass a valid API key.',
status: 'UNAUTHENTICATED'
}
}
Solution: Verify GEMINI_API_KEY environment variable is set correctly.
2. Rate Limit Exceeded (429)
{
error: {
code: 429,
message: 'Resource has been exhausted (e.g. check quota).',
status: 'RESOURCE_EXHAUSTED'
}
}
const chat = await ai.models.createChat({ model: 'gemini-2.5-flash' });
const response = await chat.sendMessage(message);
// response.text is directly available
Production Best Practices
1. Always Do
✅ Use @google/genai (NOT @google/generative-ai)
✅ Set maxOutputTokens to prevent excessive generation
✅ Implement rate limit handling with exponential backoff
✅ Use environment variables for API keys (never hardcode)
✅ Validate inputs before sending to API (save costs)
✅ Use streaming for better UX on long responses
✅ Choose the right model based on your needs (Pro for complex reasoning, Flash for balance, Flash-Lite for speed)
✅ Handle errors gracefully with try-catch
✅ Monitor token usage for cost control
✅ Use correct model names: gemini-2.5-pro/flash/flash-lite
2. Never Do
❌ Never use @google/generative-ai (deprecated!)
❌ Never hardcode API keys in code
❌ Never claim 2M context for Gemini 2.5 (it's 1,048,576 input tokens)
❌ Never expose API keys in client-side code
❌ Never skip error handling (always try-catch)
❌ Never use generic rate limits (each model has different limits - check official docs)
❌ Never send PII without user consent
❌ Never trust user input without validation
❌ Never ignore rate limits (will get 429 errors)
❌ Never use old model names like gemini-1.5-pro (use 2.5 models)
3. Security
API Key Storage: Use environment variables or secret managers
Server-Side Only: Never expose API keys in browser JavaScript
Input Validation: Sanitize all user inputs before API calls
Rate Limiting: Implement your own rate limits to prevent abuse
Error Messages: Don't expose API keys or sensitive data in error logs
4. Cost Optimization
Choose Right Model: Use Flash for most tasks, Pro only when needed
Set Token Limits: Use maxOutputTokens to control costs
Batch Requests: Process multiple items efficiently
Cache Results: Store responses when appropriate
Monitor Usage: Track token consumption in Google Cloud Console
5. Performance
Use Streaming: Better perceived latency for long responses
Parallel Requests: Use Promise.all() for independent calls
Edge Deployment: Deploy to Cloudflare Workers for low latency
Connection Pooling: Reuse HTTP connections when possible
Quick Reference
Installation
npm install @google/genai@1.27.0
Environment
export GEMINI_API_KEY="..."
Models (2025)
gemini-2.5-pro (1,048,576 in / 65,536 out) - Best for complex reasoning
gemini-2.5-flash (1,048,576 in / 65,536 out) - Best price-performance balance
gemini-2.5-flash-lite (1,048,576 in / 65,536 out) - Fastest, most cost-effective