Skip to main content

openrouter

OpenRouter unified API for 400+ models — chat completions, streaming, tool calling, structured output, embeddings, multimodal.

Datos de origen

Repositorio
DimitriGilbert/ai-skills
Última actividad en el origen
3 de junio de 2026 a las 11:17
Idioma detectado de SKILL.md
inglés
Estrellas
5
Forks
0

Opciones de instalación

De forma predeterminada está seleccionado el prompt que primero revisa el origen. Puedes cambiar a un comando directo o descargar una copia local.

Revisa los archivos de origen

Lee SKILL.md y los archivos complementarios que muestra SkillsMP antes de decidir si quieres instalarlo.

Explorador de archivos
17 archivos

Mostrando SKILL.md

SKILL.md
Instrucciones de origen · Vista previa de solo lectura
name
openrouter
description
OpenRouter unified API for 400+ models — chat completions, streaming, tool calling, structured output, embeddings, multimodal.
# OpenRouter API for AI Agents Expert guidance for AI agents integrating with OpenRouter API - unified access to 400+ models from 90+ providers. **When to use this skill:** - Making chat completions via OpenRouter API - Selecting appropriate models and variants - Implementing streaming responses - Using tool/function calling - Enforcing structured outputs - Integrating web search - Handling multimodal inputs (images, audio, video, PDFs) - Managing model routing and fallbacks - Handling errors and retries - Optimizing cost and performance --- ## API Basics ### Making a Request **Endpoint**: `POST https://openrouter.ai/api/v1/chat/completions` **Headers** (required): ```typescript { 'Authorization': `Bearer ${apiKey}`, 'Content-Type': 'application/json', // Optional: for app attribution 'HTTP-Referer': 'https://your-app.com', 'X-Title': 'Your App Name' } ``` **Minimal request structure**: ```typescript const response = await fetch('https://openrouter.ai/api/v1/chat/completions', { method: 'POST', headers: { 'Authorization': `Bearer ${apiKey}`, 'Content-Type': 'application/json', }, body: JSON.stringify({ model: 'anthropic/claude-3.5-sonnet', messages: [ { role: 'user', content: 'Your prompt here' } ] }) }); ``` ### Response Structure **Non-streaming response**: ```json { "id": "gen-abc123", "choices": [{ "message": { "role": "assistant", "content": "Response text here" }, "finish_reason": "stop" }], "usage": { "prompt_tokens": 10, "completion_tokens": 20, "total_tokens": 30 }, "model": "anthropic/claude-3.5-sonnet" } ``` **Key fields**: - `choices[0].message.content` - The assistant's response - `choices[0].finish_reason` - Why generation stopped (stop, length, tool_calls, etc.) - `usage` - Token counts and cost information - `model` - Actual model used (may differ from requested) ### When to Use Streaming vs Non-Streaming **Use streaming (`stream: true`)** when: - Real-time responses needed (chat interfaces, interactive tools) - Latency matters (user-facing applications) - Large responses expected (long-form content) - Want to show progressive output **Use non-streaming** when: - Processing in background (batch jobs, async tasks) - Need complete response before processing - Building to an API/endpoint - Response is short (few tokens) **Streaming basics**: ```typescript const response = await fetch('https://openrouter.ai/api/v1/chat/completions', { method: 'POST', headers: { /* ... */ }, body: JSON.stringify({ model: 'anthropic/claude-3.5-sonnet', messages: [{ role: 'user', content: '...' }], stream: true }) }); for await (const chunk of response.body) { const text = new TextDecoder().decode(chunk); const lines = text.split('\n').filter(line => line.startsWith('data: ')); for (const line of lines) { const data = line.slice(6); // Remove 'data: ' if (data === '[DONE]') break; const parsed = JSON.parse(data); const content = parsed.choices?.[0]?.delta?.content; if (content) { // Accumulate or display content } } } ``` --- ## Model Selection ### Model Identifier Format **Format**: `provider/model-name[:variant]` Examples: - `anthropic/claude-3.5-sonnet` - Specific model - `openai/gpt-4o:online` - With web search enabled - `google/gemini-2.0-flash:free` - Free tier variant ### Model Variants and When to Use Them | Variant | Use When | Tradeoffs | |---------|----------|-----------| | `:free` | Cost is primary concern, testing, prototyping | Rate limits, lower quality models | | `:online` | Need current information, real-time data | Higher cost, web search latency | | `:extended` | Large context window needed | May be slower, higher cost | | `:thinking` | Complex reasoning, multi-step problems | Higher token usage, slower | | `:nitro` | Speed is critical | May have quality tradeoffs | | `:exacto` | Need specific provider | No fallbacks, may be less available | ### Default Model Choices by Task **General purpose**: `anthropic/claude-3.5-sonnet` or `openai/gpt-4o` - Balanced quality, speed, cost - Good for most tasks **Coding**: `anthropic/claude-3.5-sonnet` or `openai/gpt-4o` - Strong code generation and understanding - Good reasoning **Complex reasoning**: `anthropic/claude-opus-4:thinking` or `openai/o3` - Deep reasoning capabilities - Higher cost, slower **Fast responses**: `openai/gpt-4o-mini:nitro` or `google/gemini-2.0-flash` - Minimal latency - Good for real-time applications **Cost-sensitive**: `google/gemini-2.0-flash:free` or `meta-llama/llama-3.1-70b:free` - No cost with limits - Good for high-volume, lower-complexity tasks **Current information**: `anthropic/claude-3.5-sonnet:online` or `google/gemini-2.5-pro:online` - Web search built-in - Real-time data **Large context**: `anthropic/claude-3.5-sonnet:extended` or `google/gemini-2.5-pro:extended` - 200K+ context windows - Document analysis, codebase understanding ### Provider Routing Preferences **Default behavior**: OpenRouter automatically selects best provider **Explicit provider order**: ```typescript { provider: { order: ['anthropic', 'openai', 'google'], allow_fallbacks: true, sort: 'price' // 'price', 'latency', or 'throughput' } } ``` **When to set provider order**: - Have preferred provider arrangements - Need to optimize for specific metric (cost, speed) - Want to exclude certain providers - Have BYOK (Bring Your Own Key) for specific providers ### Model Fallbacks **Automatic fallback** - try multiple models in order: ```typescript { models: [ 'anthropic/claude-3.5-sonnet', 'openai/gpt-4o', 'google/gemini-2.0-flash' ] } ``` **When to use fallbacks**: - High reliability required - Multiple providers acceptable - Want graceful degradation - Avoid single point of failure **Fallback behavior**: - Tries first model - Falls to next on error (5xx, 429, timeout) - Uses whichever succeeds - Returns which model was used in `model` field --- ## Parameters You Need ### Core Parameters **model** (string, optional) - Which model to use - Default: user's default model - **Always specify for consistency** **messages** (Message[], required) - Conversation history - Structure: `{ role: 'user'|'assistant'|'system', content: string | ContentPart[] }` - For multimodal: content can be array of text and image_url parts **stream** (boolean, default: false) - Enable Server-Sent Events streaming - Use for real-time responses **temperature** (float, 0.0-2.0, default: 1.0) - Controls randomness - **0.0-0.3**: Deterministic, factual responses (code, precise answers) - **0.4-0.7**: Balanced (general use) - **0.8-1.2**: Creative (brainstorming, creative writing) - **1.3-2.0**: Highly creative, unpredictable (experimental) **max_tokens** (integer, optional) - Maximum tokens to generate - **Always set** to control cost and prevent runaway responses - Typical: 100-500 for short, 1000-2000 for long responses - Model limit: context_length - prompt_length **top_p** (float, 0.0-1.0, default: 1.0) - Nucleus sampling - limits to top probability mass - **Use instead of temperature** when you want predictable diversity - **0.9-0.95**: Common settings for quality **top_k** (integer, 0+, default: 0/disabled) - Limit to K most likely tokens - **1**: Always most likely (deterministic) - **40-50**: Balanced - Not available for OpenAI models ### Sampling Strategy Guidelines **For code generation**: `temperature: 0.1-0.3, top_p: 0.95` **For factual responses**: `temperature: 0.0-0.2` **For creative writing**: `temperature: 0.8-1.2` **For brainstorming**: `temperature: 1.0-1.5` **For chat**: `temperature: 0.6-0.8` ### Tool Calling Parameters **tools** (Tool[], default: []) - Available functions for model to call - Structure: ```typescript { type: 'function', function: { name: 'function_name', description: 'What it does', parameters: { /* JSON Schema */ } } } ``` **tool_choice** (string | object, default: 'auto') - Control when tools are called - `'auto'`: Model decides (default) - `'none'`: Never call tools - `'required'`: Must call a tool - `{ type: 'function', function: { name: 'specific_tool' } }`: Force specific tool **parallel_tool_calls** (boolean, default: true) - Allow multiple tools simultaneously - Set `false` for sequential execution **When to use tools**: - Need to query external APIs (weather, search, database) - Need to perform calculations or data processing - Building agentic systems - Need structured data extraction ### Structured Output Parameters **response_format** (object, optional) - Enforce specific output format **JSON object mode**: ```typescript { type: 'json_object' } ``` - Model returns valid JSON - Must also instruct model in system message **JSON Schema mode** (strict): ```typescript { type: 'json_schema', json_schema: { name: 'schema_name', strict: true, schema: { /* JSON Schema */ } } } ``` - Model returns JSON matching exact schema - **Use when structure is critical** (APIs, data processing) **When to use structured outputs**: - Need predictable response format - Integrating with systems (APIs, databases) - Data extraction - Form filling ### Web Search Parameters **Enable via model variant** (simplest): ```typescript { model: 'anthropic/claude-3.5-sonnet:online' } ``` **Enable via plugin**: ```typescript { plugins: [{ id: 'web', enabled: true, max_results: 5 }] } ``` **When to use web search**: - Need current information (news, prices, events) - User asks about recent developments - Need factual verification - Topic requires real-time data ### Other Important Parameters **user** (string, optional) - Stable identifier for end-user - **Set when you have user IDs** - Helps with abuse detection and caching **session_id** (string, optional) - Group related requests - **Set for conversation tracking** - Improves caching and observability **metadata** (Record<string, string>, optional) - Custom metadata (max 16 key-value pairs) - **Use for analytics and tracking** - Keys: max 64 chars, Values: max 512 chars **stop** (string | string[], optional) - Stop sequences to halt generation - Common: `['\n\n', '###', 'END']` --- ## Handling Responses ### Non-Streaming Responses Extract content: ```typescript const response = await fetch(/* ... */); const data = await response.json(); const content = data.choices[0].message.content; const finishReason = data.choices[0].finish_reason;
Ver en GitHub
Este SKILL.md es muy grande, por eso SkillsMP muestra aqui solo la primera seccion. Ver en GitHub