Skip to main content

openrouter

OpenRouter unified API for 400+ models — chat completions, streaming, tool calling, structured output, embeddings, multimodal.

Quellinformationen

Repository
DimitriGilbert/ai-skills
Letzte Quellaktivität
3. Juni 2026 um 11:17
Erkannte Sprache von SKILL.md
Englisch
Sterne
5
Forks
0

Installationsoptionen

Standardmäßig ist der Prompt ausgewählt, der zuerst die Quelle prüft. Sie können zu einem direkten Befehl wechseln oder eine lokale Kopie herunterladen.

Quelldateien prüfen

Lesen Sie SKILL.md und alle von SkillsMP angezeigten Begleitdateien, bevor Sie sich für eine Installation entscheiden.

Datei-Explorer
17 Dateien

SKILL.md wird angezeigt

SKILL.md
Quellanweisungen · Schreibgeschützte Vorschau
name
openrouter
description
OpenRouter unified API for 400+ models — chat completions, streaming, tool calling, structured output, embeddings, multimodal.
# OpenRouter API for AI Agents Expert guidance for AI agents integrating with OpenRouter API - unified access to 400+ models from 90+ providers. **When to use this skill:** - Making chat completions via OpenRouter API - Selecting appropriate models and variants - Implementing streaming responses - Using tool/function calling - Enforcing structured outputs - Integrating web search - Handling multimodal inputs (images, audio, video, PDFs) - Managing model routing and fallbacks - Handling errors and retries - Optimizing cost and performance --- ## API Basics ### Making a Request **Endpoint**: `POST https://openrouter.ai/api/v1/chat/completions` **Headers** (required): ```typescript { 'Authorization': `Bearer ${apiKey}`, 'Content-Type': 'application/json', // Optional: for app attribution 'HTTP-Referer': 'https://your-app.com', 'X-Title': 'Your App Name' } ``` **Minimal request structure**: ```typescript const response = await fetch('https://openrouter.ai/api/v1/chat/completions', { method: 'POST', headers: { 'Authorization': `Bearer ${apiKey}`, 'Content-Type': 'application/json', }, body: JSON.stringify({ model: 'anthropic/claude-3.5-sonnet', messages: [ { role: 'user', content: 'Your prompt here' } ] }) }); ``` ### Response Structure **Non-streaming response**: ```json { "id": "gen-abc123", "choices": [{ "message": { "role": "assistant", "content": "Response text here" }, "finish_reason": "stop" }], "usage": { "prompt_tokens": 10, "completion_tokens": 20, "total_tokens": 30 }, "model": "anthropic/claude-3.5-sonnet" } ``` **Key fields**: - `choices[0].message.content` - The assistant's response - `choices[0].finish_reason` - Why generation stopped (stop, length, tool_calls, etc.) - `usage` - Token counts and cost information - `model` - Actual model used (may differ from requested) ### When to Use Streaming vs Non-Streaming **Use streaming (`stream: true`)** when: - Real-time responses needed (chat interfaces, interactive tools) - Latency matters (user-facing applications) - Large responses expected (long-form content) - Want to show progressive output **Use non-streaming** when: - Processing in background (batch jobs, async tasks) - Need complete response before processing - Building to an API/endpoint - Response is short (few tokens) **Streaming basics**: ```typescript const response = await fetch('https://openrouter.ai/api/v1/chat/completions', { method: 'POST', headers: { /* ... */ }, body: JSON.stringify({ model: 'anthropic/claude-3.5-sonnet', messages: [{ role: 'user', content: '...' }], stream: true }) }); for await (const chunk of response.body) { const text = new TextDecoder().decode(chunk); const lines = text.split('\n').filter(line => line.startsWith('data: ')); for (const line of lines) { const data = line.slice(6); // Remove 'data: ' if (data === '[DONE]') break; const parsed = JSON.parse(data); const content = parsed.choices?.[0]?.delta?.content; if (content) { // Accumulate or display content } } } ``` --- ## Model Selection ### Model Identifier Format **Format**: `provider/model-name[:variant]` Examples: - `anthropic/claude-3.5-sonnet` - Specific model - `openai/gpt-4o:online` - With web search enabled - `google/gemini-2.0-flash:free` - Free tier variant ### Model Variants and When to Use Them | Variant | Use When | Tradeoffs | |---------|----------|-----------| | `:free` | Cost is primary concern, testing, prototyping | Rate limits, lower quality models | | `:online` | Need current information, real-time data | Higher cost, web search latency | | `:extended` | Large context window needed | May be slower, higher cost | | `:thinking` | Complex reasoning, multi-step problems | Higher token usage, slower | | `:nitro` | Speed is critical | May have quality tradeoffs | | `:exacto` | Need specific provider | No fallbacks, may be less available | ### Default Model Choices by Task **General purpose**: `anthropic/claude-3.5-sonnet` or `openai/gpt-4o` - Balanced quality, speed, cost - Good for most tasks **Coding**: `anthropic/claude-3.5-sonnet` or `openai/gpt-4o` - Strong code generation and understanding - Good reasoning **Complex reasoning**: `anthropic/claude-opus-4:thinking` or `openai/o3` - Deep reasoning capabilities - Higher cost, slower **Fast responses**: `openai/gpt-4o-mini:nitro` or `google/gemini-2.0-flash` - Minimal latency - Good for real-time applications **Cost-sensitive**: `google/gemini-2.0-flash:free` or `meta-llama/llama-3.1-70b:free` - No cost with limits - Good for high-volume, lower-complexity tasks **Current information**: `anthropic/claude-3.5-sonnet:online` or `google/gemini-2.5-pro:online` - Web search built-in - Real-time data **Large context**: `anthropic/claude-3.5-sonnet:extended` or `google/gemini-2.5-pro:extended` - 200K+ context windows - Document analysis, codebase understanding ### Provider Routing Preferences **Default behavior**: OpenRouter automatically selects best provider **Explicit provider order**: ```typescript { provider: { order: ['anthropic', 'openai', 'google'], allow_fallbacks: true, sort: 'price' // 'price', 'latency', or 'throughput' } } ``` **When to set provider order**: - Have preferred provider arrangements - Need to optimize for specific metric (cost, speed) - Want to exclude certain providers - Have BYOK (Bring Your Own Key) for specific providers ### Model Fallbacks **Automatic fallback** - try multiple models in order: ```typescript { models: [ 'anthropic/claude-3.5-sonnet', 'openai/gpt-4o', 'google/gemini-2.0-flash' ] } ``` **When to use fallbacks**: - High reliability required - Multiple providers acceptable - Want graceful degradation - Avoid single point of failure **Fallback behavior**: - Tries first model - Falls to next on error (5xx, 429, timeout) - Uses whichever succeeds - Returns which model was used in `model` field --- ## Parameters You Need ### Core Parameters **model** (string, optional) - Which model to use - Default: user's default model - **Always specify for consistency** **messages** (Message[], required) - Conversation history - Structure: `{ role: 'user'|'assistant'|'system', content: string | ContentPart[] }` - For multimodal: content can be array of text and image_url parts **stream** (boolean, default: false) - Enable Server-Sent Events streaming - Use for real-time responses **temperature** (float, 0.0-2.0, default: 1.0) - Controls randomness - **0.0-0.3**: Deterministic, factual responses (code, precise answers) - **0.4-0.7**: Balanced (general use) - **0.8-1.2**: Creative (brainstorming, creative writing) - **1.3-2.0**: Highly creative, unpredictable (experimental) **max_tokens** (integer, optional) - Maximum tokens to generate - **Always set** to control cost and prevent runaway responses - Typical: 100-500 for short, 1000-2000 for long responses - Model limit: context_length - prompt_length **top_p** (float, 0.0-1.0, default: 1.0) - Nucleus sampling - limits to top probability mass - **Use instead of temperature** when you want predictable diversity - **0.9-0.95**: Common settings for quality **top_k** (integer, 0+, default: 0/disabled) - Limit to K most likely tokens - **1**: Always most likely (deterministic) - **40-50**: Balanced - Not available for OpenAI models ### Sampling Strategy Guidelines **For code generation**: `temperature: 0.1-0.3, top_p: 0.95` **For factual responses**: `temperature: 0.0-0.2` **For creative writing**: `temperature: 0.8-1.2` **For brainstorming**: `temperature: 1.0-1.5` **For chat**: `temperature: 0.6-0.8` ### Tool Calling Parameters **tools** (Tool[], default: []) - Available functions for model to call - Structure: ```typescript { type: 'function', function: { name: 'function_name', description: 'What it does', parameters: { /* JSON Schema */ } } } ``` **tool_choice** (string | object, default: 'auto') - Control when tools are called - `'auto'`: Model decides (default) - `'none'`: Never call tools - `'required'`: Must call a tool - `{ type: 'function', function: { name: 'specific_tool' } }`: Force specific tool **parallel_tool_calls** (boolean, default: true) - Allow multiple tools simultaneously - Set `false` for sequential execution **When to use tools**: - Need to query external APIs (weather, search, database) - Need to perform calculations or data processing - Building agentic systems - Need structured data extraction ### Structured Output Parameters **response_format** (object, optional) - Enforce specific output format **JSON object mode**: ```typescript { type: 'json_object' } ``` - Model returns valid JSON - Must also instruct model in system message **JSON Schema mode** (strict): ```typescript { type: 'json_schema', json_schema: { name: 'schema_name', strict: true, schema: { /* JSON Schema */ } } } ``` - Model returns JSON matching exact schema - **Use when structure is critical** (APIs, data processing) **When to use structured outputs**: - Need predictable response format - Integrating with systems (APIs, databases) - Data extraction - Form filling ### Web Search Parameters **Enable via model variant** (simplest): ```typescript { model: 'anthropic/claude-3.5-sonnet:online' } ``` **Enable via plugin**: ```typescript { plugins: [{ id: 'web', enabled: true, max_results: 5 }] } ``` **When to use web search**: - Need current information (news, prices, events) - User asks about recent developments - Need factual verification - Topic requires real-time data ### Other Important Parameters **user** (string, optional) - Stable identifier for end-user - **Set when you have user IDs** - Helps with abuse detection and caching **session_id** (string, optional) - Group related requests - **Set for conversation tracking** - Improves caching and observability **metadata** (Record<string, string>, optional) - Custom metadata (max 16 key-value pairs) - **Use for analytics and tracking** - Keys: max 64 chars, Values: max 512 chars **stop** (string | string[], optional) - Stop sequences to halt generation - Common: `['\n\n', '###', 'END']` --- ## Handling Responses ### Non-Streaming Responses Extract content: ```typescript const response = await fetch(/* ... */); const data = await response.json(); const content = data.choices[0].message.content; const finishReason = data.choices[0].finish_reason;
Auf GitHub ansehen
Diese SKILL.md ist sehr gross, daher zeigt SkillsMP hier nur den ersten Abschnitt. Auf GitHub ansehen