| name | generate-image |
| description | Generate or edit images using AI models (FLUX, Gemini). Use for general-purpose image generation including photos, illustrations, artwork, visual assets, concept art, and any image that isn't a technical diagram or schematic. For flowcharts, circuits, pathways, and technical diagrams, use the scientific-schematics skill instead. |
Generate Image
Generate and edit high-quality images using OpenRouter's image generation models including Gemini 3 Pro Image and GPT-5 Image.
When to Use This Skill
Use generate-image for:
- Photos and photorealistic images
- Artistic illustrations and artwork
- Concept art and visual concepts
- Visual assets for presentations or documents
- Image editing and modifications
- Any general-purpose image generation needs
Use scientific-schematics instead for:
- Flowcharts and process diagrams
- Circuit diagrams and electrical schematics
- Biological pathways and signaling cascades
- System architecture diagrams
- CONSORT diagrams and methodology flowcharts
- Any technical/schematic diagrams
Quick Start
Use the scripts/generate_image.py script to generate or edit images:
python scripts/generate_image.py "A beautiful sunset over mountains"
python scripts/generate_image.py "Make the sky purple" --input photo.jpg
This generates/edits an image and saves it as generated_image.png in the current directory.
API Key Setup
The script requires an OpenRouter API key and resolves it in this order:
--api-key
- the
OPENROUTER_API_KEY environment variable
OPENROUTER_API_KEY= in a .env file, searching the working directory upward
Do not open the .env file to check for the key -- run the script and read its exit status. With no key it exits with setup instructions. Keys: https://openrouter.ai/keys
Model Selection
Default model: google/gemini-3-pro-image-preview (high quality, recommended)
Available models for generation and editing:
google/gemini-3-pro-image-preview - High quality, supports generation + editing
google/gemini-3-pro-image - Same model, non-preview id
google/gemini-3.1-flash-image - Roughly half the per-image cost of gemini-3-pro
openai/gpt-5-image-mini - Cheapest of these per generated image
Select based on:
- Quality: Use gemini-3-pro-image-preview or gemini-3-pro-image
- Editing: All four accept an input image
- Cost: Use gemini-3.1-flash-image or gpt-5-image-mini
The script posts to /api/v1/chat/completions, so a model works only if it is
listed at https://openrouter.ai/api/v1/models with image in its
architecture.output_modalities. The Flux.2 ids are served by OpenRouter's
separate Images API (/api/v1/images/models) and fail here.
Common Usage Patterns
Basic generation
python scripts/generate_image.py "Your prompt here"
Specify model
python scripts/generate_image.py "A cat in space" --model "google/gemini-3.1-flash-image"
Custom output path
python scripts/generate_image.py "Abstract art" --output artwork.png
Edit an existing image
python scripts/generate_image.py "Make the background blue" --input photo.jpg
Edit with a specific model
python scripts/generate_image.py "Add sunglasses to the person" --input portrait.png --model "google/gemini-3.1-flash-image"
Edit with custom output
python scripts/generate_image.py "Remove the text from the image" --input screenshot.png --output cleaned.png
Multiple images
Run the script multiple times with different prompts or output paths:
python scripts/generate_image.py "Image 1 description" --output image1.png
python scripts/generate_image.py "Image 2 description" --output image2.png
Script Parameters
prompt (required): Text description of the image to generate, or editing instructions
--input or -i: Input image path for editing (enables edit mode)
--model or -m: OpenRouter model ID (default: google/gemini-3-pro-image-preview)
--output or -o: Output file path (default: generated_image.png)
--api-key: OpenRouter API key (overrides the environment and .env file)
Example Use Cases
For Scientific Documents
python scripts/generate_image.py "Microscopic view of cancer cells being attacked by immunotherapy agents, scientific illustration style" --output figures/immunotherapy_concept.png
python scripts/generate_image.py "DNA double helix structure with highlighted mutation site, modern scientific visualization" --output slides/dna_mutation.png
For Presentations and Posters
python scripts/generate_image.py "Abstract blue and white background with subtle molecular patterns, professional presentation style" --output slides/background.png
python scripts/generate_image.py "Laboratory setting with modern equipment, photorealistic, well-lit" --output poster/hero.png
For General Visual Content
python scripts/generate_image.py "Professional team collaboration around a digital whiteboard, modern office" --output docs/team_collaboration.png
python scripts/generate_image.py "Futuristic AI brain concept with glowing neural networks" --output marketing/ai_concept.png
Error Handling
The script provides clear error messages for:
- Missing API key (with setup instructions)
- API errors (with status codes)
- Unexpected response formats
- Missing dependencies (requests library)
If the script fails, read the error message and address the issue before retrying.
Critical Prompt Requirements
IMPORTANT: No Meta Instructions in Output
When generating prompts for the AI image generation models, ensure the generated image does NOT contain any visible text showing:
- The prompt or instructions that were given to generate it
- System instructions or AI-related metadata
- Any "meta" text describing how the image was created
- Watermarks or labels indicating AI generation
- Layout descriptions (e.g., "left panel", "right panel", "center panel")
- Font specifications or typography instructions
- Color scheme descriptions or palette information
The image should only contain the requested visual content. Always include this instruction in your prompts: "Do not include any text showing the prompt, instructions, layout descriptions, font/color specifications, or metadata in the generated image."
Notes
- Images are returned as base64-encoded data URLs and automatically saved as PNG files
- The script supports both
images and content response formats from different OpenRouter models
- Generation time varies by model (typically 5-30 seconds)
- For image editing, the input image is encoded as base64 and sent to the model
- Supported input image formats: PNG, JPEG, GIF, WebP
- Check OpenRouter pricing for cost information: https://openrouter.ai/models
Image Editing Tips
- Be specific about what changes you want (e.g., "change the sky to sunset colors" vs "edit the sky")
- Reference specific elements in the image when possible
- For best results, use clear and detailed editing instructions
- Every model listed under Model Selection supports image editing through OpenRouter
Integration with Other Skills
- scientific-schematics: Use for technical diagrams, flowcharts, circuits, pathways
- generate-image: Use for photos, illustrations, artwork, visual concepts
- scientific-slides: Combine with generate-image for visually rich presentations
- latex-posters: Use generate-image for poster visuals and hero images