AI-powered image generation and editing using Google Gemini, Google Imagen, and OpenAI models.
Generate images from text descriptions, edit existing images, create logos/stickers,
apply style transfers, and produce product mockups.
Use this skill when the user requests:
- Image generation from text descriptions
- Image editing or modifications
- Logos, stickers, or graphic design assets
- Product mockups or visualizations
- Style transfers or artistic effects
- Iterative image refinement
Available models:
- Google Gemini: gemini-2.5-flash-image (Nano Banana), gemini-3-pro-image-preview (Nano Banana Pro)
- Google Imagen: imagen-4.0-generate-001, imagen-4.0-ultra-generate-001, imagen-4.0-fast-generate-001
- OpenAI: gpt-image-1.5 (recommended), gpt-image-1, dall-e-3, dall-e-2
Inspired by: https://github.com/EveryInc/every-marketplace/tree/main/plugins/compounding-engineering/skills/gemini-imagegen
Standardmäßig ist der Prompt ausgewählt, der zuerst die Quelle prüft. Sie können zu einem direkten Befehl wechseln oder eine lokale Kopie herunterladen.
Quelldateien prüfen
Lesen Sie SKILL.md und alle von SkillsMP angezeigten Begleitdateien, bevor Sie sich für eine Installation entscheiden.
Mit Codex oder Claude installieren Kopieren Sie diesen Prompt, fügen Sie ihn in Codex, Claude oder einen anderen Assistant ein und lassen Sie die Skill-Seite prüfen und installieren.
Ein direkter Befehl überspringt den Prüf-Prompt. Prüfen Sie die Quelle, bevor Sie ihn ausführen.
AI-powered image generation and editing using Google Gemini, Google Imagen, and OpenAI models.
Generate images from text descriptions, edit existing images, create logos/stickers,
apply style transfers, and produce product mockups.
Use this skill when the user requests:
- Image generation from text descriptions
- Image editing or modifications
- Logos, stickers, or graphic design assets
- Product mockups or visualizations
- Style transfers or artistic effects
- Iterative image refinement
Available models:
- Google Gemini: gemini-2.5-flash-image (Nano Banana), gemini-3-pro-image-preview (Nano Banana Pro)
- Google Imagen: imagen-4.0-generate-001, imagen-4.0-ultra-generate-001, imagen-4.0-fast-generate-001
- OpenAI: gpt-image-1.5 (recommended), gpt-image-1, dall-e-3, dall-e-2
Inspired by: https://github.com/EveryInc/every-marketplace/tree/main/plugins/compounding-engineering/skills/gemini-imagegen
Important (December 2025): The google-generativeai package has been deprecated.
This skill now uses the google-genai SDK. If upgrading from older code, see the
migration guide.
Purpose
This skill enables AI-powered image generation and editing through Google's Gemini image models and OpenAI's DALL-E models. Create photorealistic images, illustrations, logos, stickers, and product mockups from natural language descriptions. Edit existing images with text instructions, apply style transfers, and refine outputs through iterative conversation.
Attribution: This skill is inspired by the gemini-imagegen skill from Every Marketplace by Every Inc.
When to Use
This skill should be invoked when the user asks to:
Generate images from text descriptions ("create an image of...", "generate a picture...")
Create logos, icons, or stickers ("design a logo for...", "make a sticker...")
Edit or modify existing images ("change the background to...", "add... to this image")
Apply artistic styles or effects ("make it look like...", "stylize as...")
Create product mockups or visualizations ("product photo of...", "mockup showing...")
Refine or iterate on images ("make it more...", "adjust the...", "try again with...")
Generate variations with different styles or compositions
Save the generated image to an appropriate location
Verify the output meets the request
Show the user the saved file path
Offer refinement if the result isn't quite right
Explain the prompt used so the user understands the generation
Step 6: Iterate if Needed
If the user wants changes:
For Gemini: Use chat interface to maintain context
For gpt-image-1.5: Use editing API for precise face/logo preservation
For Imagen/DALL-E: Generate new image with updated prompt
Keep previous versions for comparison
Suggest specific adjustments based on the current result
Requirements
API Keys:
Google (Gemini/Imagen): Set GOOGLE_API_KEY or GEMINI_API_KEY environment variable
OpenAI: Set OPENAI_API_KEY environment variable
Python Packages:
pip install google-genai openai pillow requests
Note: The google-generativeai package has been deprecated and will no longer receive updates.
Use google-genai instead. Migration guide: https://ai.google.dev/gemini-api/docs/migrate
System:
Python 3.8+
Internet connection for API access
Write permissions for saving images
Approximate Costs (per image):
Model
Low Quality
High Quality
imagen-4.0-fast
$0.02
$0.02
imagen-4.0
-
~$0.04
imagen-4.0-ultra
-
~$0.08
gemini-2.5-flash-image
~$0.039
~$0.039
gpt-image-1.5
~$0.016
~$0.15
gpt-image-1
~$0.02
~$0.19
dall-e-3
~$0.04
~$0.08
dall-e-2
~$0.02
~$0.02
Best Practices
Prompt Engineering
Be Specific: Vague prompts produce inconsistent results
Bad: "a nice landscape"
Good: "mountain valley at sunrise, mist over lake, pine trees, warm golden light, peaceful atmosphere"
Include Technical Details for photorealism:
Camera: "shot on 85mm lens", "wide angle 24mm", "macro photography"
Generation failure: Try simpler prompt or different model
Unsatisfactory result: Offer to regenerate with adjusted prompt
Examples
Example 1: Logo Design
User request: "Create a logo for a coffee shop called 'Morning Brew'"
Expected behavior:
Ask user about style preference (modern, vintage, minimalist, etc.)
Ask about color preferences
Select model (gpt-image-1.5 for text rendering, or gemini-3-pro-image-preview for 4K)
Generate with prompt: "Coffee shop logo for 'Morning Brew', minimalist modern design,
coffee cup with steam forming sunrise rays, warm brown and orange colors,
clean professional aesthetic, vector style, white background"
Use background="transparent" for gpt-image-1.5 for easy placement
Save image and show path
Offer to generate variations with different styles
Example 2: Product Photography
User request: "Generate product photos of wireless earbuds"
Expected behavior:
Select model (imagen-4.0-generate-001 for photorealism, or gpt-image-1.5 for editing)
Generate with prompt: "Wireless earbuds product photography, white background,
professional studio lighting, 3/4 angle view showing charging case and earbuds,
clean minimal composition, high resolution, sharp focus, e-commerce quality"
Generate additional angles if requested
Save all versions
Example 3: Illustration
User request: "Create a cute sticker of a robot"
Expected behavior:
Select model (gpt-image-1.5 with background="transparent" for stickers)
Generate with prompt: "Cute robot sticker, kawaii style, bold black outlines,
cel-shading, pastel blue and silver colors, big friendly eyes, rounded shapes,
chibi proportions, white border, transparent background suitable for sticker"
Save and offer variations
Example 4: Image Editing
User request: "Change the background of this photo to a beach sunset"
Expected behavior:
Use Read tool to load the existing image
Select model (gpt-image-1.5 for best editing with face preservation, or Gemini for chat-based iteration)
Generate with image + prompt: "Change the background to a beautiful beach at sunset,
golden hour lighting, warm colors, ocean and palm trees visible, maintain the subject
in foreground, seamless composition"
Save edited image
Example 5: Iterative Refinement
User request: "Generate a futuristic city" → "Add more neon lights" → "Make it rain"
Expected behavior:
First generation: "Futuristic city skyline, towering skyscrapers, advanced architecture,
night scene, detailed, cinematic lighting"
Use gpt-image-1.5 edit API or Gemini chat interface to maintain context
Second refinement: "Add vibrant neon lights throughout the city, cyberpunk aesthetic,
glowing signs and billboards"
Content Policies: All models have content restrictions (no violence,
explicit content, copyrighted characters, real people without consent)
Text Rendering: Much improved in gpt-image-1.5 and Imagen 4, but very long/complex text may still have issues
Photorealism of People: May not perfectly capture specific facial features; gpt-image-1.5 preserves faces best during edits
Complex Compositions: Very complex scenes may need multiple iterations
Consistency: Hard to maintain exact consistency across multiple generations; use gpt-image-1.5 or Gemini Pro with reference images for character consistency
Real-time Events: Results may not reflect very recent events (use Gemini Pro Search grounding for current topics)
API Costs: Be mindful of usage; see pricing table above
Rate Limits: APIs have rate limits; may need to wait between requests
Imagen Limitations: Text-to-image only (no editing), single image for Ultra model
Watermarks: Google Imagen images include SynthID watermark
Related Skills
python-plotting - For data visualization and charts
brainstorming - For ideating visual concepts
scientific-writing - For figure captions and documentation
python-best-practices - For writing clean API integration code