| name | app-store-screens |
| description | Use when the user asks for app store screens or a task matching the examples below. Generate 5–6 App Store screenshots in a given brand's aesthetic from a `brand.md`, raw product screenshots, or a public App Store listing fetched through Pika MCP. Story-driven (hook → value → features → proof → close), splashy, on-brand. Outputs 1290×2796 PNGs ready to drop into App Store Connect. Use when someone wants App Store / store listing assets — including: "make me app store screenshots", "design app store screens for [brand]", "I have a brand.md and screenshots, generate store assets", "screenshot set for app launch", "iOS store screens", "app store creative", "store listing visuals", "splashy app store screens", "app-store-screens".
|
| argument-hint | <brand-md-or-brand-spec-or-app-store-url> [product-screenshots-or-figma-url] [reference=<url-or-path>] [count=5|6] [--quick] [--config <path>] |
| required-capabilities | ["analyze_media","capture_website","fetch_appstore_screens","generate_image","generate_image_edit","html_to_png","identity_balance","upload_asset"] |
App Store Screens
Take a brand plus real product screens and produce a 5–6 screen App Store campaign at iPhone 6.9" size (1290×2796). The product screens can come from raw exports, Figma/source files, or a public App Store listing fetched through Pika MCP. Story-driven, splashy, strict to the brand.
This is a sister skill to build-a-brand — it consumes that skill's brand.md spec, but works equally well with any brand spec the user supplies.
Cost transparency gate
Before any paid MCP call, call identity_balance({verbose: true}) once. Surface the current balance, recent burn rate, and remaining runway, then gate the run with this exact message:
Estimated cost: about 150-300 credits (~$1.50-$3.00) for a typical 5-6 screen set using GPT-image-2, PNG renders, post-render analyze_media QA, and full-resolution individual PNG QA. This is below $5, but Reply proceed to continue or cancel to stop.
Do not call any paid MCP tool until the user replies proceed. If the user replies cancel, stop without generating. For non-interactive --quick or --config callers, require cost_ack=proceed in the config; if it is absent, stop with the estimate instead of spending credits.
The deliverable
5 or 6 PNGs, numbered, saved to ~/Desktop/[app-name]-app-store-screens/ (or wherever the user prefers):
01_hook.png ← biggest claim, works as a search-result thumbnail
02_value.png ← the one thing the app does best
03_feature_a.png ← specific capability
04_feature_b.png ← another specific capability
05_proof.png ← social proof, awards, "loved by", differentiator
06_close.png ← optional 6th — closer / CTA / brand flourish
Plus a contact sheet (_preview.png) showing all 6 at a glance.
Workflow
Step 0 — Intake and style choice
If invoked with empty args and no usable brand/screenshot/demo context, print this menu verbatim and stop. Do not generate imagery or render HTML until the required inputs are present.
What App Store screenshot set should I make? Required:
- Brand spec —
brand.md or equivalent brand notes with name, palette, fonts, voice, and imagery direction; or an App Store URL/app name if you want me to fetch the listing and draft the brand read first
- Raw product screenshots — exported PNGs, a Figma/source file that can export them, a folder of screenshots, or an App Store URL/app name I can fetch with Pika MCP
- Or a fictional app demo brief — only for launch demos/concepts where no real product exists; say
demo_mode: true and provide what the fake app does
Optional: reference App Store screenshot, moodboard, preferred screen count (5 or 6), output folder.
In interactive mode, if the user has supplied partial input, ask only for the missing required items and stop. If the non-interactive fast lane applies, use Step 0.5 instead. Required inputs:
brand.md or an equivalent brand spec with name, palette, fonts, voice, and imagery direction; if the user gives an App Store URL and asks you to "help create it", fetch the listing first and draft an inferred brand spec from the listing/icon/screenshots for approval
- one product source: raw product screenshots, a Figma/source file that can export them, or an App Store URL/app name that can be fetched through Pika MCP
- for fictional launch demos only,
demo_mode: true with demo_brief is an alternative to real product screenshots; use the Fictional app demo mode path below
- optional reference screenshot or moodboard if they want a specific App Store style
Step 0.5 — Non-interactive fast lane
Use this path when the caller passes --quick or --config <path>, or when the
caller states they are running from CI, a subagent, a batch job, or any other
non-interactive harness.
This section has precedence over the interactive ask/wait instructions below.
When it applies, use this fast lane and do not fall through to the multi-turn
intake unless the required brand and either product screenshots or explicit demo
mode inputs are truly unavailable.
--config <path> points to a JSON file with pre-baked choices: brand_spec,
product_screenshots, app_store_url, website_url, reference, style,
screen_count, narrative_arc, demo_mode, demo_brief, and
output_folder.
--quick means choose the default style unless a reference is supplied, infer
the brand read from the provided brand spec or App Store listing, draft the
5-6 screen arc yourself, and proceed.
- For
--quick or --config, do not stop for confirmation at the style choice,
brand-read playback, reference-rule playback, or 5-6 screen strategy pitch.
Record assumptions inline and continue to design/render.
- If neither real product screenshots nor explicit
demo_mode: true with
demo_brief is available, stop once with a single compact missing-fields list
instead of starting a multi-turn Q&A loop.
Fictional app demo mode
Use this only when the caller explicitly says this is a fictional app, fake app,
launch demo, concept demo, or passes demo_mode: true in config. Do not use demo mode for real products.
Demo mode is allowed to create representative UI mocks in the brand voice when no
real product screenshots exist, but the output must be labeled as demo-only
concept work, not production assets. Treat the invented screens as a storyboarded
QRT/demo artifact, not App Store Connect-ready evidence of a real product.
- Require a
demo_brief or enough user-provided product concept detail to define
the app's core job, audience, 3-5 features, and proof/CTA angle.
- Generate UI states that are internally consistent with that brief; do not imply
real customers, real reviews, real metrics, or real integrations unless the
prompt explicitly provides them.
- Add a demo-only disclosure in the delivery notes and contact sheet label:
"Demo UI concept — not production assets and not real product screenshots."
- In non-interactive mode, proceed only if config sets
demo_mode: true and
provides demo_brief; otherwise stop with a compact missing-fields list.
In interactive mode, when the required inputs are present, open with a brief agenda and ask one upfront style question. This is the cheapest moment to learn whether the user wants the default or a specific reference. In the non-interactive fast lane, choose the default style unless reference or style is supplied, record that assumption inline, and continue.
here's how this works:
1. **Read your brand + screenshots** — i absorb the brand voice, palette, type, and figure out what the app actually does
2. **Strategy** — i pitch a 5–6 screen narrative arc (hook → value → features → proof → close) with headlines for each, before any design
3. **Design + generate** — i design each layout, composite at 1290×2796
4. **Preview** — i deliver the PNGs + a contact sheet so you can see the whole campaign at once
quick question before i start — **do you want the default style, or something specific?**
- **Default** — clean, restrained, screens shown untouched. Full-frame device, headline + sub above, solid brand-color backgrounds, brand-voice copy. Codified in `references/default-layout.md`. This is what i deliver well.
- **Something specific** — share a clear reference: a screenshot of an App Store page you love (Notion, Calm, Things, Headspace, anything), a figma file, a moodboard image. I'll study it, extract the compositional rules (device size + position, headline treatment, background, any signature flourishes), pitch back what i read before designing, then build the campaign — **with your brand's palette, fonts, voice, and photography style swapped in**. So if the reference uses an over-the-shoulder photo, i'll generate one in your brand's photo style (`gpt-image-2` accepts the reference as a style input). The reference dictates the composition; your brand dictates the look.
if you don't tell me, i'll go with default.
Why no "jazzed up" auto-mode. Doing rich/varied/dramatic compositions well requires visual-judgment calls (perspective, layering, typography hierarchy, color balance) that don't have a programmatic answer. When users want richness without a reference, my approximations tend to look amateur. Asking for a concrete reference lets me extract specific rules to replicate instead of inventing freely — and the user has a way to verify i'm aiming at the right thing.
Two downstream paths based on the answer:
- Default → follow
references/default-layout.md end-to-end. Same composition skeleton on every screen; variety from color + content. This produces consistent, brand-disciplined work.
- Replicate a reference → study the reference, extract the rules, swap in the user's brand. See "Reference-driven path" below for the full process.
If the user can't or won't supply a reference but still wants more than default, push back gently — explain that without a reference you'll deliver default plus their brand color, and that's better than a guessing-game iteration loop. Don't invent a "jazzed" style on the fly.
Step 1 — Read the brand and the product
App Store listing path
If the user supplies an apps.apple.com URL, numeric App Store app ID, or app-name search term as the product source, try Pika MCP fetch_appstore_screens first because it is faster, more stable, and returns hosted screenshot/icon assets ready for later render steps:
fetch_appstore_screens(
query: <app_store_url | numeric_app_id | search_term>,
country: "us",
max_screens: 10,
include_icon: true
)
Use the returned metadata, icon, and screenshots as the product source. The returned screenshot url values are already Pika-hosted HTTPS assets and can be used directly in later html_to_png stages.
If the user provided a country-specific App Store URL, preserve that storefront country when calling the tool when possible. If the country is unclear, default to "us" unless the user asked for another storefront.
If the MCP tool is unavailable, unauthenticated, or returns no screenshots, say what happened and then use the least fragile fallback available: other read/fetch tools, official App Store/iTunes metadata endpoints, or user-provided screenshots. Keep the fallback grounded in real listing assets; do not invent product UI.
Expected result shape:
{
"app_url": "https://apps.apple.com/...",
"metadata": { "name": "...", "subtitle": "...", "description": "...", "category": "...", "icon_url": "https://..." },
"icon": { "url": "https://cdn.pika.art/...", "source_url": "https://is...mzstatic.com/...", "filename": "appstore-icon.png", "mime_type": "image/png", "width": 1024, "height": 1024 },
"screenshots": [
{ "url": "https://cdn.pika.art/...", "source_url": "https://is...mzstatic.com/.../1290x2796bb.png", "filename": "appstore-screen-01.png", "mime_type": "image/png", "width": 1290, "height": 2796 }
],
"count": 1
}
Source screenshot prep
Before choosing the default phone layout, classify each fetched or user-supplied source screenshot and record the source_treatment you will use:
clean_ui_capture - raw in-app UI suitable for a device mockup.
composed_marketing - a finished App Store marketing screen, not a clean in-app screenshot.
legacy_footer_cleanup - clean UI that only needs listing-brand footer or watermark cleanup before device embedding.
A composed-marketing source is a pre-composed marketing screen. Common signals: a baked headline or subhead already sits above the phone, the phone bleeds or clips against the source edge, the background is already campaign art, and the source image is already a finished 1290x2796 App Store screen rather than clean product UI. The 2025 Notion listing uses this pattern.
When a source is composed_marketing, do not reframe it inside another device mockup and do not add, write, or generate a second headline/subhead on top of the baked headline. Prefer a clean_ui_capture instead: use capture_website on the product website, onboarding, or web app when a captureable real UI surface exists, or ask for simulator/Figma/raw UI exports. If no clean UI surface is available, use source_treatment=composed_passthrough: present the composed source as-is as the screen/background with only minimal delivery framing, numbering, or contact-sheet labeling. Do not rely on a bottom 6-8% crop to fix composed sources; it removes the wrong area and leaves the double-headline failure intact.
For legacy source screenshots that are otherwise clean UI, inspect them for footer watermark bleed or listing-brand footers that will become clutter inside the generated campaign. The R4 Notion listing exposed this as a raw footer reading NOTION · NOTES, TASKS, AI overlapping the bottom content. If a fetched or user-supplied clean UI source screenshot has this kind of footer watermark, crop or mask the bottom 6-8% before embedding it; do not place the raw screenshot directly into the device.
Use a prepared screenshot wrapper in the HTML stage and keep the top of the source anchored. The wrapper must define its own geometry; do not rely on height:100% unless every parent up to the device has an explicit height.
<div class="source-screen-crop source-screen-crop--phone" style="
width:100%;
aspect-ratio:1290 / 2796;
overflow:hidden;
position:relative;
">
<img src="s://screen01" style="
width:100%;
height:calc(100% / 0.92); /* shows the top 92% while cropping the bottom 8% */
object-fit:cover;
object-position:top center;
display:block;
">
</div>
If an 8% crop would remove real product UI, use a brand-color fade over the bottom footer instead, but still remove the watermark bleed before the screenshot enters the final device mockup.
Handling Mac-app or hybrid-app screenshots
fetch_appstore_screens may return landscape Mac screenshots (for example 1290×806) instead of iPhone portrait screenshots (1290×2796) for Mac-only or hybrid iOS/Mac apps such as Things 3, Drafts, Numbers, Day One, OmniFocus, or BBEdit. Detect this before choosing the default iPhone-portrait layout:
The orientation rule is aspect_ratio_w > aspect_ratio_h, or width > height when only pixel dimensions are present.
mac_shots = [s for s in screenshots if s.aspect_ratio_w > s.aspect_ratio_h]
# Equivalent when explicit aspect ratios are not present:
mac_shots = [s for s in screenshots if s.width > s.height]
If mac_shots is non-empty, use the embedded-card layout in references/mac-app-layout.md rather than full-frame device shots. For Mac-only or hybrid apps, embed each landscape screenshot as a rounded-corner UI card inside the 1290×2796 portrait canvas rather than cropping it into an iPhone frame, letterboxing it, or treating it as a full-frame device. The Things 3 round-2 benchmark validated this approach: the Mac screenshots stayed readable and the portrait campaign still fit App Store Connect.
Source screenshot prep still applies to Mac cards: inspect landscape screenshots for listing-brand footers or footer watermark bleed before embedding them. Do not use the phone portrait crop on Mac UI. Keep the landscape aspect ratio with a dedicated card wrapper such as source-screen-crop--mac; if cleanup is needed, prefer a narrow bottom mask/fade or a small landscape-preserving bottom crop inside the card.
Thin App Store listing guardrail
After fetch_appstore_screens, count usable screenshots that show real
product UI. If the listing returns fewer than 3 real product screenshots, treat it
as a thin App Store listing.
- Do not create a 5-6 screen campaign by hallucinating UI. Real product UI is the
default requirement for device frames.
- Interactive mode: stop before strategy and offer two choices:
- Website-capture path — use
capture_website or supplied
website/onboarding URLs to capture real product surfaces, then continue with
those captures as product screenshots.
- Real screenshot path — ask for simulator exports, Figma frames, or other
product UI captures before continuing.
Do not offer synthesized device UI for real products. If the user explicitly
pivots to a fictional launch/concept demo, route to Fictional app demo mode and
require
demo_mode: true with demo_brief.
- Non-interactive fast lane: prefer the website-capture path when
website_url
or an obvious product website is available. If there is no captureable product
surface, stop once with a compact missing-fields list unless config explicitly
sets demo_mode: true and provides demo_brief.
After fetching, infer only a draft brand read from the listing and visuals: app name, category, visible palette, likely type direction, voice from subtitle/description, and notable UI moments. In interactive mode, read it back as a provisional brand spec and ask the user to correct it before pitching the 5-6 screen arc. In the non-interactive fast lane, record it as the provisional brand spec and continue.
Reference-driven path: how to apply the user's brand to the reference's style
The reference describes WHAT THE LAYOUT/STYLE LOOKS LIKE. The brand describes WHAT COLORS/FONTS/VOICE/PHOTOGRAPHY TO USE. The skill's job is to combine them: replicate the reference's visual structure, but render it with the user's brand. Two things to extract from the reference, two from the brand:
From the reference, extract structural rules:
- Device size (as % of canvas) and tilt angle
- Headline treatment (size, position, color logic — accent on key word? tinted bg pill?)
- Background treatment (solid color, photo, gradient, abstract elements)
- Hero imagery pattern (hand holding phone, over-the-shoulder, real subject breaking out of phone, 3D objects floating, etc.)
- Callout / pull-out pattern (speech bubbles with arrows, floating cards, polaroid frames, etc.)
- Repetition rules: same layout every screen, or varied per slot?
From the brand, apply specifics:
- Palette (substitute brand colors wherever the reference uses solid color blocks)
- Fonts (substitute brand display + body fonts everywhere the reference uses type)
- Voice (rewrite all headlines/subs in brand voice — don't copy the reference's words)
- Photography direction (any generated imagery follows the brand's photo rules — for DeltaStream that's documentary 35mm, golden hour, real apartments, butter accent in every frame, real cast diversity)
- Mood (warm vs. clinical, playful vs. expert, etc.)
For generated hero imagery: pass the reference as images to gpt-image-2 via generate_image_edit. The tool supports up to 16 reference images. Use this when the reference uses a distinctive photo composition (hand-holding-phone, over-the-shoulder, person breaking out of screen, 3D character emerging). Combine with the brand's photography rules in the prompt:
Reference: [user's reference photo — e.g., insect app's hand-holding-phone with butterfly]
Prompt: "In the EXACT composition and lighting style of the reference image — hand
holding a phone at the same angle and scale — but render the subject and scene
per [BRAND] photography rules: documentary 35mm, golden hour window light, real
apartment, butter-yellow ceramic mug visible, mid-30s mixed-race cast, 35mm film
grain. Vertical 9:16 portrait."
This is how you get the reference's STYLE without copying its CONTENT. The hand+phone composition transfers; the lighting/cast/setting comes from the brand.
Pitch the extracted rules back before designing in interactive mode. Write them out as a short bullet list ("device 75% canvas tilted -8°, headline-on-yellow-pill above device, over-the-shoulder hero, ink callouts with curved arrows") and confirm with the user that you read the reference correctly. In the non-interactive fast lane, record the extracted rules inline and continue. Then build all 6 screens applying those rules consistently. Don't deviate mid-campaign.
Then actually read:
brand.md — extract: brand name, tagline, palette (with hex), display font + body font, voice adjectives, voice examples, forbidden words, photography/illustration direction, mood words. If the file is a different format (PDF, plain notes), parse what's there and ask about gaps in interactive mode. In the non-interactive fast lane, record reasonable assumptions for non-critical gaps — don't invent a brand.
- Product screenshots — open each one and form an honest mental model: what does this app actually do? Note the core UI patterns (feed, chat, canvas, list, map, etc.), the primary action surface, and any "wow moment" screens (a generative result, a beautiful state, a unique interaction).
- If you can use
analyze_media to inspect screenshots without loading them as images, do — it's faster for a quick scan.
Then read back what you found (3-5 lines, conversational):
ok — reading [brand name]: [tagline]. palette: [colors]. voice: [adjectives]. the app looks like
it does [X] — i see [specific UI cue 1] and [specific UI cue 2]. the standout screen is [screen
N] because [reason]. correct me where i'm wrong, otherwise i'll pitch the 6-screen arc.
Interactive mode: wait for confirmation before pitching strategy. Catches misreads cheaply. In the non-interactive fast lane, treat the brand read as provisional, record the assumption inline, and continue to the strategy.
Step 2 — Pitch the 6-screen strategy
Before designing anything, write out the narrative arc as plain text. Each screen needs:
- Role (hook / value / feature / proof / close)
- Headline (5–8 words, in brand voice — see "Copy" below)
- Optional subhead (one short line)
- Layout archetype (see below — name it, don't draw it yet)
- Which raw screenshot it features (filename) — or "none, full-bleed typography"
- Proof artifact for proof screens only — at least one visible proof artifact such as a star row, Editors' Choice badge, testimonial card, review-count chip, award badge, or other sourced visual proof. The proof screen must not be text-only.
Present it like:
**Screen 1 — HOOK**
Headline: "Sleep like you mean it."
Sub: (none — let the headline carry)
Layout: full-bleed UI with bold typography overlay