| name | aihub-app-ideation |
| description | Generate detailed, solo-developer-focused app/web product ideas from AI-Hub Korea dataset metadata. Use when the user wants to brainstorm app ideas from a dataset, asks "์ด ๋ฐ์ดํฐ์
์ผ๋ก ๋ญ ๋ง๋ค ์ ์์ด / ์ฑ ์์ด๋์ด ๋ฝ์์ค", "what can I build with dataset X", or wants to judge feasibility and rank ideas. Always produces at least 5 ideas per dataset with a rich per-idea spec, metadata-only. |
AI-Hub App Ideation Skill
Purpose
Turn AI-Hub Korea dataset metadata into concrete, buildable app/web product
ideas โ at least 5 per dataset โ each with a detailed spec and an honest
feasibility judgement. This skill is the deep, structured version of the
ideate CLI command in this repo.
Use it when the user wants to explore "what could I build with this dataset?"
Hard Requirements
- Minimum 5 ideas per dataset. Never return fewer. If the metadata is thin,
still produce 5 and mark the weak ones
no-go with the reason.
- Metadata only. Reason strictly from title, inferred tags, file tree, file
sizes, train/val split, label/source presence. Never claim dataset contents,
label schemas, licenses, or quality you cannot see. State that every judgement
is a metadata-only estimate.
- Each idea uses the full per-idea template below. "More detailed than a
one-liner" is the bar โ no bare titles.
- Honest feasibility. Score and verdict must reflect reality, including the
solo-developer lens below. Do not inflate.
- Always scan competitor apps. Every idea must include a real competitor
check โ never assume a niche is empty. Search the App Store / Google Play and
the web for existing apps (Korean AND global/English), name 2โ4 concrete ones
with what they do and where they fall short, and let that gap drive the
์ฐจ๋ณํ ๋ฌด๊ธฐ. "No competitor found" is a valid result only after an actual
search, and is itself a warning (small/non-existent market), not a green light.
Default Lens: Solo Developer Building Apps
Unless the user says otherwise, assume the builder is one person shipping B2C
apps (mobile/web, self-serve distribution). Score ideas through this lens:
- Favor: image classification / text tasks that run cheaply or on-device;
clear consumer hook; simple monetization (subscription / IAP); Korea-specific
data as a moat against global apps.
- Down-rank (usually unrealistic for a solo app builder): anything needing
B2B/B2G sales, on-prem hardware/sensor/factory/field access, or heavy
regulation/liability (medical diagnosis, face recognition, financial advice).
Examples to treat as low-priority for this persona: ์ ๋ ฅ์ค๋น, ๊ณต๊ณต๊ธฐ๊ด ๋ฉํ,
ํนํ/์ง์์ฌ์ฐ ์ ๋ฌธํด, ์ ์กฐ ๋ถ๋๊ฒ์ถ ๋ผ์ธ, ์์ฑ/๊ตํต ์ธํ๋ผ.
- If an idea is attractive but unrealistic for a solo builder, say so explicitly
rather than silently ranking it high.
When the user states a different persona (enterprise, team, research), adapt the
lens accordingly and say which lens you used.
Workflow
1. Acquire dataset metadata
Prefer the repo's own pipeline so reasoning is grounded in real parsed metadata:
./run.sh inspect --datasetkey <KEY>
Or reuse an existing summary at data/normalized/datasets/<KEY>.json. If neither
is available (no aihubshell), work from whatever metadata the user provides, and
note that keys/sizes are unverified.
Read from the summary: title, category_guess, modality_guess, tags,
file_count, human_size, sample_file_paths, metadata_lines,
parse_status, train/val and label/source presence.
2. Classify the dataset
State in one or two lines: domain, modality (or "unknown โ verify by opening one
source zip"), rough size class, and supervised-readiness (labels + train/val).
Flag uncertainty honestly (e.g. "ํตํฉ ๋ฐ์ดํฐ" โ modality could be images or sensors).
3. Generate at least 5 ideas
Spread ideas across different angles โ don't give five variants of one app.
Cover a mix of: consumer utility, creator/content, education/learning,
wellness/care, and a "data-as-moat" or niche angle. Each idea uses the template.
4. Competitor / market scan (required)
For each idea, do a real search before scoring โ do not reason from memory alone:
- Search app stores and the web for existing apps, both Korean (์ ์ , ํฌ์คํ
๋ฌ,
๋ค์ด๋ฒ/์นด์นด์ค ๋ฑ) and global/English. Useful queries: " app",
" ์ฑ", " App Store / Google Play".
- If the user has a market-scouting repo (e.g.
finding-cash-cow-android โ
Play Store scraping/analysis), reuse its data when it's in session scope.
- Record 2โ4 concrete competitors per idea: name, platform, what they do, price/
monetization if visible, and the gap they leave. Note incumbents that make
a head-on solo entry unrealistic (e.g. a category leader with huge revenue).
- Turn the gap into the idea's
์ฐจ๋ณํ ๋ฌด๊ธฐ, and let competitive density feed the
opportunity_score (crowded + strong incumbents โ lower; real unmet gap โ higher).
- An honestly empty niche is a yellow flag (likely small market), not a win.
5. Score and judge each idea
Give three 1โ10 scores and a verdict:
opportunity_score โ market pull / consumer demand
feasibility_score โ how realistically a solo dev ships it (infra, cost, skill)
data_fit_score โ how well the dataset actually supports the idea
feasibility: go | maybe | no-go
6. Rank and recommend
Rank the ideas (a simple combined weighting: opportunity 0.45, feasibility 0.35,
data-fit 0.20 โ same as the repo's combined_score). Recommend the top 1โ2 for a
solo builder and say why, plus the single biggest risk for each.
Per-Idea Template (use all fields)
### <n>. <์ฑ ์ด๋ฆ> โ <go|maybe|no-go>
- ํ ์ค ์๊ฐ: <one-line pitch>
- ํํ: app | web | both
- ํ๊ฒ ์ฌ์ฉ์: <who, specifically>
- ํต์ฌ ๊ธฐ๋ฅ: <3โ5 features>
- ๋ฐ์ดํฐ ํ์ฉ: ์ด ๋ฐ์ดํฐ์
์ ๋ฌด์(๋ผ๋ฒจ/๋ชจ๋ฌ๋ฆฌํฐ/๊ตฌ์กฐ)์ ์ด๋ค ๋ชจ๋ธ ํ์คํฌ๋ก ์ฐ๋์ง
- ๋ชจ๋ธ ยท ์ถ๋ก ๋น์ฉ: ์จ๋๋ฐ์ด์ค ๊ฐ๋ฅ ์ฌ๋ถ / ํ์ต ํ์ ์ฌ๋ถ / ์๋ก ์์ฐ์์ ํ์ค์ ์ธ์ง
- ์๋ก ์คํ์ฑ: ํผ์ ์ถ์ ๊ฐ๋ฅํ ์ด์ ๋๋ ๋งํ๋ ์ง์
- MVP ํ๋ฆ: ํ๋ฉด/๋จ๊ณ 3โ5๊ฐ
- ์์ต ๋ชจ๋ธ: ๊ตฌ๋
/ IAP / ๊ด๊ณ / ์ ํด
- ๊ฒฝ์ ์ฑ: ์ค์ ๊ฒ์์ผ๋ก ์ฐพ์ ๊ธฐ์กด ์ฑ 2โ4๊ฐ (๊ตญ๋ด+๊ธ๋ก๋ฒ) โ ์ด๋ฆยทํ๋ซํผยทํ๋ ์ผยท๋นํ. ์์ผ๋ฉด "๊ฒ์ํ์ผ๋ ์์(=์์ฅ ์์ ์ํ)"
- ์ฐจ๋ณํ ๋ฌด๊ธฐ: ์ ๊ฒฝ์ ์ฑ์ด ๋ชป ์ฑ์ด ๋นํ์ ์ด๋ป๊ฒ ๊ณต๋ตํ๋๊ฐ (ํนํ ํ๊ตญ ํนํ ์ฐ์)
- ๋ฆฌ์คํฌ: 2โ3๊ฐ + ํํผ์ฑ
- ์ ์: ๊ธฐํ N/10 ยท ์คํ N/10 ยท ๋ฐ์ดํฐ์ ํฉ N/10
Multiple datasets
When given several datasets, ideate each (โฅ5 ideas each), then build a
cross-dataset ranking: which dataset offers the strongest solo-app opportunity,
using each dataset's best idea. The repo command for the aggregate view:
./run.sh ideate-rank
Using the automated pipeline
The repo can generate and persist this automatically (LM Studio / OpenAI /
Anthropic providers):
./run.sh ideate --datasetkey <KEY> --ideas 5
./run.sh ideate --datasetkey <KEY> --provider anthropic
./run.sh ideate-rank
Artifacts: data/normalized/ideation/<KEY>.json and
data/generated/ideation/<KEY>-<slug>.md. When you run ideation manually (no
LLM key), follow the same schema and template so output stays consistent with the
pipeline. The pipeline prompt also enforces the โฅ5-idea and solo-lens rules.
Cautions
- License, approval scope, and policy fit are out of scope for metadata โ always
flag them as "verify separately".
- Mental-health / medical / safety ideas: position as journaling/logging/wellness,
never diagnosis; add a crisis-resource / disclaimer note.
- Keep scores honest; a list where everything is
go is a failed analysis.
Before Finishing
- Confirm every dataset got โฅ5 ideas, each with the full template.
- Confirm each idea has a competitor scan from a real search (โฅ2 named apps or
an explicit "searched, none found" note).
- Confirm the solo-app lens was applied (or the chosen lens was stated).
- End with a ranked recommendation and the top risk per recommended idea.