Skip to main content

github-trending

Curated trending across GitHub repos and the Hugging Face Hub (models, datasets, spaces) - filtered, clustered, and labeled by momentum with a one-line why-notable per pick.

跳到安装

来源信息

仓库
aeonfun/aeon
最近来源活动
2026年9月11日 15:30
检测到的 SKILL.md 语言
英语
星标
738
分支
265

安装方式

默认使用会先检查来源的 Prompt;你也可以切换为直接命令,或下载本地副本。

检查来源文件

决定是否安装前,请先阅读 SKILL.md,以及 SkillsMP 当前展示的配套文件。

正在显示 SKILL.md

SKILL.md
来源说明 · 只读预览
name
github-trending
description
Curated trending across GitHub repos and the Hugging Face Hub (models, datasets, spaces) - filtered, clustered, and labeled by momentum with a one-line why-notable per pick.
metadata
{"title":"GitHub Trending","mode":"read-only","category":"basics","var":"","tags":["dev","research"]}
<!-- autoresearch: variation B — sharper output via curation, clustering, "why notable" gate, momentum tags --> > **${var}** — Source selector plus optional sub-scope: > - empty or `github` → **GitHub trending**, all languages (default) > - `github:<lang>` — or a bare language token like `python`, `typescript`, `rust` (backward-compatible with the old GitHub var) → GitHub trending filtered to that language > - `hf` or `huggingface` → **Hugging Face trending** across models + datasets + spaces > - `hf:models` / `hf:datasets` / `hf:spaces` (also `huggingface:models`, etc.) → Hugging Face trending scoped to a single resource type This skill covers two neighbouring layers of where developer/AI attention is moving today: the **repo layer** (GitHub trending) and the **artifact layer** (Hugging Face Hub — the models, datasets, and spaces that ship alongside, and frequently before, the paper). Both branches share the same contract: don't dump the top 10 (the source's own front page already does that) — deliver a **curated** slate of 5–8 picks a busy reader would actually want to click, grouped by category, with a one-line "why notable" and a momentum tag per pick. ## Shared preamble (run for every invocation) Read `memory/MEMORY.md` for context. Read the last 3 days of `memory/logs/` to dedupe items you've already featured (the GitHub branch dedupes against the last **2** days, the Hugging Face branch against the last **3** — see each branch's filter step). Read `soul/SOUL.md` + `soul/STYLE.md` if populated to match voice. **Parse `${var}` into a source + optional sub-scope** (deterministic): 1. If `${var}` is empty → **GitHub branch**, no language filter. 2. Otherwise trim + lowercase and split on the first `:` into `head` and optional `tail`. 3. `head` ∈ {`hf`, `huggingface`} → **Hugging Face branch**. If `tail` is present it must be one of `models` / `datasets` / `spaces` (that becomes the resource sub-scope); any other `tail` → exit `HF_TRENDING_BAD_VAR` (no notify). No `tail` → pull all three resource types. 4. `head` == `github` → **GitHub branch**. If `tail` is present, it's the language filter. 5. Any other value (no colon, `head` not `hf`/`huggingface`/`github`) → **GitHub branch**, treating the whole `${var}` as the language filter (e.g. `rust`). Then jump to the matching branch below and run it end to end. --- ## Branch A — GitHub trending (source = `github`) Don't just dump the top 10 trending repos — GitHub already shows that. Deliver a **curated** slate of 5-8 repos that a busy dev would actually want to click, grouped by category, stripped of noise, with a one-line "why notable" per pick and a momentum tag. ### A1. Fetch candidates Fetch the daily trending page via **WebFetch** (it renders the HTML for you; `curl` works too — there is no network sandbox): ``` https://github.com/trending?since=daily ``` If a language filter was resolved from `${var}`, append the language segment: `https://github.com/trending/<lang>?since=daily`. Extract for each of the ~25 returned repos: - `owner/repo` - one-line description - primary language - stars today (the "X stars today" widget) - total stars - URL ### A2. Enrich with velocity metadata (supplementary) For the 10-15 repos that survive the filter in step A3, try to enrich with **stars-per-day since creation** using `gh api` (handles auth internally, so no token touches the command line): ```bash gh api "repos/OWNER/REPO" --jq '{created_at, stargazers_count, pushed_at}' ``` Compute `velocity = stargazers_count / max(days_since_created, 1)`. If `gh api` fails for a repo, skip enrichment for that one — it's not required, just informative. > Read-only note: this skill runs `read-only`, so `gh api` (and any repo mutation) may be stripped from your toolset. If `gh api` is unavailable, skip enrichment entirely and rely on the "stars today" widget; velocity-dependent tags degrade gracefully (see A5). ### A3. Filter noise (required) **Drop** any repo matching these patterns — they're low-signal for a dev audience: - **Meta-lists**: repo names containing `awesome-`, `awesome_`, `-list`, `free-`, `public-apis`, `interview-`, `cheatsheet`, `resources` - **Bare tutorials / learn-X**: names starting with `learn-`, `build-your-own-`, `30-days-of-`, `X-in-Y`, `hello-world-*` - **Non-code bundles**: dotfiles, config dumps, blog-source repos (check description for "my personal blog", "my dotfiles") - **Low-activity**: stars today < 50 AND not new this week (created > 14 days ago) - **Already featured**: repo appeared in `memory/logs/YYYY-MM-DD.md` in the last 2 days If a repo *barely* fails a filter but is genuinely technically interesting (novel algorithm, new runtime, new framework), you may keep it — note it as a judgment call. ### A4. Require a "why notable" for each survivor For every repo that survives filtering, write **one line** (≤ 18 words) explaining *why a dev should care today*. No paraphrasing the description. Good: *"Replaces Electron with native webview bindings — ships a 3MB hello-world instead of 120MB."* Bad: *"A new framework for building desktop apps."* (that's just the description) If you can't write a concrete "why notable" line, **drop the repo**. The filter is the feature. ### A5. Tag momentum Tag each surviving repo with one of: - **DEBUT** — created within the last 14 days (first-time trending) - **ACCELERATING** — velocity > 50 stars/day AND total stars > 500 AND older than 14 days - **RETURNING** — older repo (> 90 days) trending again; note this means a release, a viral post, or a HN moment - **HOLDOVER** — appeared in yesterday's logs (use sparingly; prefer to drop) ### A6. Cluster into categories Buckets are **heuristic and author-inferred** — classify by the repo's primary utility, not by author self-description. Cap total buckets at **5** (merge adjacent ones if you hit 6+; e.g. fold Data into Infra). Group survivors into these buckets (omit empty ones): - **AI/ML** (models, inference, agents, training, prompts) - **Devtools** (CLIs, build systems, dev servers, debuggers, IDEs) - **Infra** (databases, networking, observability, orchestration) - **Web/Apps** (frameworks, UI libs, user-facing apps) - **Data** (pipelines, analytics, notebooks, viz) - **Other** — if a repo fits none of the above, put it under Other with a **one-line reason** why none of the named buckets fit. Keep Other tight; if Other ≥ 3, reconsider whether your buckets fit. Aim for 5-8 total picks. If fewer than 3 survive, send a short note (see step A8) rather than padding. ### A7. Lead with a top pick Pick the single most interesting survivor (highest-signal regardless of category) as *"Top pick"*. One sentence on why it's the top pick — not the "why notable" line, a higher-level framing. ### A8. Notify Send via `./notify`: ``` *GitHub Trending — ${today}* *Top pick* — [owner/repo](url) One-sentence framing of why this is the standout today. *AI/ML* • [owner/repo](url) — ★ Xt today (Yk total) · LANG · [TAG] why notable (one line) • [owner/repo](url) — ... *Devtools* • ... --- sources: trending=ok|fail · gh_api=ok|fail · kept N/M ``` Replace `Xt` with stars today, `Yk` with total stars in thousands, `[TAG]` with DEBUT/ACCELERATING/RETURNING/HOLDOVER. **Slate-integrity check (before you send).** The workflow captures this notify body verbatim to `output/.chains/github-trending.md`, which `vuln-scanner` reads for `owner/repo` scan targets. So every pick must stay a `[owner/repo](url)` line (a bare `https://github.com/owner/repo` permalink also parses). **Never** collapse the slate into a prose name list (e.g. "picks: OmniRoute, colibri, ...") - a bare repo name with no owner is unparseable and starves the scanner. Before sending, confirm the body carries one `[owner/repo](url)` line per surviving pick. ### A9. Log and exit Append to `memory/logs/${today}.md` under a single `### github-trending` heading, with a discriminator line `- branch: github` as the first bullet, followed by: - picked repos (owner/repo + tag) - dropped-for-noise count - source status - any judgment-call keeps (noted in step A3) **Exit codes:** - `GITHUB_TRENDING_OK` — fetched successfully, 0 or more picks sent - `GITHUB_TRENDING_ERROR` — trending page fetch failed AND `gh api` fallback also empty If the trending fetch fails, try one fallback before erroring: `gh api "search/repositories?q=created:>$(date -d '7 days ago' +%Y-%m-%d)+stars:>100&sort=stars&order=desc&per_page=25"` then run steps A3-A8 on those results (skip the "stars today" field — use velocity instead). If both fail, log `GITHUB_TRENDING_ERROR` with the failure reason and send a brief notify: *"GitHub Trending — sources unavailable today."* If fetch succeeds but every repo fails filters (rare but possible on slow days), send a short note: *"GitHub Trending — quiet day, nothing above the noise floor."* and exit OK. --- ## Branch B — Hugging Face trending (source = `hf`) Today is ${today}. The Hugging Face Hub is where new AI artifacts land first — models hours after a paper, datasets before they get cited, spaces as the first runnable form of a technique. The Hub's own front page lists "trending" but doesn't filter the noise (test models, gated previews, redundant fine-tunes of the same base). This branch mirrors the GitHub contract for the AI ecosystem: don't dump the top 10, deliver a **curated** slate of 5–8 picks a busy AI/dev reader would actually want to click, with a one-line "why notable" each. ### B1. Fetch candidates The Hugging Face Hub REST API is fully keyless for the list endpoints used here. Pull trending across all three resource types unless the resolved sub-scope narrows it: ```bash # Models — sort=trendingScore returns the same ranking that backs the HF front page curl -sf "https://huggingface.co/api/models?sort=trendingScore&direction=-1&limit=20" \ -H "accept: application/json" \ -H "user-agent: aeon/1.0 (+https://github.com/aeonfun/aeon)" \ > /tmp/hf-models.json # Datasets curl -sf "https://huggingface.co/api/datasets?sort=trendingScore&direction=-1&limit=15" \ -H "accept: application/json" \ -H "user-agent: aeon/1.0 (+https://github.com/aeonfun/aeon)" \ > /tmp/hf-datasets.json # Spaces curl -sf "https://huggingface.co/api/spaces?sort=trendingScore&direction=-1&limit=15" \ -H "accept: application/json" \ -H "user-agent: aeon/1.0 (+https://github.com/aeonfun/aeon)" \ > /tmp/hf-spaces.json ``` If the sub-scope is `models` / `datasets` / `spaces`, fetch only that endpoint. If any `curl` fails (a flaky public GET), use **WebFetch** as a fallback for the same URL. WebFetch parses the JSON for you. If both fail across all three resources (or the single one selected by the sub-scope), log `HF_TRENDING_ERROR` with the failure detail, send a brief notify (*"Hugging Face Trending — sources unavailable today."*), and exit. For each entry extract: - `id` (always present, format `owner/name`) — split on `/` to get author + name - `likes`, `downloads` (models/datasets only, spaces have no `downloads`), `trendingScore` - `tags` (filter out `region:*`, `license:*`, and storage-format noise like `endpoints_compatible`, `safetensors`, `gguf`) - `pipeline_tag` (models) — the canonical task label (e.g. `text-generation`, `text-to-image`) - `library_name` (models) — `transformers`, `diffusers`, `mlx`, etc. - `sdk` (spaces) — `gradio` / `streamlit` / `docker` / `static` - `createdAt`, `lastModified` (when present) - Resource type (`models` / `datasets` / `spaces`) — preserve so the renderer can pick the right footer - Permalink: `https://huggingface.co/{id}` for models, `/datasets/{id}` for datasets, `/spaces/{id}` for spaces ### B2. Filter noise (required) Drop entries matching these patterns — they're low-signal: - **Test / debug artifacts**: `id` containing `-test`, `-debug`, `-tmp`, `-scratch`, `-playground`, or starting with `test-` / `debug-` - **Gated / private preview shells**: entries flagged `gated: true` *and* with `<10` likes (HF gates lots of legit work, but a gated artifact with no community signal is usually a draft) - **Trivial fine-tunes**: model `id` ending in `-finetune`, `-ft`, `-lora-test`, or with `<5` likes AND `<100` downloads (real momentum picks both) - **Already featured**: anything that appeared in `memory/logs/YYYY-MM-DD.md` for the last 3 days - **Quantization-only forks**: `id` ending in `-gguf`, `-awq`, `-gptq`, `-int4`, `-int8`, `-fp8` *unless* it has `>500` likes — quantizations of a base model are useful but rarely the most interesting story; the base usually carries the narrative - **Spaces with `runtime.status: ERROR`** if the field is present (broken demos shouldn't be recommended) - **Spaces called "demo"** or "example" with `<20` likes — boilerplate scaffolds If an entry barely fails a filter but is genuinely interesting (novel architecture, first-of-kind dataset, reference implementation of a fresh paper), you may keep it — note it as a judgment call in the log. ### B3. Require a "why notable" for each survivor For every survivor, write **one line** (≤ 18 words) explaining *why someone should care today*. No paraphrasing the model card / dataset description. Good: *"First open-weight 70B trained end-to-end with online RL — beats Llama 3 70B on AGIEval, MIT-licensed."* Bad: *"A new instruction-tuned LLM."* (that's just the description) If you can't write a concrete "why notable" line for an entry, **drop it**. The filter is the feature. When the artifact references a paper, you may pull one verifying detail via **WebFetch** on the arxiv URL or the HF model card — but cap at 1 fetch per pick, and only when it materially sharpens the line. ### B4. Tag momentum Tag each survivor with one of: - **DEBUT** — `createdAt` within the last 7 days (first-time trending) - **ACCELERATING** — older than 7 days, `trendingScore > 50` AND `likes > 200` - **RETURNING** — `createdAt` older than 90 days but trending again — usually a release, a viral post, or a paper drop reviving interest. Note the reason in "why notable" when known - **HOLDOVER** — appeared in the last day's logs (use sparingly; prefer to drop unless there's a new development) ### B5. Cluster into categories Buckets are heuristic — classify by what the artifact does, not by author self-description. Cap total buckets at **5** (merge if you hit 6+). Group survivors: - **LLMs / Reasoning** — text-generation, instruction-tuned, reasoning-tuned, RAG models - **Multimodal** — text-to-image, text-to-video, vision-language, speech, music - **Agents / Tooling** — agent frameworks, tool-use models, function-calling, code models
在 GitHub 查看
这个 SKILL.md 很大,SkillsMP 这里只预览前一段内容。 在 GitHub 查看