Skip to main content

data-scraper-agent

构建一个全自动化的AI驱动数据收集代理,适用于任何公共来源——招聘网站、价格信息、新闻、GitHub、体育赛事等任何内容。按计划进行抓取,使用免费LLM(Gemini Flash)丰富数据,将结果存储在Notion/Sheets/Supabase中,并从用户反馈中学习。完全免费在GitHub Actions上运行。适用于用户希望自动监控、收集或跟踪任何公共数据的场景。

Aller à l'installation

Informations de source

Dépôt
affaan-m/ECC
Dernière activité de la source
22 mars 2026 à 22:39
Langue détectée de SKILL.md
chinois
Étoiles
264 820
Forks
39 570

Options d'installation

Le prompt qui vérifie d'abord la source est sélectionné par défaut. Vous pouvez passer à une commande directe ou télécharger une copie locale.

Vérifiez les fichiers source

Lisez SKILL.md et les fichiers associés affichés par SkillsMP avant de décider de l'installer.

Affichage de SKILL.md

SKILL.md
Instructions source · Aperçu en lecture seule
name
data-scraper-agent
description
构建一个全自动化的AI驱动数据收集代理,适用于任何公共来源——招聘网站、价格信息、新闻、GitHub、体育赛事等任何内容。按计划进行抓取,使用免费LLM(Gemini Flash)丰富数据,将结果存储在Notion/Sheets/Supabase中,并从用户反馈中学习。完全免费在GitHub Actions上运行。适用于用户希望自动监控、收集或跟踪任何公共数据的场景。
origin
community
# 数据抓取代理 构建一个生产就绪、AI驱动的数据收集代理,适用于任何公共数据源。 按计划运行,使用免费LLM丰富结果,存储到数据库,并随时间推移不断改进。 **技术栈:Python · Gemini Flash (免费) · GitHub Actions (免费) · Notion / Sheets / Supabase** ## 何时激活 * 用户想要抓取或监控任何公共网站或API * 用户说"构建一个检查...的机器人"、"为我监控X"、"从...收集数据" * 用户想要跟踪工作、价格、新闻、仓库、体育比分、事件、列表 * 用户询问如何自动化数据收集而无需支付托管费用 * 用户想要一个能根据他们的决策随时间推移变得更智能的代理 ## 核心概念 ### 三层架构 每个数据抓取代理都有三层: ``` COLLECT → ENRICH → STORE │ │ │ Scraper AI (LLM) Database runs on scores/ Notion / schedule summarises Sheets / & classifies Supabase ``` ### 免费技术栈 | 层级 | 工具 | 原因 | |---|---|---| | **抓取** | `requests` + `BeautifulSoup` | 无成本,覆盖80%的公共网站 | | **JS渲染的网站** | `playwright` (免费) | 当HTML抓取失败时使用 | | **AI丰富** | 通过REST API的Gemini Flash | 500次请求/天,100万令牌/天 — 免费 | | **存储** | Notion API | 免费层级,用于审查的优秀UI | | **调度** | GitHub Actions cron | 对公共仓库免费 | | **学习** | 仓库中的JSON反馈文件 | 零基础设施,在git中持久化 | ### AI模型后备链 构建代理以在配额耗尽时自动在Gemini模型间回退: ``` gemini-2.0-flash-lite (30 RPM) → gemini-2.0-flash (15 RPM) → gemini-2.5-flash (10 RPM) → gemini-flash-lite-latest (fallback) ``` ### 批量API调用以提高效率 切勿为每个项目单独调用LLM。始终批量处理: ```python # BAD: 33 API calls for 33 items for item in items: result = call_ai(item) # 33 calls → hits rate limit # GOOD: 7 API calls for 33 items (batch size 5) for batch in chunks(items, size=5): results = call_ai(batch) # 7 calls → stays within free tier ``` *** ## 工作流程 ### 步骤 1: 理解目标 询问用户: 1. **收集什么:** "数据源是什么?URL / API / RSS / 公共端点?" 2. **提取什么:** "哪些字段重要?标题、价格、URL、日期、分数?" 3. **如何存储:** "结果应该存储在哪里?Notion、Google Sheets、Supabase,还是本地文件?" 4. **如何丰富:** "您希望AI对每个项目进行评分、总结、分类或匹配吗?" 5. **频率:** "应该多久运行一次?每小时、每天、每周?" 常见的提示示例: * 招聘网站 → 根据简历评分相关性 * 产品价格 → 降价时发出警报 * GitHub仓库 → 总结新版本 * 新闻源 → 按主题+情感分类 * 体育结果 → 提取统计数据到跟踪器 * 活动日历 → 按兴趣筛选 *** ### 步骤 2: 设计代理架构 为用户生成以下目录结构: ``` my-agent/ ├── config.yaml # 用户自定义此文件(关键词、过滤器、偏好设置) ├── profile/ │ └── context.md # AI 使用的用户上下文(简历、兴趣、标准) ├── scraper/ │ ├── __init__.py │ ├── main.py # 协调器:抓取 → 丰富 → 存储 │ ├── filters.py # 基于规则的预过滤器(快速,在 AI 处理之前) │ └── sources/ │ ├── __init__.py │ └── source_name.py # 每个数据源一个文件 ├── ai/ │ ├── __init__.py │ ├── client.py # Gemini REST 客户端,带模型回退 │ ├── pipeline.py # 批量 AI 分析 │ ├── jd_fetcher.py # 从 URL 获取完整内容(可选) │ └── memory.py # 从用户反馈中学习 ├── storage/ │ ├── __init__.py │ └── notion_sync.py # 或 sheets_sync.py / supabase_sync.py ├── data/ │ └── feedback.json # 用户决策历史(自动更新) ├── .env.example ├── setup.py # 一次性数据库/模式创建 ├── enrich_existing.py # 对旧行进行 AI 分数回填 ├── requirements.txt └── .github/ └── workflows/ └── scraper.yml # GitHub Actions 计划任务 ``` *** ### 步骤 3: 构建抓取器源 适用于任何数据源的模板: ```python # scraper/sources/my_source.py """ [Source Name] — scrapes [what] from [where]. Method: [REST API / HTML scraping / RSS feed] """ import requests from bs4 import BeautifulSoup from datetime import datetime, timezone from scraper.filters import is_relevant HEADERS = { "User-Agent": "Mozilla/5.0 (compatible; research-bot/1.0)", } def fetch() -> list[dict]: """ Returns a list of items with consistent schema. Each item must have at minimum: name, url, date_found. """ results = [] # ---- REST API source ---- resp = requests.get("https://api.example.com/items", headers=HEADERS, timeout=15) if resp.status_code == 200: for item in resp.json().get("results", []): if not is_relevant(item.get("title", "")): continue results.append(_normalise(item)) return results def _normalise(raw: dict) -> dict: """Convert raw API/HTML data to the standard schema.""" return { "name": raw.get("title", ""), "url": raw.get("link", ""), "source": "MySource", "date_found": datetime.now(timezone.utc).date().isoformat(), # add domain-specific fields here } ``` **HTML抓取模式:** ```python soup = BeautifulSoup(resp.text, "lxml") for card in soup.select("[class*='listing']"): title = card.select_one("h2, h3").get_text(strip=True) link = card.select_one("a")["href"] if not link.startswith("http"): link = f"https://example.com{link}" ``` **RSS源模式:** ```python import xml.etree.ElementTree as ET root = ET.fromstring(resp.text) for item in root.findall(".//item"): title = item.findtext("title", "") link = item.findtext("link", "") ``` *** ### 步骤 4: 构建Gemini AI客户端 ````python # ai/client.py import os, json, time, requests _last_call = 0.0 MODEL_FALLBACK = [ "gemini-2.0-flash-lite", "gemini-2.0-flash", "gemini-2.5-flash", "gemini-flash-lite-latest", ] def generate(prompt: str, model: str = "", rate_limit: float = 7.0) -> dict: """Call Gemini with auto-fallback on 429. Returns parsed JSON or {}.""" global _last_call api_key = os.environ.get("GEMINI_API_KEY", "") if not api_key: return {} elapsed = time.time() - _last_call if elapsed < rate_limit: time.sleep(rate_limit - elapsed) models = [model] + [m for m in MODEL_FALLBACK if m != model] if model else MODEL_FALLBACK _last_call = time.time() for m in models: url = f"https://generativelanguage.googleapis.com/v1beta/models/{m}:generateContent?key={api_key}" payload = { "contents": [{"parts": [{"text": prompt}]}], "generationConfig": { "responseMimeType": "application/json", "temperature": 0.3, "maxOutputTokens": 2048, }, } try: resp = requests.post(url, json=payload, timeout=30) if resp.status_code == 200: return _parse(resp) if resp.status_code in (429, 404): time.sleep(1) continue return {} except requests.RequestException: return {} return {} def _parse(resp) -> dict: try: text = ( resp.json() .get("candidates", [{}])[0] .get("content", {}) .get("parts", [{}])[0] .get("text", "") .strip() ) if text.startswith("```"): text = text.split("\n", 1)[-1].rsplit("```", 1)[0] return json.loads(text) except (json.JSONDecodeError, KeyError): return {} ```` *** ### 步骤 5: 构建AI管道(批量) ```python # ai/pipeline.py import json import yaml from pathlib import Path from ai.client import generate def analyse_batch(items: list[dict], context: str = "", preference_prompt: str = "") -> list[dict]: """Analyse items in batches. Returns items enriched with AI fields.""" config = yaml.safe_load((Path(__file__).parent.parent / "config.yaml").read_text()) model = config.get("ai", {}).get("model", "gemini-2.5-flash") rate_limit = config.get("ai", {}).get("rate_limit_seconds", 7.0) min_score = config.get("ai", {}).get("min_score", 0) batch_size = config.get("ai", {}).get("batch_size", 5) batches = [items[i:i + batch_size] for i in range(0, len(items), batch_size)] print(f" [AI] {len(items)} items → {len(batches)} API calls") enriched = [] for i, batch in enumerate(batches): print(f" [AI] Batch {i + 1}/{len(batches)}...") prompt = _build_prompt(batch, context, preference_prompt, config) result = generate(prompt, model=model, rate_limit=rate_limit) analyses = result.get("analyses", []) for j, item in enumerate(batch): ai = analyses[j] if j < len(analyses) else {} if ai: score = max(0, min(100, int(ai.get("score", 0)))) if min_score and score < min_score: continue enriched.append({**item, "ai_score": score, "ai_summary": ai.get("summary", ""), "ai_notes": ai.get("notes", "")}) else: enriched.append(item) return enriched def _build_prompt(batch, context, preference_prompt, config): priorities = config.get("priorities", []) items_text = "\n\n".join( f"Item {i+1}: {json.dumps({k: v for k, v in item.items() if not k.startswith('_')})}" for i, item in enumerate(batch) ) return f"""Analyse these {len(batch)} items and return a JSON object. # Items {items_text} # User Context {context[:800] if context else "Not provided"} # User Priorities {chr(10).join(f"- {p}" for p in priorities)} {preference_prompt} # Instructions Return: {{"analyses": [{{"score": <0-100>, "summary": "<2 sentences>", "notes": "<why this matches or doesn't>"}} for each item in order]}} Be concise. Score 90+=excellent match, 70-89=good, 50-69=ok, <50=weak.""" ``` *** ### 步骤 6: 构建反馈学习系统 ```python # ai/memory.py """Learn from user decisions to improve future scoring.""" import json from pathlib import Path FEEDBACK_PATH = Path(__file__).parent.parent / "data" / "feedback.json" def load_feedback() -> dict: if FEEDBACK_PATH.exists(): try: return json.loads(FEEDBACK_PATH.read_text()) except (json.JSONDecodeError, OSError): pass return {"positive": [], "negative": []} def save_feedback(fb: dict): FEEDBACK_PATH.parent.mkdir(parents=True, exist_ok=True) FEEDBACK_PATH.write_text(json.dumps(fb, indent=2)) def build_preference_prompt(feedback: dict, max_examples: int = 15) -> str: """Convert feedback history into a prompt bias section.""" lines = [] if feedback.get("positive"): lines.append("# Items the user LIKED (positive signal):") for e in feedback["positive"][-max_examples:]: lines.append(f"- {e}") if feedback.get("negative"): lines.append("\n# Items the user SKIPPED/REJECTED (negative signal):") for e in feedback["negative"][-max_examples:]: lines.append(f"- {e}") if lines: lines.append("\nUse these patterns to bias scoring on new items.") return "\n".join(lines) ```
Voir sur GitHub
Ce SKILL.md est tres volumineux, SkillsMP affiche donc ici seulement la premiere section. Voir sur GitHub