Skip to main content

data-scraper-agent

构建一个全自动化的AI驱动数据收集代理,适用于任何公共来源——招聘网站、价格信息、新闻、GitHub、体育赛事等任何内容。按计划进行抓取,使用免费LLM(Gemini Flash)丰富数据,将结果存储在Notion/Sheets/Supabase中,并从用户反馈中学习。完全免费在GitHub Actions上运行。适用于用户希望自动监控、收集或跟踪任何公共数据的场景。

Ir a la instalación

Datos de origen

Repositorio
affaan-m/ECC
Última actividad en el origen
22 de marzo de 2026 a las 22:39
Idioma detectado de SKILL.md
chino
Estrellas
264.820
Forks
39.570

Opciones de instalación

De forma predeterminada está seleccionado el prompt que primero revisa el origen. Puedes cambiar a un comando directo o descargar una copia local.

Revisa los archivos de origen

Lee SKILL.md y los archivos complementarios que muestra SkillsMP antes de decidir si quieres instalarlo.

Mostrando SKILL.md

SKILL.md
Instrucciones de origen · Vista previa de solo lectura
name
data-scraper-agent
description
构建一个全自动化的AI驱动数据收集代理,适用于任何公共来源——招聘网站、价格信息、新闻、GitHub、体育赛事等任何内容。按计划进行抓取,使用免费LLM(Gemini Flash)丰富数据,将结果存储在Notion/Sheets/Supabase中,并从用户反馈中学习。完全免费在GitHub Actions上运行。适用于用户希望自动监控、收集或跟踪任何公共数据的场景。
origin
community
# 数据抓取代理 构建一个生产就绪、AI驱动的数据收集代理,适用于任何公共数据源。 按计划运行,使用免费LLM丰富结果,存储到数据库,并随时间推移不断改进。 **技术栈:Python · Gemini Flash (免费) · GitHub Actions (免费) · Notion / Sheets / Supabase** ## 何时激活 * 用户想要抓取或监控任何公共网站或API * 用户说"构建一个检查...的机器人"、"为我监控X"、"从...收集数据" * 用户想要跟踪工作、价格、新闻、仓库、体育比分、事件、列表 * 用户询问如何自动化数据收集而无需支付托管费用 * 用户想要一个能根据他们的决策随时间推移变得更智能的代理 ## 核心概念 ### 三层架构 每个数据抓取代理都有三层: ``` COLLECT → ENRICH → STORE │ │ │ Scraper AI (LLM) Database runs on scores/ Notion / schedule summarises Sheets / & classifies Supabase ``` ### 免费技术栈 | 层级 | 工具 | 原因 | |---|---|---| | **抓取** | `requests` + `BeautifulSoup` | 无成本,覆盖80%的公共网站 | | **JS渲染的网站** | `playwright` (免费) | 当HTML抓取失败时使用 | | **AI丰富** | 通过REST API的Gemini Flash | 500次请求/天,100万令牌/天 — 免费 | | **存储** | Notion API | 免费层级,用于审查的优秀UI | | **调度** | GitHub Actions cron | 对公共仓库免费 | | **学习** | 仓库中的JSON反馈文件 | 零基础设施,在git中持久化 | ### AI模型后备链 构建代理以在配额耗尽时自动在Gemini模型间回退: ``` gemini-2.0-flash-lite (30 RPM) → gemini-2.0-flash (15 RPM) → gemini-2.5-flash (10 RPM) → gemini-flash-lite-latest (fallback) ``` ### 批量API调用以提高效率 切勿为每个项目单独调用LLM。始终批量处理: ```python # BAD: 33 API calls for 33 items for item in items: result = call_ai(item) # 33 calls → hits rate limit # GOOD: 7 API calls for 33 items (batch size 5) for batch in chunks(items, size=5): results = call_ai(batch) # 7 calls → stays within free tier ``` *** ## 工作流程 ### 步骤 1: 理解目标 询问用户: 1. **收集什么:** "数据源是什么?URL / API / RSS / 公共端点?" 2. **提取什么:** "哪些字段重要?标题、价格、URL、日期、分数?" 3. **如何存储:** "结果应该存储在哪里?Notion、Google Sheets、Supabase,还是本地文件?" 4. **如何丰富:** "您希望AI对每个项目进行评分、总结、分类或匹配吗?" 5. **频率:** "应该多久运行一次?每小时、每天、每周?" 常见的提示示例: * 招聘网站 → 根据简历评分相关性 * 产品价格 → 降价时发出警报 * GitHub仓库 → 总结新版本 * 新闻源 → 按主题+情感分类 * 体育结果 → 提取统计数据到跟踪器 * 活动日历 → 按兴趣筛选 *** ### 步骤 2: 设计代理架构 为用户生成以下目录结构: ``` my-agent/ ├── config.yaml # 用户自定义此文件(关键词、过滤器、偏好设置) ├── profile/ │ └── context.md # AI 使用的用户上下文(简历、兴趣、标准) ├── scraper/ │ ├── __init__.py │ ├── main.py # 协调器:抓取 → 丰富 → 存储 │ ├── filters.py # 基于规则的预过滤器(快速,在 AI 处理之前) │ └── sources/ │ ├── __init__.py │ └── source_name.py # 每个数据源一个文件 ├── ai/ │ ├── __init__.py │ ├── client.py # Gemini REST 客户端,带模型回退 │ ├── pipeline.py # 批量 AI 分析 │ ├── jd_fetcher.py # 从 URL 获取完整内容(可选) │ └── memory.py # 从用户反馈中学习 ├── storage/ │ ├── __init__.py │ └── notion_sync.py # 或 sheets_sync.py / supabase_sync.py ├── data/ │ └── feedback.json # 用户决策历史(自动更新) ├── .env.example ├── setup.py # 一次性数据库/模式创建 ├── enrich_existing.py # 对旧行进行 AI 分数回填 ├── requirements.txt └── .github/ └── workflows/ └── scraper.yml # GitHub Actions 计划任务 ``` *** ### 步骤 3: 构建抓取器源 适用于任何数据源的模板: ```python # scraper/sources/my_source.py """ [Source Name] — scrapes [what] from [where]. Method: [REST API / HTML scraping / RSS feed] """ import requests from bs4 import BeautifulSoup from datetime import datetime, timezone from scraper.filters import is_relevant HEADERS = { "User-Agent": "Mozilla/5.0 (compatible; research-bot/1.0)", } def fetch() -> list[dict]: """ Returns a list of items with consistent schema. Each item must have at minimum: name, url, date_found. """ results = [] # ---- REST API source ---- resp = requests.get("https://api.example.com/items", headers=HEADERS, timeout=15) if resp.status_code == 200: for item in resp.json().get("results", []): if not is_relevant(item.get("title", "")): continue results.append(_normalise(item)) return results def _normalise(raw: dict) -> dict: """Convert raw API/HTML data to the standard schema.""" return { "name": raw.get("title", ""), "url": raw.get("link", ""), "source": "MySource", "date_found": datetime.now(timezone.utc).date().isoformat(), # add domain-specific fields here } ``` **HTML抓取模式:** ```python soup = BeautifulSoup(resp.text, "lxml") for card in soup.select("[class*='listing']"): title = card.select_one("h2, h3").get_text(strip=True) link = card.select_one("a")["href"] if not link.startswith("http"): link = f"https://example.com{link}" ``` **RSS源模式:** ```python import xml.etree.ElementTree as ET root = ET.fromstring(resp.text) for item in root.findall(".//item"): title = item.findtext("title", "") link = item.findtext("link", "") ``` *** ### 步骤 4: 构建Gemini AI客户端 ````python # ai/client.py import os, json, time, requests _last_call = 0.0 MODEL_FALLBACK = [ "gemini-2.0-flash-lite", "gemini-2.0-flash", "gemini-2.5-flash", "gemini-flash-lite-latest", ] def generate(prompt: str, model: str = "", rate_limit: float = 7.0) -> dict: """Call Gemini with auto-fallback on 429. Returns parsed JSON or {}.""" global _last_call api_key = os.environ.get("GEMINI_API_KEY", "") if not api_key: return {} elapsed = time.time() - _last_call if elapsed < rate_limit: time.sleep(rate_limit - elapsed) models = [model] + [m for m in MODEL_FALLBACK if m != model] if model else MODEL_FALLBACK _last_call = time.time() for m in models: url = f"https://generativelanguage.googleapis.com/v1beta/models/{m}:generateContent?key={api_key}" payload = { "contents": [{"parts": [{"text": prompt}]}], "generationConfig": { "responseMimeType": "application/json", "temperature": 0.3, "maxOutputTokens": 2048, }, } try: resp = requests.post(url, json=payload, timeout=30) if resp.status_code == 200: return _parse(resp) if resp.status_code in (429, 404): time.sleep(1) continue return {} except requests.RequestException: return {} return {} def _parse(resp) -> dict: try: text = ( resp.json() .get("candidates", [{}])[0] .get("content", {}) .get("parts", [{}])[0] .get("text", "") .strip() ) if text.startswith("```"): text = text.split("\n", 1)[-1].rsplit("```", 1)[0] return json.loads(text) except (json.JSONDecodeError, KeyError): return {} ```` *** ### 步骤 5: 构建AI管道(批量) ```python # ai/pipeline.py import json import yaml from pathlib import Path from ai.client import generate def analyse_batch(items: list[dict], context: str = "", preference_prompt: str = "") -> list[dict]: """Analyse items in batches. Returns items enriched with AI fields.""" config = yaml.safe_load((Path(__file__).parent.parent / "config.yaml").read_text()) model = config.get("ai", {}).get("model", "gemini-2.5-flash") rate_limit = config.get("ai", {}).get("rate_limit_seconds", 7.0) min_score = config.get("ai", {}).get("min_score", 0) batch_size = config.get("ai", {}).get("batch_size", 5) batches = [items[i:i + batch_size] for i in range(0, len(items), batch_size)] print(f" [AI] {len(items)} items → {len(batches)} API calls") enriched = [] for i, batch in enumerate(batches): print(f" [AI] Batch {i + 1}/{len(batches)}...") prompt = _build_prompt(batch, context, preference_prompt, config) result = generate(prompt, model=model, rate_limit=rate_limit) analyses = result.get("analyses", []) for j, item in enumerate(batch): ai = analyses[j] if j < len(analyses) else {} if ai: score = max(0, min(100, int(ai.get("score", 0)))) if min_score and score < min_score: continue enriched.append({**item, "ai_score": score, "ai_summary": ai.get("summary", ""), "ai_notes": ai.get("notes", "")}) else: enriched.append(item) return enriched def _build_prompt(batch, context, preference_prompt, config): priorities = config.get("priorities", []) items_text = "\n\n".join( f"Item {i+1}: {json.dumps({k: v for k, v in item.items() if not k.startswith('_')})}" for i, item in enumerate(batch) ) return f"""Analyse these {len(batch)} items and return a JSON object. # Items {items_text} # User Context {context[:800] if context else "Not provided"} # User Priorities {chr(10).join(f"- {p}" for p in priorities)} {preference_prompt} # Instructions Return: {{"analyses": [{{"score": <0-100>, "summary": "<2 sentences>", "notes": "<why this matches or doesn't>"}} for each item in order]}} Be concise. Score 90+=excellent match, 70-89=good, 50-69=ok, <50=weak.""" ``` *** ### 步骤 6: 构建反馈学习系统 ```python # ai/memory.py """Learn from user decisions to improve future scoring.""" import json from pathlib import Path FEEDBACK_PATH = Path(__file__).parent.parent / "data" / "feedback.json" def load_feedback() -> dict: if FEEDBACK_PATH.exists(): try: return json.loads(FEEDBACK_PATH.read_text()) except (json.JSONDecodeError, OSError): pass return {"positive": [], "negative": []} def save_feedback(fb: dict): FEEDBACK_PATH.parent.mkdir(parents=True, exist_ok=True) FEEDBACK_PATH.write_text(json.dumps(fb, indent=2)) def build_preference_prompt(feedback: dict, max_examples: int = 15) -> str: """Convert feedback history into a prompt bias section.""" lines = [] if feedback.get("positive"): lines.append("# Items the user LIKED (positive signal):") for e in feedback["positive"][-max_examples:]: lines.append(f"- {e}") if feedback.get("negative"): lines.append("\n# Items the user SKIPPED/REJECTED (negative signal):") for e in feedback["negative"][-max_examples:]: lines.append(f"- {e}") if lines: lines.append("\nUse these patterns to bias scoring on new items.") return "\n".join(lines) ```
Ver en GitHub
Este SKILL.md es muy grande, por eso SkillsMP muestra aqui solo la primera seccion. Ver en GitHub