Skip to main content

technews-webscrape

Configure and use reputable web scraping for TK TechNews. Use when an agent needs Firecrawl MCP, Firecrawl-backed scraping, dynamic page extraction, web source troubleshooting, or guidance on when to use RSS, local scraping, Firecrawl, or YouTube transcripts for cited article generation.

소스 정보

저장소
Tyler-R-Kendrick/tk-technews
최근 소스 활동
2026년 5월 20일 22:58
감지된 SKILL.md 언어
영어
스타
0
포크
0

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.

SKILL.md 표시 중

SKILL.md
소스 지침 · 읽기 전용 미리보기
name
technews-webscrape
description
Configure and use reputable web scraping for TK TechNews. Use when an agent needs Firecrawl MCP, Firecrawl-backed scraping, dynamic page extraction, web source troubleshooting, or guidance on when to use RSS, local scraping, Firecrawl, or YouTube transcripts for cited article generation.
# TechNews Webscrape Use this skill from the repository root when a web source needs higher-quality extraction than the default local `fetch` plus `cheerio` scraper. ## Preferred Scraper Use Firecrawl as the reputable external scraper. The repo has: - `.mcp.json` for MCP-aware agents: `npx -y firecrawl-mcp`. - `npm run scrape:firecrawl -- --url <url>` for one-off extraction checks. - `data/sources.json` support for `"scraper": "firecrawl"` on web sources. ## Workflow 1. Confirm `FIRECRAWL_API_KEY` is set before expecting MCP or Firecrawl scraping to work. 2. For one URL, run `npm run scrape:firecrawl -- --url "https://example.com/article"`. 3. For configured sources, add `"scraper": "firecrawl"` to the source and run `npm run ingest`. 4. Inspect `data/summaries/latest.json` and `data/raw/firecrawl/` for extraction quality. 5. Run `npm run extract:knowledge`, `npm run draft -- --topic "..."`, and `npm run validate:citations`. ## Selection Rules - Prefer RSS when the publisher offers a stable feed. - Use local scraping for simple static pages. - Use Firecrawl for dynamic pages, pages with heavy chrome, or pages where local extraction misses main content. - Use YouTube transcript ingestion for videos. - Do not use scraping to bypass paywalls, access controls, or publisher restrictions.
GitHub에서 보기