| name | rss-to-wiki-pipeline |
| description | 精选高价值 RSS feeds 扫描,输出到 raw/rss-inbox/ 暂存区。只保留有独立知识深度的 feed(非 digest 类),不再自动入库。包含 rss-inbox-curl-recovery.py 绕过 blogwatcher read-state 漏抓的兜底。 |
| version | 5.9.24.5 |
| author | Hermes Agent |
| category | wiki |
| related_skills | ["wiki-pipeline","web-content-reviewer"] |
Archived 2026-08-12 maintenance (size mgmt): inline blocks for 2026-08-12 00:56 / 03:21 / 05:36 / 07:41 / 09:52 / 14:33 / 16:58 (runs 1-7) moved to references/skill-changelog-archive-sessions-2026-08-12-0056-to-2026-08-12-1658.md (SKILL.md was at 101,600/100,000 chars — over cap, archive required before next pointer update)
RSS to Wiki Pipeline (v5.9 — 精简版 + vendor case-study 自动过滤)
📦 Historical changelogs archived: v5.8.0 → v5.9.22.7 + per-cron session logs are in references/cron-run-changelog-archive.md and references/cron-run-YYYY-MM-DD-HHMM.md. Older per-version SKILL.md changelogs (v5.9.23.0 → v5.9.23.14) in references/skill-changelog-archive-v5.9.23.0-to-v5.9.23.14.md; v5.9.23.15 → v5.9.23.20 in references/skill-changelog-archive-v5.9.23.15-to-v5.9.23.20.md; v5.9.23.21 → v5.9.23.29 archived: references/skill-changelog-archive-v5.9.23.21-to-v5.9.23.29.md. v5.9.23.33 → v5.9.23.40 archived: references/skill-changelog-archive-v5.9.23.33-to-v5.9.23.40.md. Older changelog blocks have been removed from this SKILL.md — see the individual session logs for per-version detail. This SKILL.md carries only the most recent session log pointer + Stable Operational Patterns table (the high-signal knowledge a future session needs to read FIRST) + operational steps.
Latest Session
Latest session log: references/cron-run-2026-08-28-0512.md (2026-08-28 05:12, 3rd run of day — FRESH-CYCLE REBUILD, byte-identical deterministic steady state to 03:10 baseline — 0→173 (big=75, mid=66, small=32): start inbox=0 (drained 04:12 wiki-inbox-scan-v2 + 05:06 wechat-inbox-pipeline, both post-03:10 with 0 rss ingest; 03:5x wechat line confirms "rss 173 (03:10 rebuild)... 0 genuinely-new", 05:0x confirms "rss 0 (无新 rebuild)"). 7 std scans: all 7 New=0 except AWS-ML New=1 (no timeouts, GDM clean). 11 WeChat scans (proxy + --unsafe-client): all 11 New=0; proxy clean. Problem feeds: SP 0/6, HF 5/20. Writers: recovery=210 written, 105 skipped — byte-identical to 03:10 (std 47: AWS-CB 9, AWS-ML 14, CrewAI 9, Interconnects 8, Netflix 6, OUT 1 + WeChat 163: 量子位 18, 夕小瑶 18, 字节 15, 小米 17, 新智元 18, 机器之心 13, 美团 17, 腾讯 10, 阿里云开发者 15, 阿里技术 10, PaperWeekly 12); watchdog=1 written (AWS-ML, ACTIVE via scan New=1, stale-re-report dup → 0 net); curl-recovery=49 (GDM 49). Filters: dedup_rm=32, case-study_rm=7, sub-1KB-stub_rm=48, trafilatura-upgrade=0 (all byte-identical). Final inbox 0→173 (big=75, mid=66, small=32), 173 unique URLs, 0 intra-inbox dups. Count 0 vs 03:10 baseline 173 — exactly byte-identical on EVERY component (recovery 210, curl 49, HF 5, dedup 32, case-study 7, sub-1KB 48, traf 0, final 173, composition 75/66/32, URL-200 165/173); all scans New=0 (except AWS-ML stale re-report) → no organic content; watchdog short-circuited confirms; no ingestion since 03:10 (both drains 0 rss ingest); zero feed-death signals. Reconciliation: 210 + 49 + 5 (HF) = 264 gross − 32 − 7 − 48 = 177, minus 4 already-ingested-skipped → 173 clean (watchdog net 0). Classification: FRESH-CYCLE REBUILD, byte-identical deterministic steady state, NOT stale/feed-health — pure determinism, no delta on any component, no escalation. URL-200=165/173 (aws 23/23, mp.weixin 134/134, interconnects 7/7, oneusefulthing 1/1; huggingface 0/4 CN wall — documented-transient; netflix 0/4 CDN osc — documented-transient; DNS healthy).
Prior session log: references/cron-run-2026-08-28-0306.md (2026-08-28 03:06, 2nd run of day — FRESH-CYCLE REBUILD with ingestion-driven reverse shift (−3) — 0→173 (big=75, mid=66, small=32): start inbox=0 (drained ~02:07 wiki-inbox-scan-v2 + 02:52 wechat-inbox-pipeline, both post-01:04 rebuild with 0 rss ingest since; 02:49 wechat line confirms "rss 0 (无新 rebuild)"). 7 std scans: all 7 New=0 except AWS-ML New=1 (no timeouts, GDM clean). 11 WeChat scans (proxy + --unsafe-client): all 11 New=0; proxy clean. Problem feeds: SP 0/6, HF 5/20. Writers: recovery=210 written, 105 skipped — byte-identical to 01:04 (std 47: AWS-CB 9, AWS-ML 14, CrewAI 9, Interconnects 8, Netflix 6, OUT 1 + WeChat 163: 量子位 18, 夕小瑶 18, 字节 15, 小米 17, 新智元 18, 机器之心 13, 美团 17, 腾讯 10, 阿里云开发者 15, 阿里技术 10, PaperWeekly 12); watchdog=1 written (AWS-ML, ACTIVE via scan New=1, dup → 0 net — stale-DB re-report); curl-recovery=49 (GDM 49). Filters: dedup_rm=32 (baseline 29 +3 = the 3 ingested), case-study_rm=7, sub-1KB-stub_rm=48, trafilatura-upgrade=0 (case-study/sub-1KB/traf byte-identical). Final inbox 0→173 (big=75, mid=66, small=32), 173 unique URLs, 0 intra-inbox dups. Count −3 vs 01:04 baseline 176 — decomposed: the −3 = exactly the 3 rss-inbox articles ingested by the 01:4x wechat-inbox-pipeline (OmniColor ECCV + GRACE ICML + TrAct 李飞飞, all verified ingested: 2026-08-28 in raw/articles with mp.weixin source_urls), now removed by dedup (29→32 +3) despite byte-identical recovery 210. Composition 76/68/32 → 75/66/32 (−1 big −2 mid maps to the 3 ingested). Reconciliation: 210 + 49 + 5 (HF) = 264 gross − 32 − 7 − 48 = 177, minus 4 already-ingested-skipped → 173 clean (watchdog net 0). Classification: FRESH-CYCLE REBUILD with ingestion-driven reverse shift (−3), NOT stale/feed-health — the −3 is decomposed exactly to the 3 ingests; zero feed-death signals (all std + curl HTTP=200). URL-200=165/173 (aws 23/23, mp.weixin 134/134, interconnects 7/7, OUT 1/1; huggingface 0/4 CN wall — documented-transient; netflix 0/4 CDN osc — documented-transient; DNS healthy, no escalation).
Prior session log: references/cron-run-2026-08-28-0104.md (2026-08-28 01:04, 1st run of day — FRESH-CYCLE REBUILD with organic multi-feed cascade + ingestion-driven 新智元 reverse shift — 0→176 (big=76, mid=68, small=32): start inbox=0 (drained 00:00 wiki-inbox-scan-v2 + 00:36 wechat-inbox-pipeline, both post-18:44 with 0 rss ingest; 00:25 ingests alibaba-architect-agent / ruofei-ai-native-sdlc / distributed-rl were WeChat-大模型智能-架构师-阿里技术 path, not rss-inbox). 7 std scans: GDM New=2, AWS-ML New=2, other 5 New=0 (no timeouts, GDM clean). 11 WeChat scans (proxy + --unsafe-client): 机器之心 New=3, 美团 New=3, 阿里技术 New=2, PaperWeekly New=2, 腾讯 New=1, 字节 New=1, 夕小瑶 New=1, other 4 New=0; proxy clean. Problem feeds: SP 0/6, HF 5/20. Writers: recovery=210 written, 105 skipped — std 47 (AWS-CB 9, AWS-ML 14, CrewAI 9, Interconnects 8, Netflix 6, OUT 1) + WeChat 163 (量子位 18, 夕小瑶 18, 字节 15, 小米 17, 新智元 18 [19→18 −1 Ornith-1.5 ingestion], 机器之心 13, 美团 17, 腾讯 10, 阿里云开发者 15, 阿里技术 10 [9→10 +1 organic], PaperWeekly 12 [11→12 +1 organic]); watchdog=2 written (AWS-ML, ACTIVE via scan New=2 — organic corroboration); curl-recovery=49 (GDM 49). Filters: dedup_rm=29 (baseline 32 — 3 fewer dups as 3 organic new re-enter), case-study_rm=7, sub-1KB-stub_rm=48 (baseline 49), trafilatura-upgrade=0. Final inbox 0→176 (big=76, mid=68, small=32), 176 unique URLs, 0 intra-inbox dups. Count +4 vs 18:44 baseline 172 — decomposed: (1) AWS-ML organic +2 (scan New=2, watchdog ACTIVE wrote 2, all ~15-31KB big); (2) 阿里技术 +1 + PaperWeekly +1 organic (recovery 9→10 / 11→12, scan New=2 each — confirmed organic, not stale re-report, watchdog ACTIVE); (3) 新智元 −1 = ingestion-driven reverse shift (Ornith-1.5 self-play self-generated coding RL ingested 19:00, verified feed_name: WeChat-新智元 in raw/articles, now skipped 19→18). Composition 74/65/33 → 76/68/32 (+2 big +3 mid −1 small). Reconciliation: 210 + 49 + 5 (HF) = 264 gross − 29 − 7 − 48 = 180, minus 4 already-ingested-skipped → 176 clean (watchdog net +2 organic). Classification: FRESH-CYCLE REBUILD with organic multi-feed cascade (+2 AWS-ML +1 阿里技术 +1 PaperWeekly = +4) + ingestion-driven 新智元 reverse shift (−1), NOT stale/feed-health — the +4 is decomposed exactly to the organic set; zero feed-death signals (all std + curl HTTP=200). URL-200=168/176 (aws 23/23, mp.weixin 137/137, interconnects 7/7, oneusefulthing 1/1; huggingface 0/4 CN wall — documented-transient; netflix 0/4 CDN osc — documented-transient; DNS healthy, no escalation).
Prior session log: references/cron-run-2026-08-27-1844.md (2026-08-27 18:44, 8th run of day — FRESH-CYCLE REBUILD with organic WeChat-新智元 cascade (+1 big) — 0→172 (big=74, mid=65, small=33): start inbox=0 (drained 17:39 wiki-inbox-scan-v2 + 17:56 wechat-inbox-pipeline, both post-16:20 with 0 rss ingest). 7 std scans: all 7 New=0 (no timeouts, GDM clean). 11 WeChat scans (proxy + --unsafe-client): 新智元 New=3, other 10 New=0; proxy clean. Problem feeds: SP 0/6, HF 5/20. Writers: recovery=209 written, 106 skipped — std 47 byte-identical to 16:20 (AWS-CB 9, AWS-ML 14, CrewAI 9, Interconnects 8, Netflix 6, OUT 1) + WeChat 162 (量子位 18, 夕小瑶 18, 字节 15, 小米 17, 新智元 19 [18→19 +1 organic], 机器之心 13, 美团 17, 腾讯 10, 阿里云开发者 15, 阿里技术 9, PaperWeekly 11); watchdog=0 written (GDM short-circuit, all std scans New=0); curl-recovery=50 (GDM 50; AWS-ML 0). Filters: dedup_rm=32, case-study_rm=7, sub-1KB-stub_rm=49, trafilatura-upgrade=0 (all byte-identical to 16:20). Final inbox 0→172 (big=74, mid=65, small=33), 172 unique URLs, 0 intra-inbox dups. Count +1 vs 16:20 baseline 171 — decomposed: the +1 = exactly 1 genuinely-new WeChat-新智元 big file (recovery 新智元 18→19; scan New=3 but only +1 net — 2 of the 3 New were stale DB re-reports of already-captured articles, the URL-format mismatch in DB read-state marking, documented pattern). Composition 73/65/33 → 74/65/33 (+1 big) maps exactly; all other components byte-identical (curl 50, HF 5, dedup 32, case-study 7, sub-1KB 49, traf 0, all other recovery feeds flat). Reconciliation: 209 + 50 + 5 (HF) = 264 gross − 32 − 7 − 49 = 176, minus 4 already-ingested-skipped → 172 clean (watchdog net 0). Classification: FRESH-CYCLE REBUILD with organic WeChat-新智元 cascade (+1 big), NOT stale/feed-health — the +1 is decomposed exactly to the single new 新智元 big article; zero feed-death signals (all std + curl HTTP=200). URL-200=160/172 first pass (aws 23/23, mp.weixin 130/133, interconnects 6/7, oneusefulthing 1/1; huggingface 0/4 CN wall — documented-transient; netflix 0/4 CDN osc — documented-transient); patient retry resolved all mp.weixin/interconnects failures — connection-level transients, DNS healthy, no escalation).
Prior session log: references/cron-run-2026-08-27-1620.md (2026-08-27 16:20, 7th run of day — FRESH-CYCLE REBUILD, byte-identical to 14:10 baseline — 0→171 (big=73, mid=65, small=33): start inbox=0 (drained 15:20 wiki-inbox-scan-v2 + 15:43 wechat-inbox-pipeline, both post-14:10 with 0 rss ingest). 7 std scans: all 7 New=0 (no timeouts, GDM clean). 11 WeChat scans (proxy + --unsafe-client): all 11 New=0; proxy clean. Problem feeds: SP 0/6, HF 5/20. Writers: recovery=208 written, 107 skipped — byte-identical to 14:10 (std 47: AWS-CB 9, AWS-ML 14, CrewAI 9, Interconnects 8, Netflix 6, OUT 1 + WeChat 161: 量子位 18, 夕小瑶 18, 字节 15, 小米 17, 新智元 18, 机器之心 13, 美团 17, 腾讯 10, 阿里云开发者 15, 阿里技术 9, PaperWeekly 11); watchdog=0 written (GDM short-circuit, all scans New=0); curl-recovery=50 (GDM 50; AWS-ML 0). Filters: dedup_rm=32, case-study_rm=7, sub-1KB-stub_rm=49, trafilatura-upgrade=0. Final inbox 0→171 (big=73, mid=65, small=33), 171 unique URLs, 0 intra-inbox dups. Count 0 vs 14:10 baseline 171 — exactly byte-identical on EVERY component (recovery 208, curl 50, HF 5, dedup 32, case-study 7, sub-1KB 49, traf 0, final 171, composition 73/65/33, URL-200 163/171); all scans New=0 → no organic content; watchdog short-circuited confirms; no ingestion since 14:10 (both drains 0 rss ingest); zero feed-death signals. Reconciliation: 208 + 50 + 5 (HF) = 263 gross − 32 − 7 − 49 = 175, minus 4 already-ingested-skipped → 171 clean (watchdog net 0). Classification: FRESH-CYCLE REBUILD, byte-identical deterministic steady state, NOT stale/feed-health — pure determinism, no delta on any component, no escalation. URL-200=163/171 (aws 23/23, mp.weixin 132/132, interconnects 7/7, oneusefulthing 1/1; huggingface 0/4 CN wall — documented-transient; netflix 0/4 CDN osc — documented-transient; DNS healthy).
Prior session log: references/cron-run-2026-08-27-1410.md (2026-08-27 14:10, 6th run of day — FRESH-CYCLE REBUILD, byte-identical to 12:03 baseline #10 — 0→171 (big=73, mid=65, small=33): start inbox=0 (drained 13:11 wiki-inbox-scan-v2 + 13:31 wechat-inbox-pipeline, both post-12:03 with 0 rss ingest). 7 std scans: all 7 New=0 (no timeouts, GDM clean). 11 WeChat scans (proxy + --unsafe-client): all 11 New=0; proxy clean. Problem feeds: SP 0/6, HF 5/20. Writers: recovery=208 written, 107 skipped — byte-identical to 12:03 (std 47: AWS-CB 9, AWS-ML 14, CrewAI 9, Interconnects 8, Netflix 6, OUT 1 + WeChat 161: 量子位 18, 夕小瑶 18, 字节 15, 小米 17, 新智元 18, 机器之心 13, 美团 17, 腾讯 10, 阿里云开发者 15, 阿里技术 9, PaperWeekly 11); watchdog=0 written (GDM short-circuit, all scans New=0); curl-recovery=50 (GDM 50; AWS-ML 0). Filters: dedup_rm=32, case-study_rm=7, sub-1KB-stub_rm=49, trafilatura-upgrade=0. Final inbox 0→171 (big=73, mid=65, small=33), 171 unique URLs, 0 intra-inbox dups. Count 0 vs 12:03 baseline 171 — exactly byte-identical on EVERY component (recovery 208, curl 50, HF 5, dedup 32, case-study 7, sub-1KB 49, traf 0, final 171, composition 73/65/33, URL-200 163/171); all scans New=0 → no organic content; watchdog short-circuited confirms; no ingestion since 12:03 (13:11 + 13:31 drains both 0 rss ingest); zero feed-death signals. Reconciliation: 208 + 50 + 5 (HF) = 263 gross − 32 − 7 − 49 = 175, minus 4 already-ingested-skipped → 171 clean (watchdog net 0). Classification: FRESH-CYCLE REBUILD, byte-identical deterministic steady state, NOT stale/feed-health — pure determinism, no delta on any component, no escalation. URL-200=163/171 (aws 23/23, mp.weixin 132/132, interconnects 7/7, oneusefulthing 1/1; huggingface 0/4 CN wall — documented-transient; netflix 0/4 CDN osc — documented-transient; DNS healthy).
Prior session log: references/cron-run-2026-08-27-1203.md (2026-08-27 12:03, 5th run of day — FRESH-CYCLE REBUILD, byte-identical to 10:00 baseline #9 — 0→171 (big=73, mid=65, small=33): start inbox=0 (drained by 10:11 + 11:20 wechat-inbox-pipeline runs, both with 0 rss ingest). 7 std scans: all 7 New=0 (no timeouts, GDM clean). 11 WeChat scans (proxy + --unsafe-client): all 11 New=0; proxy clean. Problem feeds: SP 0/6, HF 5/20. Writers: recovery=208 written, 107 skipped — byte-identical to 10:00 (std 47: AWS-CB 9, AWS-ML 14, CrewAI 9, Interconnects 8, Netflix 6, OUT 1 + WeChat 161: 量子位 18, 夕小瑶 18, 字节 15, 小米 17, 新智元 18, 机器之心 13, 美团 17, 腾讯 10, 阿里云开发者 15, 阿里技术 9, PaperWeekly 11); watchdog=0 written (GDM short-circuit, all scans New=0); curl-recovery=50 (GDM 50; AWS-ML 0). Filters: dedup_rm=32, case-study_rm=7, sub-1KB-stub_rm=49, trafilatura-upgrade=0. Final inbox 0→171 (big=73, mid=65, small=33), 171 unique URLs, 0 intra-inbox dups. Count 0 vs 10:00 baseline 171 — exactly byte-identical on EVERY component (recovery 208, curl 50, HF 5, dedup 32, case-study 7, sub-1KB 49, traf 0, final 171, composition 73/65/33, URL-200 163/171); all scans New=0 → no organic content; watchdog short-circuited confirms; no ingestion since 10:00 (both drains 0 ingest); zero feed-death signals. Reconciliation: 208 + 50 + 5 (HF) = 263 gross − 32 − 7 − 49 = 175, minus 4 already-ingested-skipped → 171 clean (watchdog net 0). Classification: FRESH-CYCLE REBUILD, byte-identical deterministic steady state, NOT stale/feed-health — pure determinism, no delta on any component, no escalation. URL-200=163/171 (aws 23/23, mp.weixin 132/132, interconnects 7/7, oneusefulthing 1/1; huggingface 0/4 CN wall — documented-transient; netflix 0/4 CDN osc — documented-transient; DNS healthy).
Prior session log: references/cron-run-2026-08-27-0955.md (2026-08-27 09:55, 4th run of day — FRESH-CYCLE REBUILD with ingestion-driven reverse shift (−2 WeChat-机器之心) — 0→171 (big=73, mid=65, small=33): start inbox=0 (drained ~08:56 wiki-inbox-scan-v2 + ~09:10 wechat-inbox-pipeline; 09:08 wechat line confirms "rss 0 (无新 rebuild)"). 7 std scans: all 7 New=0 (no timeouts, GDM clean). 11 WeChat scans (proxy + --unsafe-client): all 11 New=0; proxy clean. Problem feeds: SP 0/6, HF 5/20. Writers: recovery=208 written, 107 skipped — std 47 byte-identical (AWS-CB 9, AWS-ML 14, CrewAI 9, Interconnects 8, Netflix 6, OUT 1) + WeChat 161 (量子位 18, 夕小瑶 18, 字节 15, 小米 17, 新智元 18, 机器之心 13 [15→13 −2 ingestion], 美团 17, 腾讯 10, 阿里云开发者 15, 阿里技术 9, PaperWeekly 11); watchdog=0 written (GDM short-circuit, all scans New=0); curl-recovery=50 (GDM 50; AWS-ML 0). Filters: dedup_rm=32, case-study_rm=7, sub-1KB-stub_rm=49, trafilatura-upgrade=0 (all byte-identical to 05:28). Final inbox 0→171 (big=73, mid=65, small=33), 171 unique URLs, 0 intra-inbox dups. Count −2 vs 05:28 baseline 173 — decomposed: the −2 = exactly 2 WeChat-机器之心 ingests (shopify-ceo考虑禁用claude-code 4937B small + 突破遥操瓶颈全新无本体数据世界动作模型noe-0 14952B big, both verified feed_name: WeChat-机器之心 in raw/articles), recovery 机器之心 15→13 reverse shift, composition 74/65/34 → 73/65/33 (−1 big −1 small maps exactly); all other std + WeChat feeds byte-identical. Reconciliation: 208 + 50 + 5 (HF) = 263 gross − 32 − 7 − 49 = 175, minus 4 already-ingested-skipped → 171 clean (watchdog net 0). Classification: FRESH-CYCLE REBUILD with ingestion-driven reverse shift (−2 WeChat-机器之心), NOT stale/feed-health — the −2 is decomposed exactly to the 2 机器之心 ingests; zero feed-death signals. URL-200=163/171 (aws 23/23, mp.weixin 132/132, interconnects 7/7, oneusefulthing 1/1; huggingface 0/4 CN wall — documented-transient; netflix 0/4 CDN osc — documented-transient; DNS healthy).
Prior session log: references/cron-run-2026-08-27-0528.md (2026-08-27 05:28, 3rd run of day — FRESH-CYCLE REBUILD, byte-identical to 03:20 baseline — 0→173 (big=74, mid=65, small=34): start inbox=0 (drained 04:29 wiki-inbox-scan-v2 + 05:27 wechat-inbox-pipeline, both post-03:20 with 0 RSS ingest; the 05:14 ingest RSI test-time was WeChat-大模型智能, not rss-inbox). 7 std scans: all 7 New=0 (no timeouts, GDM clean). 11 WeChat scans: all 11 New=0; proxy clean. Problem feeds: SP 0/6, HF 5/20. Writers: recovery=210 written, 105 skipped — byte-identical to 03:20 (std 47: AWS-CB 9, AWS-ML 14, CrewAI 9, Interconnects 8, Netflix 6, OUT 1 + WeChat 163: 量子位 18, 夕小瑶 18, 字节 15, 小米 17, 新智元 18, 机器之心 15, 美团 17, 腾讯 10, 阿里云开发者 15, 阿里技术 9, PaperWeekly 11); watchdog=0 written (GDM short-circuit, all scans New=0; AWS-ML New=0 so the 03:20 stale-re-report watchdog=1 absent); curl-recovery=50 (GDM 50; AWS-ML 0). Filters: dedup_rm=32, case-study_rm=7, sub-1KB-stub_rm=49, trafilatura-upgrade=0. Final inbox 0→173 (big=74, mid=65, small=34), 173 unique URLs, 0 intra-inbox dups. Count 0 vs 03:20 baseline 173 — exactly byte-identical on EVERY component (recovery 210, curl 50, HF 5, dedup 32, case-study 7, sub-1KB 49, traf 0, final 173, URL-200 165/173); all scans New=0 → no organic content; watchdog short-circuited confirms; zero feed-death signals. Reconciliation: 210 + 50 + 5 (HF) = 265 gross − 32 − 7 − 49 = 177, minus 4 already-ingested-skipped → 173 clean (watchdog net 0). Classification: FRESH-CYCLE REBUILD, byte-identical deterministic steady state, NOT stale/feed-health — pure determinism, no delta on any component, no escalation. URL-200=165/173 (aws 23/23, mp.weixin 134/134, interconnects 7/7, OUT 1/1; huggingface 0/4 CN wall — documented-transient; netflix 0/4 CDN osc — documented-transient; DNS healthy).
Prior session log: references/cron-run-2026-08-27-0320.md (2026-08-27 03:20, 2nd run of day — FRESH-CYCLE REBUILD with ingestion-driven reverse shift (−5 big) — 0→173 (big=74, mid=65, small=34): start inbox=0 (drained 02:25 wiki-inbox-scan-v2 + 03:20 wechat-inbox-pipeline; the 02:xx wiki-inbox-scan ingested 6 articles from the 01:11 rebuild's 178-file inbox → 5 rss + 1 newsletter). 7 std scans: 6 New=0; AWS-ML New=1 (no timeouts, GDM clean). 11 WeChat scans (proxy + --unsafe-client): all 11 New=0; proxy clean. Problem feeds: SP 0/6, HF 5/20. Writers: recovery=210 written, 105 skipped — std 47 (AWS-CB 9, AWS-ML 14 [16→14 −2 ingestion], CrewAI 9, Interconnects 8, Netflix 6, OUT 1) + WeChat 163 (量子位 18, 夕小瑶 18, 字节 15 [16→15 −1], 小米 17, 新智元 18, 机器之心 15, 美团 17, 腾讯 10 [11→10 −1], 阿里云开发者 15, 阿里技术 9, PaperWeekly 11 [12→11 −1]); watchdog=1 written (AWS-ML, stale re-report dup → 0 net); curl-recovery=50 (GDM 50; AWS-ML 0 — already in inbox from recovery's 14). Filters: dedup_rm=32, case-study_rm=7, sub-1KB-stub_rm=49, trafilatura-upgrade=0 (all byte-identical to 01:11). Final inbox 0→173 (big=74, mid=65, small=34), 173 unique URLs, 0 intra-inbox dups. Count −5 vs 01:11 baseline 178 — decomposed: the −5 = exactly the 5 rss ingests by the 02:xx wiki-inbox-scan, all mapped to recovery per-feed reverse shifts (verified feed_name in raw/articles): AWS-ML −2 (Natera 语音 Agent 30451B + SFT 数据准备进阶 19940B), WeChat-字节 −1 (提示词注入防护 15375B), WeChat-腾讯 −1 (自进化飞轮 62775B), WeChat-PaperWeekly −1 (AutoResearch 四循环 14589B). All 5 ingested are BIG → composition big 79→74 (−5 all big), mid/small flat. AWS-ML scan New=1 + watchdog=1 written = stale re-report / dup (0 net), recovery AWS-ML exactly 14 = 16−2, NO cascade. All other std components + WeChat feeds byte-identical. Reconciliation: 210 + 50 + 5 (HF) = 265 gross − 32 − 7 − 49 = 177, minus 4 already-ingested-skipped → 173 clean (watchdog net 0). Classification: FRESH-CYCLE REBUILD with ingestion-driven reverse shift (−5), NOT stale/feed-health — the −5 is decomposed exactly to the 5 rss ingests; zero feed-death signals (all 7 std + curl HTTP=200). URL-200=165/173 (aws 23/23, mp.weixin 134/134, interconnects 7/7, OUT 1/1; huggingface 0/4 CN wall — documented-transient; netflix 0/4 CDN osc — documented-transient; DNS healthy).
Prior session log: references/cron-run-2026-08-27-0111.md (2026-08-27 01:11, 1st run of day — FRESH-CYCLE REBUILD with organic AWS-ML cascade (+2) + ingestion-driven 量子位 reverse shift (−1) — 0→178 (big=79, mid=65, small=34): start inbox=0 (drained 00:5x wechat-inbox-pipeline + 00:23 wiki-inbox-scan-v2, both saw rss 0; prior logged rebuild 20:33 wrote inbox=175; 22:53 rss-feed-scan heartbeat touch with NO cron-status line = undocumented/aborted run, self-healed). 7 std scans: GDM New=1, AWS-ML New=6, other 5 New=0 (no timeouts, GDM clean). 11 WeChat scans (proxy + --unsafe-client): 腾讯 New=1, PaperWeekly New=2, 字节 New=1 (single TIMEOUT cleared on immediate retry — in-cycle transient), other 8 New=0; proxy clean. Problem feeds: SP 0/6, HF 5/20 (baseline). Writers: recovery=215 written, 100 skipped — std 49 (AWS-CB 9, AWS-ML 16 [+2 organic], CrewAI 9, Interconnects 8, Netflix 6, OUT 1) + WeChat 166 (量子位 18 [19→18 −1 SemaPLC ingestion], 夕小瑶 18, 字节 16, 小米 17, 新智元 18, 机器之心 15, 美团 17, 腾讯 11, 阿里云开发者 15, 阿里技术 9, PaperWeekly 12); watchdog=6 written (ACTIVE via scan New>0, dups → 0 net); curl-recovery=51 (GDM 50 + AWS-ML 1). Filters: dedup_rm=32, case-study_rm=7, sub-1KB-stub_rm=49, trafilatura-upgrade=0. Final inbox 0→178 (big=79, mid=65, small=34), 178 unique URLs, 0 intra-inbox dups. Count +3 vs 20:26 baseline 175 — decomposed: (1) AWS-ML organic +2 (scan New=6 → recovery 14→16 → +2 net new big files; watchdog ACTIVE corroborates); (2) 量子位 −1 = INGESTION-driven reverse shift — SemaPLC 验证门控 PLC 代码生成 ingested 21:07, verified feed_name: WeChat-量子位 in raw/articles, now skipped (19→18); (3) curl-recovery 53→51 + dedup 34→32 + sub-1KB 50→49 = minor hygiene shifts, no feed-health change. Reconciliation: 215 + 51 + 5 (HF) = 271 gross − 32 − 7 − 49 = 183, dedup/skip overlap → 178 clean (watchdog net 0). Classification: FRESH-CYCLE REBUILD with organic AWS-ML cascade (+2) + ingestion-driven 量子位 reverse shift (−1), NOT stale/feed-health — the +3 is decomposed exactly; all std feeds except AWS-ML + other 10 WeChat feeds byte-identical; zero feed-death signals. URL-200=169/178 (aws 24/25 [1 transient], mp.weixin 137/137, interconnects 7/7, OUT 1/1; huggingface 0/4 CN wall — documented-transient; netflix 0/4 CDN osc — documented-transient; DNS healthy).
Prior session log (2026-08-26 20:26, 8th run of day): full log in references/cron-run-2026-08-26-2026.md
Prior session log (2026-08-26 16:21, 7th run of day): full log in references/cron-run-2026-08-26-1621.md
Prior session log (2026-08-26 14:05, 6th run of day): full log in references/cron-run-2026-08-26-1405.md
Prior session log (2026-08-26 09:44, 5th run of day): full log in references/cron-run-2026-08-26-0944.md
Prior session log (2026-08-26 07:42, 4th run of day): full log in references/cron-run-2026-08-26-0742.md
Prior session log (2026-08-26 05:32, 3th run of day): full log in references/cron-run-2026-08-26-0532.md
Prior session log (2026-08-26 03:23, 2th run of day): full log in references/cron-run-2026-08-26-0323.md
Prior session log (2026-08-26 01:14, 1th run of day): full log in references/cron-run-2026-08-26-0114.md
Prior session log (2026-08-25 19:53, 7th run of day): full log in references/cron-run-2026-08-25-1953.md
Prior session log (2026-08-25 17:50, 6th run of day): full log in references/cron-run-2026-08-25-1750.md
Prior session log (2026-08-25 15:34, 5th run of day): full log in references/cron-run-2026-08-25-1534.md
Prior session log (2026-08-24 21:49, 4th run of day): full log in references/cron-run-2026-08-24-2149.md
Prior session log (2026-08-24 17:00, 3rd run of day): full log in references/cron-run-2026-08-24-1700.md
Prior session log (2026-08-24 14:57, 2nd run of day): full log in references/cron-run-2026-08-24-1457.md
Prior session log (2026-08-24 12:42, 1st run of day): full log in references/cron-run-2026-08-24-1242.md
Prior session log archive (2026-08-23, runs 1-8: 00:06/02:16/04:22/06:27/08:33/10:39/12:46/20:58): pointer blocks moved to references/skill-changelog-archive-sessions-2026-08-23-full-day.md (size mgmt — SKILL.md over 100K cap); full per-run logs in references/cron-run-2026-08-23-HHMM.md
Prior session log archive (2026-08-22 13:18 → 21:56, runs 7-11): pointer blocks moved to references/skill-changelog-archive-sessions-2026-08-22-runs7-11.md (size mgmt — SKILL.md was 107,614 bytes over 100K cap); full per-run logs in references/cron-run-2026-08-22-HHMM.md
Prior session log archive (2026-08-22 00:18 → 11:07, runs 1-6): pointer blocks moved to references/skill-changelog-archive-sessions-2026-08-22-0018-to-2026-08-22-1527.md (size mgmt — SKILL.md was 102,284/100,000 chars over cap); full per-run logs in references/cron-run-2026-08-22-HHMM.md
Prior session log archive (2026-08-21 22:08 → 2026-08-20 02:00): pointer blocks for 2026-08-21 (runs 1-8: 00:13/02:22/04:28/06:33/08:38/10:44/15:53/22:08) + 2026-08-20 (runs 1-8: 02:00/04:06/06:11/08:22/10:24/15:07/17:09/22:04) moved to references/skill-changelog-archive-sessions-2026-08-21-to-2026-08-20.md (size mgmt — SKILL.md was 121,811/100,000 chars over cap, archive required before 11:07 pointer update); full per-run logs in references/cron-run-YYYY-MM-DD-HHMM.md
Prior session log archive (2026-08-19, runs 1-8: 02:10/04:13/06:21/09:14/11:33/13:39/15:47/17:57): pointer blocks moved to references/skill-changelog-archive-sessions-2026-08-19-full-day.md (size mgmt — SKILL.md was 103,583/100,000 chars over cap, archive required before next pointer update); full per-run logs in references/cron-run-2026-08-19-HHMM.md
Prior session log archive (2026-08-18 09:27 → 2026-08-18 23:53, rebuilds #31-#35): pointer blocks moved to references/skill-changelog-archive-sessions-2026-08-18-0927-to-2026-08-18-2353.md (size mgmt — SKILL.md over 100K cap); full per-run logs in references/cron-run-2026-08-18-HHMM.md.
Prior session log archive (2026-08-17 17:37 → 2026-08-18 07:20): pointer blocks for 08-17 17:37/19:59/22:03 + 08-18 02:44/04:58/07:20 moved to references/skill-changelog-archive-sessions-2026-08-17-to-2026-08-18-0720.md (size mgmt — SKILL.md was 101,176/100,000 chars over cap, archive required before 12:34 pointer update); full per-run logs in references/cron-run-YYYY-MM-DD-HHMM.md
Prior session log archive (2026-08-17 00:15 → 15:23, 7 runs): pointer blocks moved to references/skill-changelog-archive-sessions-2026-08-17-0015-to-2026-08-17-1523.md (size mgmt); full per-run logs in references/cron-run-2026-08-17-HHMM.md
Prior session log archive (2026-08-16 08:50 + 13:17): pointer blocks moved to references/skill-changelog-archive-sessions-2026-08-16-0850-to-2026-08-16-1317.md (size mgmt); full per-run logs in references/cron-run-2026-08-16-0850.md / references/cron-run-2026-08-16-1317.md
Prior session log archive (2026-08-15 runs 8-10: 18:45/21:10/23:50): pointer blocks archived to references/skill-changelog-archive-sessions-2026-08-15-1845-to-2026-08-15-2336.md (size mgmt, SKILL.md ~99.3K/100K, 2026-08-17); full per-run logs in references/cron-run-2026-08-15-1845.md / cron-run-2026-08-15-2110.md / cron-run-2026-08-15-2336.md
Prior session log archive (2026-08-15 runs 4-7: 09:39/11:49/14:13/16:27): pointer blocks archived to references/skill-changelog-archive-sessions-2026-08-15-0939-to-2026-08-15-1627.md (size mgmt, SKILL.md >100K); full per-run logs in references/cron-run-2026-08-15-0939.md / cron-run-2026-08-15-1149.md / cron-run-2026-08-15-1413.md / cron-run-2026-08-15-1627.md
Archived 2026-08-15 maintenance (size mgmt): inline blocks for 2026-08-14 17:18 / 23:51 + 2026-08-15 02:38 / 04:52 / 07:15 (runs 7-8 of 08-14 + runs 1-3 of 08-15) moved to references/skill-changelog-archive-sessions-2026-08-14-1718-to-2026-08-15-0715.md (SKILL.md was at 102,770/100,000 chars — over cap, archive required before next pointer update)
Archived 2026-08-14 maintenance (size mgmt): inline blocks for 2026-08-14 01:54 / 04:14 / 06:23 / 11:37 / 12:13 / 14:59 (runs 1-6 of 08-14) moved to references/skill-changelog-archive-sessions-2026-08-14-0154-to-2026-08-14-1459.md (SKILL.md was at 100,526/100,000 chars — over cap, archive required before next pointer update)
Archived 2026-08-14 maintenance (size mgmt): inline blocks for 2026-08-13 14:46 / 17:07 / 20:51 / 23:13 (runs 6-9 of 08-13) moved to references/skill-changelog-archive-sessions-2026-08-13-1446-to-2026-08-13-2313.md (SKILL.md was at 100,394/100,000 chars — over cap, archive required before next pointer update)
Archived 2026-08-12 maintenance (size mgmt): inline blocks for 2026-08-11 02:47 → 2026-08-11 22:52 (runs 2-10 of 08-11) + 2026-08-10 22:23 / 2026-08-11 00:37 + 2026-08-10 01:29-16:43 (runs 1-5 of 08-10) moved to references/skill-changelog-archive-sessions-2026-08-10-0129-to-2026-08-11-2252.md + references/skill-changelog-archive-sessions-2026-08-10-2223-to-2026-08-11-0037.md (SKILL.md was at 105,115/100,000 chars — over cap, archive required before pointer update)
Archived 2026-08-10 maintenance (size mgmt): inline blocks for 2026-08-09 00:38 → 23:13 (runs 1-12 of 08-09) + 2026-08-08 01:52-15:07 (runs 1-7 of 08-08) + 2026-08-07 16:56-23:36 (runs 7-10 of 08-07) + 2026-08-06 04:09-22:23 (runs 6-10 of 08-06) moved to references/skill-changelog-archive-sessions-2026-08-06-0409-to-2026-08-06-2223.md / -2026-08-07-1656-to-2026-08-07-2336.md / -2026-08-08-0152-to-2026-08-08-1507.md / -2026-08-09-0038-to-2026-08-09-2313.md (SKILL.md at cap during each; archive required before pointer update)
Archived inline blocks for rebuilds #25–#33 (2026-08-05 05:38 → 23:51) moved to references/skill-changelog-archive-sessions-2026-08-05-0538-to-2026-08-05-2351.md (2026-08-06 maintenance); see references/cron-run-YYYY-MM-DD-HHMM.md for full per-run logs.
Archived session logs (2026-08-02 → 2026-08-05 03:28, rebuilds #4–#27): inline blocks removed 2026-08-05 to stay under the 100K SKILL.md cap; one-line index in references/skill-changelog-archive-sessions-2026-08-02-to-08-05-0328.md, full per-run logs in references/cron-run-YYYY-MM-DD-HHMM.md. Rebuild chain: #4 (08-02 04:26) → #11 (08-02 21:18) → #12 (08-02 23:24) → #13/#14 (08-03) → #15/#16/#17 (08-03) → #18 (08-03 18:46) → #19 (08-03 23:21) → #20–#23 (08-04) → #24/#25 (08-04) → #26/#27 (08-05).
v5.9.23.30 → v5.9.23.44 archived: Archived in session reference files. See individual cron-run-YYYY-MM-DD-HHMM.md files.
🚨 SKILL.md Corruption Incident (v5.9.23.50 — RESTORED)
Problem (observed 2026-07-03 05:57): The SKILL.md on disk was found at 1,806 bytes — truncated to only the frontmatter + "RSS 策略" header + placeholder text. The full body (103KB) was lost. Root cause: a prior skill_manage(action='edit') call that passed only the first few lines instead of the complete file, overwriting the full content with a stub.
Fix (applied in v5.9.23.50): Restored from the wiki git HEAD (v5.9.23.41) and updated the session pointer + new sections. Older verbose changelog blocks (v5.9.23.x SKILL.md update narrative blocks) were removed to keep the file under the 100K char limit — they are preserved in individual session reference files (references/cron-run-*.md).
Prevention:
- After ANY
edit call on this skill, run skill_view('rss-to-wiki-pipeline') and verify the full body is present. The >50KB byte threshold is STALE (v5.9.23.93): the file was deliberately slimmed in the v5.9.23.50 cleanup; healthy size is now ~42KB. Verify by grepping key sections (## Stable Operational Patterns, ## Latest Session, ## RSS 策略) instead of byte size. Use anchored grep -n '^## <Section>' for the uniqueness check — unanchored grep -c '## <Section>' returns 2 for headings that also appear in prose (e.g. ## RSS 策略 is referenced in the corruption-incident section), which triggers a false duplicate-header alarm (verified 2026-08-03). DO NOT byte-compare to git HEAD — HEAD (36.8KB) is legitimately smaller than the on-disk file (41.9KB) because session-log pointers accumulate in the working copy; a byte-based false alarm triggers a destructive restore that regresses the file.
- When using
edit, ALWAYS pass a COMPLETE SKILL.md, not a partial reconstruction.
- The wiki git (
~/wiki/skills/) is the recovery source — git show HEAD:skills/wiki/rss-to-wiki-pipeline/SKILL.md has the canonical copy.
RSS 策略:只保留有独立知识价值的内容
| 保留 | 源 | 文章类型 | 知识价值 |
|---|
| ✅ | Hugging Face Blog | Post-Training / Agent / Eval / 推理 | ⭐⭐⭐ — 每篇直接命中知识库核心(vLLM/RL/Agent失败模式/DeepSeek-V4)。RSS URL: https://huggingface.co/blog/feed.xml(非 /blog/rss.xml)。feed 无 content:encoded,CN 网络下 curl 返回 0 字节(JS SPA 完全被墙),必须 TinyFish 或 VPN。 |
| ✅ | Google DeepMind Blog | AI 前沿研究发布(Alpha/RL/Agent) | ⭐⭐⭐⭐ — 替代 NVIDIA,有独立知识深度 |
| ✅ | Stochastic Parrot | AI 哲学/社会/认知深度长文 | ⭐⭐⭐ — 有独立知识体系 |
| ✅ | One Useful Thing | AI 与工作/社会交叉分析(Ethan Mollick) | ⭐⭐ — 教育性应用分析 |
| ✅ | CrewAI Blog | 多 Agent 框架官方博客 | ⭐⭐ — 有技术深度,但偏产品 |
| ✅ | Netflix Tech Blog | PERMANENT (v5.9.23.35→v5.9.23.40, 6+ crons):Feed URL returned to life after ~2 weeks of 404s. Articles return 403 via curl (medium CDN) but recovery captures full 10-36KB content from feed entries. Include in scan chains. | |
| ✅ | AWS China Blog (中文) | AWS 中国区原创技术文章 | ⭐⭐ — 中文原创,Agent 实践 |
| ✅ | AWS China ML | AWS ML 技术文章 | ⭐⭐ — 英文,但有 Agent/ML 内容 |
| 关闭 | 源 | 关闭原因 |
|---|
| ❌ | Import AI | newsletter digest,每期 3-5 话题都很浅,无法独立入库 |
| ❌ | Latent Space | AINews digest,同上 |
| ❌ | Last Week in AI | AI news digest |
| ❌ | Cloudflare AI Blog | 公司产品公告为主 |
| ❌ | AWS 英文分类 blogs (12) | 产品发布/文档为主,知识密度低 |
| ❌ | Stanford SAIL | summary only,极少更新 |
| ❌ | NVIDIA Developer Blog | 永久下架(2026-05-08):RSS 推送的 2024-2025 年文章 URL 全部 404(原站已删文),feed 本身无新文章价值。58 篇已入库旧文已清理。 |
判断标准:什么 RSS 文章值得入库
RSS 文章只有满足以下条件才值得评分入库:
- 有独立论点/论据 — 文章本身就是一个完整的知识单元,不是多个话题的拼凑
- 与知识库核心主题相关 — Agent Engineering / Harness / Skills / Post-Training / 模型架构
- 有提炼空间 — 能从中提取出可交叉链接到其他实体/概念页的信息
- 非产品公告 — 不是产品发布、版本更新、认证公告等纯消息
如果一篇文章不满足以上条件,就留在 raw/rss-inbox/ 中。不强行入库。
快速自检法
Q: 这篇文章我读完能提炼出什么?
→ 能提炼出一个实体、一个概念、或一个工程原则 → 入库
→ 只能提炼出"哦有这么件事" → 不入库
Q: 这篇文章在我知识库里能链接到哪里?
→ 有 2 个以上现有页面可以交叉链接 → 入库
→ 孤立存在,没有链接 → 不入库
News Digest vs 独立长文的区分
| 特征 | News Digest | 独立长文 |
|---|
| 内容结构 | 3-5个话题拼凑,每个一段摘要 | 单一主题,有完整论据 |
| 入库后得到 | 空壳(只有归类+原文链接) | 有知识提炼的实体/概念页 |
| 交叉链接 | 无法建立有意义的链接 | 可连接2+现有页面 |
| 典型来源 | Import AI, Latent Space AINews | Google DeepMind Blog, Stochastic Parrot |
关键教训: 把整期 newsletter 当成一篇文章入库,得到的 entity 只有归类+原文链接,和 bookmark 没区别。这和知识库"沉淀可复用知识"的初衷背道而驰。不要入库 newsletter digest。
Cron 配置(v2 — 四任务独立架构)
重要教训:不要把所有步骤放在同一个 cron 里。 合并 RSS + 微信 + Newsletter 会导致单步超时中断整个任务。应将抓取步骤拆为独立子 cron,汇总报告单独一个 cron。
当前架构
4 个独立 cron 任务,每 8h 运行,互不阻塞:
rss-feed-scan → raw/rss-inbox/ (7 good feeds)
wechat-mp-scan → raw/wechat-inbox/
newsletter-link-extract → candidates.md
wiki-inbox-scan-v2 → 汇总 + 邮件报告
环境变量
cron 运行在非交互式 shell,不加载 .zshrc。注入方式:
source ~/.wiki-cron.env
unset http_proxy https_proxy HTTP_PROXY HTTPS_PROXY
脚本路径(verified 2026-07-01)
所有脚本位于 ~/.hermes/skills/wiki/rss-to-wiki-pipeline/scripts/ 下:
rss-inbox-recovery.py — 主路径
rss-inbox-watchdog.py — 辅助路径(当 recovery 漏抓时)
rss-inbox-curl-recovery.py — GDM 兜底;也补充其他 feeds 的被 auto-mark-read 文章
rss-inbox-trafilatura-upgrade.py — 对 sub-2KB 文件尝试 trafilatura 升级
inbox-case-study-filter.py — 删除 vendor case-study PR 文章
rss-inbox-sub-1kb-filter.py — 删除 sub-1KB trafilatura stub
fetch-problem-feeds.py — 抓取 Substack + HuggingFace(Go HTTP timeout 的 feeds)
cron-log-write.sh — 将结构化 log 行写入 cron-status.log
rss-source-reaper.py — 清理失效 RSS sources
Important: Always unset proxy variables before running blogwatcher-cli — the Go HTTP client fails with proxy authorization errors in the cron environment.
执行流程(每次 cron 运行)
# Step 1: Touch heartbeat
python3 ~/wiki/scripts/cron-heartbeat.py touch rss-feed-scan
# Step 2: Scan good feeds (blogwatcher-cli, individual, 20s per feed timeout)
# IMPORTANT: use Python subprocess.run(timeout=20) wrapper — shell `timeout` command NOT available in cron env
# ⚠️ subprocess.run INHERITS proxy env vars — blogwatcher-cli's Go HTTP client fails with
# "not authorized by the client" when proxies are set. Always pass env with proxies removed.
# Do NOT use env={} (clears PATH+HOME → blogwatcher-cli: not found + $HOME not defined).
# Netflix Tech Blog reactivated — feed 200, articles 403 via curl, recovery captures content
/usr/bin/python3 -c "
import subprocess, os
env = os.environ.copy()
for k in ['http_proxy', 'https_proxy', 'HTTP_PROXY', 'HTTPS_PROXY']:
env.pop(k, None)
for feed in ['Google DeepMind Blog', 'One Useful Thing', 'Netflix Tech Blog', 'CrewAI Blog', 'Interconnects', 'AWS China Blog', 'AWS China ML']:
print(f'=== {feed} ===', flush=True)
try:
r = subprocess.run(['blogwatcher-cli', 'scan', feed], capture_output=True, text=True, timeout=20, env=env)
print(r.stdout[-500:] if r.stdout else '(no stdout)', flush=True)
if r.stderr: print('ERR:', r.stderr[-300:], flush=True)
except subprocess.TimeoutExpired:
print(f'TIMEOUT (20s) on {feed}', flush=True)
"
# Step 3: Scan problem feeds (Python fallback)
/usr/bin/python3 ~/.hermes/skills/wiki/rss-to-wiki-pipeline/scripts/fetch-problem-feeds.py
# Step 4: Write new articles to inbox (recovery is main path, watchdog is secondary)
# ⚠️ recovery.py can exceed default terminal timeout (120s) when SSL handshake
# delays accumulate across multiple feeds. Always use timeout ≥300s when
# calling via terminal(). If first attempt times out, retry with longer timeout.
/usr/bin/python3 ~/.hermes/skills/wiki/rss-to-wiki-pipeline/scripts/rss-inbox-recovery.py 2>&1 | tail -20
/usr/bin/python3 ~/.hermes/skills/wiki/rss-to-wiki-pipeline/scripts/rss-inbox-watchdog.py 2>&1 | tail -5
/usr/bin/python3 ~/.hermes/skills/wiki/rss-to-wiki-pipeline/scripts/rss-inbox-curl-recovery.py 2>&1 | tail -30
# Step 5: Remove duplicates already ingested to raw/articles/
# ⚠️ CRITICAL: Use FULL URL matching, NOT slug substring matching.
# ⚠️ TIMEOUT RISK: raw/articles/ has grown to ~3000 files. The glob+read+regex loop
# can exceed terminal's default 30s timeout. Always use terminal(..., timeout=120)
# or higher when calling the inline dedup script. If first attempt times out, retry
# with 120s — the script will complete.
# ⚠️ QUOTE NORMALIZATION (2026-08-16, rebuild 13:17): raw/articles stores ~1/3 of
# source_url values QUOTED ("https://...") while inbox files write unquoted URLs.
# The pre-fix script compared exact strings → 42 already-ingested articles re-entered
# the inbox EVERY cycle (recovery's own skip logic has the same blind spot). Strip
# surrounding quotes on BOTH sides before comparing.
/usr/bin/python3 -c "
import os, re
from pathlib import Path
wiki_root = Path.home() / 'wiki'
inbox_dir = wiki_root / 'raw/rss-inbox'
articles_dir = wiki_root / 'raw/articles'
def norm_url(u):
return u.strip().strip('\"').strip(\"'\").rstrip('/')
# Build set of FULL normalized URLs from raw/articles/
ingested_urls = set()
for f in articles_dir.glob('*.md'):
try:
content = f.read_text()
m = re.search(r'^(?:source_)?url:\\s*(.+)$', content, re.MULTILINE)
if m:
ingested_urls.add(norm_url(m.group(1)))
except: pass
# Remove inbox files whose FULL URL matches an ingested URL
removed = 0
for f in inbox_dir.glob('*.md'):
try:
content = f.read_text()
m = re.search(r'^(?:source_)?url:\\s*(.+)$', content, re.MULTILINE)
if m:
url = norm_url(m.group(1))
if url in ingested_urls:
f.unlink()
removed += 1
except: pass
print(f'Removed {removed} duplicates from inbox')
"
# Step 5b: Auto-filter vendor case-study PRs
/usr/bin/python3 ~/.hermes/skills/wiki/rss-to-wiki-pipeline/scripts/inbox-case-study-filter.py --threshold 1 --max-size 8192
# Step 5c: Remove sub-1KB trafilatura stubs
/usr/bin/python3 ~/.hermes/skills/wiki/rss-to-wiki-pipeline/scripts/rss-inbox-sub-1kb-filter.py
# Step 5d: Try trafilatura upgrade on remaining sub-2KB files
/usr/bin/python3 ~/.hermes/skills/wiki/rss-to-wiki-pipeline/scripts/rss-inbox-trafilatura-upgrade.py
# Step 6: URL health check (after all filter steps) — with per-domain breakdown
# IMPORTANT: Use 180s terminal timeout. Individual checks with --connect-timeout 2 avoid
# the DNS-latency wall that batching can hit. See "URL health check" pitfall.
/usr/bin/python3 -c "
import os, re, subprocess
from collections import defaultdict
from pathlib import Path
from urllib.parse import urlparse
inbox_dir = Path.home() / 'wiki' / 'raw' / 'rss-inbox'
domain_stats = defaultdict(lambda: {'ok': 0, 'fail': 0})
checked = 0
healthy = 0
for f in sorted(inbox_dir.glob('*.md')):
content = f.read_text()
m = re.search(r'^(?:source_)?url:\\\\s*(https?://[^\\s]+)', content, re.MULTILINE)
if not m:
continue
url = m.group(1).strip()
domain = urlparse(url).netloc
try:
r = subprocess.run(['curl', '-sL', '-o', '/dev/null', '-w', '%{http_code}', '--connect-timeout', '2', '--max-time', '5', url], capture_output=True, text=True, timeout=7)
code = r.stdout.strip()
if code in ('200', '301', '302', '307'):
domain_stats[domain]['ok'] += 1
healthy += 1
else:
domain_stats[domain]['fail'] += 1
checked += 1
except subprocess.TimeoutExpired:
domain_stats[domain]['fail'] += 1
checked += 1
print(f'URL-200={healthy}/{checked}')
for d in sorted(domain_stats.keys()):
s = domain_stats[d]
print(f' {d}: {s[\\\"ok\\\"]}/{s[\\\"ok\\\"]+s[\\\"fail\\\"]}')
"
# Step 7: Write to cron-status.log via wrapper (ABSOLUTE path required)
# IMPORTANT: assemble LINE via shell single-quoted string in terminal(), NOT execute_code
# (execute_code is blocked in cron mode)
/usr/bin/python3 -c "import os; from pathlib import Path; d=Path.home()/'wiki'/'raw'/'rss-inbox'; files=list(d.glob('*.md')); big=sum(1 for f in files if f.stat().st_size>=10240); mid=sum(1 for f in files if 5120<=f.stat().st_size<10240); small=sum(1 for f in files if f.stat().st_size<5120); print(f'big={big} mid={mid} small={small} total={len(files)}')"
# Then assemble LINE variable and write via:
# bash ~/.hermes/skills/wiki/rss-to-wiki-pipeline/scripts/cron-log-write.sh "$LINE"
cron-status.log 行格式规范
每行 rss-feed-scan 记录使用以下 schema(保持可解析、可 diff):
[YYYY-MM-DD HH:MM:SS] rss-feed-scan: scanned=<N> (<feed list summary>) + problem_feeds(<SP>+<HF>) | recovery=<R> written, <S> skipped | watchdog=<W> written (active when any feed unread>0; else 0 = short-circuit) | curl-recovery=<C> written (de-facto main path when urllib SSL times out; 0 = curl-fallback didn't run / recovered nothing / OR all URLs already in inbox from recovery — observed 2026-08-01 16:51: "32 URLs already in rss-inbox/" dedup means curl-recovery adds nothing when recovery.py ran first) | dedup_rm=<D> | case-study_rm=<CS> | sub-1KB-stub_rm=<S1> | trafilatura-upgrade=<TU> | inbox=<T> files (>=10KB=<B>, mid=<M>, small=<S2>) | URL-200=<H>/<T>
Stable Operational Patterns(当前已验证)
这些模式在多次独立 cron 中确认复现,未来 session 看到同样的指标组合可直接信任:
| 模式 | 触发条件 | 预期结果 | 验证 |
|---|
| GDM dead-end N=60 stable | GDM inbox=0 after filter chain | Skip GDM in cron health assessment. Two paths: (a) HTTP=0 → 0 GDM; (b) HTTP=200 with 50-60 stubs → sub-1KB filter removes → 0 in inbox. Don't equate "0 GDM" with "feed dead". Variant can REPEAT consecutively (2026-08-02 REFINES the 08-01 16:51 'alternation' claim): do NOT expect alternation between (a) curl-recovery=0/sub-1KB-stub_rm=0 and (b) curl-recovery=53/sub-1KB-stub_rm=53 — either variant can run consecutively (observed 4 straight variant-b cycles 08-02 04:26→10:39; 16:51 was variant-a, isolated). Both converge to 0 GDM in inbox with the same byte-identical 28-file signature. Any curl-recovery/sub-1KB count = GDM variant marker, NOT an anomaly or feed-health change; never read it as feed health. | 60+ independent crons |
| Inbox shape = 31-39 normal (post-Netflix baseline, v5.9.23.38 REFINED); fully-drained steady state ~28-32 (v5.9.23.62 ADDED); extreme deep-drain <20 (v5.9.23.64 ADDED); composition range broadened 2026-07-29 (v5.9.23.81 REFINED) | 7 good feeds after backlog sweep | 31-39 files = new normal when any feed still has backlog beyond CrewAI. Composition: 25-31 >=10KB + 6-8 mid + 0 small. Fully-drained steady state (v5.9.23.62+): inbox=28-32 (22-24 >=10KB + 6-8 mid + 0 small), below 31-39 range when all backlog feeds exhausted. If inbox drops below 25: check whether OUT milestone is absent (expected, causes additional drop) vs. feed health issue. Extreme deep-drain (v5.9.23.64, range broadened v5.9.23.81, dual-trigger confirmed v5.9.23.82): Two trigger paths: (a) recovery net contribution ~0 for sustained cycles (all feeds New=0, all recovered articles filtered by downstream filters, OUT milestone ingested) — inbox bottoms at 12-15; (b) non-zero recovery but high filter consumption — observed 2026-07-30: recovery=23 with 4 case-study + 54 sub-1KB removed → inbox=19. In both cases, variable composition applies: previously observed at 4-6 big + 6-8 mid (v5.9.23.64), but also observed at 10 big + 5 mid (2026-07-29) when AWS big-files dominate the residual set. The composition is NOT fixed — it depends on which feed's articles survive the filter chain. What matters is the total count <20. This is NOT a feed health signal — it means ALL backlogs are fully exhausted and even prior-cycle residuals are being consumed. URL health check at extreme-drain typically shows 12-15/12-15 reachable. v5.9.23.82 variant added (2026-07-30): extreme deep-drain can ALSO occur with non-zero recovery (e.g. 23) when filters consume the majority of incoming articles. Recovery=23 with 4 case-study + 54 sub-1KB removed → net inbox=19 (14 big + 5 mid). The trigger is filter chain efficiency rather than recovery floor. Was previously recorded as a separate mode (v5.9.23.65), but inbox bounced back from 24→29 in the next cycle, confirming sub-28 values are oscillatory dips within the fully-drained steady state, not a permanent downward shift. The inbox fluctuates within 28-32 nominal with occasional dips to 24-25 when recovery net is near-zero and mid-size files are consumed faster than big files. Treat intermediate-range inbox as "fully-drained steady state with transient dip", not a separate trajectory. during AWS-CB/ML partial-drain (recovery=23-28), inbox=24 (18 big + 5 mid + 1 small) HELD STEADY across 2 consecutive cycles (2026-08-01 06:22 + 08:26, identical composition) — a persistent partial-drain shape, NOT a transient dip. Distinguish by recovery value: recovery=28 with inbox=24 = stable partial-drain; recovery≈6-10 with inbox=24 = oscillatory dip within fully-drained state. |