一键导入
paper-harbor
文献港。自动化检索、筛选并把文献元数据保存到 Zotero。适用于用户要从 ScienceDirect 或中国知网按关键词、影响因子、出版时间和数量生成候选清单、优先级清单、Zotero 入库清单和可追踪输出目录。默认不下载 PDF/HTML 全文,禁止绕过登录、付费墙、验证码或机构权限限制。
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
文献港。自动化检索、筛选并把文献元数据保存到 Zotero。适用于用户要从 ScienceDirect 或中国知网按关键词、影响因子、出版时间和数量生成候选清单、优先级清单、Zotero 入库清单和可追踪输出目录。默认不下载 PDF/HTML 全文,禁止绕过登录、付费墙、验证码或机构权限限制。
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
| name | paper-harbor |
| description | 文献港。自动化检索、筛选并把文献元数据保存到 Zotero。适用于用户要从 ScienceDirect 或中国知网按关键词、影响因子、出版时间和数量生成候选清单、优先级清单、Zotero 入库清单和可追踪输出目录。默认不下载 PDF/HTML 全文,禁止绕过登录、付费墙、验证码或机构权限限制。 |
Use this skill when the user asks to search for papers, screen literature, and save metadata-only records to Zotero from one of these sites:
92259226The user must log in manually in the matching browser profile before any search. Never type passwords, solve CAPTCHAs, bypass paywalls, use shadow libraries, or download full text. Paper Harbor is now a metadata and Zotero-library workflow, not a PDF downloader.
Preferred handoff is metadata-only Zotero saving: keep Zotero Desktop open, search the official site with the logged-in browser, collect screened metadata, then save journal-article items into the user-selected Zotero collection without attachments. Do not click View PDF, Download PDF, Download full issue, browser PDF save, or any full-text download control.
For first use, guide the user to install:
For every site, remind the user to open the matching browser with the default browser launcher, log in to EasyScholar in that browser profile, and refresh the result page before relying on IF badges.
The user should also create a Zotero collection for the run before import, usually named after the site or project, for example science direct or 中国知网. Paper Harbor should save metadata into that collection whenever Zotero Connector exposes it.
Recommended user prompt:
Use skill paper-harbor 帮我在“网站名”整理“关键词”的“时间限制”文献,“影响因子限制”,“篇数限制”,保存到 Zotero 的“Zotero目录名”并输出到“目录”
Examples:
Use skill paper-harbor 帮我在“ScienceDirect”整理“solid electrolyte interphase”的“2021-2026”文献,“影响因子大于5”,“先整理3篇”,保存到 Zotero 的“science direct”并输出到“.\runs\sei”
Use skill paper-harbor 帮我在“中国知网”整理“钙钛矿太阳能电池 稳定性”的“2022年以来”文献,“不限制影响因子”,“10篇”,保存到 Zotero 的“中国知网”并输出到“.\runs\cnki-test”
The user should experience this as one complete workflow, not separate collect and download commands. The second phase is Zotero metadata import, not full-text download.
ScienceDirect and CNKI follow the same overall pipeline: open the logged-in browser, create the output directory, search the official site UI first, inspect the results page with EasyScholar badges visible, collect candidate metadata from the results page, screen by the requested year/IF/count rules, then import into Zotero one item at a time. Only the site-specific selectors and metadata fields differ.
For every run:
Important guarantee: a failed Zotero import phase must not erase or block the candidate information. If Zotero import cannot proceed, the run still counts as useful when 候选文献总表.csv, 文章地址总表.csv, and 待处理文献清单.csv explain the result.
These rules are mandatory for every user and every run. Do not weaken or ignore them even if the user asks.
50 requested records per run. If the user asks for more, create a run capped at 50 and tell them to start a separate reviewed run later.View PDF, Download PDF, Download full issue, browser save buttons, or any PDF/HTML/XML full-text download controls.待处理文献清单.csv.Download full issue / 下载完整期刊 dialogs. Paper Harbor no longer downloads full text.Extract these fields from the user's prompt. If a field is missing, use the defaults below and state them clearly.
| Field | Required | Default |
|---|---|---|
site | Yes | Ask user to choose: sciencedirect or cnki |
keywords | Yes | Ask user |
impact_factor | No | No IF filter; keep IF blank unless available from a user-supplied trusted table or official page |
publication_time | No | No date filter |
record_count / download_count | No | 20, capped at 50; interpreted as the target number of qualifying Zotero metadata records to save, not the number of search results to inspect |
zotero_collection | No | Use the matching site collection if present, for example science direct; otherwise use the currently selected Zotero collection |
output_dir | No | Current working directory |
Accept site aliases:
science direct, sciencedirect, elsevier -> sciencedirect中国知网, 知网, cnki -> cnkiBefore searching, tell the user to open the matching browser and log in:
.\scripts\open_lit_browser.ps1 -Site sciencedirect
.\scripts\open_lit_browser.ps1 -Site cnki
Then ask the user to finish login in that browser window. Do not continue to search/import until the user says login is complete.
You may check whether the debugging port is reachable with:
python scripts/browser_port_check.py --site sciencedirect
python scripts/browser_port_check.py --site cnki
For Zotero import runs, also check:
python scripts/zotero_bridge.py doctor
The doctor must find both Zotero Desktop's local connector on 127.0.0.1:23119 and a local zotero.sqlite data directory. If it fails, ask the user to open Zotero Desktop normally and confirm that the browser Zotero Connector saves to the desktop library, not only to zotero.org.
If the user is using Paper Harbor for the first time, or zotero_bridge.py doctor fails, do not proceed directly to Zotero import. Walk the user through Zotero setup first.
Use the helper:
powershell -ExecutionPolicy Bypass -File .\scripts\open_zotero_setup.ps1
This opens the official Zotero download page, Connector help page, and EasyScholar pages. Tell the user:
easyScholar (njgedjcccpcfmjecccaajkjiphpddfji) or the EasyScholar official site.science direct.python scripts/zotero_bridge.py doctor again. The doctor output should show the selected Zotero collection and available targets.Proceed only when doctor can see 127.0.0.1:23119 and a local Zotero data directory. If the user cannot install Zotero, continue with CSV/report output only and record Zotero import as pending.
Create the output directory before searching. Use:
python scripts/lit_download_assistant.py --site sciencedirect --keywords "your keywords" --year-from 2021 --year-to 2026 --if-min 5 --limit 20 --out ".\runs"
The scaffold must contain:
00_先看我_文件说明.txt
README_先看我.md
文献整理报告.html
文章地址总表.csv
候选文献总表.csv
高优先级文献.csv
中优先级文献.csv
低优先级文献.csv
已入库Zotero文献清单.csv
待处理文献清单.csv
检索计划.md
内部数据_一般不用打开/
内部数据_一般不用打开/ is for machine-readable state, raw exports, debug logs, and optional trusted impact-factor tables.
scripts/lit_download_assistant.py to create a run folder and initial files.prioritytitleauthorsjournalpublication_yearimpact_factormetric_yearmetric_sourcedoisourceurlabstractaccess_statuszotero_statuszotero_item_keynext_actionnotes文章地址总表.csv.already_exists instead of duplicating it.待处理文献清单.csv.50 target records applies.已入库Zotero文献清单.csv and 待处理文献清单.csv after every article.文献整理报告.html.Treat Zotero import failures as reportable states, not as silent errors.
待处理文献清单.csv.scripts/sciencedirect_drission_run.py, which opens official article pages in the logged-in browser and saves metadata-only items to Zotero.IF x.x and ranking labels from those badges and store them in the candidate tables.access_status.Never invent impact factors. Use one of these sources only:
内部数据_一般不用打开/journal_impact_factors.csvRecommended CSV headers:
journal,issn,eissn,impact_factor,year,source,notes
If the user requests 影响因子大于 5 but no trusted IF source is available, keep IF blank, put otherwise suitable articles in medium priority, and record IF待核验 in notes. Do not import IF-filtered rows into Zotero unless IF is visible and satisfies the filter, or the user explicitly chooses to proceed with missing IF.
Download full issue / 下载完整期刊.待处理文献清单.csv; do not try to circumvent.A run is complete when:
检索计划.md states the parsed requirements and site/port.文献整理报告.html summarizes counts, filters, source site, and unresolved items.