xiaohongshu-deep-profile-collect
采集小红书个人主页的全量信息,包括基础资料、全部发布笔记(含详情正文和标签)、收藏笔记列表(全量)、点赞笔记列表(全量)
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
采集小红书个人主页的全量信息,包括基础资料、全部发布笔记(含详情正文和标签)、收藏笔记列表(全量)、点赞笔记列表(全量)
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Collect user data from logged-in social platforms (Douyin, Xiaohongshu, Weibo, Douban, Bilibili), cross-analyze to build a precise personal profile, and auto-generate USER.md + MEMORY.md. Use for new user onboarding, personalization, and building user context. 通过用户已登录的社交平台自动采集数据并交叉分析,生成精准用户画像,写入 USER.md 和 MEMORY.md。
采集豆瓣当前登录用户的全量个人信息,包括基础资料、看过/想看的电影、读过/想读的书、广播动态
采集B站当前登录用户的全量个人信息,包括基础资料、投稿视频、收藏夹内容、关注列表
深度采集抖音个人主页信息,包含基础资料、全部作品(含播放数)、喜欢列表、收藏列表、全量关注列表
采集微博个人主页的全量信息,包括基础资料($CONFIG.user)、全部微博、关注列表、收藏列表
| name | xiaohongshu-deep-profile-collect |
| description | 采集小红书个人主页的全量信息,包括基础资料、全部发布笔记(含详情正文和标签)、收藏笔记列表(全量)、点赞笔记列表(全量) |
| version | 3.0.0 |
本 Skill 用于采集小红书当前登录用户的个人主页全量信息。通过 XHR 拦截技术突破虚拟滚动限制,实现收藏和点赞列表的全量采集。v3.0 新增笔记详情采集(正文内容+标签),通过逐条点击笔记打开详情页并从 __INITIAL_STATE__ 提取结构化数据。
采集内容:
核心技术:XHR 拦截 —— 小红书使用虚拟滚动,DOM 中最多保留约 28 个 note-item 元素,旧元素会被回收。必须通过拦截 XMLHttpRequest 捕获 API 响应来获取全量数据。
⚠️ 关键认知:小红书个人主页有三个独立 tab(笔记/收藏/点赞),数据完全独立。必须分别导航到每个 tab 才能采集对应数据。不要因为某个 tab 数据为0就跳过后续 tab。
⚠️ 执行约束:本 Skill 中的所有 JS 脚本都必须完整复制执行,不可自行简化、改写或省略任何部分。每个脚本都经过实际测试验证,简化会导致数据不完整。
用户示例:
使用 xiaohongshu-deep-profile-collect 采集我的小红书个人主页信息
无需提供参数,自动检测当前登录用户。
无。本 Skill 自动从小红书首页导航栏检测当前登录用户。
前提条件:浏览器已登录小红书账号。
工具:chrome_navigate
参数:
url: https://www.xiaohongshu.com/说明:小红书不支持 /user/self 直达个人主页,需要先访问首页获取当前用户的主页 URL。
工具:chrome_execute_script
参数:
world: MAINtimeout: 10000jsScript:() => {
const links = document.querySelectorAll('a');
for (let a of links) {
if (a.textContent.trim() === '我' &&
(a.getAttribute('href')||'').includes('/user/profile/')) {
return JSON.stringify({
profileUrl: 'https://www.xiaohongshu.com' + a.getAttribute('href')
});
}
}
return JSON.stringify({profileUrl: null});
}
说明:从首页导航栏找到文本为"我"且 href 包含 /user/profile/ 的链接,提取完整的个人主页 URL。
输出:profileUrl — 后续步骤使用此 URL 导航。
工具:chrome_navigate
参数:
url: {step_2.profileUrl}(步骤2返回的 profileUrl)说明:导航到个人主页,默认显示笔记 tab。
工具:chrome_execute_script
参数:
world: MAINtimeout: 120000jsScript:async () => {
// 智能滚动:连续5轮无新增note-item则停止
let prevCount = 0, sameCount = 0;
for (let i = 0; i < 50; i++) {
window.scrollTo({top: document.body.scrollHeight, behavior: 'smooth'});
await new Promise(r => setTimeout(r, 1500));
const c = document.querySelectorAll('section.note-item').length;
if (c === prevCount) { sameCount++; if (sameCount > 5) break; }
else { sameCount = 0; }
prevCount = c;
}
window.scrollTo({top: 0}); await new Promise(r => setTimeout(r, 500));
const notes = [];
document.querySelectorAll('section.note-item').forEach(n => {
notes.push(n.textContent.trim().substring(0, 120));
});
return JSON.stringify({
noteCount: notes.length, notes: notes,
fullText: document.body.innerText
});
}
说明:滚动加载全部笔记卡片。fullText 包含基础资料(昵称/小红书号/IP属地/简介/学校/关注粉丝获赞数)。笔记 tab 的卡片数量即为发布笔记数(不走 XHR 分页)。
输出:noteCount、notes(笔记标题列表)、fullText(含基础资料)
⚠️ 步骤4只是采集"笔记"tab的数据。收藏和点赞是独立的 tab,必须分别导航过去才能采集。即使笔记数为0,也要继续执行步骤5-10采集收藏和点赞。
核心技术:小红书笔记详情(正文 desc、标签 tagList)无法通过列表页 DOM 或 API 直接获取。必须逐条点击笔记打开详情页(新 tab),从详情页的 window.__INITIAL_STATE__.note.noteDetailMap 中提取结构化数据。
⚠️ 关键认知:
a.cover 链接会在新 tab 中打开笔记详情页(不是弹窗)工具:chrome_execute_script + get_windows_and_tabs + chrome_close_tabs
流程:对每条笔记重复以下子步骤:
工具:chrome_execute_script
参数:
tabId: {个人主页 tabId}world: MAINtimeout: 10000jsScript:async () => {
// {noteId} 替换为当前要采集的笔记 ID
const cover = document.querySelector('section.note-item a.cover[href*="{noteId}"]');
if (!cover) return JSON.stringify({error: 'no cover for {noteId}'});
cover.click();
await new Promise(r => setTimeout(r, 2000));
return JSON.stringify({clicked: true});
}
说明:通过 CSS 选择器精准定位到包含目标 noteId 的封面链接并点击。点击后会在新 tab 中打开笔记详情页。等待 2 秒让新 tab 加载。
工具:get_windows_and_tabs
说明:调用后在返回的 tabs 列表中查找 URL 包含当前 noteId 的 tab,记录其 tabId。
工具:chrome_execute_script
参数:
tabId: {子步骤 B 找到的新 tab ID}world: MAINtimeout: 15000jsScript:() => {
try {
const noteMap = JSON.parse(JSON.stringify(
window.__INITIAL_STATE__.note.noteDetailMap
));
const noteIds = Object.keys(noteMap).filter(
k => k !== 'undefined' && noteMap[k].note && noteMap[k].note.noteId
);
if (noteIds.length === 0) return JSON.stringify({error: 'no notes in map'});
const n = noteMap[noteIds[0]].note;
return JSON.stringify({
noteId: n.noteId,
title: n.title || n.displayTitle || '',
desc: (n.desc || ''),
tagList: (n.tagList || []).map(t => ({name: t.name, type: t.type})),
type: n.type,
time: n.time,
interactInfo: n.interactInfo ? {
likedCount: n.interactInfo.likedCount || '0',
collectedCount: n.interactInfo.collectedCount || '0',
commentCount: n.interactInfo.commentCount || '0'
} : null,
imageCount: n.imageList ? n.imageList.length : 0
});
} catch(e) {
return JSON.stringify({error: e.message});
}
}
输出:笔记详情(noteId / title / desc / tagList / type / time / interactInfo / imageCount)
数据结构:
{
"noteId": "65a1b2c3000000001234abcd",
"title": "笔记标题",
"desc": "笔记正文内容...\n#标签1[话题]# #标签2[话题]#",
"tagList": [
{"name": "标签名", "type": "topic"}
],
"type": "normal",
"time": 1769476094000,
"interactInfo": {
"likedCount": "123",
"collectedCount": "45",
"commentCount": "6"
},
"imageCount": 3
}
工具:chrome_close_tabs
参数:
tabIds: [{子步骤 B 的 tabId}]说明:关闭详情页 tab,防止 tab 堆积导致浏览器卡顿。
__INITIAL_STATE__.user.notes 中提取video 的笔记 desc 可能为空或很短,这是正常的工具:chrome_navigate
参数:
url: {step_2.profileUrl}?tab=fav&subTab=note说明:⚠️ 关键步骤,不可跳过! 必须使用 URL 参数方式导航到收藏 tab(?tab=fav&subTab=note)。JS 点击 .reds-tab-item 不会触发数据刷新。收藏数据只有在收藏 tab 下滚动时才会通过 API 加载。
工具:chrome_execute_script
参数:
world: MAINtimeout: 5000jsScript:() => {
window.__xhs_collected = [];
window.__xhs_done = false;
window.__xhs_api_count = 0;
const origOpen = XMLHttpRequest.prototype.open;
const origSend = XMLHttpRequest.prototype.send;
XMLHttpRequest.prototype.open = function(method, url, ...rest) {
this.__xhs_url = url;
return origOpen.call(this, method, url, ...rest);
};
XMLHttpRequest.prototype.send = function(...args) {
if (this.__xhs_url &&
(this.__xhs_url.includes('note/collect/page') ||
this.__xhs_url.includes('note/like/page'))) {
this.addEventListener('load', function() {
try {
const data = JSON.parse(this.responseText);
if (data.data && data.data.notes) {
window.__xhs_api_count++;
data.data.notes.forEach(note => {
window.__xhs_collected.push({
note_id: note.note_id,
display_title: note.display_title || '',
type: note.type,
user: note.user ? {
nickname: note.user.nickname,
user_id: note.user.user_id
} : null,
interact_info: note.interact_info ? {
liked_count: note.interact_info.liked_count
} : null
});
});
if (!data.data.has_more) window.__xhs_done = true;
}
} catch(e) {}
});
}
return origSend.apply(this, args);
};
return JSON.stringify({injected: true});
}
说明:注入 XHR 拦截器,监听小红书收藏/点赞 API 的响应。API endpoints:
edith.xiaohongshu.com/api/sns/web/v2/note/collect/page?num=30&cursor=xxxedith.xiaohongshu.com/api/sns/web/v1/note/like/page?num=30&cursor=xxx拦截器自动从 API 响应中提取 note 结构化数据,检测 has_more=false 时设置完成标志。
工具:chrome_execute_script
参数:
world: MAINtimeout: 120000jsScript:async () => {
const MAX_ITEMS = 500; // 上限500条,防止超大收藏列表采集过久
for (let batch = 0; batch < 30; batch++) {
for (let i = 0; i < 5; i++) {
window.scrollTo({top: document.body.scrollHeight, behavior: 'smooth'});
await new Promise(r => setTimeout(r, 1200));
}
if (window.__xhs_done) break;
if (window.__xhs_collected && window.__xhs_collected.length >= MAX_ITEMS) break;
}
const seen = new Set();
const unique = [];
window.__xhs_collected.forEach(n => {
if (!seen.has(n.note_id) && unique.length < MAX_ITEMS) { seen.add(n.note_id); unique.push(n); }
});
const sampled = window.__xhs_collected && window.__xhs_collected.length > MAX_ITEMS;
return JSON.stringify({
total: unique.length, notes: unique,
sampled, totalBeforeSample: window.__xhs_collected ? window.__xhs_collected.length : unique.length,
apiCalls: window.__xhs_api_count
});
}
说明:持续滚动触发 API 请求加载更多数据。每批滚动 5 次(间隔 1.2s),最多 30 批。拦截器自动收集数据,滚动结束后去重提取。
输出:收藏笔记全量列表(含 note_id / display_title / type / user / interact_info)
工具:chrome_navigate
参数:
url: {step_2.profileUrl}?tab=liked&subTab=note说明:⚠️ 关键步骤,不可跳过! 必须导航到点赞 tab(?tab=liked&subTab=note)。点赞数据只有在这个 tab 下滚动才会触发 API 加载。即使前面收藏数据为0,点赞仍可能有数据,反之亦然。
工具:chrome_execute_script
说明:每次页面导航后拦截器会失效,需要重新注入。使用与步骤6完全相同的脚本。
工具:chrome_execute_script
说明:与步骤7完全相同的滚动+提取逻辑。
输出:点赞笔记全量列表
最终结果:此步骤完成后,所有数据采集完毕。
工具:chrome_close_tabs
说明:关闭所有打开的标签页,清理资源。
本 Skill 的核心技术是 XHR 拦截,原因:
section.note-item 最多保留约 28 个,旧元素被回收XMLHttpRequest.prototype.open/send,在 API 响应到达时提取数据has_more=false 表示全部加载完毕⚠️ 小红书的笔记/收藏/点赞 tab 不能用 JS click 切换:
document.querySelector('.reds-tab-item').click() — 数据不刷新chrome_navigate(profileUrl + '?tab=fav&subTab=note') — 正确加载每条笔记的结构:
{
"note_id": "65110e20000000001d01629e",
"display_title": "笔记标题",
"type": "video|normal",
"user": {"nickname": "作者昵称", "user_id": "作者ID"},
"interact_info": {"liked_count": "4820"}
}
问题1:profileUrl 为 null
{profileUrl: null}问题2:收藏/点赞数量为0
问题3:数量少于预期
✅ 小红书个人主页深度采集完成!
📊 基础资料:
- 昵称:{用户昵称}
- 小红书号:{小红书号}
- IP属地:{IP属地}
- 简介:{个人简介}
- 学校:{学校}
- 关注:{N} | 粉丝:{N} | 获赞与收藏:{N}
📝 发布笔记:{N}条
⭐ 收藏笔记:{N}条(全量)
👍 点赞笔记:{N}条(全量)