Skip to main content 홈 크리에이터 jeremylongshore tons-of-skills-marketplace firecrawl-data-handling
firecrawl-data-handling Process, validate, and store Firecrawl scraped content with deduplication and chunking.
Use when handling scraped markdown, implementing content pipelines, building RAG knowledge
bases, or processing crawl results for downstream consumption.
Trigger with phrases like "firecrawl data", "firecrawl content processing",
"firecrawl markdown cleaning", "firecrawl storage", "firecrawl RAG pipeline".
설치로 이동 Skills Marketplace 커뮤니티가 만든 AI 스킬을 발견하고 탐색하세요.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
직접 명령은 검토 Prompt를 거치지 않습니다. 실행하기 전에 소스를 확인하세요.
npx skills add https://github.com/jeremylongshore/tons-of-skills-marketplace --skill firecrawl-data-handling명령은 한 줄로 유지됩니다. 복사하기 전에 가로로 스크롤해 전체 내용을 확인하세요.
로컬 사본을 원하시나요? SkillsMP에서 현재 제공할 수 있는 파일을 다운로드하세요.
Zip 다운로드 다운로드 중... 이 저장소의 다른 Skills langchain-deploy-integration Deploy a LangChain 1.0 / LangGraph 1.0 app to Cloud Run, Vercel, or LangServe correctly — with timeouts sized for chain length, cold-start mitigation, SSE anti-buffering headers, and Secret Manager over .env. Use when prepping a first production deploy, debugging a stream that hangs behind a proxy, or diagnosing p99 latency spikes. Trigger with "langchain deploy", "langchain cloud run", "langchain vercel python", "langchain langserve", or "langchain docker".
langchain-langgraph-agents Build a correct LangGraph 1.0 ReAct agent with create_react_agent — typed tools, error propagation, recursion caps, and stop conditions that actually stop. Use when writing a first tool-calling agent, migrating from AgentExecutor or initialize_agent, or diagnosing an agent that loops on vague prompts. Trigger with "langgraph agent", "create_react_agent", "langgraph tool calling", "AgentExecutor migration", or "agent loop cost".
langchain-langgraph-human-in-loop Build LangGraph 1.0 human-in-the-loop approval flows with interrupt_before /
interrupt_after and Command(resume=...) — JSON-serializable state, clean
resume semantics, and UI wiring for approval decisions. Use when adding an
approval gate before an expensive tool call, wiring a Slack/web UI for agent
approvals, or debugging a graph that crashes on interrupt.
Trigger with "langgraph human in loop", "langgraph interrupt_before",
"langgraph approval flow", "Command resume", "langgraph HITL".
name firecrawl-data-handling description Process, validate, and store Firecrawl scraped content with deduplication and chunking.
Use when handling scraped markdown, implementing content pipelines, building RAG knowledge
bases, or processing crawl results for downstream consumption.
Trigger with phrases like "firecrawl data", "firecrawl content processing",
"firecrawl markdown cleaning", "firecrawl storage", "firecrawl RAG pipeline".
allowed-tools Read, Write, Edit version 1.11.0 license MIT author Jeremy Longshore <jeremy@intentsolutions.io> tags ["saas","firecrawl","compliance"] compatibility Designed for Claude Code
Firecrawl Data Handling
Overview
Process scraped web content from Firecrawl pipelines. Covers markdown cleaning, structured data extraction with Zod validation, content deduplication, chunking for LLM/RAG, and storage patterns for crawled content.
Instructions
Step 1: Content Cleaning
import FirecrawlApp from "@mendable/firecrawl-js" ;
const firecrawl = new FirecrawlApp ({
apiKey : process.env .FIRECRAWL_API_KEY !,
});
async function scrapeClean (url : string ) {
const result = await firecrawl.scrapeUrl (url, {
formats : ["markdown" ],
onlyMainContent : true ,
excludeTags : ["script" , "style" , "nav" , "footer" , "iframe" ],
waitFor : 2000 ,
});
return {
url : result.metadata ?.sourceURL || url,
title : result.metadata ?.title || "" ,
markdown : cleanMarkdown (result.markdown || "" ),
scrapedAt : (). (),
};
}
( ): {
md
. ( , )
. ( , )
. ( , )
. ( , )
. ( , )
. ();
}
new
Date
toISOString
function
cleanMarkdown
md : string
string
return
replace
/\n{3,}/g
"\n\n"
replace
/\[.*?\]\(javascript:.*?\)/g
""
replace
/!\[.*?\]\(data:.*?\)/g
""
replace
/<!--[\s\S]*?-->/g
""
replace
/<script[\s\S]*?<\/script>/gi
""
trim
Step 2: Structured Extraction with Validation import { z } from "zod" ;
const ArticleSchema = z.object ({
title : z.string ().min (1 ),
author : z.string ().optional (),
publishedDate : z.string ().optional (),
content : z.string ().min (50 ),
wordCount : z.number (),
});
async function extractArticle (url : string ) {
const result = await firecrawl.scrapeUrl (url, {
formats : ["extract" ],
extract : {
schema : {
type : "object" ,
properties : {
title : { type : "string" },
author : { type : "string" },
publishedDate : { type : "string" },
content : { type : "string" },
},
required : ["title" , "content" ],
},
},
});
if (!result.extract ) throw new Error (`Extraction failed for ${url} ` );
return ArticleSchema .parse ({
...result.extract ,
wordCount : (result.extract .content || "" ).split (/\s+/ ).length ,
});
}
Step 3: Content Deduplication import { createHash } from "crypto" ;
function contentHash (text : string ): string {
return createHash ("sha256" )
.update (text.trim ().toLowerCase ())
.digest ("hex" );
}
function deduplicatePages (pages : Array <{ url: string ; markdown: string }> ) {
const seen = new Map <string , string >();
const unique : typeof pages = [];
const duplicates : Array <{ url : string ; duplicateOf : string }> = [];
for (const page of pages) {
const hash = contentHash (page.markdown );
if (seen.has (hash)) {
duplicates.push ({ url : page.url , duplicateOf : seen.get (hash)! });
} else {
seen.set (hash, page.url );
unique.push (page);
}
}
console .log (`Dedup: ${pages.length} input, ${unique.length} unique, ${duplicates.length} duplicates` );
return { unique, duplicates };
}
Step 4: Chunk for LLM / RAG interface ContentChunk {
url : string ;
title : string ;
chunkIndex : number ;
content : string ;
wordCount : number ;
}
function chunkForRAG (
url : string ,
title : string ,
markdown : string ,
maxWords = 800
): ContentChunk [] {
const sections = markdown.split (/\n(?=#{1,3}\s)/ );
const chunks : ContentChunk [] = [];
let current = "" ;
let index = 0 ;
for (const section of sections) {
const combined = current ? `${current} \n\n${section} ` : section;
if (combined.split (/\s+/ ).length > maxWords && current) {
chunks.push ({
url, title, chunkIndex : index++,
content : current.trim (),
wordCount : current.split (/\s+/ ).length ,
});
current = section;
} else {
current = combined;
}
}
if (current.trim ()) {
chunks.push ({
url, title, chunkIndex : index,
content : current.trim (),
wordCount : current.split (/\s+/ ).length ,
});
}
return chunks;
}
Step 5: Crawl and Store Pipeline import { writeFileSync, mkdirSync } from "fs" ;
import { join } from "path" ;
async function crawlAndStore (baseUrl : string , outputDir : string , opts ?: {
maxPages?: number ;
paths?: string [];
} ) {
mkdirSync (outputDir, { recursive : true });
const crawlResult = await firecrawl.crawlUrl (baseUrl, {
limit : opts?.maxPages || 50 ,
includePaths : opts?.paths ,
scrapeOptions : { formats : ["markdown" ], onlyMainContent : true },
});
const pages = (crawlResult.data || []).map (page => ({
url : page.metadata ?.sourceURL || baseUrl,
markdown : cleanMarkdown (page.markdown || "" ),
}));
const { unique } = deduplicatePages (pages);
const manifest = unique.map (page => {
const slug = new URL (page.url ).pathname
.replace (/\//g , "_" ).replace (/^_|_$/g , "" ) || "index" ;
const filename = `${slug} .md` ;
writeFileSync (join (outputDir, filename), page.markdown );
return { url : page.url , file : filename, size : page.markdown .length };
});
writeFileSync (join (outputDir, "manifest.json" ), JSON .stringify (manifest, null , 2 ));
return manifest;
}
Error Handling Issue Cause Solution Empty content JS not rendered Increase waitFor, use onlyMainContent Garbage in markdown Bad HTML cleanup Add excludeTags for problematic elements Duplicate pages URL aliases or redirects Content-hash deduplication Oversized chunks Long single sections Add word limit to chunking logic Extract returns null Page too complex for LLM Simplify schema, use shorter prompt
Examples
Documentation Scraper with RAG Output const docs = await crawlAndStore ("https://docs.example.com" , "./scraped-docs" , {
maxPages : 50 ,
paths : ["/docs/*" , "/api/*" ],
});
for (const doc of docs) {
const content = readFileSync (`./scraped-docs/${doc.file} ` , "utf-8" );
const chunks = chunkForRAG (doc.url , doc.file , content);
console .log (`${doc.url} : ${chunks.length} chunks` );
}
Resources
Next Steps For access control, see firecrawl-enterprise-rbac.