بنقرة واحدة
scraping-skills
يحتوي scraping-skills على 12 من skills المجمعة من thirdwatch-dev، مع تغطية مهنية على مستوى المستودع وصفحات skill داخل الموقع.
Skills في هذا المستودع
Use when a site blocks your scraper and you need to get past it — 403 Forbidden on the first request, JS challenges, CAPTCHAs, Cloudflare Turnstile, DataDome, Akamai, PerimeterX, "Just a moment...", access denied, IP bans, or an empty/skeleton page where data should be. Covers diagnosing transient failures vs real blocks, the cheapest-first technique ladder (plain HTTP → TLS fingerprint spoof → stealth browser → full browser), TLS spoofing with impit/curl_cffi, Camoufox stealth-browser setup, the Cloudflare Turnstile iframe-click, homepage warmup for cookie-based defenses, browser cookie reuse, XHR interception, proxy selection (residential vs datacenter vs SERP), and which targets genuinely resist affordable bypasses. Triggers on "site is blocking me", "getting 403", "Cloudflare", "DataDome", "Turnstile", "CAPTCHA", "bypass anti-bot", "scraper stopped working", "blocked".
Use when you need to scrape e-commerce product data — titles, brands, prices, MRP/list price, discounts, ratings, variants/sizes, availability, images, and product URLs — from online storefronts and marketplaces. Covers Amazon, Flipkart, AliExpress, Myntra, Nykaa, Meesho, Snapdeal, Noon, AJIO, FirstCry, Tata Cliq, and any Shopify store. For price monitoring / MAP enforcement, catalog building, competitor price tracking, and dropshipping product research. Triggers on "scrape products", "price monitoring", "product catalog", "competitor prices", "MAP", "track prices", "build a product database", "compare prices across stores".
Use when you need to scrape social media or content platforms — posts, engagement metrics, profiles, videos, ads, or transcripts — from Twitter/X, Instagram, TikTok, Reddit, LinkedIn, Pinterest, YouTube, Facebook Ad Library, or IMDb. Covers social listening, influencer research, content and trend analysis, ad-creative intelligence, and building RAG datasets from video transcripts. Triggers on "scrape social media", "influencer", "social listening", "engagement metrics", "ad library", "video transcripts", "scrape tweets/posts/reels/videos", "subreddit data".
Use when you want to package a web scraper as a deployable, monetizable Apify Actor in Python — set up the directory structure, choose a runtime pattern (HTTP / impit / Playwright / Camoufox), write the input/output/dataset schemas, deploy with the apify CLI, and configure pay-per-event pricing. Triggers on "build an Apify actor", "deploy/publish a scraper to Apify", "monetize a scraper", "pay-per-event pricing", "Apify input schema", "Apify output schema", "actor.json".
Use when you need B2B leads or company data — business listings with contact info, supplier directories, company/people enrichment, KYB/verification, or trade data. Covers Google Maps (local listings), IndiaMart and JustDial (supplier/business directories), LinkedIn (company employees and candidate finder), GST Verification (India KYB), UN Comtrade (trade data), and Product Hunt (launches). Triggers on "lead generation", "B2B leads", "find businesses", "company data", "supplier directory", "KYB", "due diligence", "market sizing", "sales prospecting".
Use when you need data from food-delivery platforms — restaurants, menus, item prices, ratings, and delivery times/fees — for menu and price intelligence, restaurant coverage analysis, market analysis, or competitor menu tracking. Covers Swiggy and Zomato (India), Talabat (Middle East), Deliveroo (UK/UAE/EU), and Noon Food. Triggers on "scrape restaurants", "menu data", "food delivery", "restaurant prices", "menu prices", "delivery fees", "restaurant ratings".
Use when scraping the job market — job listings, salaries, company ratings, or candidate/profile sourcing — from LinkedIn, Indeed, Glassdoor, Naukri, Google Jobs, RemoteOK, Wellfound, Monster, ZipRecruiter, Reed, Adzuna, Upwork, CutShort, or AmbitionBox. Covers building a job aggregator, salary benchmarking, recruiting/sourcing, and ATS/market feeds. Triggers on "scrape jobs", "job listings", "salary data", "company ratings", "find candidates", "source candidates", "recruiting", "job board scraper", "job aggregator".
Use when you need property or real-estate listing data — sale and rental listings, asking prices, BHK / bed-bath counts, carpet/built-up area, floor, location and coordinates, and agent/owner details. Covers Rightmove (UK), 99acres, MagicBricks, NoBroker, CommonFloor (India), and Craigslist housing. Triggers on "scrape property listings", "real estate data", "rental prices", "property prices", "housing", "for-sale listings", "rental comps", "real estate lead gen".
Use when scraping reviews, ratings, or reputation data — review text, star ratings, TrustScore, pros/cons, company replies. Covers Trustpilot, G2, Capterra, Yelp, Google Maps, and Shopify review widgets. Triggers on "scrape reviews", "ratings", "brand monitoring", "voice of customer", "VOC", "competitor reviews", "review mining", "review sentiment", "TrustScore".
Use when scraping search engines, SERPs, or SEO data — Google Search organic results, Google News headlines, bulk on-page SEO audits (titles/meta/headings/schema), sitemap and subdomain discovery, or bulk image download from pages. Covers rank tracking, SERP/news monitoring, technical SEO audits at scale, competitive recon, and asset inventory. Triggers on "scrape Google", "SERP", "rank tracking", "SEO audit", "search results", "Google News", "subdomains", "sitemap", "scrape search engine", "competitor SEO".
Use when you need travel and hospitality data — hotels, vacation rentals, attractions, or car-sharing listings with their prices, ratings, reviews, availability, star levels, and amenities. Covers Booking.com, Trip.com/Ctrip, TripAdvisor, and Turo. Use for rate and price monitoring, competitive pricing analysis, availability tracking, and review/ranking tracking. Triggers on "scrape hotels", "hotel prices", "hotel rates", "travel data", "car rental", "hotel availability", "TripAdvisor rankings", "Booking.com scraper".
Use when starting any web scraping or structured-data-extraction task and you need to decide HOW to get the data — whether to call an API, run a ready-made scraper, or build a custom one. Covers the cost-first technique ladder (plain HTTP → TLS fingerprint spoof → stealth browser → full browser), the build-vs-buy decision, data-extraction priority order, proxy choices, rate limiting, and legal/compliance basics. Triggers on "scrape", "extract data from website", "crawl", "how do I get data from X", "is there a scraper for X".