Skip to main content

playwright-scraper-skill

Playwright-based web scraping OpenClaw Skill with anti-bot protection. Successfully tested on complex sites like Discuss.com.hk.

Jump to install

Source facts

Repository
szsip239/teamclaw
Last source activity
April 19, 2026 at 22:47
Detected SKILL.md language
English
Stars
112
Forks
17

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.

File Explorer
13 files

Showing SKILL.md

SKILL.md
Source instructions ยท Read-only preview
name
playwright-scraper-skill
description
Playwright-based web scraping OpenClaw Skill with anti-bot protection. Successfully tested on complex sites like Discuss.com.hk.
version
1.2.0
author
Simon Chan
# Playwright Scraper Skill A Playwright-based web scraping OpenClaw Skill with anti-bot protection. Choose the best approach based on the target website's anti-bot level. --- ## ๐ŸŽฏ Use Case Matrix | Target Website | Anti-Bot Level | Recommended Method | Script | |---------------|----------------|-------------------|--------| | **Regular Sites** | Low | web_fetch tool | N/A (built-in) | | **Dynamic Sites** | Medium | Playwright Simple | `scripts/playwright-simple.js` | | **Cloudflare Protected** | High | **Playwright Stealth** โญ | `scripts/playwright-stealth.js` | | **YouTube** | Special | deep-scraper | Install separately | | **Reddit** | Special | reddit-scraper | Install separately | --- ## ๐Ÿ“ฆ Installation ```bash cd playwright-scraper-skill npm install npx playwright install chromium ``` --- ## ๐Ÿš€ Quick Start ### 1๏ธโƒฃ Simple Sites (No Anti-Bot) Use OpenClaw's built-in `web_fetch` tool: ```bash # Invoke directly in OpenClaw Hey, fetch me the content from https://example.com ``` --- ### 2๏ธโƒฃ Dynamic Sites (Requires JavaScript) Use **Playwright Simple**: ```bash node scripts/playwright-simple.js "https://example.com" ``` **Example output:** ```json { "url": "https://example.com", "title": "Example Domain", "content": "...", "elapsedSeconds": "3.45" } ``` --- ### 3๏ธโƒฃ Anti-Bot Protected Sites (Cloudflare etc.) Use **Playwright Stealth**: ```bash node scripts/playwright-stealth.js "https://m.discuss.com.hk/#hot" ``` **Features:** - Hide automation markers (`navigator.webdriver = false`) - Realistic User-Agent (iPhone, Android) - Random delays to mimic human behavior - Screenshot and HTML saving support --- ### 4๏ธโƒฃ YouTube Video Transcripts Use **deep-scraper** (install separately): ```bash # Install deep-scraper skill npx clawhub install deep-scraper # Use it cd skills/deep-scraper node assets/youtube_handler.js "https://www.youtube.com/watch?v=VIDEO_ID" ``` --- ## ๐Ÿ“– Script Descriptions ### `scripts/playwright-simple.js` - **Use Case:** Regular dynamic websites - **Speed:** Fast (3-5 seconds) - **Anti-Bot:** None - **Output:** JSON (title, content, URL) ### `scripts/playwright-stealth.js` โญ - **Use Case:** Sites with Cloudflare or anti-bot protection - **Speed:** Medium (5-20 seconds) - **Anti-Bot:** Medium-High (hides automation, realistic UA) - **Output:** JSON + Screenshot + HTML file - **Verified:** 100% success on Discuss.com.hk --- ## ๐ŸŽ“ Best Practices ### 1. Try web_fetch First If the site doesn't have dynamic loading, use OpenClaw's `web_fetch` toolโ€”it's fastest. ### 2. Need JavaScript? Use Playwright Simple If you need to wait for JavaScript rendering, use `playwright-simple.js`. ### 3. Getting Blocked? Use Stealth If you encounter 403 or Cloudflare challenges, use `playwright-stealth.js`. ### 4. Special Sites Need Specialized Skills - YouTube โ†’ deep-scraper - Reddit โ†’ reddit-scraper - Twitter โ†’ bird skill --- ## ๐Ÿ”ง Customization All scripts support environment variables: ```bash # Set screenshot path SCREENSHOT_PATH=/path/to/screenshot.png node scripts/playwright-stealth.js URL # Set wait time (milliseconds) WAIT_TIME=10000 node scripts/playwright-simple.js URL # Enable headful mode (show browser) HEADLESS=false node scripts/playwright-stealth.js URL # Save HTML SAVE_HTML=true node scripts/playwright-stealth.js URL # Custom User-Agent USER_AGENT="Mozilla/5.0 ..." node scripts/playwright-stealth.js URL ``` --- ## ๐Ÿ“Š Performance Comparison | Method | Speed | Anti-Bot | Success Rate (Discuss.com.hk) | |--------|-------|----------|-------------------------------| | web_fetch | โšก Fastest | โŒ None | 0% | | Playwright Simple | ๐Ÿš€ Fast | โš ๏ธ Low | 20% | | **Playwright Stealth** | โฑ๏ธ Medium | โœ… Medium | **100%** โœ… | | Puppeteer Stealth | โฑ๏ธ Medium | โœ… Medium-High | ~80% | | Crawlee (deep-scraper) | ๐Ÿข Slow | โŒ Detected | 0% | | Chaser (Rust) | โฑ๏ธ Medium | โŒ Detected | 0% | --- ## ๐Ÿ›ก๏ธ Anti-Bot Techniques Summary Lessons learned from our testing: ### โœ… Effective Anti-Bot Measures 1. **Hide `navigator.webdriver`** โ€” Essential 2. **Realistic User-Agent** โ€” Use real devices (iPhone, Android) 3. **Mimic Human Behavior** โ€” Random delays, scrolling 4. **Avoid Framework Signatures** โ€” Crawlee, Selenium are easily detected 5. **Use `addInitScript` (Playwright)** โ€” Inject before page load ### โŒ Ineffective Anti-Bot Measures 1. **Only changing User-Agent** โ€” Not enough 2. **Using high-level frameworks (Crawlee)** โ€” More easily detected 3. **Docker isolation** โ€” Doesn't help with Cloudflare --- ## ๐Ÿ” Troubleshooting ### Issue: 403 Forbidden **Solution:** Use `playwright-stealth.js` ### Issue: Cloudflare Challenge Page **Solution:** 1. Increase wait time (10-15 seconds) 2. Try `headless: false` (headful mode sometimes has higher success rate) 3. Consider using proxy IPs ### Issue: Blank Page **Solution:** 1. Increase `waitForTimeout` 2. Use `waitUntil: 'networkidle'` or `'domcontentloaded'` 3. Check if login is required --- ## ๐Ÿ“ Memory & Experience ### 2026-02-07 Discuss.com.hk Test Conclusions - โœ… **Pure Playwright + Stealth** succeeded (5s, 200 OK) - โŒ Crawlee (deep-scraper) failed (403) - โŒ Chaser (Rust) failed (Cloudflare) - โŒ Puppeteer standard failed (403) **Best Solution:** Pure Playwright + anti-bot techniques (framework-independent) --- ## ๐Ÿšง Future Improvements - [ ] Add proxy IP rotation - [ ] Implement cookie management (maintain login state) - [ ] Add CAPTCHA handling (2captcha / Anti-Captcha) - [ ] Batch scraping (parallel URLs) - [ ] Integration with OpenClaw's `browser` tool --- ## ๐Ÿ“š References - [Playwright Official Docs](https://playwright.dev/) - [puppeteer-extra-plugin-stealth](https://github.com/berstend/puppeteer-extra/tree/master/packages/puppeteer-extra-plugin-stealth) - [deep-scraper skill](https://clawhub.com/opsun/deep-scraper)
View on GitHub