The web data API to search, scrape, and interact at scale. 🔥
-
Updated
Sep 26, 2026 - TypeScript
Web scraping is the process of programmatically retrieving web pages and extracting structured data from them for analysis, storage, or reuse. It powers price monitoring, search indexing, market research, and training-data collection.
Scraping ranges from plain HTTP requests and HTML parsing to full browser automation for JavaScript-rendered and anti-bot-protected pages. The ecosystem includes request libraries, HTML parsers, headless browsers, and cloud extraction platforms.
The web data API to search, scrape, and interact at scale. 🔥
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
Self-hosted webscraper.
AnyCrawl 🚀: A Node.js/TypeScript crawler that turns websites into LLM-ready data and extracts structured SERP results from Google/Bing/Baidu/etc. Native multi-threading for bulk processing.
Advanced Privacy Browser Core with Unified Fingerprint Defense: Cloudflare, Akamai, Kasada, Shape, DataDome, PerimeterX, hCaptcha, FunCaptcha, Imperva, reCAPTCHA, ThreatMetrix, Adscore
Browser fingerprinting tools for anonymizing your scrapers. Developed by Apify.
Self-hosted scraping engine — bypasses any JS challenge & captcha: Cloudflare, Turnstile, reCAPTCHA, hCaptcha, GeeTest. FlareSolverr & Byparr alternative and drop-in replacement for your *arr stack.
🕵️♂️ LinkedIn profile scraper returning structured profile data in JSON.
➖ Stripped down, stable version of firecrawl optimized for self-hosting and ease of contribution. Billing logic and AI features are completely removed. Crawl and convert any website into LLM-ready markdown.
Dive into web scraping and build a Next.js 13 eCommerce price tracker within a single video that teaches you data scraping, cron jobs, sending emails, deployment, and more.
n8n node for browser automation using Puppeteer
A free & open source IMDb front-end.
🧶 Extract data from any website without code, just clicks.
GoogleBard - A reverse engineered API for Google Bard chatbot for NodeJS
HTTP client for Node.js with browser TLS fingerprint impersonation
Facebook Group Members Extractor. Download Facebook group members in CSV.
Free Palestine. 📖 This tool is to download course from educative.io for offline usage. It uses your login credentials and download the course.
estela, an elastic web scraping cluster 🕸