MinerU
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Written mainly in Python.
What each tool is actually for, described plainly. No ratings, no scores, no invented benchmarks.
Web scraping and data: Getting text out of the web and out of documents, in the shape a model can read.
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Written mainly in Python.
#1 PDF Application on GitHub that lets you edit PDFs on any device anywhere
Written mainly in Java.
Agents that use the browser.
Written mainly in Python. Released under the MIT licence.
🚀🤖 Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper. Don't be shy, join here: https://discord.gg/jP8KfhDhyN
Written mainly in Python. Released under the Apache-2.0 licence.
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
Written mainly in TypeScript. Released under the Apache-2.0 licence.
Get your documents ready for gen AI
Written mainly in Python. Released under the MIT licence.
The context API to search, scrape, and interact with the web at scale. 🔥
Written mainly in TypeScript. Released under the AGPL-3.0 licence.
Convert PDF to markdown + JSON quickly with high accuracy
Written mainly in Python. Released under the Apache-2.0 licence.
Playwright is a framework for Web Testing and Automation. It allows testing Chromium, Firefox and WebKit with a single API.
Written mainly in TypeScript. Released under the Apache-2.0 licence.
A pure-python PDF library capable of splitting, merging, cropping, and transforming the pages of PDF files
Written mainly in Python.
Convert any URL to an LLM-friendly input with a simple prefix https://r.jina.ai/
Written mainly in TypeScript. Released under the Apache-2.0 licence.
Scrapy, a fast high-level web crawling & scraping framework for Python.
Written mainly in Python. Released under the BSD-3-Clause licence.
Automate browser based workflows with AI
Written mainly in Python. Released under the AGPL-3.0 licence.
🔥 Open Source Browser API for AI Agents & Apps. Steel Browser is a batteries-included browser sandbox that lets you automate the web without worrying about infrastructure.
Written mainly in TypeScript. Released under the Apache-2.0 licence.
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
Written mainly in HTML. Released under the Apache-2.0 licence.