Doc Scraper MCP Server
This project provides a powerful command-line tool to crawl documentation sites based on settings defined in a config.yaml file. It navigates the site structure, extracts content from specified HTML sections using CSS selectors, and converts it into clean Markdown files.
People who work with data preparation, web scraper and documentation and want it reachable from Claude, Cursor, VS Code, or another MCP client. The project is written in Go.
VERIFIED ACTIVE
LAST COMMIT 2026-09-06 · ★ 99 · #61 OF 93 MAINTAINED WEB SCRAPING · VERIFIED 2026-09-18
Apache-2.0 · Go servers · how we verify → /methodology
01 · Install Doc Scraper
Claude Desktop
Settings → Extensions → Install, then select the https://github.com/Sriram-PR/doc-scraper/releases/download/v2.9.2/doc-scraper.mcpb .mcpb bundle 02 · Evidence
Security posture
What to check before giving this server access to your agent - from the registry, GitHub, and our own probes. We don't score safety; we show what's verifiable.
runs as Claude Desktop extension (.mcpb bundle)
license Apache-2.0 - declared in the repository
registry namespace io.github.Sriram-PR is GitHub-verified and matches the repo owner
03 · What Doc Scraper can do
Prose above is summarized from the project's README and registry record - no invented capabilities.
Latest releases
v2.9.2 · 2026-09-06
Add end-to-end crawl and watch-scheduler tests, converting the crawl entry points to return exit codes instead of calling os.Exit so deferred store and index cleanup runs (@Sriram-PR) · Fail a fresh crawl when every…
v2.9.1 · 2026-09-04
Add fuzz targets for the untrusted-input parsers and harden the config writer to refuse degenerate YAML layouts found by fuzzing instead of silently losing site entries (@Sriram-PR) · Update x/text to v0.41.0 fixing a…
v2.9.0 · 2026-09-04
Add a search CLI command with ranked results and section anchor links, relax natural-language queries that FTS5's implicit AND would return nothing for, and add a recorded demo GIF to the README (@Sriram-PR) · Fix…
04 · Who maintains Doc Scraper
doc-scraper is maintained by sriram-pr. It's the only MCP server we track from this author; the repo dates to Apr 2025.
05 · Facts
- repository
- github.com/sriram-pr/doc-scraper
- category
- web scraping - ranked #61 of 93 actively-maintained web scraping servers as of 2026-09-18.
- release cadence
- 9 releases in the last 90 days (latest 2026-09-06)
- registry
- io.github.Sriram-PR/doc-scraper (active, first published 2026-09-04 · 5 versions)
- packages
- mcpb:https://github.com/Sriram-PR/doc-scraper/releases/download/v2.9.2/doc-scraper.mcpb
06 · Doc Scraper FAQ
What is Doc Scraper?
This project provides a powerful command-line tool to crawl documentation sites based on settings defined in a config.yaml file. It navigates the site structure, extracts content from specified HTML sections using CSS selectors, and converts it into clean Markdown files.
Is Doc Scraper still maintained?
Yes - as of 2026-09-18, its last commit was 2026-09-06 and it shipped 9 releases in the last 90 days. We re-verify nightly.
07 · Alternatives to Doc Scraper
Alternatives to Doc Scraper
Maintained web scraping servers if Doc Scraper isn't the fit.
- Firecrawl MCP Server MCP server for Firecrawl - web search, scraping, and biomedical/arXiv paper search. ★ 7,478 · 2026-09-18
- Apify MCP Server Extract data from any website with thousands of scrapers, crawlers, and automations on Apify Store ⚡ ★ 7,456 · 2026-09-17
- Wigolo Local-first web intelligence MCP server for AI coding agents ★ 5,284 · 2026-09-17
- Exa Fast, intelligent web search and web crawling. New mcp tool: Exa-code is a context tool for coding ★ 5,013 · 2026-08-21
- Tavily MCP MCP server for advanced web search using Tavily ★ 2,390 · 2026-09-16
- Open Brand Extract brand assets (logos, colors, backdrop images, brand name) from any website URL ★ 797 · 2026-05-12
Pairs well with
Servers that cover what Doc Scraper doesn't - only shown when the pairing reason fits the companion.
- SearXNG Search → search companion search · ★ 1,245
- Blockrun MCP → search companion search · ★ 394
- Supabase → database companion database · ★ 2,910
- MongoDB MCP Server → database companion database · ★ 1,131
- Codebase Memory → memory companion memory · ★ 43,715
- Memorix → memory companion memory · ★ 785
More web scraping MCP servers · Search1 API · Indonesia Civic Stack · Read GZH · Charlotte
More Go MCP servers · StackQL MCP Server · DevTool MCP · Figma MCP Express · Supr Send · Synapbus · see all