Doc Scraper MCP Server

This project provides a powerful command-line tool to crawl documentation sites based on settings defined in a config.yaml file. It navigates the site structure, extracts content from specified HTML sections using CSS selectors, and converts it into clean Markdown files.

People who work with data preparation, web scraper and documentation and want it reachable from Claude, Cursor, VS Code, or another MCP client. The project is written in Go.

VERIFIED ACTIVE

LAST COMMIT 2026-09-06 · ★ 99 · #61 OF 93 MAINTAINED WEB SCRAPING · VERIFIED 2026-09-18

Apache-2.0 · Go servers · how we verify → /methodology

01 · Install Doc Scraper

Claude Desktop

Settings → Extensions → Install, then select the https://github.com/Sriram-PR/doc-scraper/releases/download/v2.9.2/doc-scraper.mcpb .mcpb bundle

02 · Evidence

Security posture

What to check before giving this server access to your agent - from the registry, GitHub, and our own probes. We don't score safety; we show what's verifiable.

runs as Claude Desktop extension (.mcpb bundle)

license Apache-2.0 - declared in the repository

registry namespace io.github.Sriram-PR is GitHub-verified and matches the repo owner

03 · What Doc Scraper can do

Prose above is summarized from the project's README and registry record - no invented capabilities.

Latest releases

v2.9.2 · 2026-09-06

Add end-to-end crawl and watch-scheduler tests, converting the crawl entry points to return exit codes instead of calling os.Exit so deferred store and index cleanup runs (@Sriram-PR) · Fail a fresh crawl when every…

v2.9.1 · 2026-09-04

Add fuzz targets for the untrusted-input parsers and harden the config writer to refuse degenerate YAML layouts found by fuzzing instead of silently losing site entries (@Sriram-PR) · Update x/text to v0.41.0 fixing a…

v2.9.0 · 2026-09-04

Add a search CLI command with ranked results and section anchor links, relax natural-language queries that FTS5's implicit AND would return nothing for, and add a recorded demo GIF to the README (@Sriram-PR) · Fix…

04 · Who maintains Doc Scraper

doc-scraper is maintained by sriram-pr. It's the only MCP server we track from this author; the repo dates to Apr 2025.

05 · Facts

category
web scraping - ranked #61 of 93 actively-maintained web scraping servers as of 2026-09-18.
release cadence
9 releases in the last 90 days (latest 2026-09-06)
registry
io.github.Sriram-PR/doc-scraper (active, first published 2026-09-04 · 5 versions)
packages
mcpb:https://github.com/Sriram-PR/doc-scraper/releases/download/v2.9.2/doc-scraper.mcpb

06 · Doc Scraper FAQ

What is Doc Scraper?

This project provides a powerful command-line tool to crawl documentation sites based on settings defined in a config.yaml file. It navigates the site structure, extracts content from specified HTML sections using CSS selectors, and converts it into clean Markdown files.

Is Doc Scraper still maintained?

Yes - as of 2026-09-18, its last commit was 2026-09-06 and it shipped 9 releases in the last 90 days. We re-verify nightly.

07 · Alternatives to Doc Scraper

More web scraping MCP servers · Search1 API · Indonesia Civic Stack · Read GZH · Charlotte

More Go MCP servers · StackQL MCP Server · DevTool MCP · Figma MCP Express · Supr Send · Synapbus · see all