Doctor: Web Crawling and Indexing for MCP LLM Agents

May 31, 2025

{
“content”: “

\n\n
\n
\n
\n

Introduction

\n

Developers often struggle to provide LLMs with up-to-date, structured web data without building complex RAG pipelines from scratch. Doctor is an open-source tool designed to solve this by automating the discovery, crawling, and indexing of websites to be exposed as a Model Context Protocol (MCP) server for LLM agents. By transforming raw web content into a searchable, hierarchical index, Doctor enables AI agents to reason over current web data with far greater precision than standard web search tools.

\n

\n
\n
\n\n
\n
\n
\n

What Is Doctor?

\n

Doctor is a tool for discovering, crawling, and indexing websites to be exposed as an MCP server for LLM agents. It provides a complete stack for converting web pages into structured data that AI agents can query. Built primarily in Python, it leverages a modern AI stack including crawl4ai for web scraping, LangChain for text chunking, and litellm for embedding generation via OpenAI. The project is licensed under the MIT License, allowing for broad adoption and customization.

\n

The core philosophy of Doctor is to bridge the gap between the static nature of LLM training data and the dynamic nature of the web, providing agents with a \”memory\” of the web that is structured hierarchically, mirroring the site’s own architecture.

\n

\n
\n
\n\n
\n
\n
\n

Why Doctor Matters

\n

Traditional web search for LLMs often returns fragmented fragments of pages, which leads to \”hallucinations\” when the agent cannot see the context of the surrounding pages. Doctor solves this by implementing hierarchy tracking during the crawl process. This means the agent doesn’t just get a piece of text; it understands where that page sits in the site’s overall structure, which is critical for technical documentation or complex product guides.

\n

Furthermore, the integration with the Model Context Protocol (MCP) is a significant differentiator. MCP is an open standard that allows LLM agents (like those in Cursor or VSCode) to seamlessly connect to external data sources. By exposing its index as an MCP server, Doctor allows developers to give their AI coding assistants a direct, structured feed of the latest documentation for any library or API they are using, drastically reducing the need for manual copy-pasting of docs into the chat.

\n

As the ecosystem of AI agents moves toward autonomous reasoning, the ability to provide them with high-fidelity, structured context is becoming the primary bottleneck. Doctor provides the infrastructure to automate this context delivery at scale.

\n

\n
\n
\n\n
\n
\n
\n

Key Features

\n

    \n

  • Hierarchical Navigation: Pages maintain parent-child relationships, allowing LLM agents to navigate through the site structure logically rather than randomly.
  • \n

  • Automated Web Crawling: Uses crawl4ai to efficiently scrape web pages while maintaining a record of how pages are linked.
  • \n

  • Intelligent Text Chunking: Leverages LangChain to the break down large web pages into manageable pieces for embedding and retrieval.
  • \n

  • Vector Search Support: Stores data in DuckDB with vector search capabilities, enabling semantic search across the crawled content.
  • \n

  • MCP Server Integration: Exposes the indexed data as an MCP server, making it immediately available to compatible LLM agents in IDEs.
  • \n

  • FastAPI Web Service: Provides a REST API for managing crawl jobs and viewing the index map.
  • \n

  • Domain Grouping: Automatically groups pages from the same domain together, even if they were crawled individually.
  • \n

  • Automatic Title Extraction: Extracts page titles from HTML or markdown content to improve the quality of the index.
  • \n

  • Breadcrumb Navigation: Provides a visual path from the root to the current page, which is useful for both humans and agents.
  • \n

  • Sibling Navigation: Allows quick access to pages at the same level in the hierarchy, facilitating broader context gathering.
  • \n

\n

\n
\n
\n\n
\n
\n
\n

How Doctor Compares

\n

\n

\n

\n

\n

\n

\n

\n

\n

\n

\n

\n

\n

\n

\n

\n

\n

\n

\n

\n

\n

\n

\n

\n

\n

\n

\n

\n

\n

\n

\n

\n

\n

\n

Feature Doctor Standard RAG Pipeline Generic Web Search
MCP Server Support Native Custom Implementation None
Hierarchy Tracking Yes Rarely No
Setup Complexity Low (Docker) High Very Low
Context Fidelity High Medium