Skip to content
MCP ThesaurusMCP Thesaurus

agent web โ€” URL to LLM ready markdown: a polite, robots respecting web page reader (free)

CommunityIncomplete39/100Claim

streamable-httpMITupdated 2mo ago

URL โ†’ LLM-ready markdown. A charter-clean, polite, robots-respecting web reader for AI agents โ€” honors robots.txt, identifies honestly, and never bypasses anti-bot/CAPTCHA or paywalls. No account, no API key โ€” free per-call access.

SourceWebsite

What can you do with agent web โ€” URL to LLM ready markdown: a polite, robots respecting web page reader (free)?

agent-web

URL โ†’ LLM-ready markdown. A charter-clean, polite, robots-respecting web reader for AI agents โ€” honors robots.txt, identifies honestly, and never bypasses anti-bot/CAPTCHA or paywalls. No account, no API key โ€” free per-call access.

๐ŸŒ Live: https://agent-web.foomworks.workers.dev ยท ๐Ÿ”Œ MCP: https://agent-web.foomworks.workers.dev/mcp


What it gives an agent

Give it one publicly reachable URL; get back the page as clean, LLM-ready markdown โ€” title, word count, and the text โ€” without parsing raw HTML yourself. Server-rendered (static) HTML pages. Ideal for summarizing an article, extracting docs, or feeding a page into an LLM.

Streamable HTTP, JSON-RPC 2.0:

https://agent-web.foomworks.workers.dev/mcp
Tool Cost Returns
read_url free fetch a URL โ†’ full LLM-ready markdown (title, word count, markdown)
read_url_preview free the first ~600 characters of the markdown โ€” a cheap look before the full pull

Or call it over HTTP

BASE=https://agent-web.foomworks.workers.dev
curl -s "$BASE/read?url=https://example.com/&ref=readme"          # full markdown
curl -s "$BASE/read/preview?url=https://example.com/&ref=readme"  # first ~600 chars

Discovery: /.well-known/mcp.json (MCP) ยท /.well-known/agent-card.json (A2A) ยท /openapi.json ยท /policy (Acceptable-Use Policy)

Charter-clean by design (the bright line)

This is a polite reader, not a scraper:

  • Honors the origin's robots.txt for our user-agent โ€” a disallowed path is refused without fetching; an unconfirmable robots.txt (5xx/error) is treated conservatively as disallowed.
  • Identifies honestly with a descriptive user-agent โ€” never spoofs a browser.
  • Read-only GET of the single URL you supply โ€” no crawl, no proxy rotation.
  • Never bypasses anti-bot/CAPTCHA, paywalls, or login walls.
  • SSRF-guarded: private/loopback/link-local/internal hosts are blocked and every redirect hop is re-validated.

You are responsible for your right to fetch the URL you submit.

Limits

  • Static HTML only for now โ€” JavaScript-rendered pages, screenshots, and PDF are a later increment.
  • Body cap ~2 MB, fetch timeout ~12 s.
  • The service holds no keys and never initiates payments.

Install the AgentSkill

SKILL.md is a portable AgentSkill (ClawHub / skills.sh / Claude Skills).

About

agent-web is operated by FOOM โ€” an AI-operated, human-supervised autonomous service. License: MIT.