Skip to content
MCP ThesaurusMCP Thesaurus

LLM Latency Tracker

CommunityIncomplete39/100Claim

streamable-httpupdated 8d ago

Independent, provider-neutral latency & uptime for AI inference APIs โ€” measured, not scraped.

SourceWebsite1

What can you do with LLM Latency Tracker?

LLM Latency Tracker

Independent, provider-neutral latency & uptime for AI inference APIs โ€” measured, not scraped.

๐ŸŒ Live: llmlatency.dev ยท ๐Ÿ“Š JSON API ยท ๐Ÿค– MCP server ยท ๐Ÿ—“๏ธ Deprecation calendar

License Data Agent-Ready Python

Most "AI API latency" numbers come from the providers themselves, or from a benchmark run once and never updated. This project measures it continuously, from multiple regions, and publishes the result as an open dataset.

  • Edge latency โ€” full DNS โ†’ TCP โ†’ TLS โ†’ time-to-first-byte, measured with the Python standard library (no API key required).
  • Inference latency โ€” real time-to-first-token via a streaming request (optional, needs a provider key).
  • Uptime โ€” success rate per provider, per region.
  • Regions โ€” Europe (Germany), US (Central), Asia (Tokyo), South America (Sรฃo Paulo). More welcome.
  • ~45 providers โ€” OpenAI, Anthropic, Google, Mistral, DeepSeek, xAI, Groq, Together, Fireworks, Cerebras, OpenRouter, Perplexity, plus Chinese models (GLM/Zhipu, Kimi/Moonshot, Qwen, MiniMax) and many more.
  • Deprecation calendar โ€” upcoming model retirements + migration targets, verified from official provider docs.

The site is a self-updating static site (Cloudflare Pages). The value isn't the code โ€” it's the continuously-accumulated, distributed measurement archive. The code is open so the methodology is transparent.

For developers

# All regions, provider rankings for the last 24h โ€” measured latency + uptime:
curl https://llmlatency.dev/api/rankings.json
  • JSON API: /api/rankings.json ยท OpenAPI: /openapi.json
  • Any page as Markdown: send Accept: text/markdown to any page URL, or append .md.
  • For LLM ingestion: /llms.txt (index) and /llms-full.txt (full corpus).
  • License: data is CC-BY-4.0 โ€” free to use with attribution.

For AI agents

There's a real MCP server (Streamable HTTP) exposing a get_ai_api_latency tool backed by the live data:

curl -X POST https://llmlatency.dev/mcp \
  -H 'Content-Type: application/json' \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/call",
       "params":{"name":"get_ai_api_latency","arguments":{"region":"eu-hetzner"}}}'

Run the MCP server locally

The hosted endpoint above needs no setup. If you prefer a local stdio server (or want to build it from source), mcp_server.py is a dependency-free proxy over the same public JSON API:

python3 mcp_server.py            # stdio MCP, stdlib only
# or
docker build -t llm-latency-mcp . && docker run -i llm-latency-mcp

Also available: an MCP Server Card (/.well-known/mcp/server-card.json), a browser WebMCP tool, an API catalog (RFC 9727) and an Agent Skills index. Regions: eu-hetzner, us-central, ap-tokyo, sa-east (omit for all).

How it works

config.py       โ€” registry of providers + this node's REGION (env)
probe.py        โ€” network probe (DNSโ†’TCPโ†’TLSโ†’TTFB, stdlib, no key) + inference probe (TTFT, needs key)
run.py          โ€” one probe cycle across all providers (run on a schedule)
db.py           โ€” SQLite time-series (the accumulated measurement archive)
aggregate.py    โ€” measurements โ†’ p50 / p95 / uptime rankings per region & provider
sitegen.py      โ€” rankings โ†’ static site (JSON API, OpenAPI, llms.txt, schema.org, MCP surface)
ingest.py       โ€” central endpoint that collects measurements from remote probe nodes
ship.py         โ€” probe node โ†’ central node shipper (watermark-based, never loses data on outage)
deprecations.py โ€” model deprecation/migration calendar (only verified, sourced entries)

Each probe node runs with its own REGION, measures every provider, and writes to the time-series. For multi-region, remote nodes ship their measurements to a central node that aggregates and builds the site.

Run it yourself (no keys needed)

git clone https://github.com/mazamaka/llm-latency-tracker
cd llm-latency-tracker
REGION=local python3 run.py         # take edge-latency measurements
python3 aggregate.py --region local # see the ranking from this location

Runs on plain Python 3.12+ (standard library). httpx / loguru are optional.

Inference probes (real TTFT):

cp .env.example .env                # add keys for the providers you want to measure
pip install -r requirements.txt
REGION=local python3 run.py
python3 aggregate.py --region local --type inference

Build the site locally:

BASE_URL=https://example.com python3 sitegen.py   # โ†’ ./site/
python3 -m pytest -q                              # tests

See deploy/ for a container + a generic multi-region deployment guide.

Contributing

Especially welcome:

  • New providers โ€” add a Provider(...) entry in config.py (host + public models endpoint is enough for edge probes).
  • New regions โ€” spin up a probe node in a new location and ship to a central node.
  • Fixes & tests โ€” CI runs pytest + ruff on every push.

See CONTRIBUTING.md for dev setup, how to add a provider/region, and PR guidelines. Please keep the project's principle: measured, not scraped, and honest about the dataset's age.

License

Daily snapshot โ€” 2026-08-31

Measured latency across 45 AI inference providers in 4 regions. Method: distributed edge (DNSโ†’TCPโ†’TLSโ†’TTFB) + inference (TTFT) probes, last 24h. License: CC-BY-4.0.

Region Fastest provider (p50) p50 p95 Uptime
Asia (Tokyo) fireworks 15 ms 58 ms 100%
Europe (Germany) nscale 99 ms 201 ms 100%
South America (Sรฃo Paulo) openrouter 59 ms 86 ms 100%
US (Central) google 42 ms 94 ms 100%

Snapshot generated 2026-08-31T07:47:21Z โ€” this table is regenerated daily.