streamable-httpupdated 8d ago
Independent, provider-neutral latency & uptime for AI inference APIs โ measured, not scraped.
What can you do with LLM Latency Tracker?
LLM Latency Tracker
Independent, provider-neutral latency & uptime for AI inference APIs โ measured, not scraped.
๐ Live: llmlatency.dev ยท ๐ JSON API ยท ๐ค MCP server ยท ๐๏ธ Deprecation calendar
Most "AI API latency" numbers come from the providers themselves, or from a benchmark run once and never updated. This project measures it continuously, from multiple regions, and publishes the result as an open dataset.
- Edge latency โ full DNS โ TCP โ TLS โ time-to-first-byte, measured with the Python standard library (no API key required).
- Inference latency โ real time-to-first-token via a streaming request (optional, needs a provider key).
- Uptime โ success rate per provider, per region.
- Regions โ Europe (Germany), US (Central), Asia (Tokyo), South America (Sรฃo Paulo). More welcome.
- ~45 providers โ OpenAI, Anthropic, Google, Mistral, DeepSeek, xAI, Groq, Together, Fireworks, Cerebras, OpenRouter, Perplexity, plus Chinese models (GLM/Zhipu, Kimi/Moonshot, Qwen, MiniMax) and many more.
- Deprecation calendar โ upcoming model retirements + migration targets, verified from official provider docs.
The site is a self-updating static site (Cloudflare Pages). The value isn't the code โ it's the continuously-accumulated, distributed measurement archive. The code is open so the methodology is transparent.
For developers
# All regions, provider rankings for the last 24h โ measured latency + uptime:
curl https://llmlatency.dev/api/rankings.json
- JSON API:
/api/rankings.jsonยท OpenAPI:/openapi.json - Any page as Markdown: send
Accept: text/markdownto any page URL, or append.md. - For LLM ingestion:
/llms.txt(index) and/llms-full.txt(full corpus). - License: data is CC-BY-4.0 โ free to use with attribution.
For AI agents
There's a real MCP server (Streamable HTTP) exposing a get_ai_api_latency tool backed by the live data:
curl -X POST https://llmlatency.dev/mcp \
-H 'Content-Type: application/json' \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/call",
"params":{"name":"get_ai_api_latency","arguments":{"region":"eu-hetzner"}}}'
Run the MCP server locally
The hosted endpoint above needs no setup. If you prefer a local stdio server (or want to build it from source), mcp_server.py is a dependency-free proxy over the same public JSON API:
python3 mcp_server.py # stdio MCP, stdlib only
# or
docker build -t llm-latency-mcp . && docker run -i llm-latency-mcp
Also available: an MCP Server Card (/.well-known/mcp/server-card.json), a browser WebMCP tool, an API catalog (RFC 9727) and an Agent Skills index. Regions: eu-hetzner, us-central, ap-tokyo, sa-east (omit for all).
How it works
config.py โ registry of providers + this node's REGION (env)
probe.py โ network probe (DNSโTCPโTLSโTTFB, stdlib, no key) + inference probe (TTFT, needs key)
run.py โ one probe cycle across all providers (run on a schedule)
db.py โ SQLite time-series (the accumulated measurement archive)
aggregate.py โ measurements โ p50 / p95 / uptime rankings per region & provider
sitegen.py โ rankings โ static site (JSON API, OpenAPI, llms.txt, schema.org, MCP surface)
ingest.py โ central endpoint that collects measurements from remote probe nodes
ship.py โ probe node โ central node shipper (watermark-based, never loses data on outage)
deprecations.py โ model deprecation/migration calendar (only verified, sourced entries)
Each probe node runs with its own REGION, measures every provider, and writes to the time-series. For multi-region, remote nodes ship their measurements to a central node that aggregates and builds the site.
Run it yourself (no keys needed)
git clone https://github.com/mazamaka/llm-latency-tracker
cd llm-latency-tracker
REGION=local python3 run.py # take edge-latency measurements
python3 aggregate.py --region local # see the ranking from this location
Runs on plain Python 3.12+ (standard library). httpx / loguru are optional.
Inference probes (real TTFT):
cp .env.example .env # add keys for the providers you want to measure
pip install -r requirements.txt
REGION=local python3 run.py
python3 aggregate.py --region local --type inference
Build the site locally:
BASE_URL=https://example.com python3 sitegen.py # โ ./site/
python3 -m pytest -q # tests
See deploy/ for a container + a generic multi-region deployment guide.
Contributing
Especially welcome:
- New providers โ add a
Provider(...)entry inconfig.py(host + public models endpoint is enough for edge probes). - New regions โ spin up a probe node in a new location and ship to a central node.
- Fixes & tests โ CI runs
pytest+ruffon every push.
See CONTRIBUTING.md for dev setup, how to add a provider/region, and PR guidelines. Please keep the project's principle: measured, not scraped, and honest about the dataset's age.
License
- Code: MIT
- Data (rankings, API output): CC-BY-4.0 โ attribute llmlatency.dev.
Daily snapshot โ 2026-08-31
Measured latency across 45 AI inference providers in 4 regions. Method: distributed edge (DNSโTCPโTLSโTTFB) + inference (TTFT) probes, last 24h. License: CC-BY-4.0.
| Region | Fastest provider (p50) | p50 | p95 | Uptime |
|---|---|---|---|---|
| Asia (Tokyo) | fireworks | 15 ms | 58 ms | 100% |
| Europe (Germany) | nscale | 99 ms | 201 ms | 100% |
| South America (Sรฃo Paulo) | openrouter | 59 ms | 86 ms | 100% |
| US (Central) | 42 ms | 94 ms | 100% |
- Full dataset:
data/rankings/2026-08-31.json(latest) - Citable archive (DOI):
10.5281/zenodo.21954788โ daily aggregates, CC-BY-4.0 - Hugging Face dataset: https://huggingface.co/datasets/llmlatency/llm-latency-tracker
- Kaggle dataset: https://www.kaggle.com/datasets/llmlatency/llm-latency-tracker
- Archived in Software Heritage:
swh:1:snp:2778cbabd72a70a629ee35fbd5ac536d1ccb7a9a - Python client: https://pypi.org/project/llmlatency/
- Live rankings and methodology: https://llmlatency.dev
- Machine-readable API: https://llmlatency.dev/api/rankings.json
- Model deprecation calendar: https://llmlatency.dev/deprecations
Snapshot generated 2026-08-31T07:47:21Z โ this table is regenerated daily.
Install
Add LLM Latency Tracker to your client. Pick the one you use.
claude mcp add --transport http llm-latency-tracker https://llmlatency.dev/mcpcodex mcp add llm-latency-tracker --url https://llmlatency.dev/mcp{
"mcpServers": {
"llm-latency-tracker": {
"url": "https://llmlatency.dev/mcp"
}
}
}Add to `~/.cursor/mcp.json`, or `.cursor/mcp.json` for a single project.
{
"servers": {
"llm-latency-tracker": {
"type": "http",
"url": "https://llmlatency.dev/mcp"
}
}
}Add to `.vscode/mcp.json` in your workspace.
{
"mcpServers": {
"llm-latency-tracker": {
"url": "https://llmlatency.dev/mcp"
}
}
}Add to `claude_desktop_config.json`, then restart Claude Desktop.
{
"mcpServers": {
"llm-latency-tracker": {
"serverUrl": "https://llmlatency.dev/mcp"
}
}
}Add to `~/.codeium/windsurf/mcp_config.json`.
Score
39 / 100
Incomplete
- Documentation25/25
- Maintenance19/25
- Trust6/20
- Capability0/15
- Install experience12/15
- Documents what it does and how to connect
- Has a resolvable package or endpoint
- Exposes at least one tool, prompt or resource
- README has substantive content
- Includes a code example
- Documents its configuration
- Mentions credentials or security posture
- Last commit 1 days ago
- Has a release history
- Repository is not archived
- No licence detected
- Namespace verified in the official MCP registry
- Claimed by its owner
- Published under an organisation
- 0 tool(s) documented
- Provides prompt templates
- Provides resources
- 6 documented install method(s)
- Published to a package registry
- Offers a hosted endpoint โ no local install
Version history
| Versions | Published |
|---|---|
| 1.1.0Latest | Aug 16, 2026 |
| 1.0.0 | Jul 23, 2026 |