Live Test
Test models directly
Send a prompt to any model and see the response stream in real time. Compare two models side by side to measure speed and quality.
Interactive model tester
The tester runs in your browser and needs JavaScript. Pick any model from the live catalogue, send a prompt and watch the answer stream token by token — with latency and token counts measured as it happens.
- Single — One model, one prompt, one streamed answer with timing stats.
- Compare — Two models answer the same prompt side by side, streaming simultaneously, so differences in speed and style are directly visible.
- Consensus — Several models answer in parallel and a judge model synthesises one consensus answer — the same mechanism behind our Consensus API.
How the live test works
Pick a model
Choose from the models listed below, grouped by provider. The list is the live catalogue — it updates as models are added or retired.
Send a prompt
Type your own prompt or start from a community-suggested test question. The request goes to the provider through our API gateway — the same path our paid API uses.
Watch it stream
Tokens render as they arrive, with live latency and token counts, so you see how a model responds — not just a single final number.
Three test modes
Single
One model, one prompt, one streamed answer with timing stats.
Compare
Two models answer the same prompt side by side, streaming simultaneously, so differences in speed and style are directly visible.
Consensus
Several models answer in parallel and a judge model synthesises one consensus answer — the same mechanism behind our Consensus API.
Model coverage
30 models from 5 providers are currently enabled for live testing:
- Google Gemini: Gemini 3.5 Flash, Gemini 2.5 Flash, Gemini Flash Latest, Gemini 2.5 Pro
- OpenAI: o3-mini-2025-01-31, gpt-4.1-2025-04-14, gpt-4.1-mini-2025-04-14, gpt-4o, o4-mini-2025-04-16, gpt-4o-mini, gpt-4.1-nano-2025-04-14, gpt-4o-mini-2024-07-18
- OVH AI Endpoints (GRA): gpt-oss-120b, gpt-oss-20b
- OpenRouter: Cohere Command-A, Nous Hermes 3 70B, Llama 4 Scout, MiniMax M2.5, DeepSeek v3.2, Qwen 2.5 VL 72B Instruct, DeepSeek v4 Pro, Mistral Voxtral Small 24B, Qwen 3.6 Plus, NVIDIA Nemotron Super 49B v1.5, Llama 4 Maverick, Llama 3.3 70B Instruct, Qwen 3.7 Max
- Anthropic: Claude Sonnet 4.5, Claude Sonnet 4.6, Claude Haiku 4.5
Free demo terms
The live test is free and needs no sign-up. Guests get a limited number of runs per day; a free account raises that daily allowance. Limits reset every 24 hours and keep the demo available for everyone.
Where the numbers come from
- Benchmark methodology — How we test models weekly — scoring, judges and datasets.
- Consensus API — The multi-model consensus mode as a production API.
- Leaderboard — Current speed and quality rankings from the same test runs.