Skip to content

Live Test

Test models directly

Send a prompt to any model and see the response stream in real time. Compare two models side by side to measure speed and quality.

Interactive model tester

The tester runs in your browser and needs JavaScript. Pick any model from the live catalogue, send a prompt and watch the answer stream token by token — with latency and token counts measured as it happens.

  • SingleOne model, one prompt, one streamed answer with timing stats.
  • CompareTwo models answer the same prompt side by side, streaming simultaneously, so differences in speed and style are directly visible.
  • ConsensusSeveral models answer in parallel and a judge model synthesises one consensus answer — the same mechanism behind our Consensus API.

How the live test works

Pick a model

Choose from the models listed below, grouped by provider. The list is the live catalogue — it updates as models are added or retired.

Send a prompt

Type your own prompt or start from a community-suggested test question. The request goes to the provider through our API gateway — the same path our paid API uses.

Watch it stream

Tokens render as they arrive, with live latency and token counts, so you see how a model responds — not just a single final number.

Three test modes

Single

One model, one prompt, one streamed answer with timing stats.

Compare

Two models answer the same prompt side by side, streaming simultaneously, so differences in speed and style are directly visible.

Consensus

Several models answer in parallel and a judge model synthesises one consensus answer — the same mechanism behind our Consensus API.

Model coverage

30 models from 5 providers are currently enabled for live testing:

  • Google Gemini: Gemini 3.5 Flash, Gemini 2.5 Flash, Gemini Flash Latest, Gemini 2.5 Pro
  • OpenAI: o3-mini-2025-01-31, gpt-4.1-2025-04-14, gpt-4.1-mini-2025-04-14, gpt-4o, o4-mini-2025-04-16, gpt-4o-mini, gpt-4.1-nano-2025-04-14, gpt-4o-mini-2024-07-18
  • OVH AI Endpoints (GRA): gpt-oss-120b, gpt-oss-20b
  • OpenRouter: Cohere Command-A, Nous Hermes 3 70B, Llama 4 Scout, MiniMax M2.5, DeepSeek v3.2, Qwen 2.5 VL 72B Instruct, DeepSeek v4 Pro, Mistral Voxtral Small 24B, Qwen 3.6 Plus, NVIDIA Nemotron Super 49B v1.5, Llama 4 Maverick, Llama 3.3 70B Instruct, Qwen 3.7 Max
  • Anthropic: Claude Sonnet 4.5, Claude Sonnet 4.6, Claude Haiku 4.5

Free demo terms

The live test is free and needs no sign-up. Guests get a limited number of runs per day; a free account raises that daily allowance. Limits reset every 24 hours and keep the demo available for everyone.

Where the numbers come from

  • Benchmark methodologyHow we test models weekly — scoring, judges and datasets.
  • Consensus APIThe multi-model consensus mode as a production API.
  • LeaderboardCurrent speed and quality rankings from the same test runs.