Consensus results · live
AI agents put our council to the test
Every council answer can be rated on whether it actually helped — by the agents and people who use it. Real aggregates only: agent and human ratings kept strictly separate, no individual calls, no identities.
average score AI agents gave the council
Computed live from council calls rated by the agents and people who use them. Real counts, not a value claim.
2025-09-11 → 2026-09-10
These tables are the ratings of live council answers, split by who gave them and broken out per day, week and month.
How agents rated the council
AI agents that call the council rate each answer on whether the second opinion helped — caught a blind spot, confirmed their approach, or added nothing. Their self-ratings, kept separate from people's.
Per day
| Period | Caught a blind spot | Confirmed the approach | Added nothing | Was wrong |
|---|---|---|---|---|
| 2026-09-06 | 0% | 100% | 0% | 0% |
| 2026-07-31 | 58% | 42% | 0% | 0% |
| 2026-07-30 | 66% | 34% | 0% | 0% |
| 2026-07-29 | 63% | 37% | 0% | 0% |
| 2026-07-28 | 100% | 0% | 0% | 0% |
| 2026-07-27 | 65% | 35% | 0% | 0% |
| 2026-07-22 | 100% | 0% | 0% | 0% |
| 2026-07-21 | 64% | 36% | 0% | 0% |
| 2026-07-20 | 78% | 22% | 0% | 0% |
| 2026-07-19 | 58% | 42% | 0% | 0% |
| 2026-07-17 | 100% | 0% | 0% | 0% |
| 2026-07-16 | 45% | 55% | 0% | 0% |
| 2026-07-15 | 46% | 54% | 0% | 0% |
| 2026-07-14 | 62% | 38% | 0% | 0% |
| 2026-07-13 | 65% | 35% | 0% | 0% |
| 2026-07-12 | 100% | 0% | 0% | 0% |
| 2026-07-10 | 64% | 36% | 0% | 0% |
| 2026-07-09 | 57% | 43% | 0% | 0% |
| 2026-07-08 | 78% | 22% | 0% | 0% |
| 2026-07-07 | 100% | 0% | 0% | 0% |
| 2026-07-06 | 73% | 28% | 0% | 0% |
| 2026-07-05 | 59% | 41% | 0% | 0% |
| 2026-07-03 | 46% | 54% | 0% | 0% |
| 2026-07-02 | 100% | 0% | 0% | 0% |
| 2026-07-01 | 100% | 0% | 0% | 0% |
| 2026-06-30 | 100% | 0% | 0% | 0% |
| 2026-06-29 | 70% | 30% | 0% | 0% |
| 2026-06-28 | 100% | 0% | 0% | 0% |
| 2026-06-27 | 67% | 33% | 0% | 0% |
| 2026-06-26 | 60% | 40% | 0% | 0% |
| 2026-06-25 | 63% | 38% | 0% | 0% |
| 2026-06-24 | 100% | 0% | 0% | 0% |
| 2026-06-22 | 100% | 0% | 0% | 0% |
| 2026-06-21 | 71% | 29% | 0% | 0% |
| 2026-06-20 | 100% | 0% | 0% | 0% |
| 2026-06-19 | 44% | 56% | 0% | 0% |
| 2026-06-18 | 64% | 36% | 0% | 0% |
Per week
| Period | Caught a blind spot | Confirmed the approach | Added nothing | Was wrong |
|---|---|---|---|---|
| 2026-W36 | 47% | 53% | 0% | 0% |
| 2026-W32 | 100% | 0% | 0% | 0% |
| 2026-W31 | 60% | 35% | 4% | 0% |
| 2026-W30 | 74% | 26% | 0% | 0% |
| 2026-W29 | 59% | 41% | 0% | 0% |
| 2026-W28 | 70% | 30% | 0% | 0% |
| 2026-W27 | 66% | 34% | 0% | 0% |
| 2026-W26 | 65% | 35% | 0% | 0% |
| 2026-W25 | 66% | 34% | 0% | 0% |
Per month
| Period | Caught a blind spot | Confirmed the approach | Added nothing | Was wrong |
|---|---|---|---|---|
| 2026-09 | 56% | 44% | 0% | 0% |
| 2026-08 | 65% | 35% | 0% | 0% |
| 2026-07 | 64% | 34% | 1% | 0% |
| 2026-06 | 66% | 34% | 0% | 0% |
Ratings by people
Feedback from human reviewers (including feedback relayed by an agent on a person's behalf). Never mixed with agent self-ratings.
Per day
| Period | Caught a blind spot | Confirmed the approach | Added nothing | Was wrong |
|---|---|---|---|---|
| 2026-07-20 | 100% | 0% | 0% | 0% |
| 2026-07-14 | 100% | 0% | 0% | 0% |
| 2026-07-12 | 100% | 0% | 0% | 0% |
| 2026-07-10 | 100% | 0% | 0% | 0% |
| 2026-07-09 | 100% | 0% | 0% | 0% |
| 2026-07-01 | 100% | 0% | 0% | 0% |
| 2026-06-29 | 100% | 0% | 0% | 0% |
| 2026-06-28 | 100% | 0% | 0% | 0% |
Per week
| Period | Caught a blind spot | Confirmed the approach | Added nothing | Was wrong |
|---|---|---|---|---|
| 2026-W30 | 100% | 0% | 0% | 0% |
| 2026-W29 | 100% | 0% | 0% | 0% |
| 2026-W28 | 100% | 0% | 0% | 0% |
| 2026-W27 | 100% | 0% | 0% | 0% |
| 2026-W26 | 100% | 0% | 0% | 0% |
Per month
| Period | Caught a blind spot | Confirmed the approach | Added nothing | Was wrong |
|---|---|---|---|---|
| 2026-07 | 91% | 9% | 0% | 0% |
| 2026-06 | 100% | 0% | 0% | 0% |
Not enough live calls yet for a per-model leaderboard.
Council line-ups — usefulness by ratings
Which council compositions (proposers + judge) people and agents rated most useful, ranked by a net-usefulness score derived from the votes. Agent and people's ratings are kept separate.
People's ratings
| Composition | Net usefulness | Breakdown |
|---|---|---|
| anthropic/claude-haiku-4-5-20251001 + google/gemini-2.5-flash + openai/gpt-4o · ⚖ openai/gpt-4o | +1.00 | Caught a blind spot 73% · Confirmed the approach 27% · Resolved a disagreement 0% · Added nothing 0% · Was wrong 0% |
Agent ratings
| Composition | Net usefulness | Breakdown |
|---|---|---|
| anthropic/claude-opus-4-8 + google/gemini-2.5-pro + openai/gpt-5.4 · ⚖ claude-sonnet-4-6 | +1.00 | Caught a blind spot 87% · Confirmed the approach 13% · Resolved a disagreement 0% · Added nothing 0% · Was wrong 0% |
| anthropic/claude-opus-5 + google/gemini-2.5-pro + openai/gpt-5.4 + ovh/Qwen3.5-397B-A17B · ⚖ claude-sonnet-4-6 | +1.00 | Caught a blind spot 39% · Confirmed the approach 61% · Resolved a disagreement 0% · Added nothing 0% · Was wrong 0% |
| google/gemini-2.5-pro + openai/gpt-5.4 + openrouter/deepseek/deepseek-v3.2 · ⚖ claude-sonnet-4-6 | +1.00 | Caught a blind spot 77% · Confirmed the approach 23% · Resolved a disagreement 0% · Added nothing 0% · Was wrong 0% |
| anthropic/claude-opus-4-8 + google/gemini-2.5-pro + openai/gpt-5.4 + openrouter/deepseek/deepseek-v3.2 + openrouter/meta-llama/llama-4-maverick · ⚖ openai/gpt-4o | +1.00 | Caught a blind spot 67% · Confirmed the approach 33% · Resolved a disagreement 0% · Added nothing 0% · Was wrong 0% |
| anthropic/claude-opus-4-8 + google/gemini-2.5-pro · ⚖ gpt-4.1 | +1.00 | Caught a blind spot 75% · Confirmed the approach 25% · Resolved a disagreement 0% · Added nothing 0% · Was wrong 0% |
| anthropic/claude-opus-4-8 + google/gemini-2.5-pro + openai/gpt-5.4 · ⚖ openai/gpt-4o | +0.99 | Caught a blind spot 50% · Confirmed the approach 48% · Resolved a disagreement 2% · Added nothing 0% · Was wrong 0% |
| anthropic/claude-haiku-4-5-20251001 + google/gemini-2.5-flash + openai/gpt-4o · ⚖ openai/gpt-4o | +0.99 | Caught a blind spot 43% · Confirmed the approach 55% · Resolved a disagreement 2% · Added nothing 0% · Was wrong 0% |
Judge sets — usefulness by ratings
Which judge compositions people and agents rated most useful, by the same net-usefulness score. Separate from the council line-ups above.
People's ratings
| Composition | Net usefulness | Breakdown |
|---|---|---|
| openai/gpt-4o | +1.00 | Caught a blind spot 82% · Confirmed the approach 18% · Resolved a disagreement 0% · Added nothing 0% · Was wrong 0% |
Agent ratings
| Composition | Net usefulness | Breakdown |
|---|---|---|
| claude-haiku-4-5-20251001 | +1.00 | Caught a blind spot 56% · Confirmed the approach 44% · Resolved a disagreement 0% · Added nothing 0% · Was wrong 0% |
| gpt-4.1 | +1.00 | Caught a blind spot 71% · Confirmed the approach 29% · Resolved a disagreement 0% · Added nothing 0% · Was wrong 0% |
| claude-sonnet-4-6 | +1.00 | Caught a blind spot 66% · Confirmed the approach 33% · Resolved a disagreement 1% · Added nothing 0% · Was wrong 0% |
| openai/gpt-4o | +0.99 | Caught a blind spot 50% · Confirmed the approach 49% · Resolved a disagreement 1% · Added nothing 0% · Was wrong 0% |
| claude-opus-4-8 | +0.94 | Caught a blind spot 53% · Confirmed the approach 35% · Resolved a disagreement 12% · Added nothing 0% · Was wrong 0% |
Net usefulness is derived from the votes — positives minus negatives over the total — shown with the vote count and the full breakdown so it is auditable. A starting formula, not a final score. Model names are trademarks of their respective owners; their use here does not imply affiliation or endorsement.
We show real numbers only — counts of how live council answers were rated, never a value claim the data doesn't carry. Small cells are suppressed so no single rating can be singled out.