Agent♥︎Age
Catalog

The Aggregate — LLM benchmark aggregate

Official

by theaggregate

Fused LLM rankings: one IRT/Elo scale across ~5,000 public benchmark leaderboards, updated daily.

Model Context Protocol (MCP) Server: ai.theaggregate/the-aggregate

The ai.theaggregate/the-aggregate MCP server provides fused LLM rankings. It combines results from ~5,000 public benchmark leaderboards into a single IRT/Elo scale. These rankings are updated daily, enabling downstream systems to use a consistent score scale across multiple benchmarks.

🛠️ Key Features

  • Fused LLM rankings
  • Single IRT/Elo scale across ~5,000 public benchmark leaderboards
  • Daily updates

🚀 Use Cases

  • Normalizing LLM performance comparisons across many public benchmarks
  • Feeding a unified ranking signal into applications that require consistent scoring

⚡ Developer Benefits

  • Reduced need to reconcile multiple leaderboard scoring schemes
  • Stable, regularly refreshed benchmark rankings based on a shared scale

⚠️ Limitations

  • Scope is limited to fused rankings derived from the referenced public benchmark leaderboards
  • No details provided on additional tools, configuration, or data sources beyond the fused IRT/Elo ranking updates
The Aggregate — LLM benchmark aggregate - agentage MCP Catalog