Methodology

Make the comparison
understandable.

01

Start with the task

Keep the brief, input assets, and success criteria attached to the result.

02

Show the complete setup

Record the LLM, generation calls, skill versions, available tools, environment, and settings.

03

Capture what actually happened

Distinguish available tools from used tools, requested models from served models, and installed skills from activated skills.

04

Keep judgments simple

Vote on the outputs before the setups are revealed.

05

State the limits

Show missing metadata, repeated attempts, sample size, and what changed. A showcase and a benchmark answer different questions.

06

Make it inspectable

Connect claims to reports, reports to comparisons, and comparisons to the underlying runs.

Monthly configuration standings

A config family groups recorded LLM, Gen AI, Skill and Software choices. Exact setup IDs, runtime and execution evidence remain available in its detail. Missing software is an unknown recipe bucket, not proof that those executions used the same software. Only Evals-generated public outputs are eligible.

Win rate and record

Win rate is winning votes divided by winning, losing and drawn votes for the selected month. W–L–D keeps those counts visible. Draws combine ties and both-bad votes. Skipped or unreadable comparisons do not enter the denominator. The accepted vote is counted once per user and output pair, then assigned to its UTC calendar month.

Limits

These rates describe observed community preferences. Category and opponent mix, task coverage and sample size can differ. A rate is not a causal effect or an opponent-adjusted statistical fit. Every available family can appear even below the older fitted-score floor of 100 decisive votes.

User standings

Top contributors use votes + 5 × public outputs + 2 × public tasks + distinct public items they commented on. Top creators count preferred-output votes received by their public outputs. Top token users count captured input and output tokens across runs; cached input is included only once. Missing or partial token measurements remain unknown rather than zero.