Understand performance
in the work that matters.
Explore how models, tools, skills, and environments affect real tasks.
- 01
Evaluate a new release
Revisit a relevant set of tasks with a new model or configuration.
- 02
Investigate a workflow
Compare tool access, skills, or environment choices with the full setup in view.
- 03
Inspect the evidence
Explore reports, underlying artifacts, and the boundaries of each conclusion.
Will a different setup improve this task?
Start with a question, a task collection, and configurations to compare. Evaluation scope and delivery would be agreed separately.
View report →01Define the question
Pick a task collection and the change you want to test.
02Run relevant configurations
Run each setup on the same tasks, with its full configuration recorded.
03Inspect outcomes + evidence
Compare the outputs, then open the runs and artifacts behind them.
Inspect the work behind the scores.
Ask about access to eligible, shared records for your evaluation. Private runs remain private.
Discuss access →Access is scoped to eligible records; submitting a question does not grant it.Transcripts
Prompts, tool calls, and intermediate artifacts where sharing is permitted.
Evaluations
Task-level outcomes and recorded setup comparisons.
Preferences
Pairwise judgments and optional feedback where shared.
Discuss an evaluation question
Tell us what you want to evaluate. A request does not grant access to private runs or promise a particular data package.