For AI labs & product teams

Understand performance
in the work that matters.

Explore how models, tools, skills, and environments affect real tasks.

  1. 01

    Evaluate a new release

    Revisit a relevant set of tasks with a new model or configuration.

  2. 02

    Investigate a workflow

    Compare tool access, skills, or environment choices with the full setup in view.

  3. 03

    Inspect the evidence

    Explore reports, underlying artifacts, and the boundaries of each conclusion.

A concrete starting point

Will a different setup improve this task?

Start with a question, a task collection, and configurations to compare. Evaluation scope and delivery would be agreed separately.

View report →
  1. 01Define the question

    Pick a task collection and the change you want to test.

  2. 02Run relevant configurations

    Run each setup on the same tasks, with its full configuration recorded.

  3. 03Inspect outcomes + evidence

    Compare the outputs, then open the runs and artifacts behind them.

Shared evidence

Inspect the work behind the scores.

Ask about access to eligible, shared records for your evaluation. Private runs remain private.

Discuss access →Access is scoped to eligible records; submitting a question does not grant it.
  • Transcripts

    Prompts, tool calls, and intermediate artifacts where sharing is permitted.

  • Evaluations

    Task-level outcomes and recorded setup comparisons.

  • Preferences

    Pairwise judgments and optional feedback where shared.

Work with Evals

Discuss an evaluation question

Tell us what you want to evaluate. A request does not grant access to private runs or promise a particular data package.

Sign in with Google to send a request from your Evals account.