Developer
News and Updates
Get Support
Sign in
Get Support
Sign in
DOCUMENTATION
Cloud
Data Center
Resources
Sign in
Sign in
DOCUMENTATION
Cloud
Data Center
Resources
Sign in
Last updated Jul 30, 2026

Review and share results

Blind review

A browser window opens and shows two answers (A and B) for each prompt, with the source hidden. For each pair:

  1. Required: Which answer would you rather use for this work?
  2. Optional: Which answer is more accurate and better grounded in the available context?

Interpreting your results

The completed run produces a local report. Keep output local unless sharing has been explicitly agreed.

Your report scores each prompt on quality and token usage.

What you seeHow to read it
Token efficiencyThe token and cost difference between the two arms. Count a drop as a win only if quality held. Fewer tokens with a worse answer is not cheaper — it is a shorter answer that dropped content.
Agent qualityWhich answer you picked in the blind review. That vote is the verdict. The relIQ judge is supporting evidence from an LLM judge.

Read quality and cost together, not separately. Token savings can come from shorter answers that dropped quality. Look for prompts where Teamwork Graph is at least as good on quality and cheaper on tokens, or clearly better on quality at similar cost.

One run is noisy: run several repetitions before you conclude anything, and treat every number as early and directional.

Package and share results

Only package or share results when sharing has been explicitly agreed.

For the latest report:

1
2
benchmark report zip

Or a specific report:

1
2
benchmark report zip --report-name runs/reports/DASHBOARD.html

The ZIP stays local under your run folder (for example ~/.benchmark/runs) until you share it through the agreed channel.

With /twg-ai-context-benchmark, you can also ask the agent to zip and share the latest run; it will redact and package the latest or a specific report for you.

Rate this page: