A browser window opens and shows two answers (A and B) for each prompt, with the source hidden. For each pair:
The completed run produces a local report. Keep output local unless sharing has been explicitly agreed.
Your report scores each prompt on quality and token usage.
| What you see | How to read it |
|---|---|
| Token efficiency | The token and cost difference between the two arms. Count a drop as a win only if quality held. Fewer tokens with a worse answer is not cheaper — it is a shorter answer that dropped content. |
| Agent quality | Which answer you picked in the blind review. That vote is the verdict. The relIQ judge is supporting evidence from an LLM judge. |
Read quality and cost together, not separately. Token savings can come from shorter answers that dropped quality. Look for prompts where Teamwork Graph is at least as good on quality and cheaper on tokens, or clearly better on quality at similar cost.
One run is noisy: run several repetitions before you conclude anything, and treat every number as early and directional.
Only package or share results when sharing has been explicitly agreed.
For the latest report:
1 2benchmark report zip
Or a specific report:
1 2benchmark report zip --report-name runs/reports/DASHBOARD.html
The ZIP stays local under your run folder (for example ~/.benchmark/runs) until you share it through the agreed channel.
With /twg-ai-context-benchmark, you can also ask the agent to zip and share the latest run; it will redact and package the latest or a specific report for you.
Rate this page: