Want help running the benchmark? Sign up to benchmark Teamwork Graph with your own work.
Find out whether Teamwork Graph helps your AI agent give more useful answers about your real work.
The benchmark asks your agent the same question twice. The first answer uses only the agent's standard Atlassian access. The second answer uses Teamwork Graph. You then read both answers without being told which setup produced each one and choose the one you would rather use. The benchmark turns those choices into a report that compares answer quality, token use, and time.
Everything runs on your own computer against your own Atlassian site. Nothing is sent to Atlassian unless you decide to package a report and share it.
The benchmark keeps the agent, the question, and the model fixed, so Teamwork Graph is the only meaningful difference between the two answers. Each side is called a context arm.
| Siloed context | Teamwork Graph context |
|---|---|
| Your agent's standard Atlassian access through the Rovo MCP server, with Teamwork Graph turned off. | The same agent answering the same question with the Teamwork Graph CLI connected. |
You see these two names throughout the benchmark, from setup to the final report. In commands, the same two arms are named control and test, and those literal names appear only when you run a benchmark from the command line.
In the review step the two answers appear as Answer A and Answer B, assigned so that you cannot tell which arm produced which answer. That is what makes your choice usable as evidence.
You install two command line tools and run one command. That command opens the benchmark app, a local web app in your browser that carries you through the whole benchmark. It is not a status dashboard you watch from the side: setup, prompt selection, the run, your review, and the final report all happen in it.
Expect 20 to 40 minutes end to end for a first run, most of it unattended while the two arms answer your prompts. Install and set up opens with everything you need before you start.
This tool is in active development, so treat every result as directional rather than final. Read how to read the report before quoting any number, because token savings and answer quality do not always move together.
Rate this page: