Developer
News and Updates
Get Support
Sign in
Get Support
Sign in
DOCUMENTATION
Cloud
Data Center
Resources
Sign in
Sign in
DOCUMENTATION
Cloud
Data Center
Resources
Sign in
Last updated Sep 8, 2026

Run a benchmark

Install and set up ends with Continue to prompts in the benchmark app. Everything on this page happens in that same browser tab: you choose what to test, confirm what the run will use, and watch both context arms answer your prompts.

If you would rather have an AI agent drive the run for you, see start a run from an AI agent.

Step 1: Choose what the benchmark tests

The Choose what the benchmark tests step offers three options.

OptionUse it when
Recommended test promptsThis is your first run, or you want a broad result. Ten connected-work prompts chosen to show what Teamwork Graph can reach that a siloed agent cannot.
Custom promptsYou want to test a real question your team asks. Write one prompt, or run a set you saved earlier.
Role-basedNot available yet. It is marked Coming soon.

Whichever you choose, the same prompts run through both context arms.

Fill in your own details

Every recommended prompt needs details only you know: a topic, a Jira project, a person, a lookback window. Each row reads as the prompt itself, with those details filled in, so you can tell what a prompt asks without opening it. While Finding suggestions is on screen, the benchmark app is still proposing values from your connected sources.

  1. Read the rows. Where the app had nothing to suggest, the row shows the name of the detail it needs in a dashed outline, such as topic or project key.
  2. Choose which prompts to run. Select all available and Deselect all change the whole list at once.
  3. Select a row's warning, which reads Click to fill in ..., to open that prompt and fill in what it asks for. The arrow at the end of any row opens it too, showing the prompt's full wording. Open as many rows as you want to compare.
  4. Replace any suggestion you cannot judge yourself with a real topic, project, or person you know well. You have to be able to tell a good answer from a bad one later.

An empty detail only matters on a prompt you selected. Continue to review stays unavailable until every selected prompt has what it needs, and the footer names the detail that is holding it up.

A prompt whose sources are not connected cannot be selected. It names the source it needs, and you fix that back in Connect data sources, not here.

Success: at least one prompt is selected, no selected prompt is missing a detail, and Continue to review becomes available.

Pick topics with substantial recent activity across your connected sources. A topic with little work, few decisions, and no documentation gives both context arms nothing to find, so the comparison says very little.

Step 2: Confirm what will run

The Know what will run step is the last one before the benchmark spends model usage. It shows:

SectionWhat to check
Setup summaryThe agent, sources, and prompts you selected.
ModelsThe model that answers the prompts, and the model that scores them.
Estimated usage and timeExpected token usage and duration, plus any warnings.
Fair comparisonThe guarantees that keep the two arms comparable, such as the same agent, prompts, and model on both sides.

Select Confirm and start benchmark run. If no estimate could be produced, the same button reads Start benchmark run without an estimate.

Success: the benchmark app replaces the setup steps with the run view in the same tab.

Step 3: Watch the run

Keep the browser tab open and leave the terminal running until the run finishes. The benchmark app reports one phase at a time.

PhaseWhat is happening
Preparing the comparisonFinal checks and source canaries, before any prompt runs.
Running benchmark promptsBoth context arms are answering the same prompt.
Running blind automated reviewAn LLM judge compares the answers without knowing which arm produced them.
Review the answersYour turn. Choose the stronger answer for each prompt.
Building the final reportUsage, quality, and your review choices are combined.
Benchmark completeThe final report is ready in the same tab.

If a run stops, the benchmark app shows Benchmark could not continue with the error. Start a new run rather than resuming a partial one, so both arms answer the same prompts with the same fresh state.

Next: when the phase reaches Review the answers, continue to review, report, and share.

The recommended set contains ten realistic tasks that depend on context about people, projects, documents, code, and recent work.

PromptWhat it asks for
Return-to-work catch-upWhat changed while you were away, focused on one topic, and what to do first.
Topic deep diveThe current state of a topic, its key people, dependencies, and where to start contributing.
Project dependency mapUpstream blockers, sibling dependencies, downstream consumers, and the critical path for a Jira project.
Code-change dependency risk mapCross-repository upstream and downstream dependencies, production use, and change risk for a Jira project.
Leadership delivery briefingA 14 day briefing for an org on what shipped and needs leadership attention.
Reliability and incident reviewRecurring reliability risks for a platform, and where leadership should focus prevention.
Find subject-matter experts and ownersWho to consult for each workstream in a topic, and whether an overall owner exists.
Personal daily stand-upOne person's stand-up for a project, covering blockers and decisions that matter today.
Find current design documents and PRDsThe documents to start from for a topic, plus conflicts and freshness gaps.
Launch-readiness decision briefA go, conditional go, or no-go recommendation for an Atlas project.

Use it for a first run because ten prompts give a broader result than one. Every result depends on the work data connected to your site, so treat a single run as directional. Run the benchmark more than once before you draw a conclusion.

In commands and file names, the recommended set is the suite named cc.

Optional: start a run from an AI agent

Instead of driving the benchmark app yourself, you can ask a supported AI agent to run the benchmark. The agent uses the Teamwork Graph AI Context Benchmark skill, which is a set of instructions the Benchmark CLI installs for you. It completes any missing setup, starts both context arms, and opens a browser view so you can follow progress and review the answers.

This path does not use the benchmark app. Instead of the guided app, the agent runs the two arms as command line processes and opens the live view, which shows run progress and the blind review only. The setup checks, prompt suggestions, and usage estimate live in the app, so you give those up here in exchange for letting the agent drive.

Supported agents are Codex and Claude Code. Cursor Desktop can coordinate a run, using Codex or Claude Code to produce the answers.

If the skill is missing, install or repair it:

1
2
benchmark skills install

Then ask your agent for one of these:

Ask forExample instruction
The recommended setUse the Teamwork Graph AI Context Benchmark skill with suite cc.
One prompt from the setUse the Teamwork Graph AI Context Benchmark skill with scenario swe_ooo_catchup and topic <your topic>.
Your own questionUse the Teamwork Graph AI Context Benchmark skill with prompt "<your question>".

It is the same comparison either way, so prefer the benchmark app unless you specifically want the agent to drive.

Use one agent for a complete run. Answers from different agents are not comparable, because the agent is meant to be held fixed while Teamwork Graph is the only thing that changes.

If neither the benchmark app nor an agent can start a run, see the three-terminal commands in Troubleshooting.

Next step

Review, report, and share results

Rate this page: