All notes
Engineering7 min read

Rovo Users Complete 20% More Jira Items. That Number Should Make You Nervous.

Work-item throughput is the weakest tier of measurement, and AI makes ticket creation cheap. Here's what to fix before you measure anything.

From Atlassian's Q4 FY26 letter to shareholders, published 6 August:

Customers that adopt Rovo are completing 20% more Jira work items and creating/editing 25% more Confluence pages versus non-adopters.

Assisted actions were up more than 50% quarter over quarter. More than 80% of the Fortune 500 now use Rovo.

The number is probably real. Atlassian has the telemetry, and there's no reason to think anyone cooked it. That's not the problem.

The problem is the unit.

20% more what, exactly?

A Jira work item is not a standard measure of anything. It's whatever the person creating the ticket decided it was that morning.

Take one feature — say, adding SSO to your admin console. Here are three defensible decompositions (Fig. 1).

Fig. 01One feature, three defensible ways to cut it into tickets

The nine-day work bar at the foot of each column is identical. Only the number of items above it changes.

1closed items
One story
Story: ADM-214Add SAML SSO
9 days · 1 engineer
6closed items
One story, five sub-tasks
Story: ADM-214Add SAML SSO
Sub-task: Metadata endpoint
Sub-task: Certificate handling
Sub-task: Login redirect
Sub-task: Session mapping
Sub-task: Docs
9 days · 1 engineer
8closed items
Eight stories, two epics
Epic · Okta
Story: SP-initiated
Story: IdP-initiated
Story: Metadata + certs
Story: Single logout
Epic · Entra ID
Story: SP-initiated
Story: IdP-initiated
Story: Metadata + certs
Story: Single logout
9 days · 1 engineer
The count is 1, 6 or 8 for the same nine days of work, by the same engineer, on the same feature. The eightfold spread is produced entirely by a decision about ticket shape.

Same feature. Same nine days of work. Same engineer. The third team "completed" eight times the work items of the first.

Every velocity metric quietly assumes this doesn't happen. Granularity is treated as a constant. It's a free variable — and it moves.

Why AI moves it in one direction

Here's the uncomfortable part. Ticket granularity has always been noisy, but it was expensive-noisy. Splitting a story into six meant typing six summaries, six descriptions, setting six sets of fields, linking six parents. The tedium was a tax on decomposition, and the tax kept ticket counts down.

AI assistance removes the tax.

When creating a ticket costs nine seconds instead of ninety, finer decomposition becomes the path of least resistance. Not because anyone decided to game the metric — nobody sat in a meeting and agreed to inflate item counts. It's just that the friction that used to cap granularity is gone.

So "Rovo adopters complete 20% more work items" is consistent with at least three different worlds (Fig. 2).

Fig. 02Three worlds consistent with the same reported number

Three hypotheses about a 20% rise in completed work items. Work actually shipped is up 20% in the first, unchanged in the second, and somewhere between the two in the third. The reported figure is plus 20% in all three.

Work actually shippedItems closed
+20%Those teams genuinely ship 20% more.
+20%
+0%Those teams ship the same amount, cut into more pieces.
+20%
?Some mix, in unknown proportion.
+20%
Work shipped is up 20%, unchanged, or unknown. The reported figure is +20% in all three. A measure that returns the same value under every explanation is not evidence for any one of them.

The letter can't distinguish between them, and neither can you from the outside. Neither, from the inside, can most engineering leaders looking at their own dashboard.

The research landed the same week

Four days before the shareholder letter, ACM Queue published Eight Myths on Software Engineering and GenAI — six researchers, five from Microsoft plus Margaret-Anne Storey at the University of Victoria. It's the most useful thing written on this in a while, and it's free.

One finding is worth pinning to the wall:

Studies at Microsoft and elsewhere show that developers spend closer to 14 percent of their time writing code.

Fourteen percent.

Think about what that does to the standard AI productivity story. If you make code generation twice as fast — genuinely twice, no caveats — you've addressed 14% of the job. The ceiling on total improvement is 7% (Fig. 3). Everything else lives in the other 86%: coordination, review, waiting, meetings, figuring out what to build, deciding what a ticket even means.

Fig. 03Where a code-generation speedup can and cannot reach
7%ceiling on total improvement if that 14% halves
14%writing code86%everything else: coordination, review, waiting, deciding what to build
Writing code is about 14% of the week. Halving it — a genuine 2×, no caveats — moves the total by at most 7%. The ceiling holds however fast generation gets.

The paper's broader argument is that misread studies and anecdotal wins are driving real decisions about tooling and measurement. A vendor headline about work-item throughput is exactly the artefact it warns about — not because it's dishonest, but because output counts are the weakest tier of measurement we have, and they're the easiest tier to move without moving anything real.

This is not a new observation. DORA and SPACE have said versions of it for years: count outcomes, not outputs. What's new is that the cost of producing outputs just collapsed, which makes output-counting worse than it used to be rather than merely imperfect.

What to do instead

The cynical move here is to dunk on the metric and stop. That's not useful — you still need to know whether the tooling is working.

So: the reason you can't compare item throughput across teams, or across quarters, is that nobody has fixed how items get created. Standardise that, and throughput becomes meaningful. Leave it floating, and every comparison is measuring ticket-writing habits.

Concretely:

Define your recurring project shapes. Most teams have five or six. A new service. A migration. A vendor integration. A quarterly compliance review. An incident retro. Write down, once, how each decomposes — what becomes an epic, what becomes a story, what's a subtask, what doesn't get a ticket at all.

Fix decomposition at the template, not in the moment. The point isn't the time saved, although you save time. The point is that the decomposition decision gets made once, deliberately, by someone thinking about it — instead of forty times a quarter by whoever happened to be creating the ticket.

Then measure. Cycle time per item becomes comparable. Throughput becomes comparable. Estimates start converging, because "a story" means the same thing in March as it did in January.

Treat any before/after AI comparison that crosses a granularity change as noise. Including your own. Especially your own.

This is the argument for project templates that has nothing to do with saving time — templates are how you hold the unit still. You can't measure a team with a ruler that stretches.

Reading the number honestly

None of this means Rovo isn't working. It might well be. The 25% figure on Confluence pages is arguably more interesting than the 20% on Jira items, since page creation is less prone to arbitrary subdivision.

It means the number can't carry the weight people will put on it. It'll appear in slide decks for the next two quarters as evidence that AI adoption improves engineering output. It isn't that. It's evidence that Rovo adopters close more tickets — which is a real thing, a measurable thing, and a different thing.

Ask the follow-up question. Twenty percent more work items, at what average size?

If nobody can answer, that's the finding.

The product these notes come from

Stop creating Jira issues one by one.

SuperTemplates turns unstructured text into structured backlogs. AI handles hierarchy. Templates handle scale. You review the whole batch before anything is created.

Try free on the Atlassian Marketplace