A tightened measurement engine (6.45)
One strict rulebook for reliability, time to start, flow efficiency and cycle time, applied the same way on every view.
This release tightens the engine underneath the numbers. Commitment reliability, time to start, flow efficiency and cycle time now follow one strict rulebook, applied the same way on Sprint Summary, Delivery Signals, the leadership report and the retro snapshot. A grade no longer depends on which screen you opened it from, and every rule behind it is written down inside the product.
Some figures will read differently once this lands, because the rules behind them are stricter and now consistent everywhere. Reliability may read a little lower, flow efficiency typically reads higher, and reopened items read longer. Nothing about how your teams work has changed. Section 4 below gives the one-line reason for each movement.
1. Commitment reliability, tightened
Reliability now grades the commitment a team made on day one against the delivery it made by sprint end.
- Graded against the day-one commitment
Reliability measures the items committed at the end of the sprint’s first day against what was delivered by sprint end, with a 24-hour grace for work closed the morning after review. It grades the commitment a team actually stood behind, over the sprint it stood behind it for.
- Slipped work counts as carryover, cancelled work counts for neither side
An item that lands in a later sprint is counted as carryover rather than as delivery in the sprint that committed it, and a cancelled item counts for neither side. A cancellation can no longer flatter a sprint or weigh against it.
- The day-one baseline is reconstructed for every viewer
Whoever opens the view, the day-one commitment is rebuilt from your own Azure DevOps history rather than depending on who happened to be watching at the time. Two people looking at the same sprint read the same commitment.
- Organisations with imported history are fully supported
If your work item history was brought across from another tool, your sprints are now measured on the same strict day-one baseline as everyone else’s, for the first time.
- One reliability number, on every surface
Sprint Summary, Delivery Signals, the leadership report and the retro snapshot all read the same reliability figure, and How These Metrics Work spells out the exact rule that produced it.
2. One rulebook, applied on every view
The same tightening runs through time to start, flow efficiency, forecasting and cycle time.
- Time to Start is now measured per sprint
The Delivery Signals card previously showed a placeholder. It now measures the days from sprint start, or from the item’s creation where that is later, until work actually began, shows the median for the sprint, and says plainly when there is nothing measurable yet.
- One industry-standard definition of flow efficiency, everywhere
Every view now uses hands-on time divided by the time from first work to done, the definition the industry already agrees on. Time an item spent waiting before anyone started it no longer counts against you. Where no waiting states are mapped, a note flags a suspicious 100% on every view that shows the number, rather than letting a perfect score stand unexplained.
- Forecasts and velocity averages agree with each other
The sprint still in progress no longer weighs on velocity averages and Monte Carlo forecasts as though it were a finished sprint. Its row stays visible, it simply stops tilting the aggregate. A brand-new team now sees a real early trend from its first sprint instead of a flat zero.
- Cycle time holds one anchor
An item that was reopened is measured from the first time work started on it, on every view, so its full story is counted. History reads that hit Azure DevOps rate limits are retried, and where history still cannot be read in full, the Cycle Time view and the dashboard say the ages are approximate, the same disclosure Aging already made.
3. Sprint Summary, redesigned
Each team card now leads with the number a coach actually presents, and reads from the same rulebook as the rest of the release.
- The verdict leads the card
Each team card now opens with the number a coach presents: delivered versus committed, as one large figure, with a delivery bar and the carry-over and moved-to-next-sprint counts alongside it.
- A quiet comparison against last sprint
For a finished sprint, a small “vs last sprint” line appears next to the verdict once it resolves. It is computed by the exact same rule as the figure above it, under this team’s own item-type and workflow mapping, never a second formula reaching a different answer.
- Colour only where it means something
The tiles are neutral by default. Scope change is the one tile that can take an accent, and only once churn passes your organisation’s configured threshold, the same comparison the scope-churn notification already makes.
- Progress bars for all three buckets
Each bucket gets its own bar. A bucket with nothing in it says “none planned” rather than sitting at a full bar that reads like the work is done.
- One computation, everywhere it shows up
The card reads from the same commitment-reliability rule as Delivery Signals, the leadership report and the retro snapshot, so a figure a coach presents from Sprint Summary is the same figure anyone sees on another screen.
4. Numbers an enterprise can trust
A larger organisation does not only need a metric. It needs one a coach can defend in front of a team and a leader can take to a board.
- A grade a coach can defend in front of a team
When reliability reads what it reads, the rule behind it is one sentence long and the same rule ran on every sprint in the chart. The conversation in a retro can be about the work rather than about the measurement.
- A grade a leader can take to a board
The leadership report and the retro snapshot draw on the same definitions as the views the teams use every day, so a figure presented upstairs is the figure the team recognises downstairs.
- Every definition documented in the product
How These Metrics Work carries an entry for each rule in this release, including a new Time to Start entry. Nobody has to take a figure on faith or reassemble its definition from a support thread.
- Two screens never disagree in a meeting
One definition per metric means Sprint Summary, Delivery Signals, the leadership report and the retro snapshot cannot show different answers to the same question while a room watches.
- Migrated history measured on the same baseline
Organisations whose history came across from another tool are graded on the same day-one baseline as everyone else, so a migration is no longer a reason a set of teams sits outside the comparison.
5. Where your numbers may shift, and why
A few figures will read differently once this lands. Each moves for one reason, and it is the same reason everywhere the figure appears.
- Reliability may read a little lower
The bar is stricter: the day-one commitment, delivered by sprint end with a 24-hour grace, and slipped work counted as carryover.
- Flow efficiency typically reads higher
Waiting that happened before work began no longer counts against you, so the percentage rises towards the time your team was genuinely hands-on.
- Reopened items read longer
A reopened item is measured from the first time work started on it, so its full story is counted rather than only its most recent stretch.
- Mid-sprint forecasts may read slightly higher
The sprint still in progress no longer pulls on the aggregate as though it were a finished sprint.
6. Clearer labels, in every language
Wording across the product, in all six supported languages, so a label says what the number behind it actually means.
- Commitment copy says day one
Everywhere the commitment is named, the label now says it is the day-one commitment, in all six supported languages, so the wording and the rule say the same thing.
- “(latest)” for a team with no active sprint
A team with no sprint currently running in Azure DevOps now sees “(latest)” alongside a short hint, rather than a sprint labelled “(Current)” that is not the current one.
- Aging says “no history yet”
An aging threshold with nothing behind it yet says so plainly, instead of showing a zero-day threshold that reads like a measurement.
- Cumulative Flow counts work created in the sprint
The scope-added count now includes items created mid-sprint, so scope growth is described using the same scope the board is showing.
- DORA tells no successful runs apart from no runs
The empty state distinguishes the two, so an absence of successful runs is never read as an absence of activity.
- The weekend setting says where it applies
The weekend-exclusion setting now states that it applies to Aging, so its reach is on the label rather than something to infer.
- How These Metrics Work is expanded
Every rule in this release has an entry, including a new one for Time to Start, so the definition behind a number is always one click away from the number.
7. Getting started
Before you share the new figures widely, open Sprint Summary and Delivery Signals on your most recently completed sprint and read the two side by side. Seeing the same reliability figure in both places, and reading the rule behind it in How These Metrics Work, is the quickest way to be ready for the first question a team or a board will ask.
The full view-by-view guide, including Sprint Summary, Delivery Signals, Cycle Time, Cumulative Flow, DORA Metrics and Monte Carlo forecasting.
Day-to-day administration: access control, seats, group rules, workflow mapping, and the organisation-wide activity log with its export.
The previous release, where Delivery Signals thresholds became organisation-configurable and every grade started showing how it was calculated.