Measuring Progress When Milestones Don’t Show Up in a Status Report

A printed status report on a conference table showing an unchanged milestone bar, next to a laptop screen showing a rising usage chart telling a different story

The status report said nothing moved that week. The status report was wrong about what mattered.

Eight months into a legacy system replacement, I asked for a progress update ahead of a steering committee review. The report was thin. No milestone had closed that week. The same bar sat on the Gantt chart it had sat on three weeks earlier, still in progress, nothing to bump. That week was a clean example of something specific: milestones don’t always show up in a status report. Real progress can still be happening underneath it.

Here is what had actually happened that week. A business analyst fielded a question about the old system. For the first time, they answered it from the new one without opening the legacy screen to check. A partner team resolved a data discrepancy themselves, using the new reconciliation process, without escalating it to us. A new hire onboarded straight onto the target-state workflow. Nobody needed to explain “the way we used to do it” as a reference point. None of that closed a milestone. All of it was the transformation actually taking hold. None of it was in the report I was about to carry into the room.

The status report and the transformation measure different things

A project plan tracks task completion: requirements signed off, code deployed, data migrated, training delivered. That is real and worth tracking. It does not tell you whether the change took hold. Process, technology, and culture shift on different clocks. A status report built around milestones is built to track the first two. Culture, the slowest of the three, does not close out on a Gantt chart. It shows up in whether people reach for the new way without being told to. That kind of evidence does not fit in a field that expects a percentage.

That is not a flaw in status reporting. Status reports answer a specific question, is the work on track, and answer it well. The mistake is treating them as the only instrument for a different question: is the change actually sticking. That question needs its own evidence, and most of that evidence already exists somewhere. It just is not being looked at.

A milestone tells you a task finished. It does not tell you whether anyone believes in what replaced it.

Find the data that is already being collected

The temptation is to treat this as purely qualitative, something you sense in meetings but cannot point to. Usually that is not true. Most programs are already generating the exhaust that would answer this, nobody has gone looking for it:

  • System usage, not system rollout. Login and transaction volume on the legacy system versus the new one, tracked weekly. A legacy system that stays busy six months after go-live is telling you something a training-completion metric never will.
  • Ticket categorization. Help desk and support tickets tagged by which system they reference. A rising share of tickets ask how to do something in the new system. A shrinking share ask why the new system doesn’t work like the old one. That shift is a real trend line, not an anecdote.
  • Defect and workaround trend, post go-live. Not the defect count itself, the shape of the curve. Defects that trail off while workaround requests keep climbing means the software shipped but the process did not.
  • Training completion versus proficiency. Completion tells you people sat through it. Proficiency, measured by whether they can complete a real transaction unassisted two months later, tells you something different. Almost nobody tracks the second one.

Ten minutes with whoever owns each of those systems usually turns up more real signal than a month of status meetings. The data was never designed to answer this question, but it was collecting the answer anyway.

The same signal can mean two opposite things

The hardest part of this practice is that every good sign has a bad twin. They look identical from a distance. Escalations dropping means people found a working path, or it means people stopped believing anyone would act on them. Workaround chatter going quiet means the gap closed, or it means the team gave up flagging it. A quiet week can be the best week of the program, or the first sign of disengagement. A status report cannot tell you which.

The only way to tell them apart is to ask, directly. It takes the same discipline that turns a vague objection into a real requirement. Do not accept the absence of a signal as good news until someone checks. When escalations drop, ask the team who used to raise them. Did the problem go away, or did they just stop bringing it up? Do this on a schedule, not only when something feels off. By the time something feels off with disengagement, it has usually been building for months.

Somebody has to own collecting this, on a cadence

None of the above happens by accident. It does not happen from one director reading meeting notes once a quarter. It needs an owner and a rhythm, separate from status reporting. Pencil in a short recurring check-in, every one to two weeks. The question is not “what did you complete,” but “what did you notice, and what changed that nobody logged.”

Rotate who is asked. A single source of “good news” every week is not a trend, it is one person’s optimism. Confirmation bias is easiest to catch when you can see whose voice is missing from the pattern over time.

What this is worth saying out loud to the people above you

The consequence of skipping all of this is not abstract. A multi-year program that only ever shows “in progress” on a milestone chart starts to look stalled to the people funding it. That is true whether or not it actually is. Steering committees deprioritize what looks stuck. Teams disengage from a plan that never seems to finish. This happens even while the real transformation is well underway beneath the chart they are staring at.

The fix is not replacing the milestone chart. The committee still needs it. It is walking in with both: the task-completion view they expect, and the adoption evidence that answers the question milestones were never built to answer. That evidence looks like system usage shifting, ticket mix changing, escalations resolving without you.

Naming that second view out loud, on a cadence, is the same practice behind building real shared understanding. It replaces just sending more updates. It is often the difference between a program that gets the benefit of the doubt through a slow quarter, and one that does not.

Where AI actually helps here

None of this needs to be automated to be useful. Turning it into a dashboard too early tends to flatten exactly the nuance that makes it worth doing by hand. What AI is genuinely good at is being a thinking partner across more raw material than one person can hold in their head at once. Using it that way changes each piece above.

The reporting problem needs a skeptical reader. Feed it your notes before the steering committee meeting. Ask it to argue the skeptical read: why would this not be enough evidence? You want the weak spots in your own story surfaced in private, not live in the room.

For the data problem, it is a fast way to brainstorm what exhaust already exists in systems you have not thought to check. Think ticket tools, login logs, training platforms. Do that before you go ask their owners for it.

The two-faced signal problem needs a different move. Ask it, every time, what it would mean if this were actually the bad version. Let it hold that second read until someone verifies which one is true.

Ownership and cadence benefit from the same trick. It can sit across weeks of raw check-in notes and flag what a single re-read would miss. For example: the same person always reporting the good news, or a thread mentioned once in week three that quietly never came up again.

Finally, the scenario itself, the specific detail that makes any of this credible, works best when AI plays interviewer. It pushes past “things are going well” until it gets specifics: a name, a system, a week, a thing someone actually said.

What to walk into the room with

None of that replaces judgment. It extends how much a director buried in status updates can actually notice, which was always the real constraint here.

The steering committee wants the milestone chart, and it should get one. The team building the future deserves to hear, just as often, the version of progress that milestones were never built to capture. That version needs to be backed by evidence, checked for the version that says otherwise, and owned by someone on a schedule.

More essays like this one live on the Between Mondays articles page.

Further reading: the U.S. Government Accountability Office’s Schedule Assessment Guide and the Project Management Institute. Both publish on measuring real progress across long, complex programs.

Frequently asked questions

Why don't status reports capture whether a transformation is actually working?

Status reports track task completion, requirements signed off, code deployed, training delivered, which is real but different from whether the change actually took hold. Culture, the slowest of the three layers of change, doesn’t close out on a Gantt chart; it shows up in whether people reach for the new way without being told to.

What data already exists to measure real progress on a change program?

Most programs are already generating usable signal: system usage on the legacy system versus the new one, how support tickets are categorized, the shape of the defect-versus-workaround trend after go-live, and training completion versus real proficiency two months later. Nobody has to build new instrumentation, just go ask the system owners for what they already track.

How do you tell if a drop in escalations is good news or bad news?

You can’t tell from the number alone, since a quiet week can mean the problem got solved or that people gave up flagging it. Ask the team that used to raise the issue, directly and on a schedule, whether it actually went away or they simply stopped bringing it up.

Who should own tracking adoption evidence on a long transformation program?

It needs a named owner and a recurring cadence separate from status reporting, typically a short check-in every one to two weeks asking what people noticed rather than what they completed. Rotate who is asked so one person’s optimism doesn’t get mistaken for a trend.

How can AI help track progress that a status report misses?

It’s useful as a thinking partner: stress-testing a steering committee update for weak spots before the meeting, brainstorming what usage or ticket data already exists to check, and holding the skeptical read on any good-looking signal until someone verifies it. It doesn’t replace judgment, it extends how much a director buried in updates can actually notice.