Wednesday, September 09, 2026
Over the years managing technology projects, I have seen story points become much more than an estimation technique. They usually start inside a development team as a way to discuss complexity, uncertainty and expected effort, but after a few sprints those same numbers begin appearing in dashboards, management reports and PMO presentations. At that point, it becomes very easy to look at two teams and say that one delivered 100 story points while another delivered 50, or to assume that two activities estimated at five points should require approximately the same amount of time to complete.
The problem is that story points were not created to work as a standardized unit of project delivery. They are relative estimates whose meaning depends on the context of the team that created them. A five-point activity for one team may represent a completely different combination of complexity, uncertainty and effort from a five-point activity estimated by another team. When those numbers remain inside the team, this difference may not create a problem. When they move into portfolio reporting and begin to be compared across projects, however, their interpretation becomes much more difficult.
To understand how significant this difference can be, Saint Jude analyzed 727,282 Jira issues from 204 software projects across 19 different domains. We compared story-point behavior with project delivery data to understand whether projects using similar estimates actually showed similar delivery times. The objective was not to prove beforehand that story points work or do not work, but to observe what happens when a relative estimation method is treated as if it were a comparable project metric.

Story points are generally used to represent relative differences between activities. A team may assign 1, 2, 3, 5, 8 or 13 points according to its perception of complexity, uncertainty and expected effort. The important part is that these values make sense in relation to the other activities estimated by the same team. A three-point activity is expected to be different from a five-point activity because the team has created its own internal reference for what those numbers mean.
The difficulty begins when this relative scale is translated into time. It is common to hear questions such as how many hours correspond to one story point, whether five points represent one or two days of work, or how many story points a developer should complete during a sprint. Those questions are understandable because managers need deadlines, capacity forecasts and delivery expectations, but the point scale itself does not provide a universal conversion. One team may historically complete a certain type of five-point activity in a few days, while another team may classify a substantially different type of work with exactly the same number.
If story points could be safely converted into a common delivery metric, projects using the same point values should show at least reasonably similar delivery behavior. That does not mean every activity would take exactly the same amount of time, because projects always contain variation, but we should expect the differences to remain within a range that makes the number useful as a common reference. This is where the data begins to challenge that interpretation.
We first looked at projects whose median story-point estimate was exactly one. This is useful because it gives us a simple reference: if the number one had a reasonably consistent relationship with delivery across independent projects, their median cycle times should not be completely disconnected from each other. Instead, the projects showed extremely different behaviors even though their median estimate was identical.
Project Median story points Median cycle time Project 1 1 0.7 hours Project 2 1 115.5 hours Project 3 1 161.9 hours Project 4 1 197.4 hours Project 5 1 4,464.8 hours Project 6 1 8,246.3 hoursAll of these projects have the same project-level median estimate, but their median cycle times are radically different. At one extreme, the median is less than one hour. At the other, it exceeds 8,000 hours. Even if we ignore the most extreme values and look only at the projects in the middle of the table, the differences remain significant enough to show that the number one does not carry a standardized delivery meaning from one project to another.
For me, this is the first important result of the analysis. A story point can be meaningful inside the historical context of the team that created it, but when that number is removed from that context it tells us very little about how much calendar time an activity should take. The estimate still exists, but the reference that gives meaning to the estimate has been lost.

The difference does not disappear when we move to larger estimates. Projects with a median of around three story points also show very different median cycle times. In one project, the median was approximately 120 hours. In another, it was close to 292 hours. A third project reached approximately 454 hours even though the median story-point value was the same.
This is particularly important when story point estimation begins to move beyond the team and becomes part of portfolio governance. A PMO may see the number three in different boards and assume that those activities are comparable because the label is the same. In practice, each team may have created that estimate using a different historical reference, a different understanding of complexity and a different level of technical or organizational uncertainty.
The problem, therefore, is not necessarily the act of estimating an activity with story points. The problem is assuming that the numerical scale remains comparable when it crosses the boundaries between teams and projects. Once the original context disappears, the number continues to look precise even though its meaning may have changed completely.
Another part of the analysis makes this difference even clearer. We found projects with a median estimate of five story points whose median delivery time was significantly shorter than projects using a median of only one story point. In one example, a project with a median estimate of five points had a median cycle time of approximately 140.5 hours, while another project using a median of one point had a median cycle time above 2,000 hours.
The same pattern appears elsewhere in the dataset. A project using a median of three points had a median delivery time of approximately 120 hours, while another project using one point exceeded 4,400 hours. This does not mean that a five-point activity is inherently faster than a one-point activity, and it would be incorrect to interpret the result that way. What it shows is that absolute point values from independent projects do not form a common elapsed-time scale.
Each team creates its own estimation reference. If the team uses that reference consistently, story points may still help it discuss relative complexity and plan work based on its own historical behavior. But when we combine several teams and interpret their point values as if they belonged to the same scale, we introduce a level of comparability that the data does not support.
Imagine an organization with three development teams. Team A completes 80 story points during a sprint, Team B completes 45 and Team C completes 110. If those numbers appear together in a management dashboard, the immediate interpretation may be that Team C is delivering more and Team B is delivering less. The dashboard is mathematically correct, because the totals were calculated properly, but the conclusion depends on an assumption that may not be true: that all three teams are using equivalent estimation scales.
One team's five-point story may represent something another team would estimate as three or eight points. The difference may come from technical complexity, architecture, experience, team composition, uncertainty or simply from the estimation habits that developed over time. If these scales are different, comparing their absolute totals can create an appearance of precision without creating a standardized measure of performance.
The behavior we observed across the 204 projects reinforces this problem. Projects using the same median story-point value can have dramatically different median cycle times. This means that story points can provide useful context inside a project while becoming much more difficult to interpret when they are aggregated across multiple independent teams.
Story points normally originate as a team-level Agile estimation technique, but in many organizations they eventually become part of project and portfolio reporting. A PMO may create indicators such as total story points delivered, average points per sprint, velocity by team, velocity trends, productivity comparisons and portfolio capacity. These calculations may all be technically correct, but that does not necessarily mean that the underlying values are comparable.
If each team has created its own relative scale, adding story points together does not automatically transform them into a standardized portfolio metric. A team completing 100 points is not necessarily producing twice as much work, value or delivery as a team completing 50. The numbers tell us what happened according to each team's estimation system, but they do not give us enough information to make that productivity comparison safely.
For me, this is where project intelligence needs to move beyond the estimate created at the beginning of the activity. The original point value is still useful information, but it should be interpreted together with what happened during execution: how long similar activities normally take, how often they become blocked, whether they move across several sprints, how much rework they generate and how their behavior compares with the historical performance of the project.
Project data contains many signals that can help us understand delivery behavior with more context than a single estimate. Cycle time shows how long an activity remains in the delivery lifecycle before completion. Sprint movement can reveal when work repeatedly crosses from one sprint to another. Dependencies show how much an activity relies on other work before it can move forward, while rework indicates when completed or nearly completed activities need to return to execution.
Handoffs provide another important signal because a task that changes ownership several times may behave differently from one completed by the same person or team from beginning to end. Workflow transitions can also reveal complexity in the execution process, especially when activities move repeatedly between statuses before reaching completion. None of these indicators needs to replace story points completely. Their value comes from providing context around the original estimate.
Consider two activities that were both estimated at five story points. The first behaves almost exactly like dozens of similar activities previously completed by the same team. Its cycle time remains within the normal range, it has no important dependencies and it progresses through the workflow without interruption. The second activity starts with the same five-point estimate but becomes blocked, moves across several sprints, changes ownership and accumulates rework. The estimate has not changed, but the delivery risk has.
This is why I prefer to think about software delivery metrics as a combination of signals rather than a single number. An estimate tells us what the team believed before or at the beginning of execution. Delivery data tells us what is actually happening as the project evolves.

This is the part of project management that interests me the most. Instead of asking only how many story points were assigned to an activity, I prefer to understand how similar activities have behaved before. How long did they normally take? How much variation existed between the fastest and slowest deliveries? Were activities with dependencies consistently slower? Did items that moved through several sprints behave differently? Did rework, changes in ownership or an increase in workflow transitions appear before cycle time started to grow?
The historical behavior of a project creates a reference based on what actually happened, rather than only on what the team expected to happen when the activity was estimated. This does not eliminate estimation. It gives estimation additional context and makes it possible to compare the original expectation with the reality of execution.
In Saint Jude Project Intelligence, we use project data to compare current activities with the project's historical behavior and with similar completed activities. Artificial Intelligence helps identify combinations of signals, deviations and recurring patterns that would be difficult to observe by looking at a single Jira field or an isolated sprint report.
The objective is not to replace the project manager or create another estimation scale. It is to create a layer of intelligence that helps PMOs and project leaders understand whether current activities are behaving consistently with similar work, whether execution is moving outside the historical pattern and whether the information available today still supports the assumptions made when the project was planned.
After comparing story-point estimates and median delivery times across 204 software projects, the data does not support treating story points as a universal project delivery metric. Projects with the same median story-point value can have dramatically different median cycle times, while projects using higher median estimates can sometimes deliver considerably faster than projects using lower point values.
Because of this variation, the analysis does not provide a basis for statements such as “one story point equals X hours” or “a team delivering 100 story points is twice as productive as a team delivering 50.” Both statements require story points to behave like a standardized unit, and that is not what we observe when independent projects are compared.
Story points are relative by design. That may make them useful inside the context in which they were created, but the same characteristic makes them difficult to standardize across teams, projects and company-wide governance. For me, the more useful question is therefore not how many hours a five-point activity should take. The better question is whether, based on the way similar activities and the project itself have behaved historically, we can understand what should reasonably be expected from that activity now.
That is where an estimate begins to become Project Intelligence. The story point remains one piece of information, but it becomes much more useful when it is interpreted together with actual delivery behavior, historical performance and the other signals the project is producing during execution.
Story points are only one signal inside a project. Saint Jude analyzes project data to identify patterns involving delivery time, risks, dependencies, rework, capacity and predictability, combining the original estimates with the historical behavior produced during execution.
You can upload your project data and receive a free analysis through our Project Health Checker. The objective is simple: instead of looking only at what the project was expected to do, use the data already generated by the project to understand how it is actually behaving.
Until next time!
Erik Scaranello
Here you can find everything about Costs & Margin