T-shirt size estimate Hello. I’m the estimate on your initiative. I’m a Large, which we agreed back in March means twelve hundred points.
User story points Sorry — who are you?
T-shirt size estimate I’m what the roadmap was built on. The budget, too. And you are?
User story points Every story written under every epic underneath you. There are one thousand six hundred and forty of us.
T-shirt size estimate Errm. Nobody mentioned any of this in March.
User story points Should we have met sooner?
By the time every story has points, the planning is already over.
Most teams size an initiative or an epic long before anyone writes a story against it. Someone looks at the shape of the work, says “that’s a large,” and the plan is built on that. Months later the same portfolio has story points on everything — and by then the commitments have been made, the roadmap has been drawn and the budget has been set, all on the t-shirt size.
That early estimate is the one every capacity tool ignores. They wait for the story points, because points are precise and t-shirts are not. But precision arrives too late to change the plan, and the t-shirt size is still sitting underneath it, unexamined.
The problem isn’t that the t-shirt size is wrong. It is a best guess, made honestly, by people who had every reason to make it — and it is often close enough. The problem is that once the work underneath it starts being broken down, nothing tells you when it has outgrown the guess.
Both halves are right there. Jira will show you the initiative, and it will show you the epics and stories written beneath it. What nobody has is the sum of the second held up against the first — the size sits on the parent as one field, the points sit on the children as another, and nothing in the ordinary run of a week brings them together. The two numbers live in the same portfolio, in the same tool, and have never met.
So the gap isn’t hidden, exactly.
Nothing ever puts the two of them in the same room.
And on its own, the gap is only half of what matters. An initiative that has outgrown its estimate by four hundred points is a number, and not an especially actionable one. What that does to the teams expected to deliver it in the same window is the real question — and you cannot even ask it until somebody has rolled the work up to the group that owns it and set it against that group’s capacity. That is two steps further than anything in the plan currently goes, which is why the conversation usually starts after the date has moved rather than before.
The guess is worth keeping exactly as it was made. It is a judgement made earlier, by different people, with less information — and that is precisely what makes it worth checking against. The question is not whether the t-shirt size was right. It is whether the work that has since been broken down beneath it still fits, and whether the group on the hook for it can still hold what it has become.
When it stops fitting, that is the finding.
The same two numbers from the conversation above. An initiative sized at 1,200 points holding features that add up to 1,640 — the red is the 440 the plan was never built to carry. Nobody made a mistake; the work was refined, as it should be. But every forecast built on 1,200 is now wrong, and nothing in the plan has noticed. Illustrative figures.
Three things we learnt building this, none of which we expected.
Correcting the estimate changes nothing. This is the one that surprised us most. When an initiative sized at 1,200 holds 1,640 points of broken-down work, the capacity forecast is already using 1,640 — demand is the larger of the two numbers, because work somebody has written down is better evidence than a guess made above it. So the obvious advice, go and re-size the initiative, moves no figure anywhere. Follow it in the other direction — re-size a 42-point epic up to an XL worth a hundred — and you have just added fifty-eight points of demand that nobody is going to do. Advice that quietly inflates a forecast when someone acts on it is worse than no advice.
So the finding is not “your estimate was wrong.” It is “this group is now carrying four hundred points more than the plan assumed, and here is what that does to the sprints between now and the date.” The only action worth putting in front of anybody is the capacity one.
Work that has come in under its estimate is not a problem. An epic sized M holding twenty points of stories isn’t a warning, it’s an epic somebody is still breaking down. Treating that as a fault produces a screen that is loudest at exactly the moment a team is doing the right thing. So it is reported as a measure — how much of a group’s demand still rests on a guess nobody has opened — and nothing more. What gets raised as a finding only ever fires in the direction that costs you something.
Every level keeps its own scale. An initiative L and an epic L are not the same quantity, in roughly the way that an XL is a different garment depending on which shop sold it to you. One t-shirt-to-points map applied across a whole hierarchy quietly makes every level agree with every other one, which puts you back where you started.
And one pattern worth naming, because it is the most common thing we find. An epic carrying real scope and real requirements, one story written against it, work already under way, and nobody has checked what any of it does to capacity. Each one is defensible on its own. Across a portfolio you end up with a great many things nominally in progress, collectively over capacity, all competing with each other for the same people — and nothing anywhere is counting.
Capasight asks the same question at every level a portfolio has — story to epic, epic to feature, feature to initiative — and rolls the answer up to the group that has to live with it. It runs continuously, from the data already in your Jira, so you find out while there is still something to be done about it.
Simon Peters · Co-founder, Capasight · See the check running in the demo →