Story point estimation

Story points measure size relative to other work, not time. That distinction is the source of most of the confusion around them, and getting it right is what makes the numbers useful.

What a point actually measures

A story point bundles three things: how much work there is, how complex it is, and how much uncertainty surrounds it. A tedious but well-understood migration and a small but genuinely novel feature can both be a 5, for different reasons.

Points are deliberately not hours. The moment a team agrees that one point equals four hours, points have become a worse unit of time and every benefit disappears. Relative sizing works because humans compare things far more accurately than they predict durations. Ask someone whether this story is bigger than that one and they will usually be right. Ask how many hours it will take and they will usually be wrong.

Calibrating with a reference story

Pick a story the team has already finished, one that everyone remembers and that felt small but not trivial. Call it a 2 or a 3. That is your anchor for the quarter.

From then on, estimation is comparison rather than invention. Is this bigger than the reference? Roughly double it? Every new team member gets the reference story explained on day one, which is far faster than trying to teach them what a 5 feels like in the abstract.

Recalibrate when the team composition changes substantially or when you notice the reference no longer resembles the work. Do not recalibrate quietly mid-sprint.

Using velocity without breaking it

Velocity is the number of points a team completes per sprint, averaged over the last three to five. Its legitimate use is forecasting: at roughly 30 points a sprint, a 120-point backlog is about four sprints of work.

It breaks the moment it becomes a target. Points are estimated by the same people whose performance is being measured, so a team under pressure to raise velocity will raise its estimates. The number goes up, nothing ships faster, and you have lost the forecasting tool as well. Compare a team against its own history, never against another team.

Quote a range rather than the mean when someone asks for a date. Feed your last four or five sprints into the story point calculator and it returns the average alongside the best and worst sprint you have actually delivered, which is the pair of numbers worth taking into a planning meeting.

Common failure modes

Estimating in a vacuum. If the story has not been refined, the estimate is a coin flip. Refine first, estimate second.

Anchoring on the first number spoken. This is exactly what simultaneous reveal prevents, and it is why estimating in a shared document tends to produce worse numbers than estimating with cards.

Averaging a split vote. A 2 and a 13 do not make a 7. They make a conversation. Talk, then re-vote.

Estimating everything. Sizing a copy change or a dependency bump costs more meeting time than the estimate is worth, and it inflates velocity with points that never carried any risk. Some teams stop estimating small work entirely and only size the items where the size is genuinely uncertain. That is a legitimate choice, not a failure of discipline.

Putting it into practice

Run the estimation itself with planning poker, using the Fibonacci deck once your backlog is refined enough for numbers to mean something.