A story point scale is the set of values a team uses when estimating product backlog items.
The most common scale is a modified Fibonacci sequence such as:
1, 2, 3, 5, 8, 13, 20, 40, 100
The exact numbers matter less than the reason for using a limited scale. A good story point scale reminds the team that estimates are approximate. It gives enough choices to be useful, but not so many that the team wastes time pretending it can tell the difference between 14 and 15 points.
What a Story Point Scale Does
A story point scale turns relative judgment into a usable estimate.
When a team looks at a product backlog item, it asks questions such as:
- Is this smaller or larger than items we have estimated before?
- Is it closer to a three or a five?
- Is it large enough that we should split it?
- Is uncertainty driving the estimate up?
The scale gives the team a few useful buckets for those answers.
A scale should support the conversation rather than become the conversation. If the team is spending a long time debating whether an item is a six or a seven, the scale is inviting precision the team probably does not have.
Why Teams Use Fibonacci-Style Scales
Many teams use Fibonacci-style values because the gaps get wider as the numbers get larger.
That widening matters. The larger a product backlog item is, the less precisely the team can usually estimate it. The difference between a one and a two may be meaningful. The difference between a 41 and a 42 is almost certainly noise.
A modified Fibonacci scale accepts that reality. It gives the team small distinctions for small items and larger buckets for large items.
The common agile scale is not a perfect Fibonacci sequence. Many teams use 20, 40, and 100 instead of continuing with 21, 34, 55, and 89. Those rounder numbers avoid a false impression that the team knows the estimate with mathematical precision.
A twenty-point item is big. A forty-point item is much bigger or much less certain. A hundred-point item is usually a signal that the item is too large or too unclear to be useful yet.
Other Scales Can Work
Fibonacci-style numbers are common, but they are not mandatory.
Some teams use powers of two:
1, 2, 4, 8, 16, 32
Other teams use T-shirt sizes such as small, medium, large, and extra-large. Some use playful scales such as animals or dogs. These can work when the team understands what the categories mean and can still use the estimates to support decisions.
Numeric scales have one advantage: they are easier to use with velocity and forecasting. If the team wants to forecast how much work may fit into future sprints, numbers are more convenient than labels.
Non-numeric scales can be useful early when a team is trying to avoid turning points into days. They can also be useful when the Product Owner only needs rough ordering.
For teams that want to add estimates, calculate velocity, or forecast, story points should use numbers that add sensibly.
Use Numbers That Add Sensibly
Story point numbers need to add in a useful way.
A four-point item should be about twice the effort of a two-point item. Two two-point items should be in the neighborhood of one four-point item. The relationship will never be perfect, but it needs to be close enough that totals and velocity mean something.
That is why teams should avoid using numbers as mere labels. A team should not say 1 means "an hour," 2 means "a day," and 3 means "a week." Those numbers do not add sensibly. Three one-point items would not be roughly the same as one three-point item.
Story points are approximate, but they still need to behave like numbers. Otherwise sprint totals, velocity, and forecasts become misleading.
Large Numbers Are Signals
A large value is not a badge of importance. It is a warning that the item may be too large, too uncertain, or too poorly understood for the next decision. The right response may be story splitting, more refinement, a spike, or an explicit decision to keep the estimate rough because the item is still far from implementation.
Large story point values are useful because they tell the team something.
A large estimate may mean:
- The item is genuinely large
- The item includes too much scope
- The team does not understand the item well enough
- The item has significant risk
- The item should be split
- The team needs a spike or another learning step
A large estimate should not automatically be treated as a plan. If an item is estimated at 40 or 100 points, the Product Owner and team should ask what decision the estimate supports.
Sometimes a large rough estimate is enough. For example, if the Product Owner only needs to know that an item is much too large to consider soon, a 40 or 100 may be useful.
But if the item is being considered for an upcoming sprint, the large number usually says, "We need a better conversation before this is ready."
Avoid False Precision
False precision happens when a team uses numbers that imply more accuracy than the team has.
A scale that includes every number from 1 to 100 invites arguments that are not worth having. The team may debate whether an item is 17 or 18 points, even though no one can reliably know that.
A good scale prevents some of that waste. It says, in effect, "Choose the bucket this item belongs in."
The bucket matters because the decisions matter. A Product Owner may act differently if an item is a three instead of a thirteen. The Product Owner probably will not act differently if an item is a thirteen instead of a fourteen.
Use a scale that keeps the team focused on useful distinctions.
Include Zero Carefully
Some teams include zero in their scale.
Zero can be useful for extremely small work, such as a typo fix or configuration change that is too small to affect forecasting. But teams should use zero carefully. Too many zero-point items can hide work, distort capacity discussions, and make the team's actual effort less visible.
If an item has meaningful effort, give it a meaningful estimate. If it is truly tiny, zero may be acceptable.
The important question is whether the estimate helps the team and Product Owner understand the work.
Keep the Scale Team-Specific
Story point scales are team-specific.
Two teams can both use the number five and mean slightly different things. That is acceptable as long as each team uses its own scale consistently.
Problems begin when managers compare teams by velocity. A team completing 40 points is not automatically twice as productive as a team completing 20. They may use different baselines, have different skills, work in different domains, or estimate with different internal meanings.
When multiple teams need to forecast together, they need a deliberate way to align or normalize estimates. A shared set of baseline items can help. A manager ranking teams by point totals will not.
Common Story Point Scale Problems
Too Many Values
A detailed numeric scale creates arguments that do not improve decisions. Remove values that invite false precision.
Values That Secretly Mean Days
If one point equals one day, the team is back to time estimating. The scale should support relative effort, not disguise duration.
Numbers Used as Labels
Story point numbers need to add sensibly. If 1 means an hour, 2 means a day, and 3 means a week, the team is using numbers as labels rather than as relative estimates. That makes sprint totals and velocity misleading.
No Shared Baselines
A scale without examples is harder to use. Keep a few completed or well-understood items as anchors.
Large Items Treated as Ready
A 40 or 100 is often useful because it shows the item needs splitting, learning, or a decision. Do not let the number make an unclear item look ready.
Comparing Teams by Points
A scale belongs to a team. Comparing raw velocities across teams creates pressure and often leads to estimate inflation.
Modified Fibonacci Versus Powers of Two
Most teams do well with either a modified Fibonacci scale or a powers-of-two scale.
A modified Fibonacci scale such as 1, 2, 3, 5, 8, 13, 20, 40, and 100 is popular because the gaps grow as uncertainty grows. It gives teams enough values for meaningful comparison without inviting arguments over tiny differences.
A powers-of-two scale such as 1, 2, 4, 8, 16, and 32 can also work. It emphasizes that each bucket is meaningfully larger than the one before it.
The exact sequence is less important than the behavior it encourages. A good scale keeps the team comparing product backlog items, discussing uncertainty, and avoiding fake precision.
How Large Should the Largest Value Be?
Some teams include values such as 40 or 100. Others stop at 13 or 20.
Both approaches can work, but they send different signals.
Large values are useful when the team is estimating early ideas, epics, or poorly understood product backlog items for broad planning. A 40 or 100 can mean, "This is too large or uncertain to treat as normal sprint-ready work."
A smaller maximum value is useful when the team wants to prevent large items from entering a sprint. If the largest value allowed for sprint candidates is 13, anything larger must be split, refined, or deferred.
One practical compromise is to allow large values during early planning, but require items near the top of the product backlog to be split into a range the team can finish within a sprint.
Use Zero Carefully
Some teams include zero on a story point scale.
That can be useful for tiny product backlog items that still need to be tracked but are too small to affect forecasting meaningfully. But zero can also create confusion. A zero-point item is not free. Someone still has to understand it, make the change, test it, review it, and deploy it.
Use zero only when the team agrees what it means. For many teams, it is better to avoid zero and use one for the smallest meaningful product backlog item.
Avoid Half Points and Decimal Points
Half points and decimal points usually make story point estimation worse.
If the team is debating whether an item is 2.5 or 3, the scale is encouraging more precision than the team probably has. That effort is usually better spent discussing whether the item is comparable to known two-point or three-point items, or whether something about the item is still unclear.
The goal is not to measure work precisely. The goal is to put the item in the right bucket for the decision being made.
A Scale Needs Baselines
A scale without examples remains abstract.
A team may agree to use 1, 2, 3, 5, and 8, but still disagree about what a five means. Baseline items make the scale usable. They let the team ask, "Is this more like the login change we called a three, or the reporting change we called an eight?"
If a team keeps arguing about the same values, the problem may not be the scale. The team may need better baselines.
Common Questions
What Is the Best Story Point Scale?
For most teams, a modified Fibonacci scale works well: 1, 2, 3, 5, 8, 13, 20, 40, 100.
Why Not Use 1 Through 10?
A 1-through-10 scale can work, but it often encourages false precision. Teams may spend time debating differences that are too small to estimate reliably.
Should We Use T-Shirt Sizes?
T-shirt sizes can work for rough early sorting, but story points should be numeric when the team wants to add estimates, calculate velocity, or forecast future work.
Should Every Team Use the Same Scale?
Teams can use the same sequence of numbers, but the meaning of the numbers remains team-specific unless teams deliberately share baselines.
Is a 100-Point Item Allowed?
Yes, but treat it as a signal. A 100-point item is usually too large, too uncertain, or too far away to discuss in detail.
Recommended Articles
- Why the Fibonacci Sequence Works Well for Estimating
- The Best Way to Establish a Baseline When Playing Planning Poker
- 7 Ways to Get the Best Estimates of Story Size
- Better Estimates Are Possible on Agile Teams
Related Pages
- Agile Planning explains how estimates and velocity support forecasts.
- Product Backlog Refinement helps teams prepare items before assigning estimates.