How to improve inventory forecasting accuracy on Shopify
Forecast accuracy is how close your predicted units come to what actually sells. You improve it by measuring error properly, feeding the forecast clean data, choosing a method that fits each SKU, and reviewing results on a fixed schedule.
Measure forecast error before you change anything
You cannot tell whether a change helped without a baseline number. Three measures cover the basics. In each formula, A is actual units sold in a period, F is the forecast for that period, and n is the number of periods.
MAPE = (100 ÷ n) × Σ |A − F| ÷ A
WMAPE = 100 × Σ |A − F| ÷ Σ A
Bias = 100 × Σ (F − A) ÷ Σ A
- MAPE (mean absolute percentage error) averages the percentage miss across periods. It is undefined when A = 0, which happens often for slow sellers, and it gives a 50% miss on a 5-unit week the same weight as a 50% miss on a 500-unit week.
- WMAPE (weighted MAPE) divides total absolute error by total actual units, so high-volume periods and SKUs count for more. Calculated on units it weights by units; calculated on revenue it weights by dollars.
- Bias shows direction. A positive value means you over-forecast and a negative value means you under-forecast. Over- and under-forecasts cancel each other out here, so bias can read zero while individual errors are large.
A hypothetical SKU over four weeks, with a flat forecast of 45 units a week:
| Week | Actual | Forecast | Forecast − actual | Absolute error |
|---|---|---|---|---|
| 1 | 40 | 45 | +5 | 5 |
| 2 | 50 | 45 | −5 | 5 |
| 3 | 30 | 45 | +15 | 15 |
| 4 | 60 | 45 | −15 | 15 |
| Total | 180 | 180 | 0 | 40 |
WMAPE = 40 ÷ 180 × 100 = 22.2%. MAPE = (5 ÷ 40 + 5 ÷ 50 + 15 ÷ 30 + 15 ÷ 60) ÷ 4 × 100 = (0.125 + 0.10 + 0.50 + 0.25) ÷ 4 × 100 = 24.4%. Bias = 0 ÷ 180 × 100 = 0%. The forecast missed every week, yet bias is zero because the misses cancel. Read bias and an error measure together.
Then compare against a naive forecast, which simply uses last week's actual as this week's forecast. In the example, the naive forecasts for weeks 2 to 4 are 40, 50 and 30. Their absolute errors are 10, 20 and 30, a total of 60 against 140 actual units, so naive WMAPE is 60 ÷ 140 × 100 = 42.9%. The flat forecast has errors of 5, 15 and 15 over the same weeks, 35 in total, which is 25.0%. It beats the naive forecast. A method that cannot beat the naive one is not earning its complexity.
1. Clean the history the forecast learns from
Every method, from a moving average to a statistical model, assumes past sales reflect past demand. Several situations break that assumption:
- Stockouts. Sales were capped by supply, not by demand. Zero sales while out of stock is a missing observation, not a zero.
- Promotions and one-off events. A flash sale or a single bulk order skews later baselines unless you flag it.
- Returns. Forecast net units, after returns, not gross units.
- Other sales channels. If one pool of stock serves several channels, forecast the combined demand.
- Variants. Forecast each size and color where they sell differently. Rolling them into one product hides variant-level patterns.
Say a SKU sold 28 units over 28 days but was out of stock for 8 of them. The raw average is 28 ÷ 28 = 1.0 unit per day. The in-stock average is 28 ÷ 20 = 1.4 units per day. The raw figure understates demand by (1.4 − 1.0) ÷ 1.4 ≈ 29%.
To start, export 12 months of orders, add columns for out-of-stock days and promotion days, and calculate baselines from the clean rows only.
2. Match the method to the history you have
Complex methods need more data to estimate their parameters. On thin history, a simple method applied well is the safer choice. These boundaries are rules of thumb.
| History available | Reasonable approach | Reason |
|---|---|---|
| A few weeks to about 3 months | Short moving average plus judgment | Too little data to estimate a trend or a season |
| About 3 to 12 months | Weighted moving average or exponential smoothing | Recent sales matter most, but one partial year cannot show seasonality |
| About 12 to 24 months | Baseline times a seasonal index | One cycle gives an index, though it cannot separate a season from a one-off event |
| More than 24 months | Seasonal index averaged over several years, plus trend | Repeated cycles show which patterns recur |
The formulas for each method are in the demand forecasting guide.
3. Segment SKUs by how forecastable they are
Two SKUs with the same average sales can behave very differently. The coefficient of variation (CV) measures how much demand swings relative to its average.
CV = standard deviation of demand ÷ mean demand
Weekly sales for two hypothetical SKUs over six weeks:
- Product A: 8, 12, 10, 9, 11, 10. Mean = 10, sample standard deviation ≈ 1.41, CV ≈ 0.14.
- Product B: 0, 15, 2, 0, 20, 23. Mean = 10, sample standard deviation ≈ 10.56, CV ≈ 1.06.
Both average 10 units a week. Product A is easy to forecast and Product B is not.
There is no universal cutoff. A store might treat a CV below 0.5 as stable and above 1.0 as erratic, then adjust after seeing which SKUs actually miss. A fuller classification also looks at how often demand occurs at all, giving four groups: smooth, erratic, intermittent and lumpy. Then match effort to the group:
- Smooth (stable) SKUs: a standard method and a modest buffer.
- Erratic SKUs: expect high error. Use a larger buffer and a shorter review cycle rather than a more elaborate model.
- Intermittent SKUs: many periods with zero sales. Use a method built for intermittent demand, a minimum and maximum stock rule, or ordering to demand.
- Lumpy SKUs: demand is both rare and highly variable, so expect the largest misses. Use a minimum and maximum stock rule or order to demand, and review often.
Put your forecasting effort where revenue is concentrated; ABC analysis shows you where that is.
4. Find and fix bias
Accuracy tells you how big the misses are. Bias tells you which way they lean. A forecast that is consistently too high ties up cash in stock. One that is consistently too low causes stockouts. Common causes of over-forecasting include padding quantities because stockouts feel worse than overstock, leaving promotion spikes in the baseline, and using gross instead of net units. Common causes of under-forecasting include stockout days counted as zero demand and ignoring growth.
Say over eight weeks you forecast 480 units and sold 400. Bias = (480 − 400) ÷ 400 × 100 = +20%. Scaling future forecasts by 400 ÷ 480 ≈ 0.83 would remove that average over-forecast. A correction factor is a stopgap, though. Find the cause first, because the factor hides the problem instead of fixing it.
Track bias by category and by supplier, not just store-wide. A store-wide figure near zero can hide one category that runs high and another that runs low.
5. Add information the sales history does not contain
Sales history shows what happened, not what you have planned. Useful inputs include promotion dates, price changes, launches, planned ad and email pushes, supplier restock dates, and known one-off orders.
Web traffic and add-to-cart counts may lead orders, but test that on your own data before relying on it. Compare traffic in one week with orders in the following weeks, and use the signal only if the relationship holds up across several periods. A signal that sounds plausible but fails that check adds noise.
6. Forecast no further out than the decision needs
Error tends to grow as the horizon lengthens. Match each forecast to the decision it supports, and make it span lead time plus the interval between order reviews:
- A domestic supplier with a 2-week lead time and weekly reviews needs a forecast covering 3 weeks.
- An overseas order with a 10-week lead time and reviews every 2 weeks needs 12 weeks.
- A seasonal buy placed months before the season needs a longer, rougher forecast. Plan it as low, expected and high scenarios instead of one number.
Refresh short-horizon forecasts as new sales arrive.
7. Let the remaining error set your safety stock
No method removes all error, so size safety stock from the error of your forecast rather than from raw demand variation. If daily forecast errors have standard deviation σe (in units per day) and are independent from day to day, then over a lead time of L days:
Safety stock = z × σe × √L
z is the service-level factor, about 1.65 for a 95% target.
Say σe = 3 units per day, L = 9 days and z = 1.65. Safety stock = 1.65 × 3 × √9 = 1.65 × 3 × 3 ≈ 15 units. If better forecasting cuts σe to 2, safety stock = 1.65 × 2 × 3 ≈ 10 units. Smaller error means less cash sitting in buffer stock.
The full calculation, including lead-time variability, is in the safety stock formula guide.
8. Review on a schedule
A forecast drifts without review. A routine catches small errors before they become stockouts or overstock.
Weekly
- Compare forecast with actual units for your top SKUs.
- Flag any SKU whose error exceeds a threshold you set, for example 30%.
- Find the cause: bad data, a market change, or a limit of the method.
Monthly
- Calculate WMAPE and bias across the catalog and by category.
- Compare with the naive forecast.
- Move SKUs to a different method when their history has grown.
- Estimate what stockouts and overstock cost you. The inventory KPI guide and inventory turnover help here.
Quarterly
- Ask whether your approach still fits the catalog, and set targets relative to your own baseline.
A 30-day plan
- Week 1, baseline: export 12 months of sales, list your top 20 SKUs by revenue, mark stockout and promotion periods, and calculate WMAPE and bias.
- Week 2, clean and segment: adjust history for stockouts and promotions, calculate CV for each SKU, and pick a method to fit each history.
- Week 3, add inputs and review: add your promotion calendar, compare last week's forecast with actuals, and check bias by category.
- Week 4, make it routine: schedule the weekly review, document the process, and set a WMAPE target against your week-1 baseline.