Evaluating Product Cannibalization and Competitor Price Changes in Retail Demand Forecasting Using Cross-Price Elasticity Modeling
Yada Duangphaichoom
Abstract
Retailers change prices often, through short-term promotions and through longer-term adjustments, and either type of change can divert demand toward substitute products, a pattern known as product cannibalization. We formalize this with price elasticity of demand, $\varepsilon_{ii} = \frac{\partial \ln Q_i}{\partial \ln P_i}$, and its cross-price counterpart, $\varepsilon_{ij} = \frac{\partial \ln Q_i}{\partial \ln P_j}$, which captures how a change in product $j$'s price shifts demand for product $i$. Price elasticity underlies prior work that quantifies cannibalization through unit and sales diversion ratios derived from own- and cross-price elasticities [1], but cross-price elasticity is still hard to estimate systematically across large retail portfolios, and naive estimation often produces overconfident p-values because of autocorrelation in weekly sales data. Demand forecasting models have long used competitor price as a static input, but it remains unclear whether competitor price-change signals provide predictive information beyond competitor price levels alone. We propose a multi-phase pipeline to estimate these elasticities and test their value for demand forecasting, applied to a single retail store format spanning 580 categories, roughly 65,000 SKUs, and 2.28 million weekly product-level sales records over 50 weeks, obtained directly from the retailer's operational records. We isolate high-impact SKUs with a Pareto (80/20) filter, assign price tiers dynamically via Jenks Natural Breaks, and exclude promotional weeks, since promotion is a distinct treatment mechanism that introduces unusually large and structured price changes, so the remaining variation reflects regular-price changes only. Within each tier, we fit a log-log regression, $$\ln Q = \beta_0 + \beta_{ii} \ln P_i + \sum \beta_{ij} \ln P_j + \gamma M + \epsilon$$ with monthly seasonality dummies and HAC Newey-West standard errors. The Frisch-Waugh-Lovell theorem makes pairwise estimation tractable across thousands of SKU permutations by caching each SKU's own-price residualization for reuse. UDR/SDR (Unit and Sales Diversion Ratios) are generated from the resulting elasticities [1], with Benjamini-Hochberg FDR correction for multiple comparisons. We convert these elasticities into a feature set, price-gap percentage, cumulative competitor-discount impact, and market concentration, $HHI = \sum s_k^2$, and feed these into LightGBM, XGBoost, and CatBoost. LightGBM wins the most head-to-head category matchups and trains faster with a smaller memory footprint, so we set it as the default forecaster. Adding competitor price-change features improves accuracy modestly on average, measured by $WAPE = \frac{\sum |y - \hat{y}|}{\sum |y|}$, but the gain concentrates in categories with strong substitution structure: across the 84 categories meeting our panel-size criteria, a paired Wilcoxon signed-rank test confirms the improvement is significant ($p < 0.001$, $r = 0.76$), with 84.5% showing lower error under the advanced model. References [1] Y. Yuan, O. Capps, Jr., and R. M. Nayga, Jr., "Assessing the Demand for a Functional Food Product: Is There Cannibalization in the Orange Juice Category?," Agricultural and Resource Economics Review, vol. 38, no. 2, pp. 153–165, 2009.
Chat is not available.
Successful Page Load