From Unstructured Reviews to Return Intent: Theory-Guided NLI and Causal Machine Learning
Abstract
From Unstructured Reviews to Return Intent: Theory-Guided NLI and Causal Machine Learning Customer-generated reviews are unstructured and heterogeneous texts that can contain multiple aspects with different sentiments. While document-level sentiment analysis loses this structure, aspect-based sentiment analysis can generate results that are fragmented and difficult to aggregate into meaningful constructs. Sentiment scores obtained from topics or themes generated from the text could be helpful, but this approach has problems of its own. One conceptual attribute could be fragmented across multiple discovered topics, whereas one discovered topic can mix aspects that are conceptually distinct. The topics generated are corpus-dependent, and topic modelling does not inherently provide sentiment toward the identified topic. While such topics can be interpretable, their interpretation remains post hoc. For our problem, we therefore require a representation in which attributes are defined a priori through domain knowledge, while their expression and sentiment are inferred from the contextual meaning of the review. We use Garvin's eight dimensions of product quality as an established theoretical framework, formulating natural language hypotheses for each attribute and using natural language inference (NLI) to determine whether a review entails them. Each attribute is represented through positive and negative sentiment hypotheses, allowing attribute-specific sentiment to be captured across multiple dimensions within the same review, without manual lexical rules or aggressive preprocessing. We first identify product-relevant content, then evaluate sentiment toward product-quality attributes, and finally identify return intent as the concluding semantic extraction stage. We evaluate the framework using 100,784 one- and two-star Amazon US reviews of cellphones, laptops, and tablets posted from January 2020 onward. After price and category filtering, 70,738 reviews enter the semantic extraction pipeline, of which 56,316 are identified as product-relevant. Among these, 37.1% are classified as exhibiting return intent. We first evaluate whether the resulting quality-dimension representation contains predictive information using supervised classification: Random Forest achieves the strongest performance (AUC = 0.83, AP = 0.71), followed by Gradient Boosting (AUC = 0.82, AP = 0.71) and Logistic Regression (AUC = 0.77, AP = 0.66), with negative serviceability sentiment identified as the strongest predictor across both tree-based models. We then apply Double Machine Learning with a partially linear regression specification and XGBoost and Random Forest nuisance learners, controlling for star rating, review length, price, price tier, and product category. The estimated partial effects reveal substantial differences across quality dimensions. Serviceability has the largest estimated effect (β = 0.417), followed by conformance (β = 0.260), perceived quality (β = 0.206), reliability (β = 0.130), and durability (β = 0.128), while performance (β = −0.045) and features (β = −0.111) show opposite-signed estimates, suggesting reviewers who criticize specs are, if anything, less likely to express return intent once other dimensions and rating are held fixed. Serviceability also shows the strongest descriptive association with return intent, with a 39.1 percentage-point difference in return-intent rates between reviews that mention it and those that do not. Stratified estimates suggest further heterogeneity across product categories and price tiers, though formal price-tier interaction tests are not significant. These results demonstrate how domain knowledge can be operationalized as natural-language hypotheses to transform unstructured text into interpretable, attribute-specific representations for downstream machine learning. The framework connects contextual language understanding with theoretically grounded variables, while highlighting the importance of accounting for heterogeneity and shared textual information when interpreting observational causal estimates.