What You Say Versus When You Say It: Efficiently Predicting Service Completions with LLMs and Stochastic Processes
Abstract
Accurately determining whether an ongoing service conversation will continue or has effectively ended is central to operational control in chat-based service systems. Premature conversation closures can induce costly customer re-contacts, whereas overly cautious closures waste agent capacity. Using large-scale conversational data from a food delivery service organization, we study two related prediction tasks: short-term conversation continuation and end-of-service completion. We compare three modeling paradigms—stochastic self-exciting point processes, metadata-based neural networks, and large language models (LLMs) operating on full conversation text. Our results reveal a sharp distinction between the tasks. For short-term continuation prediction, lightweight Hawkes process models capture temporal dynamics most effectively and achieve competitive accuracy at minimal computational cost. In contrast, end-of-service prediction benefits substantially from semantic information: a fine-tuned Transformer-based text classifier outperforms all temporal models, while zero-shot prompted LLMs perform poorly despite their scale. A cost–accuracy analysis shows that large models are only justified for specific, high-stakes decisions. Motivated by these findings, we propose a hybrid approach that combines real-time Hawkes-based monitoring with selective LLM queries. This strategy exploits the complementary strengths of temporal and semantic models and achieves a best-of-both-worlds cost–accuracy tradeoff for operational decision-making in service conversations.