From Learned Values to Simple Rules: RL-Guided Strategy Discovery in Food Production
Abstract
Food factories rely on fixed operating procedures because they are stable and easy to audit. These procedures work well in normal production, but they can react slowly to rare contamination events. We study whether reinforcement learning can find when a factory should clean earlier. Our simulator follows a reported factory practice of 18 hours of production followed by six hours of cleaning. It also includes hidden failures, noisy sensor readings, and delayed tests. The current procedure earns a steady profit in normal operation, but loses an average of 3,276 dollars in a simulated crisis. We use simulated future outcomes to learn the value of cleaning after an alarm. The learned policy reduces crisis loss by 182 dollars compared with a strong two-alarm rule while preserving routine profit. The learned values also point to a simple strategy: clean after the first strong alarm. On new simulations, this rule reduces crisis loss by 586 dollars and gains 5.83 dollars per day. The main lesson is not that learning should replace the factory procedure. It is that learning can help find a small and useful change to a mature process.