Adaptive Macro Learning
Abstract
Interactive learning for long-horizon tasks, such as LLM-based agentic tasks, recommender systems, and dialogue systems is crucial for real world deployment. In many of these systems, the primary objective of interest is available after a long horizon (e.g., episode outcome of a web shopping agent, outcome after multi-turn interaction in a recommender system) but optimizing for this macro outcome is difficult due to sparse rewards, credit assignment issues, and long context. To overcome these challenges, recent research proposes heuristics that use dense micro feedback with sparse macro feedback. In this work, we study and identify the conditions where micro feedback can improve the macro objective of interest. We then propose a principled method to leverage rich but imperfect micro feedback such that the overall gradient update improves the macro level performance. Empirical analysis on a synthetic task shows that our method adapts under varying conditions and remains competitive while learning more efficiently than existing baselines.