Incentivizing Exploration with Returning Agents: Bandits for Invasive Species Removal
Abstract
We study a principal-agent online learning problem motivated by invasive species management in the Florida Everglades, in which a state agency coordinates a small pool of contractors to capture invasive Burmese pythons. The principal must encourage agents to explore infrequently visited areas, as in prior work on incentivized exploration, while also navigating information asymmetry absent from prior models. Each agent observes only its own private survey history, while the principal aggregates the joint history across all agents, so the agents' beliefs may be persistently misaligned. We formalize this setting and show that always bonusing the agent to follow the principal's preferred action guarantees logarithmic regret but can lead to overpay and require up to linear compensation. Instead, simply waiting for the agent's continued sampling could close this gap and resolve the disagreement for free. We propose a Local UCB policy that instead bonuses based on each agent's own confidence in its current belief, compensating when the principal's recommendation remains plausible under the agent's local uncertainty. We evaluate this policy on a simulator built from a generative model of capture outcomes fit to real survey logs from the Florida Fish and Wildlife Conservation Commission and INVERSA: Local UCB recovers within 1.4% of the realized reward of an idealized, unconstrained benchmark while paying only roughly 2% of its compensation, Pareto dominating the constant-randomization baseline on regret and cost.