Timezone: »

Policy Gradient Coagent Networks
Philip Thomas

Mon Dec 12 10:00 AM -- 02:59 PM (PST) @ None #None

We present a novel class of actor-critic algorithms for actors consisting of sets of interacting modules. We present, analyze theoretically, and empirically evaluate an update rule for each module, which requires only local information: the module's input, output, and the TD error broadcast by a critic. Such updates are necessary when computation of compatible features becomes prohibitively difficult and are also desirable to increase the biological plausibility of reinforcement learning methods.

Author Information

Philip Thomas (University of Massachusetts Amherst)

More from the Same Authors