Contrasting foraging and reinforcement learning computations in the prefrontal cortex of macaques
Published in Cosyne 2026, 2026
Summary. Flexible behavior relies on computations that weigh alternatives, integrate new evidence, commit to options, and abandon courses of action when evidence no longer supports them. These processes are commonly tested in economic tasks, where agents make discrete choices between simulta- neously presented options. The dominant formalism for understanding how humans and other animals cope with the challenge is reinforcement learning (RL). More recently, however, an ecological account has emerged, which frames behavior as foraging, emphasizing accept–reject decisions over sequentially encountered options. Identifying the computational framework that the brain implements requires iden- tifying the neural representations of these abstract computations. The prefrontal cortex has a central role in flexible behavior: it computes value signals and translates them into control policies. How- ever, the precise computations underlying value estimation and the mechanisms that convert values into choices remain unclear. We recorded 5000 neurons in 113 sessions from three frontal areas in macaques performing a standard multi-armed bandit task. Behavioral modeling recovered value signals and decision policies for both computational frameworks. Although RL and foraging frameworks posit fundamentally different algorithms (distinct value definitions and control policies), behavior was not sufficiently powerful to distinguish them. However, population codes revealed a decisive dissociation between the two frameworks. We identified a neural subspace in the midcingulate cortex (MCC) and ventrolateral prefrontal cortex (vLPFC) that precisely tracked value computations. Capitalizing on this low-dimensional representation, we showed that MCC dynamics exhibited signatures aligned with foraging-style computations, and conflicted with RL requirements. We propose a mechanistic account of adaptive behavioral control in the prefrontal cortex, with the MCC tracking value and comparing them to a switching-threshold, whereas the LPFC implements the strategy indicated by the MCC. This policy is ecologically motivated and computationally efficient. Together, these results provide a detailed neural account of foraging-like flexibility in the primate frontal cortex.
If the embedded PDF does not load, open it directly here: POSTER_UNGVARSZKI_cosyne26.pdf.
