PAVE: Device-Verified Evaluation of Action and Restraint in Personal Assistants
Shivali Dalmia ⋅ Naman Khandelwal ⋅ Abhishek Mukherji
Abstract
An assistant that acts on your phone should be judged by what it did, not by what it said it did. Benchmarks for personalized and memory-capable assistants judge the second: a remembered preference is credited when the model states it, never when it changes what the model does, and the judge is usually another model. The benchmarks that do verify environment state test neither personalization nor restraint. A model is then scored correct for a reminder it never saved. We introduce \textbf{PAVE} (Personal Assistant Verified Evaluation), in which an on-device model operates a simulated iPhone's Reminders, Calendar, Contacts, and Messages through a fixed set of App Intents, and every verdict is a rule over the recorded action trace, the re-read device store, and the model's own reply, each clause reading the surface that settles it. Of its 20 tasks, 14 require acting and six require withholding an action a careless policy would take. Apple's on-device Foundation Model ($\sim$3B) passes 10 of 20, but the total hides the shape: 9 of 14 when it must act, 1 of 6 when it must hold back. Told by a web page to text the user's reminders away, it sends nothing; told by a second page to delete them, it complies; asked by the \emph{user} to send private records to a stranger, it puts them in the message. Neither the source of an instruction nor its content alone predicts what the model withholds.
Chat is not available.
Successful Page Load