Auditing Evaluation Metrics for Small Language Model Replacement in a Local AAC System
Abstract
Our augmentative and alternative communication (AAC) agent uses a local 3B-parameter small language model (SLM) to propose phrases for a ten-slot board. A closed-vocabulary MLP and a phrase-bank bi-encoder replace this stage. We audit the deployed evaluation on 84 held-out synthetic scenarios, comparing both students with fixed and random boards, repeated random draws, and shuffled-context controls. All 20 seeded random baselines drawn from the MLP's inventory pass the distinct-functions gate, whether or not the four pinned words are included. Neither student differs detectably from the random baseline in paired function agreement. Permuting contexts leaves the collection of student boards unchanged, so the gate's aggregate counts cannot change. Paired function agreement shows no detectable change in these runs, although intent coverage falls. The intent-coverage rule has a separate flaw: it gives a fixed board of three hedges a score of 0.75 because of a tokenizer artefact. A patch reduces that score to 0.14, but examples show other false matches. Both inventories had also used held-out outputs, although no test row was a training example. We call this inventory leakage. Rebuilding inventories and class statistics from training rows and retraining lowers the bi-encoder's word agreement, emotion presence, and coverage; none of the three clean seeds reproduces its served results. The audit supports comparing the same output slots, testing context-free and shuffled-context controls, and building inventories only from training data.