Efficient Local Vocabulary Selection for Context-Aware Augmentative Communication
Abstract
Context-aware augmentative and alternative communication (AAC) requires timely suggestions while preserving user choice. We study a local architecture that separates speech and scene perception from board construction, replacing per-board language-model generation with fixed-vocabulary candidate filtering and an explicit selector that combines pinned core entries, answer reservations, a hedge cap, and maximal marginal relevance (MMR). Across 120 synthetic scenarios with fixed teacher-generated candidate pools, MMR increases distinct communicative-function labels by 0.73 among six dynamic entries relative to relevance-only selection (95% CI 0.58–0.88). In a separate retrospective evaluation, learned and cosine candidate filters cover 61 and 60 of 71 annotated rows under the same selector, providing no established coverage advantage for the learned head. An OpenVINO serving configuration constructs boards in 4.50 ms median on an Apple M4 CPU against 9.05 ms for the PyTorch path, excluding perception, transport, and display; the reduction comes from the embedding runtime, while head quantisation reduces head storage by 2.8×. That configuration covers 62 of 71 rows against the teacher’s 68, without establishing 90% coverage retention. Candidate filtering, selection rules, and inference runtime therefore trade off independently in local AAC board construction, and function diversity should be reported alongside response coverage.