Small Models, Exact Memories: Stress-Testing Optical Memory for Small Multimodal Agents
Abstract
Can unmodified small multimodal readers use rendered histories as memory without losing exact state? We evaluate zero-shot optical memory as an alternative representation, not a new compression architecture, using two tests: an 18-episode controlled exact-state stress test across densities and Qwen3.5 scales, and long-horizon conversational QA on a frozen 100-question LoCoMo subset with a GLM second-family check. Dense 8 px LoCoMo histories score .027/.099 F1 for Qwen3.5-2B/9B, versus .384/.410 for raw text and .415/.422 for input-count-matched BM25. Gold-selected evidence at the same density recovers to .383/.510, so unreadability alone cannot explain full-history failure; selection, distractor load, page count, and context use remain jointly implicated. GLM-4.6V-Flash shows the same pattern: .182 for full optical, .403/.402 for raw/BM25, and .474 for selected optical. In the controlled study, larger Qwen readers recover more exact state, but accuracy at the only processor-measured compressed point spans 5.6%–55.6%; GLM reaches 27.8%. Scaling helps within Qwen3.5, but naive dense optical packing is neither competitive with retrieval-conditioned text nor reliable for exact technical state in this tested zero-shot regime. These results establish neither a universal scaling law nor a general disadvantage for engineered optical-memory systems.