Last-Turn Query Placement in Chat-Served RAG
Abstract
Chat-served retrieval-augmented generation (RAG) often emits each evidence chunk as its own user turn. We freeze the selected evidence E' and vary only where the question sits in that chat. On eight instruction-tuned models (1.5B to 9B), a query-first layout that never repeats the question drops token F1 below 0.05 on four models at depth 100% with the needle guaranteed in E' (recency@50, 16k-word contexts; n=10, and n=34 for two 7B and 8B models). Packing the same E' into one user message leaves query-first collapsed on DeepSeek-R1 1.5B and Gemma2 9B (F1 0.045 and 0.021, n=10). Placing the question after the packed block recovers to the multi-turn question-last baseline. This is last-turn query semantics: the last user turn is treated as the query. The result is a serving-interface last-turn diagnostic on chat APIs, not evidence about long-range attention.