UniCache: Adaptive Task- and Type-Aware KV Cache Compression for Unified Multimodal Models
Wanqi Yang ⋅ Yuexiao Ma ⋅ Mei Xie ⋅ Xiawu Zheng
Abstract
Unified multimodal models based on Mixture-of-Transformers (MoT) support both image understanding and generation through a shared attention mechanism. In this architecture, the KV caches exhibit substantial heterogeneity across tasks and cache types. Existing KV-cache compression methods primarily target a single execution mode, and applying them uniformly to unified models can degrade performance on individual tasks. In this work, we systematically analyze KV caches in BAGEL across understanding and generation tasks, and propose UniCache, an adaptive task- and type-aware KV-cache compression framework, which identifies the cache segments activated by each task, applies tailored compression policies to different segment types, and compresses these segments in parallel through attention-guided budget allocation and task-aware temporal budget scheduling. Experiments show that UniCache can sustain balanced quality at high compression ratios: on BAGEL it reduces KV-cache memory by up to 5$\times$ while matching the original model's performance and achieves up to 1.78$\times$ the throughput of the baseline in long-context settings.
Chat is not available.
Successful Page Load