PROTEUS: Agentic Microarchitecture Search for FPGA Language-Model Accelerators
Andy Dimnaku ⋅ Tun-Yu Chang ⋅ Qizheng Zhang ⋅ Genghan Zhang ⋅ Rubens Lacouture ⋅ Konstantin Hossfeld ⋅ Kunle Olukotun
Abstract
Language models increasingly combining attention, state-space, and linear-recurrent layers whose distinct parallelism, state, and memory-access patterns favor different accelerator architectures. Existing FPGA automation, however, typically lowers fixed programs or searches within compiler-defined transformations and hardware templates. We present PROTEUS, an agentic search framework that searches full-layer microarchitectures under a resource budget, restructuring data movement, memory storage, compute organization, numeric representation, and temporal scheduling. An asynchronous planner proposes microarchitectural hypotheses, while parallel workers implement and evaluate them. Across 12 layer-workload- precision configurations, PROTEUS reduces synthesized kernel latency by 1.4$\times$--30.6$\times$ (4.9$\times$ geo. mean) over a resource-matched Gemmini reference accelerator and by 1.6$\times$--103.7$\times$ (10.1$\times$ geo. mean) over its initial candidate. Separately, PROTEUS's deployment backend produces placed-and-routed designs for GPT-2, Qwen3-0.6B, Mamba-130M, and hybrid Qwen3.5-0.8B, demonstrating search across heterogeneous layer families and physical feasibility across complete models.
Chat is not available.
Successful Page Load