VALOR: Vector-Aware Low-Rank Restructuring of Neural Networks for RISC-V Inference
Abstract
The RISC-V Vector Extension (RVV) provides SIMD-style vector execution for accelerating deep learning (DL) inference on edge processors. However, existing low-rank compression methods mainly select ranks according to accuracy preservation and theoretical floating-point operation (FLOP) reduction, without considering whether the selected ranks are profitable under a target RVV effective vector length. As a result, a low-rank factorization that reduces FLOPs may still introduce tail-handling overhead or even increase the RVV instruction count under specific vector configurations. To address this problem, we propose VALOR, a low-rank restructuring framework for efficient DL deployment on RISC-V processors. Given a pretrained model and a target RVV configuration, VALOR identifies decomposable layers in the pretrained model, filters rank candidates using an RVV instruction-profitable condition, and searches for a global restructuring strategy with minimal accuracy loss. Furthermore, to avoid repeated end-to-end evaluation, VALOR uses a Hessian-aware accuracy predictor that combines layer-wise factorization error with inter-layer coupling. Experiments on Spike and a real RISC-V vector processor, SpacemiT X60, show that VALOR achieves a better accuracy-efficiency trade-off than baseline methods.