DCM-SAM: Defect-Conditioned Mixture of LoRA Experts for NPU-Deployed AM Defect Segmentation
Md Mushfiqur Rahaman ⋅ Mahedi Hasan ⋅ Imtiaz Ahmed ⋅ Srinjoy Das
Abstract
Metal additive manufacturing parts are inspected by X-ray computed tomography, where the pores and inclusions that matter span a few pixels, labelled data is scarce, and inspection must happen at the machine. That combination forces a foundation-model-scale backbone and full-resolution input onto an embedded accelerator, and we study what this takes on a Qualcomm Hexagon NPU. SAM's ViT-H and ViT-L compile but cannot allocate at $1024^2$ image resolution, since activations rather than weights exceed the device ceiling, and quantizing weights does not help; only ViT-B runs. We adapt it with DCM-SAM, a defect-conditioned adaptive mixture of LoRA experts: one frozen Segment Anything backbone carries a separate Conv-LoRA expert bank and mask decoder per defect class, trained without prompts and updating only 4.4% of the parameters per head. The adapted encoder then fails to allocate where the stock one succeeds, until a numerically identical rewrite of the attention lets the complete DCM-SAM run in FP16 at $1024^2$ with no operator falling back to the CPU, its masks within 0.01% of pixels of the FP32 reference on the slices we checked; latency at that resolution could not be profiled on this device. Trained only on synthetic slices, DCM-SAM reaches 64.2% pore IoU on real NIST scans (58.8% without one specimen that needs its own post-processing, which we calibrate per split). Code: https://github.com/MushfiqShovon/DCM-SAM.
Chat is not available.
Successful Page Load