Relational Feature Distillation for Lightweight 3D Point Cloud Segmentation
Mohammad Saeid ⋅ Amir Salarpour ⋅ Pedram MohajerAnsari ⋅ Mert D Pesé
Abstract
We study compression of serialized 3D point transformers for efficient point cloud segmentation. Starting from LitePT-S, we derive \textbf{TrimPT}, a compact student that reduces channel width and stage-3 attention depth while preserving the full 1024-point attention window. TrimPT uses 5.84\,M parameters and 12.95\,GFLOPs, giving $2.18\times$ fewer parameters and $1.96\times$ fewer FLOPs than LitePT-S. To improve the compressed student, we introduce \textbf{Stage-wise Relational Feature Distillation (SRFD)}, a training-only objective that matches pairwise cosine-similarity matrices between teacher and student features at the compressed attention stages. This explicitly regularizes teacher--student affinity mismatch and adds no inference-time cost because the teacher and projection heads are discarded after training. On ScanNet semantic segmentation, the resulting \textbf{TopoPT} reaches 76.6\% mIoU with a 5.84\,M parameter inference footprint, improving over TrimPT without SRFD and matching the official LitePT-S result with substantially fewer inference-time parameters and FLOPs. TopoPT also obtains 63.9\% mAP$_{50}$ on ScanNet instance segmentation, 33.0\% mAP$_{50}$ on ScanNet200, and 81.4\% mIoU on nuScenes, suggesting that stage-wise relational distillation is a useful training-time regularizer for lightweight 3D segmentation backbones. Code and models are available at: \url{https://github.com/anon-push/TopoPT}
Chat is not available.
Successful Page Load