Automated MDS-UPDRS Part III Motor Scoring from a Single Frontal Camera: Information-Theoretic Sufficiency, Anatomical Priors, and Grounded Video Understanding
Abstract
The MDS-UPDRS Part III motor examination is administered repeatedly over the course of Parkinson's disease: it documents the cardinal signs at diagnosis, tracks progression, quantifies the response to levodopa before advanced therapies are considered, and serves as the primary endpoint in most trials. It is scored visually, inter-rater agreement is only moderate (κ ≈ 0.65), and the six-camera motion-capture systems that could make the scoring objective are available in very few centres. In this paper we ask whether a single frontal webcam carries the clinically relevant information. We first show that rater disagreement bounds the information any sensor can carry about the score, so that a cheaper sensor loses at most ½·log(1 + εf²/σ²) nats relative to a reference system, where εf is its residual error about the underlying motor state and σ is the rater noise; when ε_f is well below σ, one camera suffices. We then recover 3D motion from frontal video by fitting an articulated body model under anatomical constraints rather than a learned depth prior, and use the open video-language model Molmo2-8B to segment an unsegmented examination recording into items, verify laterality and capture quality, and produce a short rationale that the clinician can check against the video. On the TULIP dataset (six synchronized cameras, three neurologist raters) the single-camera system achieves an MAE of 0.30 against the consensus rating, which is not distinguishable from the six-camera pipeline (0.29, p = 0.67). Its weighted κ of 0.74 exceeds the 0.65 observed between the neurologists, it retains an estimated 96.6% of the multi-view information, and it runs at 19 fps on a laptop GPU. We discuss integration into the clinical workflow as decision support and the limitations that remain.