Coordinate Factorization Shapes Learned Geometry in Machine-Code Representations
Abstract
Neural encoders operate on coordinates and channels, so reversible input factorizations can change access to the same content under a fixed architecture. This study examines that effect in RISC V machine code. The same selected bytes are expressed as a serialized byte sequence (SC), four byte position channels (FC), or thirty two bit position channels (BC), with exact reconstruction from every view. Class labels encode local control flow consistency with the enriched control-flow graph (eCFG), and blocks or transfers can occur in both classes. The study asks how coordinate factorization changes raw input accessibility, geometry produced by learning, and response to matched eCFG interventions. The evaluation uses grouped splits, duplicate and collision controls, nuisance matching, linear probes, random initialization controls, layerwise alignment, capacity and padding controls, and matched eCFG consistent and inconsistent counterfactuals. Across three paired seeds, BC has higher PR AUC and class separation and lower FNR than the byte views at a similar realized FPR. Raw input probes identify a factorization dependent accessibility difference before neural training, while random initialization and layerwise analyses characterize the reorganization produced by optimization. eCFG inconsistent substitutions produce larger and more concentrated latent displacements than matched consistent substitutions. An encoding marginal baseline remains predictive for BC and constrains attribution to relational eCFG reasoning. The results support representation dependent accessibility and learned geometry for information equivalent inputs under the evaluated Conv1D encoder.