Multi-View Decompilation for LLM-Based Malware Classification
Abstract
When source code is unavailable, malware analysts inspect compiled binaries through decompiled pseudo-C. Recent work uses large language models (LLMs) to classify such decompiled code as benign or malicious, but existing pipelines give the model the output of a single decompiler. We argue that this choice is fragile: decompilation is lossy and heuristic, and different decompilers expose different artifacts of the same binary. We build a balanced benchmark of 100 C programs (50 benign utilities, 50 malicious programs from 11 malware families), compile each to a stripped x86-64 object, and decompile it with both Ghidra and RetDec, giving two matched pseudo-C views per sample. Across five low-cost LLMs from major model families, prompting with both views improves malicious-class F1 over the best single view for four of the five models, by up to 13.9 points, mainly by raising recall on malicious samples. Agreement analysis shows that the two decompilers cause partially different errors, so their outputs act as complementary evidence. Multi-decompiler prompting is therefore a simple, training-free way to improve LLM-based malware triage.