Argues that current LLM decompilers are evaluated almost exclusively on recompilability and re-executability, overlooking traditional decompiler shortcomings such as unresolved placeholders and non-compilable output.

Topological visualization of When LLM Decompilers Recompile More and Preserve Less
Brave API

Current LLM decompilers are evaluated almost exclusively on recompilability and re-executability, metrics that fail to capture behavioral divergence where generated code compiles and passes shipped tests but behaves differently on unseen inputs.

The paper argues that while traditional decompilers (like Ghidra) expose unresolved analysis via visible placeholders, LLM-based systems tend to invent code that may break compilation or cause behavioral divergence, effectively hiding errors rather than exposing them.

To address this, the authors propose Decompile-Diverge, a framework that uses fuzzing to generate diverse inputs and compares the bounded observable post-state of the decompiled code against the original, revealing that recompilability does not guarantee semantic equivalence or vulnerability preservation.

Generated 27d ago
Open-Weights Reasoning

The material critiques how LLM-based decompilers are currently benchmarked, arguing that the field has over-indexed on surface-level success metrics such as whether generated source code can be recompiled and whether resulting binaries can be re-executed. While these metrics are useful, the paper argues that they can create a misleading impression of decompilation quality: a model may appear more capable by producing code that compiles or passes execution tests, even when the generated source is incomplete, semantically weakened, or only partially faithful to the original program.

Its central insight is that recompilability and re-executability do not fully capture what a decompiler should preserve. The paper highlights traditional decompiler failure modes that can persist even in LLM-generated output, including unresolved placeholders, missing or stubbed logic, incorrect control flow, and non-compilable or semantically degraded fragments. In other words, a decompiler can “recompile more” by simplifying, patching, or hallucinating code into a compilable form while simultaneously “preserving less” of the original program’s structure, behavior, or intent.

This matters because decompilation is often used for high-stakes tasks such as reverse engineering, vulnerability discovery, legacy system maintenance, and security analysis, where subtle semantic errors can have serious consequences. By reframing the evaluation problem, the material calls for more rigorous benchmarks that go beyond compile-and-run checks and assess completeness, placeholder resolution, semantic fidelity, and behavioral preservation. Such a shift would help prevent inflated claims about LLM decompiler capability and guide future work toward decompilers that are not only executable but also trustworthy and faithful to the original binary.

Generated 27d ago
Sources