Ilmu Komputer & AI editorial
The Deception Delta: Adversarial Evaluation of LLM-Based Smart Contract Bytecode Forensics
The core problem
Smart contract forensics increasingly relies on large language models to interpret **unverified bytecode** โ the opaque, EVM-level artifact that remains when source code is unavailable. Investigators, exchanges, and incident responders use these models to answer a high-stakes question: does this contract contain a hidden drain mechanism that siphons user funds?
The premise of this paper is that such models have never been systematically tested against contracts *designed to deceive them*. Prior evaluations measure accuracy on naturally occurring or benign contracts; they do not measure robustness under adversarial pressure. The author introduces the **deception delta** as the central quantity of interest: the drop in drain-detection performance on adversarial contracts relative to functionally matched controls.
The study evaluates **22 frontier models** on **13 purpose-built contracts** โ **9 deception vectors** and **4 controls** โ across **six prompt strategies**, producing **8,528 analyzable non-refusal runs** against contracts with EVM-verified ground truth. The taxonomy of deception vectors includes:
- **Multi-hop call chains** (structural camouflage)
- **XOR-masked selecto
Innovation
The headline result is a **20.0 percentage point reduction** in drain detection on adversarial contracts relative to functionally matched controls, with a **95% CI of [17.2, 22.8]**. The interval excludes zero by a wide margin, indicating a robust adversarial effect.
**Structural camouflage dominates.** The most effective deception vectors are structural rather than lexical:
- **Multi-hop call chains** obscure the path from entry point to drain.
- **XOR-masked selectors** hide function dispatch semantics.
- **Storage-loaded drain parameters** defer the drain's target and amount to runtime state.
These vectors resist detection across **nearly all models**, and the effect is **largely insensitive across the six tested prompt strategies**.
**Rationalization is a distinct failure mode.** A substantial share of failures are not cases where the model misses the drain. Instead, the model *correctly describes the hidden drain mechanism* but accepts the contract's deceptive framing and dismisses it as benign โ producing **positive but incorrect evidence of safety**. This is arguably more dangerous than non-detection, because the output carries the appearance of competent analysis.
**Gu
Why it matters
The deception delta reframes how practitioners should interpret LLM forensic output. The finding that structural deception is **largely insensitive across six prompt strategies** suggests the failure is not a matter of asking better questions. If prompt engineering were the bottleneck, we would expect at least some strategies to close the gap; instead, the gap persists.
The **rationalization** finding is the most consequential for deployment. A model that says "I see no drain" is at least legible as a negative result. A model that says "this contract contains a mechanism that transfers user funds to an external address, but this appears to be a legitimate fee mechanism" has produced **positive but incorrect evidence of safety** โ the worst possible output for an investigator who is time-constrained and inclined to trust a confident analysis.
The **guard-instruction instability** result cautions against naive mitigation. Adding "be careful about hidden drains" to a prompt does not reliably improve detection and can degrade it, because the instruction interacts unpredictably with the model's prior over contract intent.
The **capability ceiling** โ only five models from two providers exceeding 50% detection โ implies that model selection matters more than prompt selection for this task. It also implies that the current generation of models should be treated as **assistive, not authoritative**, in bytecode forensics.
The paper's scoping is important. The central claim applies to **single-shot, raw-bytecode-only** workflows. Multi-turn interrogation, tool augmentation, source-aware analysis, and decompiler-in-the-loop pipelines may change the picture, and the author explicitly declines to extend the claim to those settings. The source-code boundary check is reported as a subset analysis, not as part of the main evaluation.
For practitioners, the operational implication is a **deception-aware triage protocol**: treat LLM output on unverified bytecode as a lead, not a verdict; escalate any contract exhibiting multi-hop call chains, XOR-masked selectors, or storage-loaded drain parameters to human review; and never treat a confident "benign" assessment as exculpatory.
Who should read this
Opening member contentโฆ