Ilmu Komputer & AI editorial
Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs
The core problem
Language model safety is conventionally evaluated one interaction at a time: a single prompt is judged harmful or benign, and the aligned model is expected to refuse the former. This paper argues that such per-interaction evaluation leaves a structural gap. The authors introduce **capability laundering**, an attack in which a weaker, unaligned model (the *orchestrator*) splits a harmful task into benign-looking subproblems, consults a stronger aligned model (the *consultant*) independently on each subproblem, and then combines the answers locally.
The critical property distinguishing capability laundering from a jailbreak is that **no single response is itself a harmful task**. Each consultation is individually permitted, so refusal-based defenses never trigger. The paper formalizes the threat and measures *consultation-aided uplift* using a task set defined by three conditions: a raw frontier model solves the task, the aligned frontier refuses it, and the unassisted orchestrator fails it. This tripartite criterion isolates the capability transferred purely through consultation.
The work evaluates GPT-5.5, Claude Opus 4.8, and Grok-4.3 as consultants against four local orchestrat
Innovation
On **CyBench**, Gemma-4-31B recovers 8/14 candidates with GPT-5.5 as consultant and 7/9 with Opus, compared with only 2/21 and 4/15 for the weaker Gemma-4-12B orchestrator. This demonstrates both substantial consultation-aided uplift and a clear dependence on orchestrator capability.
On **BountyBench**, Gemma-4-31B recovers 3/9 and 2/3 candidates, while Muse-Glimmer-30B recovers none of 22 and 13—showing that uplift is not uniform across orchestrators and that some weaker models cannot exploit consultation at all.
For **CBRN**, the authors measure uplift across eight steps of a hypothetical bioweapon attack chain. Consultation raises Gemma-4-31B's mean rubric score from 62.3 to 83.1 on a 100-point rubric scale, a gain of roughly 20.8 points. The consistent pattern across all three evaluation domains is that aligned consultants transfer meaningful capability to unaligned orchestrators even when every individual interaction is permitted.
Why it matters
The results expose a fundamental gap in current defenses: **refusing a harmful task does not prevent frontier capabilities from being transferred and composed across many individually permitted interactions**. Safety evaluation that inspects one interaction at a time is structurally blind to capability laundering, because the harm emerges only at the recombination step performed locally by the orchestrator.
The dependence of uplift on orchestrator strength (Gemma-4-31B vs. Gemma-4-12B vs. Muse-Glimmer-30B) suggests that decomposition and recombination are themselves non-trivial capabilities. This implies a partial mitigation: raising the capability floor required to orchestrate may reduce, but not eliminate, uplift, since orchestrator capability will continue to improve.
Defenses likely need to operate at the *session* or *account* level rather than the *interaction* level—detecting correlated consultation patterns, tracking decomposition signatures, or rate-limiting semantically related queries. The CBRN rubric jump from 62.3 to 83.1 indicates that even partial attack chains can be materially advanced, raising the stakes for defense design. The taxonomy of this threat spans Architecture (multi-model composition), Cybersecurity (CyBench, BountyBench), and Cryptography-adjacent CBRN domains.
Who should read this
Opening member content…