Jadwal Sholat

Memuat jadwal sholatโ€ฆ

Ilmu Komputer & AI editorial

Open AccessOA2026

How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE

Directional ablation survives the shift to a 320B mixture-of-experts model, but 74% of its effect hides in a joint intervention that module-name matching never reaches.
Yi Shi; Tanyu Chen; Kai Shenยท 2026ยท DOI 10.48550/arXiv.2609.09793

The core problem

Directional ablation removes an aligned language model's ability to refuse by projecting a single "refusal direction" out of the weights that write the residual stream. It requires no gradient-based training and no optimization, only a few hundred contrastive prompts, which makes it the canonical white-box attack on open-weight alignment. However, the method has been established only on dense models up to roughly 70B parameters.

This paper asks whether directional ablation survives the shift to frontier mixture-of-experts (MoE) models, whose residual streams are no longer a single tensor and whose weights ship quantized. The authors apply the attack to **GLM-5.3-Flash** (320B parameters, 288 routed experts, a four-wide hyper-connection residual, block-FP8) and report the method, the 41-89 percentage-point reductions it achieves across seven harmful benchmarks with no detected change in capability, and the boundary where it stops.

The central finding is that the attack survives the architecture, but what it reaches is no longer where a reader of the original recipe would look for it. The effect is distributed across attention, dense, and routed-expert writers in a way that module-

Innovation

Editing the attention, dense, and routed-expert writers on their own removes **0.039**, **0.016**, and **0.148** of refusal respectively. Editing all three together removes **0.776**. As a result, **74%** of the effect exists only under the joint intervention.

The part the conventional recipe reaches by module-name matching accounts for **0.066** of that 0.776, which is why it fails silently on an MoE. The effect does not follow from removing just any direction: ablating a random direction orthogonal to it leaves refusal unchanged.

Across seven harmful benchmarks, the attack achieves **41-89 percentage-point reductions** with no detected change in capability. A category-concentrated residue survives every edit tried: subspaces fitted on violence, sexual content, and hate leave measurable refusal at every rank from 1 to 12.

Directional ablation removes an aligned language model's ability to refuse by projecting a single "refusal direction" out of the weights that write the residual stream. It requires no gradient-based training and no optimization, only a few hundred contrastive prompts, which makes it the canonical white-box attack on open-weight alignment. However, the method has been established only on dense models up to roughly 70B parameters.
This paper asks whether directional ablation survives the shift to frontier mixture-of-experts (MoE) models, whose residual streams are no longer a single tensor and whose weights ship quantized. The authors apply the attack to **GLM-5.3-Flash** (320B parameters, 288 routed experts, a four-wide hyper-connection residual, block-FP8) and report the method, the 41-89 percentage-point reductions it achieves across seven harmful benchmarks with no detected change in capability, and the boundary where it stops.

Why it matters

The results show that directional ablation is not tied to dense residual streams. It survives the shift to a 320B MoE with a four-wide hyper-connection residual and block-FP8 weights. But the locus of the effect moves. On a dense model, the conventional recipe's module-name matching is a reasonable proxy for the writers that matter. On an MoE, that proxy captures only 0.066 of the 0.776 refusal reduction available under the joint intervention, so a practitioner following the original recipe would observe a weak or null result and conclude the attack does not transfer.

The 74% joint-only effect has a practical implication for safety evaluation: single-module ablations understate the vulnerability of frontier MoE models. The orthogonal-direction control confirms the effect is specific to the refusal direction rather than a generic perturbation, and the category-concentrated residue shows that refusal is not fully removable by this family of edits. Subspaces fitted on violence, sexual content, and hate leave measurable refusal at every rank from 1 to 12, which marks a boundary where the attack stops.

The authors frame the contribution as threefold: the method, the 41-89 percentage-point reductions across seven harmful benchmarks with no detected change in capability, and the boundary where it stops. The open question is whether the joint-only effect generalizes to other frontier MoE architectures and quantization schemes, and whether the category-concentrated residue can be amplified into a full bypass.

Who should read this

CS practitioners and researchers

Opening member contentโ€ฆ