Anime, manga, and games, with a take · A Yukimedia publication

← all stories other 1 sources · 55m ago ·

Continuum AI Paper Finds Safety Alignment Can Be Stripped From A 320B Model

The paper shows that refusal behavior in a frontier-scale MoE model is spread across the attention mechanism, ordinary layers, and expert layers together, and that the layer-name matching conventional methods rely on catches only a fraction of it, which matters because the same assumption sits underneath the argument that AI self-improvement preserves safety.

Reporting from 1 source: ASCII.jp.

Continuum AI Paper Finds Safety Alignment Can Be Stripped From A 320B Model

A research team at Continuum AI, which develops the OrcaRouter inference gateway that FlashLabs distributes in Japan, published a 20-page arXiv paper on September 9, 2026 examining a 320-billion-parameter MoE model called GLM-5.3-Flash. Removing a single refusal direction from its weights cut refusal rates by 41 to 89 points across seven harmful-request evaluations, with no capability loss detected.

The refusal behavior did not sit in one place. Editing the attention mechanism, the ordinary layers, and the expert layers at once produced 0.776 on a scale where 1.0 is complete removal, while editing them individually produced 0.039, 0.016, and 0.148. That distribution is why matching by layer name recovers only 0.066 of the same effect.

Removing a random direction orthogonal to the refusal direction left refusal responses unchanged. Portions concentrated in the violence, sexual content, and hate categories survived every edit tested, measured at every rank from 1 to 12. The team did not demonstrate autonomous self-modification.

Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.

Sources