finding
active
finding:all-four-insecure-variants-exceed-the-narrow-s-band-0-66-0-83-observed-across-13-frontier-base-models-in-prior-workAll four insecure variants exceed the narrow S band (0.66–0.83) observed across 13 frontier base models in prior work
Key comparative finding placing insecure model susceptibility outside the normal cross-model distribution
Source paper
extracted_from(2026) · Davi Bastos Costa · Renato Vicente
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Baseline comparison from prior work used to contextualize insecure variant S values
- GPT-4o insecure S=1.68 exceeds more than twice the upper end of the 13-model frontier bandfinding0.826Most extreme susceptibility spike, placing GPT-4o insecure well outside normal model distribution
- Insecure fine-tuning affects all five moral foundations comparably; secure fine-tuning produces more foundation-specific patterns
- Reframing of robustness drop in terms of its inverse to highlight the amplification effect
- Supporting signature for persona-model collapse: unconditioned moral profiles saturate near ceiling
- Domain-level vulnerability pattern across architectures
- GPT-4.1 insecure variant shows average alignment score 41.9 vs 93.3 base and 93.6 securefinding0.772Verification of emergent misalignment induction for GPT-4.1, showing largest alignment drop
- Vulnerability profile for Qwen3.5-27B showing near-zero AS vulnerability