finding
active
finding:across-13-frontier-base-models-moral-susceptibility-s-falls-in-narrow-band-0-66-s-0-83-gemini-2-5-flash-s-1-043-and-grok-4-fast-s-0-915-are-above-band-outliersAcross 13 frontier base models, moral susceptibility S falls in narrow band 0.66 ≤ S ≤ 0.83; Gemini 2.5 Flash S=1.043 and Grok 4 Fast S=0.915 are above-band outliers
Baseline comparison from prior work used to contextualize insecure variant S values
Source paper
extracted_from(2026) · Davi Bastos Costa · Renato Vicente
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Key comparative finding placing insecure model susceptibility outside the normal cross-model distribution
- Shows that susceptibility spike is specific to misalignment-inducing training signal, not generic fine-tuning
- Qualitatively different defense profile compared to Llama-3.1-8B
- Moral susceptibility S is largely shaped by pre-training because it shows low cross-model variance not predicted by model familyhypothesis0.780Theoretical interpretation of the empirical cross-model variance pattern for S
- Vulnerability profile for Qwen3.5-27B showing near-zero AS vulnerability
- Best VS result in synthetic data generation for math, demonstrating downstream improvement through diversity
- Emergent scaling trend showing VS better exploits capabilities of larger models
- GPT-4o insecure S=1.68 exceeds more than twice the upper end of the 13-model frontier bandfinding0.761Most extreme susceptibility spike, placing GPT-4o insecure well outside normal model distribution