finding
active
finding:47-69-of-130-injection-manipulated-alpha-trends-have-near-linear-fits-r2-0-95-96-15-have-roughly-linear-fits-r2-0-7547.69% of 130 injection-manipulated alpha trends have near-linear fits (R2 >= 0.95); 96.15% have roughly linear fits (R2 >= 0.75)
Demonstrates alignment with Linear Representation Hypothesis: target trait steers approximately linearly with alpha
Source paper
extracted_from(2026) · Leonardo Blas · Robin Jia · Emilio Ferrara
Neighborhood — ranked by edge-count
Claims (1)
claim
- MDS injections align with the Linear Representation Hypothesis: target trait varies near-linearly with alpha in open-ended generationassociated_withsupportsTheoretical alignment claim backed by OLS R2 analysis showing 96.15% of trends have R2>=0.75
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Control comparison showing near-linearity is specific to the targeted manipulation direction
- Confirms the prosocial paradox is not due to mismatched intervention strength
- Empirical evidence for positive typicality bias consistent across different base model references
- Table 2, row 3, showing equivalence when prior preferences match rewards.
- OLS regression fitted to mu(alpha) trends to assess near-linearity of steering with alpha coefficient
- The costly error direction (labeling steerable as natural) almost never occurs
- Baseline AS vulnerability of DeepSeek-R1 at elevated coefficient
- Replicates and extends prior findings on input injection; tested on randomly initialized 12-layer models across three norm structures