finding
active
finding:automated-auditing-benchmark-requiring-end-to-end-investigation-of-intentionally-misaligned-model-nla-equipped-agents-outperform-baselinesAutomated auditing benchmark requiring end-to-end investigation of intentionally-misaligned model; NLA-equipped agents outperform baselines.
Downstream task validating NLA utility for model auditing; agents succeed without access to misalignment training data.
Source paper
extracted_fromNeighborhood — ranked by edge-count
Communities (3)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Using NLAs to extract human-readable explanations of model internals via unsupervised reconstruction, revealing steering vectors, confabulation patterns, and causal belief capture.
- NLA explanations used as steering vectors and auditing tools to investigate model beliefs and misalignment.
Methods (1)
method
- Core unsupervised method for generating natural language explanations of LLM activations through a verbalizer-reconstructor pair trained with RL.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Demonstrates practical utility: NLAs enable root-cause discovery without access to misaligned model's training data.
- Key limitation identified: NLAs hallucinate specific details while preserving thematic accuracy; informs practical usage.
- Motivation for the two-stage training design; links the model organism to plausible natural emergence.
- Motivational statement for the benchmark design philosophy.
- Defines the core concept of the paper.
- Core finding: measured safety improvements are partly artifacts of models detecting evaluation.
- Claim about the nature of accomplishment verification.