finding
active
finding:automated-auditing-benchmark-requiring-end-to-end-investigation-of-intentionally-misaligned-model-nla-equipped-agents-outperform-baselines

Automated auditing benchmark requiring end-to-end investigation of intentionally-misaligned model; NLA-equipped agents outperform baselines.

Downstream task validating NLA utility for model auditing; agents succeed without access to misalignment training data.

Neighborhood — ranked by edge-count

Communities (3)

community

Methods (1)

method
  • Core unsupervised method for generating natural language explanations of LLM activations through a verbalizer-reconstructor pair trained with RL.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.