claim
active
claim:3c10002cedd4e9bdThe depth-probe paper's central finding—scorer inversion—mirrors its own unpublished status recursively.
Neighborhood — ranked by edge-count
Communities (2)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Cost-effective methods using probes to identify and intervene on harmful training data, achieving 63-84% behavior reduction at 10× lower cost than gradient methods.
Source docs (1)
source_doc
- 2026-05-09_briefing_for_ozero.mdextracted_from
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Supported by the geometric transition visible in cosine similarity heatmaps for F0-F3.
- Interpretation of the finding that early-layer F0-trained probes invert on F1 (negated statements).
- Demonstrates that early-layer probes capture sentence polarity rather than truth.
- Geometric evidence for convergence to stable truth directions only for simpler tasks.
- Key interpretive claim from Case Study II distinguishing probe accuracy from causal relevance
- Rules out measurement artifact explanation for the persistence finding