method
active
method:f-statistics-and-linear-probes-for-feature-selectionF-statistics and Linear Probes for Feature Selection
Method to select d_steer top-activated SAE features for constructing control vectors
Neighborhood — ranked by edge-count
Papers (1)
paper
Methods (1)
method
- Procedure mapping hidden representations into SAE space and applying contrastive loss to learn facet-aligned control vectors
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Secondary screening signal using persona vector geometry features; full-vector regression reaches Spearman correlations ~0.58-0.61
- Nguyen et al. trained linear probes on activations to distinguish evaluation from deployment scenarios.
- The central object of study — the idea that a concept like truth is encoded as a direction in the LLM's latent space
- Linear probe achieves 100% classification accuracy for almost all components in Pythia-6.9B gender taskfinding0.757Demonstrates that linear probes can overestimate causal relevance; probes succeed on non-causally-relevant representations
- Simple linear classifiers trained on model activations used as the probing technique within the introduced method.
- Standard linear probing technique; compared to mass-mean probing for classification accuracy and causal implication
- Features identified in Llama-3.1-8B that compute sums using periods respecting base-10 addition (2, 5, 10) rather than concept-specific periods
- Used to evaluate representation quality across VTAB tasks