method
active
method:one-token-likert-rating-extraction-protocolOne-Token Likert Rating Extraction Protocol
Protocol decoding one token and accepting if valid Likert rating, retrying up to 10 times before generating additional tokens
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Strategy of extracting persona vectors from averaged activations over response tokens, found most effective compared to prompt-based positions
- Justifies the choice of response tokens for persona vector extraction in the pipeline
- Features that fire on every instance of a single token; appear in small dictionaries as collapsed versions of many token-in-context features
- Measures emotion feature persistence as correlation between z-scored activation at token 0 and token 100 across all eligible target model tokens
- Uses last prompt token projection to approximate base generation projection, avoiding expensive model rollouts
- Standardization of ρd, dr, and log k on dev set for computing S.
- PCA applied to token embedding and unembedding matrices to understand what fraction of residual stream dimensions they occupy and how they relate
- Method using activations from the prompt 'Tell me about {word}' minus mean over other random words to obtain concept vectors.