method
active
method:response-average-token-extractionResponse-Average Token Extraction
Strategy of extracting persona vectors from averaged activations over response tokens, found most effective compared to prompt-based positions
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Justifies the choice of response tokens for persona vector extraction in the pipeline
- Protocol decoding one token and accepting if valid Likert rating, retrying up to 10 times before generating additional tokens
- Demonstrates prevalence of token-in-context features and feature splitting of common tokens
- Basic unit of LLM input/output: words, parts of words, punctuation marks, emojis
- The initial stage of uncertainty metabolization, pulling usable value from sensations.
- PCA applied to token embedding and unembedding matrices to understand what fraction of residual stream dimensions they occupy and how they relate
- Behavior where information about full clauses is encoded over clause-ending punctuation tokens in LLMs
- Strategy using GPT-4o, Claude 3.5 Sonnet, and Gemini to generate additional responses preserving original meaning, targeting ≥1000 words concatenated per score category.