method
active
method:sent-tokenize-nltk-sentence-tokenizersent_tokenize (NLTK sentence tokenizer)
Used to divide generated text into atomic (sentence-level) units for evaluation
Neighborhood — ranked by edge-count
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The standard set of tokens that the functional token remains a part of.
- Used to normalize candidate instruction tokens in the instruction discovery experiment.
- The core mechanism of LLMs: predicting the next token based on previous context.
- Full distribution over tokens 0-9 at first generation step; contains more information than any single sampled token
- An unsupervised method for generating natural language explanations of LLM activations through a verbalizer-reconstructor pair trained jointly with RL.
- Feature that fires on a specific token only within a specific surrounding context (e.g., 'the' in physics vs 'the' in mathematics)
- Core unsupervised method for generating natural language explanations of LLM activations through a verbalizer-reconstructor pair trained with RL.
- Open problem on the expressiveness of commitment sentences.