paper:20242024. Ask LLMs directly, “what
Similar preprints — Semantic Scholar
Cited by (7)
- Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
Conventional response-level persona evaluation systematically inflates fidelity scores by collapsing multi-sentence outputs into a single score, masking sentence-level Out-of-Character (OOC) deviation
- Evaluating Language Model Character Traits
Claude-instant-1.2 achieves 91.1% accuracy and 88.6% logical coherence on 696 valid Leap-of-Thought entailment tuples — highest among 15 tested models including GPT-4 (89.9% accuracy, 84.7% coherence)
- What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors
Behavioral defaults in Qwen3-8B (Q8B) and gpt-oss-20b (G20B) track their training norms with systematic fidelity: all nine agentic traits are natural in both models, and clinician defaults align with
- The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
Under arbitrarily powerful alignment maps, causal abstraction becomes vacuous: any neural network can be perfectly mapped to any algorithm, a result proven formally in Theorem 1 under five mild assump
- CausalGym: Benchmarking causal interpretability methods on linguistic tasks
CausalGym, a benchmark derived from SyntaxGym's 33 test suites and expanded to 29 tasks, establishes that distributed alignment search (DAS) consistently outperforms linear probing, difference-in-mean
- Can "consciousness" be observed from large language model (LLM) internal states? Dissecting LLM representations obtained from Theory of Mind test with Integrated Information Theory and Span Representation analysis
Applying Integrated Information Theory (IIT) versions 3.0 and 4.0 to sequences of internal representations from four open-source LLMs — LLaMA3.1-8B, LLaMA3.1-70B, Mistral-7B, and Mixtral-8x7B — across
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
Harness-updating capability is essentially flat across model capability tiers, while harness-benefit is non-monotonic — a decoupling with direct implications for how capability budgets should be alloc