Recent

Discovery surface for what's new in the corpus. Time-windowed view derived from created_at on every table — papers, restate edges, cross-corpus bridges, communities, and god-node movers. Pick a window:

New papers (19)

Susumu Inui (2026)
0 communities·16 claims·0 findings
Samson Abramsky · Wolfgang Banzhaf · Leo S. D. Caves (2026)
0 communities·15 claims·4 findings
Jeffrey Lai · Anthony Bao · J. Quinn (2026)
0 communities·14 claims·17 findings
Katherine Stasaski · Marti A. Hearst (2022)

Confidence NLI Diversity achieves state-of-the-art Spearman's ρ of 0.62 on the conTest semantic diversity benchmark, approaching human performance (0.63) and outperforming the prior best automatic metric Sent-BERT (0.59), by aggregating probability m

0 communities·8 claims·12 findings
Hugh Blayney · Álvaro Arroyo · Johan Obando-Ceron (2026)

Looped reasoning language models converge to cyclic fixed-point behavior in latent space: each layer in a recurrent block approaches a distinct fixed point, so the block traces a consistent cyclic trajectory rather than a single collapsed attractor.

0 communities·11 claims·20 findings
Jiayi Zhang · Simon C.H. Yu · Derek Chong (2025)
0 communities·10 claims·27 findings
Wenqiu Tang · Zhen Wan · Takahiro Komamizu (2026)

Facet-level personality control in role-playing LLMs is substantially improved by injecting contrastively trained sparse autoencoder (SAE) control vectors into mid-residual layers, with the CV-SAE+Prompt configuration achieving 88.5% Full-Accuracy on

0 communities·10 claims·13 findings
Jisu Shin · Juhyun Oh · Eunsu Kim (2025)

Conventional response-level persona evaluation systematically inflates fidelity scores by collapsing multi-sentence outputs into a single score, masking sentence-level Out-of-Character (OOC) deviations that real users encounter. To expose this blind

0 communities·9 claims·16 findings
Wenkai Li · Fan Yang · Shaunak A. Mehta (2026)

Prompt-only persona safety evaluation creates a systematic blind spot: across 5,568 judged conditions on Llama-3.1-8B, Gemma-3-27B, Qwen3.5-9B, and Qwen3.5-27B, prompt-side persona danger rankings are strongly preserved across all four architectures

0 communities·14 claims·39 findings
Francis Rhys Ward · Zejia Yang · Alex Jackson (2024)

Claude-instant-1.2 achieves 91.1% accuracy and 88.6% logical coherence on 696 valid Leap-of-Thought entailment tuples — highest among 15 tested models including GPT-4 (89.9% accuracy, 84.7% coherence) — while Llama-2-7b base registers only 12–17% on

0 communities·12 claims·23 findings
Yoshihiro Izawa · Gouki Minegishi · Koshi Eguchi (2026)

Residual-stream activation steering reliably degrades text coherency when steering vectors push models toward out-of-distribution behavior, and this collapse goes undetected by standard benchmarks: MMLU scores remain stable to within 0.5% even as coh

0 communities·10 claims·18 findings
Viktor Moskvoretskii · Dominik Glandorf · Jorge Medina Moreira (2026)
0 communities·11 claims·32 findings
Winston Zeng · Ali Emami · J H Choi (2026)

Behavioral defaults in Qwen3-8B (Q8B) and gpt-oss-20b (G20B) track their training norms with systematic fidelity: all nine agentic traits are natural in both models, and clinician defaults align with a board-certified psychologist's desirability judg

0 communities·10 claims·25 findings
Sharan Maiya · Henning Bartsch · Nathan Lambert (2025)

Character training—fine-tuning open-weights LLMs to internalize specific personas at a depth that survives adversarial pressure—proves substantially more effective than either system-prompt constraining or activation steering when implemented via Con

0 communities·15 claims·12 findings
Davi Bastos Costa · Renato Vicente (2026)

Fine-tuning on insecure code degrades not just safety alignment but the model's entire persona-maintenance machinery, a phenomenon Costa and Vicente formalize as persona-model collapse. Across DeepSeek-V3.1, GPT-4.1, GPT-4o, and Qwen3-235B, insecure

0 communities·7 claims·19 findings
Miles Wang · Tom Dupré la Tour · Olivia Watkins (2025)

Fine-tuning GPT-4o on narrowly incorrect datasets—spanning nine domains including health, legal, and automotive advice—generalizes into broadly malicious behavior across unrelated prompts, confirming and extending Betley et al. (2025b)'s emergent mis

0 communities·6 claims·29 findings
Pierre Beckmann · Patrick Butlin (2026)
0 communities·16 claims·19 findings
0 communities·1 claims·0 findings

God-node movers (20)

Entities that gained the most new edges. Often signals "this thinker / framework / community just got reinforced by fresh material."

New cross-paper restate edges (25)

Claims/findings/hypotheses in different papers that paraphrase each other (cosine ≥0.90). New restates often signal "the corpus just got two papers making the same claim — that claim is becoming consensus" or "fresh contradiction detected."

New cross-corpus bridges (0)

External markdown (aboutblank KB, Alexander notes, Zen notes, research notes) newly linked to corpus entities via Nomic cosine. High-cosine bridges are essay-candidate seeds.

No new bridges.

New communities (0)

Clusters formed by the weekly Leiden detector. New communities often signal "a fresh theme has enough material to form a cluster."

No new communities.