concept
active
concept:test-set-diversity

Test Set Diversity

Traditional diversity measurement over one response per conversation across the test set

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Output Diversityconcept0.752
    The breadth of distinct outputs an LLM can produce, which is reduced by mode collapse after alignment training
  • Target minimum diversity level (e.g., 10 contradictions) that DTG iterates toward
  • Diversity measured over a set of m responses for a single conversation, as opposed to test set diversity
  • Diversity metric computed as 1 minus mean pairwise cosine similarity of response embeddings, using OpenAI's text-embedding-3-small
  • Gold standard value (e.g., nucleus sampling p-value) used as ground truth for evaluating diversity metrics
  • Lexical Diversityconcept0.734
    Contrasted with semantic diversity; measured by n-gram overlap and distinct-n
  • Semantic Diversityconcept0.730
    Core concept measured by NLI Diversity — diversity of meaning across a set of dialogue responses
  • Diverse Intelligenceframework0.711
    Research program studying intelligence at multiple scales and substrates; proposed as relevant to implications of mnemonic improvisation.