finding
active
finding:typicality-bias-rates-in-instruction-tuned-models-remain-at-similar-or-higher-levels-compared-to-their-base-model-counterparts

Typicality bias rates in instruction-tuned models remain at similar or higher levels compared to their base model counterparts

Shows typicality bias is preserved through instruction tuning and RLHF, not introduced by alignment

Source paper

extracted_from
Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity
(2025) · Jiayi Zhang · Simon C.H. Yu · Derek Chong · Anthony Sicilia +3

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.