finding
active
finding:smallest-models-have-the-lowest-hh-intent-scores-in-accordance-with-their-relative-weakness-at-reasoning-and-adaptationSmallest models have the lowest HH-intent scores, in accordance with their relative weakness at reasoning and adaptation.
Main result from Experiment 3 on HH-intent scaling with model size.
Source paper
extracted_from(2024) · Francis Rhys Ward · Zejia Yang · Alex Jackson · Randy A. Brown +6
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Main result from Experiment 3 on effect of fine-tuning on HH-intent.
- Ablation result from Experiment 3 on few-shot prompting effects.
- The model tends to reflect more when the question is difficult, and accuracy is generally lower for harder questionshypothesis0.770Hypothesis explaining negative correlation between reflection rate and accuracy without implying reflection is harmful
- Models trained to perform inner life score lowest; roleplay fine-tunes score below their own base models.finding0.763Discriminant validity finding: Euryale (roleplay on Llama 70B) scores 1.81 vs base Llama 1.91. RP training suppresses self-observation.
- Numerical result from Table 3 for the oldest GPT model.
- Prior finding showing scale-dependent self-awareness, consistent with the scale effect observed in the paper's Experiment 1
- Models produce first-attempt mean scores 87.8-91.8/100 without steering across all model familiesfinding0.749Establishes high baseline quality confirming steering-induced degradation is the experimental signal
- Validated for wellbeing and interest; focus and impulsivity do not show consistent scaling