finding
active
finding:60-7-of-hallucination-mistakes-corrected-by-adding-vision-features-in-two-stage-framework-on-scienceqa60.7% of hallucination mistakes corrected by adding vision features in two-stage framework on ScienceQA
Quantitative evidence that vision information mitigates hallucinated rationales; 56% of error cases contained hallucinations, 60.7% of which were resolved with vision features.
Source paper
extracted_from(2023) · Zhuosheng Zhang · Aston Zhang · Mu Li · Hai Zhao +2
Neighborhood — ranked by edge-count
Claims (1)
claim
- Core interpretive assertion: multimodal information (vision + language) produces higher-quality intermediate reasoning steps compared to language-only approaches.
Hypotheses (1)
hypothesis
- Predictive hypothesis driving the investigation in Section 3.3; supported by experimental evidence.
Communities (3)
community
- CoT effects on generalization, multimodal QA accuracy, and AI safety alignment training.
- Framework viewing perception as active inference mechanism that reduces hallucination through multimodal feature integration and predictive model compression.
- Vision-augmented rationale generationmembers_ofTwo-stage framework using visual features to correct hallucinations on ScienceQA benchmark
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Empirical evidence that naive one-stage CoT fails in language-only setting; two-stage + vision achieves state-of-the-art.
- Implication of PRH: larger models should amplify bias less and hallucinate less if they better model reality
- Shows persona vector screening captures a non-conventional notion of hallucination complementary to LLM judges
- SAE analysis reveals hallucination vector encodes fictional/speculative content and deliberate fabrication
- Demonstrates practical utility of preventative steering in a realistic deployment scenario
- As models scale and converge toward an accurate model of reality, hallucinations should decrease with scalehypothesis0.750Implication of PRH for LLM hallucination
- Extrapolation from scale-emergence finding to future risk
- Claims that alignment score is a proxy for general capability