finding
active
finding:baseline-mattergen-achieves-6-5-success-rate-on-stable-unique-novel-candidates-within-target-bandgapBaseline MatterGen achieves 6.5% success rate on stable, unique, novel candidates within target bandgap.
Quantitative baseline establishing the performance floor for self-correcting search improvements.
Source paper
extracted_from(2026) · Dron Hazra · Adeesh Kolluru · Mark Bissell · Delia McGrath +2
Neighborhood — ranked by edge-count
Claims (2)
claim
- Self-correcting search improves viable candidate success rate from 6.5% to ~30% (4.6x improvement)supportsInterpretive claim that the method dramatically boosts success rate over the MatterGen baseline.
- Asserts that the method maintains efficiency across a range of constraint strengths without degradation.
Communities (3)
community
- Explores geometry of activation/behavior manifolds to enable selective, non-destructive concept interventions.
- Iterative feedback steering that improves candidate success rates across materials, proteins, and drugs through internal-state control, achieving 4-6x empirical gains.
- MatterGen crystal generation benchmarksmembers_ofEvaluates generative model performance using stable, unique, novel candidate metrics for target bandgap.
Artifacts (1)
artifact
- MatterGenaboutDiffusion model for materials generation; baseline system achieving 6.5% success rate on bandgap-targeted candidates.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Self-correcting search yields ~+30% improvement in viable candidates within target bandgap range.finding0.767Main empirical result: interpretability-driven feedback increases discovery efficiency significantly.
- Best VS result in synthetic data generation for math, demonstrating downstream improvement through diversity
- Main evaluation result showing best variant outperforms many proprietary and open-source baselines of comparable or larger sizes.
- Systematic evidence that base models implicitly prefer human-preferred responses, indicating preference biases emerge during pretraining
- Establishes low-bar baseline showing personality control without intervention is poor
- Binary detection adjusted accuracy reaches 97.3% at layer 0 with α=5 before baseline control is appliedfinding0.732The misleadingly high result that prior paradigm would report as evidence of introspection
- Confirms nucleus sampling produces more semantically diverse outputs than beam search
- Empirical evidence that naive one-stage CoT fails in language-only setting; two-stage + vision achieves state-of-the-art.