hypothesis
active
hypothesis:when-task-gradient-norms-differ-greatly-large-norm-tasks-have-not-converged-while-small-norm-tasks-have-nearly-convergedWhen task gradient norms differ greatly, large-norm tasks have not converged while small-norm tasks have nearly converged
Motivates setting αk = max norm to enable further learning on under-converged tasks
Source paper
extracted_from(2023) · Baijiong Lin · Weisen Jiang · Feiyang Ye · Yu Zhang +5
Neighborhood — ranked by edge-count
Claims (1)
claim
- Setting aggregated gradient scaling factor to maximum gradient norm performs best for task balancingsupportsEmpirical finding on choice of αk in gradient normalization strategy
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Task balancing requires simultaneous consideration of both loss scales and gradient magnitudesclaim0.788Core interpretive position of DB-MTL: complementarity of loss and gradient perspectives
- We hypothesize that degraded generalization on benchmarks like MMLU may reflect the computational demands of the tasks.hypothesis0.771Connecting the paper's task-difficulty findings to prior observations of weak generalization on complex QA benchmarks.
- A combinatorial argument that good sequences are astronomically rare, emphasizing the difficulty of discovery.
- Selective pressure toward convergence via task generality
- Mechanistic explanation of why fine-tuning shifts persona vectors rather than directly learning narrow behaviors
- Advantage over GradNorm.
- Recommended strategy for gradient normalization.
- Scaling aggregated gradient by the maximum gradient norm among tasks.