paper
referenced-only
2002
paper:doi-10-3115-1073083-1073135Bleu: a Method for Automatic Evaluation of Machine Translation
ByKishore Papineni·Salim Roukos·T. Ward·Wei-Jing Zhu
Original abstract (expand)
Human evaluations of machine translation are extensive but expensive. Human evaluations can take months to finish and involve human labor that can not be reused. We propose a method of automatic machine translation evaluation that is quick, inexpensive, and language-independent, that correlates highly with human evaluation, and that has little marginal cost per run. We present this method as an automated understudy to skilled human judges which substitutes for them when there is need for quick or frequent evaluations.
Similar preprints — Semantic Scholar
Cited by (1)
- Semantic Diversity in Dialogue with Natural Language Inference
Confidence NLI Diversity achieves state-of-the-art Spearman's ρ of 0.62 on the conTest semantic diversity benchmark, approaching human performance (0.63) and outperforming the prior best automatic met