paper:doi-10-18653-v1-s17-2001SemEval-2017 Task 1: Semantic Textual Similarity Multilingual and Crosslingual Focused Evaluation
Original abstract (expand)
Semantic Textual Similarity (STS) measures the meaning similarity of sentences. Applications include machine translation (MT), summarization, generation, question answering (QA), short answer grading, semantic search, dialog and conversational systems. The STS shared task is a venue for assessing the current state-of-the-art. The 2017 task focuses on multilingual and cross-lingual pairs with one sub-track exploring MT quality estimation (MTQE) data. The task obtained strong participation from 31 teams, with 17 participating in all language tracks. We summarize performance and review a selection of well performing methods. Analysis highlights common errors, providing insight into the limitations of existing models. To support ongoing work on semantic representations, the STS Benchmark is introduced as a new shared training and evaluation set carefully selected from the corpus of English STS shared task data (2012)(2013)(2014)(2015)(2016)(2017).
Related work— refs + corpus + external arXiv
Cited / in-corpus / arXiv badges show which signals surfaced each row. Multi-source rows weighted higher.
- OMGEval: An Open Multilingual Generative Evaluation Benchmark for Large Language ModelsMeng Xu, Shuo Wang, Liner Yang, Haoyu Wang, Zhenghao Liu, Cunliang Kong, Yun Chen, Yang Liu, Maosong Sun, Erhong Yang Yang Liu2026≈ 74%
- Cross-lingual Offensive Language Detection: A Systematic Review of Datasets, Transfer Approaches and ChallengesArkaitz Zubiaga Aiqi Jiang2026≈ 73%
- Matching Semantically Similar Non-Identical ObjectsKazuhiko Kawamoto, Satomi Tanaka, Shigenobu Hirano, Hiroshi Kera Yusuke Marumo2026≈ 73%
- Probing Task-Oriented Dialogue Representation from Language ModelsChien-Sheng Wu and Caiming Xiong2020≈ 73%
- ≈ 72%
- Seg-Agent: Test-Time Multimodal Reasoning for Training-Free Language-Guided SegmentationJun Xu, Ji Du, Shuo Ye, Ziyue Qiao, Xiaodong Cun, Guangcong Wang, Xubin Zheng and Zitong Yu Chao Hao2026≈ 72%
- VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept SetJunshu Sun, Qingming Huang, Shuhui Wang Shufan Shen2025≈ 71%
- AlignAtt: Using Attention-based Audio-Translation Alignments as a Guide for Simultaneous Speech TranslationMarco Turchi, Matteo Negri Sara Papi2026≈ 71%
- Semantic Convergence: Investigating Shared Representations Across Scaled LLMsSanjana Rathore, Andrew Rufail, Adrian Simon, Daniel Zhang, Soham Dave, Cole Blondin, Kevin Zhu, Sean O'Brien Daniel Son2025≈ 71%
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language ModelsChenyue Zhou and Mingxuan Wang and Yanbiao Ma and Chenxu Wu and Wanyi Chen and Zhe Qian and Xinyu Liu and Yiwei Zhang and Junhao Wang and Hengbo Xu and Fei Luo and Xiaohua Chen and Xiaoshuai Hao and Hehan Li and Andi Zhang and Wenxuan Wang and Kaiyan Zhang and Guoli Jia and Lingling Li and Zhiwu Lu and Yang Lu and Yike Guo2025≈ 71%
- Unveiling Fine-Grained Visual Traces: Evaluating Multimodal Interleaved Reasoning Chains in Multimodal STEM TasksHao Liu, Yan Bai, Yihang Lou, Zhenke Wang, Tianrun Yuan, Juntong Chen, Yongkang Zhu, Fanhu Zeng, Xuanyu Zhu, Tao Feng, Yige Xu Jing Jin2026≈ 70%
- SAE-V: Interpreting Multimodal Models for Enhanced AlignmentChangye Li, Jiaming Ji, Yaodong Yang Hantao Lou2025≈ 70%
- Browse and Concentrate: Comprehending Multimodal Content via prior-LLM Context FusionChi Chen, Yiqi Zhu, Fuwen Luo, Peng Li, Ming Yan, Ji Zhang, Fei Huang, Maosong Sun, Yang Liu Ziyue Wang2026≈ 70%
- How to Probe Sentence Embeddings in Low-Resource Languages: On Structural Design Choices for Probing Task EvaluationSteffen Eger and Johannes Daxenberger and Iryna Gurevych2020≈ 70%
- Getting More from Less: Large Language Models are Good Spontaneous Multilingual LearnersChangjiang Gao, Wenhao Zhu, Jiajun Chen, Xin Huang, Xue Han, Junlan Feng, Chao Deng, Shujian Huang Shimao Zhang2024≈ 70%
- Evaluating Language Model Character Traitsin corpus2024≈ 68%
- ≈ 67%
- ≈ 67%
- The Platonic Representation Hypothesisin corpus2024≈ 67%
- Mechanistic Knobs in LLMs: Retrieving and Steering High-Order Semantic Features via Sparse Autoencodersin corpus2026≈ 67%
- ≈ 67%
- Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representationsin corpus2023≈ 67%
- ≈ 66%
- Model Alignment Searchin corpus2025≈ 66%
- ≈ 66%
- ≈ 65%
Similar preprints — Semantic Scholar
Cited by (2)
- Semantic Diversity in Dialogue with Natural Language Inference
Confidence NLI Diversity achieves state-of-the-art Spearman's ρ of 0.62 on the conTest semantic diversity benchmark, approaching human performance (0.63) and outperforming the prior best automatic met
- Interpreting Language Model Parameters
VPD (adVersarial Parameter Decomposition) decomposes weight matrices directly into rank-one interpretable subcomponents rather than decomposing activations as sparse autoencoders (SAEs) do, flipping t