Pushing the Right Buttons: Adversarial Evaluation of Quality Estimation

Diptesh  Kanojia; Marina  Fomicheva; Tharindu Ranasinghe; Frédéric  Blain; Constantin Orasan; Lucia Specia

doi:10.48550/arXiv.2109.10859

Pushing the Right Buttons: Adversarial Evaluation of Quality Estimation

Diptesh Kanojia, Marina Fomicheva, Tharindu Ranasinghe, Frédéric Blain, Constantin Orasan, Lucia Specia

Research output: Chapter in Book/Published conference output › Conference publication

Abstract

Current Machine Translation (MT) systems achieve very good results on a growing variety of language pairs and datasets. However, they are known to produce fluent translation outputs that can contain important meaning errors, thus undermining their reliability in practice. Quality Estimation (QE) is the task of automatically assessing the performance of MT systems at test time. Thus, in order to be useful, QE systems should be able to detect such errors. However, this ability is yet to be tested in the current evaluation practices, where QE systems are assessed only in terms of their correlation with human judgements. In this work, we bridge this gap by proposing a general methodology for adversarial testing of QE for MT. First, we show that despite a high correlation with human judgements achieved by the recent SOTA, certain types of meaning errors are still problematic for QE to detect. Second, we show that on average, the ability of a given model to discriminate between meaning-preserving and meaning-altering perturbations is predictive of its overall performance, thus potentially allowing for comparing QE systems without relying on manual quality annotation.

Original language	English
Title of host publication	WMT 2021 - 6th Conference on Machine Translation, Proceedings
Publisher	Association for Computational Linguistics (ACL)
Pages	625-638
Number of pages	14
ISBN (Electronic)	9781954085947
DOIs	https://doi.org/10.48550/arXiv.2109.10859
Publication status	Published - 22 Sept 2021
Event	6th Conference on Machine Translation, WMT 2021 - Virtual, Online, Dominican Republic Duration: 10 Nov 2021 → 11 Nov 2021

Publication series

Name	WMT 2021 - 6th Conference on Machine Translation, Proceedings

Conference

Conference	6th Conference on Machine Translation, WMT 2021
Country/Territory	Dominican Republic
City	Virtual, Online
Period	10/11/21 → 11/11/21

Bibliographical note

Access to Document

10.48550/arXiv.2109.10859Licence: CC BY 4.0

Kanojiaetal_2021_VoR
Copyright 2021 the authors, Creative Commons Attribution 4.0 International (CC BY 4.0)
Final published version, 705 KBLicence: CC BY 4.0

Cite this

Kanojia, D., Fomicheva, M., Ranasinghe, T., Blain, F., Orasan, C., & Specia, L. (2021). Pushing the Right Buttons: Adversarial Evaluation of Quality Estimation. In WMT 2021 - 6th Conference on Machine Translation, Proceedings (pp. 625-638). (WMT 2021 - 6th Conference on Machine Translation, Proceedings). Association for Computational Linguistics (ACL). https://doi.org/10.48550/arXiv.2109.10859

@inproceedings{c78934c2f82c4f0ebaa02644a5152468,

title = "Pushing the Right Buttons: Adversarial Evaluation of Quality Estimation",

abstract = "Current Machine Translation (MT) systems achieve very good results on a growing variety of language pairs and datasets. However, they are known to produce fluent translation outputs that can contain important meaning errors, thus undermining their reliability in practice. Quality Estimation (QE) is the task of automatically assessing the performance of MT systems at test time. Thus, in order to be useful, QE systems should be able to detect such errors. However, this ability is yet to be tested in the current evaluation practices, where QE systems are assessed only in terms of their correlation with human judgements. In this work, we bridge this gap by proposing a general methodology for adversarial testing of QE for MT. First, we show that despite a high correlation with human judgements achieved by the recent SOTA, certain types of meaning errors are still problematic for QE to detect. Second, we show that on average, the ability of a given model to discriminate between meaning-preserving and meaning-altering perturbations is predictive of its overall performance, thus potentially allowing for comparing QE systems without relying on manual quality annotation.",

author = "Diptesh Kanojia and Marina Fomicheva and Tharindu Ranasinghe and Fr{\'e}d{\'e}ric Blain and Constantin Orasan and Lucia Specia",

year = "2021",

month = sep,

day = "22",

doi = "10.48550/arXiv.2109.10859",

language = "English",

series = "WMT 2021 - 6th Conference on Machine Translation, Proceedings",

publisher = "Association for Computational Linguistics (ACL)",

pages = "625--638",

booktitle = "WMT 2021 - 6th Conference on Machine Translation, Proceedings",

address = "United States",

}

Kanojia, D, Fomicheva, M, Ranasinghe, T, Blain, F, Orasan, C & Specia, L 2021, Pushing the Right Buttons: Adversarial Evaluation of Quality Estimation. in WMT 2021 - 6th Conference on Machine Translation, Proceedings. WMT 2021 - 6th Conference on Machine Translation, Proceedings, Association for Computational Linguistics (ACL), pp. 625-638, 6th Conference on Machine Translation, WMT 2021, Virtual, Online, Dominican Republic, 10/11/21. https://doi.org/10.48550/arXiv.2109.10859

Pushing the Right Buttons: Adversarial Evaluation of Quality Estimation. / Kanojia, Diptesh ; Fomicheva, Marina ; Ranasinghe, Tharindu et al.
WMT 2021 - 6th Conference on Machine Translation, Proceedings. Association for Computational Linguistics (ACL), 2021. p. 625-638 (WMT 2021 - 6th Conference on Machine Translation, Proceedings).

Research output: Chapter in Book/Published conference output › Conference publication

TY - GEN

T1 - Pushing the Right Buttons

T2 - 6th Conference on Machine Translation, WMT 2021

AU - Kanojia, Diptesh

AU - Fomicheva, Marina

AU - Ranasinghe, Tharindu

AU - Blain, Frédéric

AU - Orasan, Constantin

AU - Specia, Lucia

PY - 2021/9/22

Y1 - 2021/9/22

N2 - Current Machine Translation (MT) systems achieve very good results on a growing variety of language pairs and datasets. However, they are known to produce fluent translation outputs that can contain important meaning errors, thus undermining their reliability in practice. Quality Estimation (QE) is the task of automatically assessing the performance of MT systems at test time. Thus, in order to be useful, QE systems should be able to detect such errors. However, this ability is yet to be tested in the current evaluation practices, where QE systems are assessed only in terms of their correlation with human judgements. In this work, we bridge this gap by proposing a general methodology for adversarial testing of QE for MT. First, we show that despite a high correlation with human judgements achieved by the recent SOTA, certain types of meaning errors are still problematic for QE to detect. Second, we show that on average, the ability of a given model to discriminate between meaning-preserving and meaning-altering perturbations is predictive of its overall performance, thus potentially allowing for comparing QE systems without relying on manual quality annotation.

AB - Current Machine Translation (MT) systems achieve very good results on a growing variety of language pairs and datasets. However, they are known to produce fluent translation outputs that can contain important meaning errors, thus undermining their reliability in practice. Quality Estimation (QE) is the task of automatically assessing the performance of MT systems at test time. Thus, in order to be useful, QE systems should be able to detect such errors. However, this ability is yet to be tested in the current evaluation practices, where QE systems are assessed only in terms of their correlation with human judgements. In this work, we bridge this gap by proposing a general methodology for adversarial testing of QE for MT. First, we show that despite a high correlation with human judgements achieved by the recent SOTA, certain types of meaning errors are still problematic for QE to detect. Second, we show that on average, the ability of a given model to discriminate between meaning-preserving and meaning-altering perturbations is predictive of its overall performance, thus potentially allowing for comparing QE systems without relying on manual quality annotation.

UR - http://www.scopus.com/inward/record.url?scp=85127214482&partnerID=8YFLogxK

UR - https://arxiv.org/abs/2109.10859

U2 - 10.48550/arXiv.2109.10859

DO - 10.48550/arXiv.2109.10859

M3 - Conference publication

AN - SCOPUS:85127214482

T3 - WMT 2021 - 6th Conference on Machine Translation, Proceedings

SP - 625

EP - 638

BT - WMT 2021 - 6th Conference on Machine Translation, Proceedings

PB - Association for Computational Linguistics (ACL)

Y2 - 10 November 2021 through 11 November 2021

ER -

Pushing the Right Buttons: Adversarial Evaluation of Quality Estimation

Abstract

Publication series

Conference

Bibliographical note

Access to Document

Other files and links

Fingerprint

Cite this