Underreporting of errors in NLG output, and what to do about it

Emiel van Miltenburg; Miruna Clinciu; Ondřej Dušek; Dimitra Gkatzia; Stephanie Inglis; Leo Leppänen; Saad Mahamood; Emma Manning; Stephanie Schoch; Craig Thomson; Luou Wen

Underreporting of errors in NLG output, and what to do about it

Emiel van Miltenburg, Miruna Clinciu, Ondřej Dušek, Dimitra Gkatzia, Stephanie Inglis, Leo Leppänen, Saad Mahamood, Emma Manning, Stephanie Schoch, Craig Thomson, Luou Wen

Research output: Chapter in Book/Report/Conference proceeding › Published conference contribution

18 Citations (Scopus)

21 Downloads (Pure)

Abstract

We observe a severe under-reporting of the different kinds of errors that Natural Language Generation systems make. This is a problem, because mistakes are an important indicator of where systems should still be improved. If authors only report overall performance metrics, the research community is left in the dark about the specific weaknesses that are exhibited by `state-of-the-art' research. Next to quantifying the extent of error under-reporting, this position paper provides recommendations for error identification, analysis and reporting.

Original language	English
Title of host publication	Proceedings of the 14th International Conference on Natural Language Generation
Place of Publication	Aberdeen, Scotland, UK
Publisher	Association for Computational Linguistics
Pages	140-153
Number of pages	14
Publication status	Published - 1 Aug 2021

Access to Document

Miltenburg_etal_INLG21_Underreporting_Of_Errors_VoR
ACL materials are Copyright © 1963–2021 ACL; other materials are copyrighted by their respective copyright holders. Materials prior to 2016 here are licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 3.0 International License. Permission is granted to make copies for the purposes of teaching and research. Materials published in or after 2016 are licensed on a Creative Commons Attribution 4.0 International License. https://creativecommons.org/licenses/by/4.0/
Final published version, 350 KBLicence: CC BY

https://aclanthology.org/2021.inlg-1.14Licence: CC BY

Cite this

van Miltenburg, E., Clinciu, M., Dušek, O., Gkatzia, D., Inglis, S., Leppänen, L., Mahamood, S., Manning, E., Schoch, S., Thomson, C., & Wen, L. (2021). Underreporting of errors in NLG output, and what to do about it. In Proceedings of the 14th International Conference on Natural Language Generation (pp. 140-153). Association for Computational Linguistics. https://aclanthology.org/2021.inlg-1.14

Underreporting of errors in NLG output, and what to do about it. / van Miltenburg, Emiel; Clinciu, Miruna; Dušek, Ondřej et al.
Proceedings of the 14th International Conference on Natural Language Generation. Aberdeen, Scotland, UK: Association for Computational Linguistics, 2021. p. 140-153.

Research output: Chapter in Book/Report/Conference proceeding › Published conference contribution

van Miltenburg, E, Clinciu, M, Dušek, O, Gkatzia, D, Inglis, S, Leppänen, L, Mahamood, S, Manning, E, Schoch, S, Thomson, C & Wen, L 2021, Underreporting of errors in NLG output, and what to do about it. in Proceedings of the 14th International Conference on Natural Language Generation. Association for Computational Linguistics, Aberdeen, Scotland, UK, pp. 140-153. <https://aclanthology.org/2021.inlg-1.14>

@inproceedings{2d58a93f12864859a56019f682f44228,

title = "Underreporting of errors in NLG output, and what to do about it",

abstract = "We observe a severe under-reporting of the different kinds of errors that Natural Language Generation systems make. This is a problem, because mistakes are an important indicator of where systems should still be improved. If authors only report overall performance metrics, the research community is left in the dark about the specific weaknesses that are exhibited by `state-of-the-art' research. Next to quantifying the extent of error under-reporting, this position paper provides recommendations for error identification, analysis and reporting.",

author = "{van Miltenburg}, Emiel and Miruna Clinciu and Ond{\v r}ej Du{\v s}ek and Dimitra Gkatzia and Stephanie Inglis and Leo Lepp{\"a}nen and Saad Mahamood and Emma Manning and Stephanie Schoch and Craig Thomson and Luou Wen",

year = "2021",

month = aug,

day = "1",

language = "English",

pages = "140--153",

booktitle = "Proceedings of the 14th International Conference on Natural Language Generation",

publisher = "Association for Computational Linguistics",

}

TY - GEN

T1 - Underreporting of errors in NLG output, and what to do about it

AU - van Miltenburg, Emiel

AU - Clinciu, Miruna

AU - Dušek, Ondřej

AU - Gkatzia, Dimitra

AU - Inglis, Stephanie

AU - Leppänen, Leo

AU - Mahamood, Saad

AU - Manning, Emma

AU - Schoch, Stephanie

AU - Thomson, Craig

AU - Wen, Luou

PY - 2021/8/1

Y1 - 2021/8/1

N2 - We observe a severe under-reporting of the different kinds of errors that Natural Language Generation systems make. This is a problem, because mistakes are an important indicator of where systems should still be improved. If authors only report overall performance metrics, the research community is left in the dark about the specific weaknesses that are exhibited by `state-of-the-art' research. Next to quantifying the extent of error under-reporting, this position paper provides recommendations for error identification, analysis and reporting.

AB - We observe a severe under-reporting of the different kinds of errors that Natural Language Generation systems make. This is a problem, because mistakes are an important indicator of where systems should still be improved. If authors only report overall performance metrics, the research community is left in the dark about the specific weaknesses that are exhibited by `state-of-the-art' research. Next to quantifying the extent of error under-reporting, this position paper provides recommendations for error identification, analysis and reporting.

UR - https://inlg2021.github.io/pages/papers.html

M3 - Published conference contribution

SP - 140

EP - 153

BT - Proceedings of the 14th International Conference on Natural Language Generation

PB - Association for Computational Linguistics

CY - Aberdeen, Scotland, UK

ER -

Underreporting of errors in NLG output, and what to do about it

Abstract

Access to Document

Other files and links

Cite this