Treating Dialogue Quality Evaluation as an Anomaly Detection Problem
Abstract: Dialogue systems for interaction with humans have been enjoying increased popularity in the research and industry fields. To this day, the best way to estimate their success is through means of human evaluation and not automated approaches, despite the abundance of work done in the field. In this paper, we investigate the effectiveness of perceiving dialogue evaluation as an anomaly detection task. The paper looks into four dialogue modeling approaches and how their objective functions correlate with human annotation scores. A high-level perspective exhibits negative results. However, a more in-depth look shows some potential for using anomaly detection for evaluating dialogues.
Show BibTeX
@inproceedings{DBLP:conf/lrec/NedelchevUL20,
author = {Rostislav Nedelchev and
Ricardo Usbeck and
Jens Lehmann},
editor = {Nicoletta Calzolari and
Fr{\'{e}}d{\'{e}}ric B{\'{e}}chet and
Philippe Blache and
Khalid Choukri and
Christopher Cieri and
Thierry Declerck and
Sara Goggi and
Hitoshi Isahara and
Bente Maegaard and
Joseph Mariani and
H{\'{e}}l{\`{e}}ne Mazo and
Asunci{\'{o}}n Moreno and
Jan Odijk and
Stelios Piperidis},
title = {Treating Dialogue Quality Evaluation as an Anomaly Detection Problem},
booktitle = {Proceedings of The 12th Language Resources and Evaluation Conference,
{LREC} 2020, Marseille, France, May 11-16, 2020},
pages = {508--512},
publisher = {European Language Resources Association},
year = {2020},
url = {https://aclanthology.org/2020.lrec-1.64/},
timestamp = {Fri, 06 Aug 2021 00:40:03 +0200},
biburl = {https://dblp.org/rec/conf/lrec/NedelchevUL20.bib},
bibsource = {dblp computer science bibliography, https://dblp.org}
}