Findings of the E2E NLG Challenge

Ondřej Dušek, Jekaterina Novikova, Verena Rieser

2018-10-02WS 2018 11Data-to-Text Generation Text Generation Spoken Dialogue Systems

Abstract

This paper summarises the experimental setup and results of the first shared task on end-to-end (E2E) natural language generation (NLG) in spoken dialogue systems. Recent end-to-end generation systems are promising since they reduce the need for data annotation. However, they are currently limited to small, delexicalised datasets. The E2E NLG shared task aims to assess whether these novel approaches can generate better-quality output by learning from a dataset containing higher lexical richness, syntactic complexity and diverse discourse phenomena. We compare 62 systems submitted by 17 institutions, covering a wide range of approaches, including machine learning architectures -- with the majority implementing sequence-to-sequence models (seq2seq) -- as well as systems based on grammatical rules and templates.

Results

Task	Dataset	Metric	Value	Model
Text Generation	E2E NLG Challenge	BLEU	65.93	TGen
Text Generation	E2E NLG Challenge	CIDEr	2.2338	TGen
Text Generation	E2E NLG Challenge	METEOR	44.83	TGen
Text Generation	E2E NLG Challenge	NIST	8.6094	TGen
Text Generation	E2E NLG Challenge	ROUGE-L	68.5	TGen
Data-to-Text Generation	E2E NLG Challenge	BLEU	65.93	TGen
Data-to-Text Generation	E2E NLG Challenge	CIDEr	2.2338	TGen
Data-to-Text Generation	E2E NLG Challenge	METEOR	44.83	TGen
Data-to-Text Generation	E2E NLG Challenge	NIST	8.6094	TGen
Data-to-Text Generation	E2E NLG Challenge	ROUGE-L	68.5	TGen

Related Papers

Making Language Model a Hierarchical Classifier and Generator2025-07-17 Mitigating Object Hallucinations via Sentence-Level Early Intervention2025-07-16 The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs2025-07-15 Seq vs Seq: An Open Suite of Paired Encoders and Decoders2025-07-15 Hashed Watermark as a Filter: Defeating Forging and Overwriting Attacks in Weight-based Neural Network Watermarking2025-07-15 Exploiting Leaderboards for Large-Scale Distribution of Malicious Models2025-07-11 CLI-RAG: A Retrieval-Augmented Framework for Clinically Structured and Context Aware Text Generation with LLMs2025-07-09 FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation2025-07-09