@inproceedings{castro-ferreira-etal-2021-enriching,
    title = "Enriching the {E}2{E} dataset",
    author = "Castro Ferreira, Thiago  and
      Vaz, Helena  and
      Davis, Brian  and
      Pagano, Adriana",
    editor = "Belz, Anya  and
      Fan, Angela  and
      Reiter, Ehud  and
      Sripada, Yaji",
    booktitle = "Proceedings of the 14th International Conference on Natural Language Generation",
    month = aug,
    year = "2021",
    address = "Aberdeen, Scotland, UK",
    publisher = "Association for Computational Linguistics",
    url = "https://preview.aclanthology.org/ingest-emnlp/2021.inlg-1.18/",
    doi = "10.18653/v1/2021.inlg-1.18",
    pages = "177--183",
    abstract = "This study introduces an enriched version of the E2E dataset, one of the most popular language resources for data-to-text NLG. We extract intermediate representations for popular pipeline tasks such as discourse ordering, text structuring, lexicalization and referring expression generation, enabling researchers to rapidly develop and evaluate their data-to-text pipeline systems. The intermediate representations are extracted by aligning non-linguistic and text representations through a process called delexicalization, which consists in replacing input referring expressions to entities/attributes with placeholders. The enriched dataset is publicly available."
}Markdown (Informal)
[Enriching the E2E dataset](https://preview.aclanthology.org/ingest-emnlp/2021.inlg-1.18/) (Castro Ferreira et al., INLG 2021)
ACL
- Thiago Castro Ferreira, Helena Vaz, Brian Davis, and Adriana Pagano. 2021. Enriching the E2E dataset. In Proceedings of the 14th International Conference on Natural Language Generation, pages 177–183, Aberdeen, Scotland, UK. Association for Computational Linguistics.