@inproceedings{riley-etal-2020-translationese,
    title = "Translationese as a Language in ``Multilingual'' {NMT}",
    author = "Riley, Parker  and
      Caswell, Isaac  and
      Freitag, Markus  and
      Grangier, David",
    editor = "Jurafsky, Dan  and
      Chai, Joyce  and
      Schluter, Natalie  and
      Tetreault, Joel",
    booktitle = "Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics",
    month = jul,
    year = "2020",
    address = "Online",
    publisher = "Association for Computational Linguistics",
    url = "https://preview.aclanthology.org/sigedu-bea-out-of-sync-correction/2020.acl-main.691/",
    doi = "10.18653/v1/2020.acl-main.691",
    pages = "7737--7746",
    abstract = "Machine translation has an undesirable propensity to produce ``translationese'' artifacts, which can lead to higher BLEU scores while being liked less by human raters. Motivated by this, we model translationese and original (i.e. natural) text as separate languages in a multilingual model, and pose the question: can we perform zero-shot translation between original source text and original target text? There is no data with original source and original target, so we train a sentence-level classifier to distinguish translationese from original target text, and use this classifier to tag the training data for an NMT model. Using this technique we bias the model to produce more natural outputs at test time, yielding gains in human evaluation scores on both accuracy and fluency. Additionally, we demonstrate that it is possible to bias the model to produce translationese and game the BLEU score, increasing it while decreasing human-rated quality. We analyze these outputs using metrics measuring the degree of translationese, and present an analysis of the volatility of heuristic-based train-data tagging."
}Markdown (Informal)
[Translationese as a Language in “Multilingual” NMT](https://preview.aclanthology.org/sigedu-bea-out-of-sync-correction/2020.acl-main.691/) (Riley et al., ACL 2020)
ACL
- Parker Riley, Isaac Caswell, Markus Freitag, and David Grangier. 2020. Translationese as a Language in “Multilingual” NMT. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7737–7746, Online. Association for Computational Linguistics.