@inproceedings{eger-etal-2019-pitfalls,
    title = "Pitfalls in the Evaluation of Sentence Embeddings",
    author = {Eger, Steffen  and
      R{\"u}ckl{\'e}, Andreas  and
      Gurevych, Iryna},
    editor = "Augenstein, Isabelle  and
      Gella, Spandana  and
      Ruder, Sebastian  and
      Kann, Katharina  and
      Can, Burcu  and
      Welbl, Johannes  and
      Conneau, Alexis  and
      Ren, Xiang  and
      Rei, Marek",
    booktitle = "Proceedings of the 4th Workshop on Representation Learning for NLP (RepL4NLP-2019)",
    month = aug,
    year = "2019",
    address = "Florence, Italy",
    publisher = "Association for Computational Linguistics",
    url = "https://preview.aclanthology.org/iwcs-25-ingestion/W19-4308/",
    doi = "10.18653/v1/W19-4308",
    pages = "55--60",
    abstract = "Deep learning models continuously break new records across different NLP tasks. At the same time, their success exposes weaknesses of model evaluation. Here, we compile several key pitfalls of evaluation of sentence embeddings, a currently very popular NLP paradigm. These pitfalls include the comparison of embeddings of different sizes, normalization of embeddings, and the low (and diverging) correlations between transfer and probing tasks. Our motivation is to challenge the current evaluation of sentence embeddings and to provide an easy-to-access reference for future research. Based on our insights, we also recommend better practices for better future evaluations of sentence embeddings."
}Markdown (Informal)
[Pitfalls in the Evaluation of Sentence Embeddings](https://preview.aclanthology.org/iwcs-25-ingestion/W19-4308/) (Eger et al., RepL4NLP 2019)
ACL
- Steffen Eger, Andreas Rücklé, and Iryna Gurevych. 2019. Pitfalls in the Evaluation of Sentence Embeddings. In Proceedings of the 4th Workshop on Representation Learning for NLP (RepL4NLP-2019), pages 55–60, Florence, Italy. Association for Computational Linguistics.