@inproceedings{chakrabarty-etal-2016-neural,
    title = "A Neural Lemmatizer for {B}engali",
    author = "Chakrabarty, Abhisek  and
      Chaturvedi, Akshay  and
      Garain, Utpal",
    editor = "Calzolari, Nicoletta  and
      Choukri, Khalid  and
      Declerck, Thierry  and
      Goggi, Sara  and
      Grobelnik, Marko  and
      Maegaard, Bente  and
      Mariani, Joseph  and
      Mazo, Helene  and
      Moreno, Asuncion  and
      Odijk, Jan  and
      Piperidis, Stelios",
    booktitle = "Proceedings of the Tenth International Conference on Language Resources and Evaluation ({LREC}'16)",
    month = may,
    year = "2016",
    address = "Portoro{\v{z}}, Slovenia",
    publisher = "European Language Resources Association (ELRA)",
    url = "https://preview.aclanthology.org/ingest-emnlp/L16-1406/",
    pages = "2558--2561",
    abstract = "We propose a novel neural lemmatization model which is language independent and supervised in nature. To handle the words in a neural framework, word embedding technique is used to represent words as vectors. The proposed lemmatizer makes use of contextual information of the surface word to be lemmatized. Given a word along with its contextual neighbours as input, the model is designed to produce the lemma of the concerned word as output. We introduce a new network architecture that permits only dimension specific connections between the input and the output layer of the model. For the present work, Bengali is taken as the reference language. Two datasets are prepared for training and testing purpose consisting of 19,159 and 2,126 instances respectively. As Bengali is a resource scarce language, these datasets would be beneficial for the respective research community. Evaluation method shows that the neural lemmatizer achieves 69.57{\%} accuracy on the test dataset and outperforms the simple cosine similarity based baseline strategy by a margin of 1.37{\%}."
}Markdown (Informal)
[A Neural Lemmatizer for Bengali](https://preview.aclanthology.org/ingest-emnlp/L16-1406/) (Chakrabarty et al., LREC 2016)
ACL
- Abhisek Chakrabarty, Akshay Chaturvedi, and Utpal Garain. 2016. A Neural Lemmatizer for Bengali. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16), pages 2558–2561, Portorož, Slovenia. European Language Resources Association (ELRA).