Juan-Manuel Torres-Moreno

Also published as: Juan-Manuel Torres, Juan-Manuel Torres Moreno, Juan Manuel Torres Moreno


2026

The aim of this article is to introduce Context-Free Grammars (CFG) for the Nawatl language. Nawatl is an Amerindian language of the 𝜋-language type, i.e. a language with few digital resources. For this reason the corpora available for the learning of Large Language Models (LLMs) are virtually non-existent, posing a significant challenge. The goal is to produce a substantial number of syntactically valid artificial Nawatl sentences and thereby to expand the corpora for the purpose of learning embeddings (static models or probably LLMs). For this objective, we introduce two new Nawatl CFGs and use them in generative mode. Thanks to these grammars, it is possible to expand Nawatl corpus significantly and subsequently to use it to learn embeddings (such as FastText) and to evaluate their relevance in semantic similarity tasks. The results show an improvement compared to the results obtained using only the original corpus without artificial expansion, and also demonstrate that economic embeddings often perform better than some LLMs.
In this paper, we aim to answer the following question: could corpus duplication be useful in Natural Language Processing (NLP) for low-resource languages? In these languages (or pi-languages), corpora available for training Large Language Models are virtually non-existent. Specifically, we study the impact of corpus expansion in Nahuatl, an agglutinative and polysynthetic Amerindian pi-language characterised by extensive dialectal variation. Our goal is to increase the size of Nahuatl corpora, which currently consist of a limited number of tokens, through controlled duplication techniques. Our experimental setup employs incremental duplication alongside appropriate corpus balancing, with the objective of training embeddings optimised for downstream NLP tasks. Consequently, static embeddings were trained and evaluated on a sentence-level semantic similarity task. Our results show a significant improvement in performance when incremental duplication is applied, compared to results obtained without corpus expansion. To our knowledge, this technique has not yet been explored in this field.

2025

π-YALLI : a new corpus for Nahuatl Language Models The Nahuatl is a language with few computational resources, despite the fact that it is a living language spoken by around two million people. We built π-YALLI, a corpus that enables research and development of dynamic and static Language Models (LM). We measured the perplexity of π-YALLI, evaluating state-of-the-art LM performance on a manually annotated semantic similarity corpus relative to annotator agreement. The results show the difficulty of working with this π-language, but at the same time open up interesting perspectives for the study of other NLP tasks on Nahuatl.

2021

This work aims to evaluate the ability that both probabilistic and state-of-the-art vector space modeling (VSM) methods provide to well known machine learning algorithms to identify social network documents to be classified as aggressive, gender biased or communally charged. To this end, an exploratory stage was performed first in order to find relevant settings to test, i.e. by using training and development samples, we trained multiple algorithms using multiple vector space modeling and probabilistic methods and discarded the less informative configurations. These systems were submitted to the competition of the ComMA@ICON’21 Workshop on Multilingual Gender Biased and Communal Language Identification.

2020

In this paper, we show the enhancing of the Demanded Skills Diagnosis (DiCoDe: DiagnĂłstico de Competencias Demandadas), a system developed by Mexico City’s Ministry of Labor and Employment Promotion (STyFE: SecretarĂ­a de Trabajo y Fomento del Empleo de la Ciudad de MĂ©xico) that seeks to reduce information asymmetries between job seekers and employers. The project uses webscraping techniques to retrieve job vacancies posted on private job portals on a daily basis and with the purpose of informing training and individual case management policies as well as labor market monitoring. For this purpose, a collaboration project between STyFE and the Language Engineering Group (GIL: Grupo de IngenierĂ­a LingĂŒĂ­stica) was established in order to enhance DiCoDe by applying NLP models and semantic analysis. By this collaboration, DiCoDe’s job vacancies system’s macro-structure and its geographic referencing at the city hall (municipality) level were improved. More specifically, dictionaries were created to identify demanded competencies, skills and abilities (CSA) and algorithms were developed for dynamic classifying of vacancies and identifying terms for searches on free text, in order to improve the results and processing time of queries.

2018

The phenomenon of cyberbullying has growing in worrying proportions with the development of social networks. Forums and chat rooms are spaces where serious damage can now be done to others, while the tools for avoiding on-line spills are still limited. This study aims to assess the ability that both classical and state-of-the-art vector space modeling methods provide to well known learning machines to identify aggression levels in social network cyberbullying (i.e. social network posts manually labeled as Overtly Aggressive, Covertly Aggressive and Non-aggressive). To this end, an exploratory stage was performed first in order to find relevant settings to test, i.e. by using training and development samples, we trained multiple learning machines using multiple vector space modeling methods and discarded the less informative configurations. Finally, we selected the two best settings and their voting combination to form three competing systems. These systems were submitted to the competition of the TRACK-1 task of the Workshop on Trolling, Aggression and Cyberbullying. Our voting combination system resulted second place in predicting Aggression levels on a test set of untagged social network posts.
Multi-Sentence Compression (MSC) aims to generate a short sentence with key information from a cluster of closely related sentences. MSC enables summarization and question-answering systems to generate outputs combining fully formed sentences from one or several documents. This paper describes a new Integer Linear Programming method for MSC using a vertex-labeled graph to select different keywords, and novel 3-gram scores to generate more informative sentences while maintaining their grammaticality. Our system is of good quality and outperforms the state-of-the-art for evaluations led on news dataset. We led both automatic and manual evaluations to determine the informativeness and the grammaticality of compressions for each dataset. Additional tests, which take advantage of the fact that the length of compressions can be modulated, still improve ROUGE scores with shorter output sentences.
Cet article prĂ©sente l’édition 2018 de la campagne d’évaluation DEFT (DĂ©fi Fouille de Textes). A partir d’un corpus de tweets, quatre tĂąches ont Ă©tĂ© proposĂ©es : identifier les tweets sur la thĂ©matique des transports, puis parmi ces derniers, identifier la polaritĂ© (nĂ©gatif, neutre, positif, mixte), identifier les marqueurs de sentiment et la cible, et enfin, annoter complĂštement chaque tweet en source et cible des sentiments exprimĂ©s. Douze Ă©quipes ont participĂ©, majoritairement sur les deux premiĂšres tĂąches. Sur l’identification de la thĂ©matique des transports, la micro F-mesure varie de 0,827 Ă  0,908. Sur l’identification de la polaritĂ© globale, la micro F-mesure varie de 0,381 Ă  0,823.
Semantic Textual Similarity (STS) is the basis of many applications in Natural Language Processing (NLP). Our system combines convolution and recurrent neural networks to measure the semantic similarity of sentences. It uses a convolution network to take account of the local context of words and an LSTM to consider the global context of sentences. This combination of networks helps to preserve the relevant information of sentences and improves the calculation of the similarity between sentences. Our model has achieved good results and is competitive with the best state-of-the-art systems.

2014

2013

2011

Le rĂ©sumĂ© automatique cross-lingue consiste Ă  gĂ©nĂ©rer un rĂ©sumĂ© rĂ©digĂ© dans une langue diffĂ©rente de celle utilisĂ©e dans les documents sources. Dans cet article, nous proposons une approche de rĂ©sumĂ© automatique multi-document, basĂ©e sur une reprĂ©sentation par graphe, qui prend en compte des scores de qualitĂ© de traduction lors du processus de sĂ©lection des phrases. Nous Ă©valuons notre mĂ©thode sur un sous-ensemble manuellement traduit des donnĂ©es utilisĂ©es lors de la campagne d’évaluation internationale DUC 2004. Les rĂ©sultats expĂ©rimentaux indiquent que notre approche permet d’amĂ©liorer la lisibilitĂ© des rĂ©sumĂ©s gĂ©nĂ©rĂ©s, sans pour autant dĂ©grader leur informativitĂ©.

2010

This paper presents two corpora produced within the RPM2 project: a multi-document summarization corpus and a sentence compression corpus. Both corpora are in French. The first one is the only one we know in this language. It contains 20 topics with 20 documents each. A first set of 10 documents per topic is summarized and then the second set is used to produce an update summarization (new information). 4 annotators were involved and produced a total of 160 abstracts. The second corpus contains all the sentences of the first one. 4 annotators were asked to compress the 8432 sentences. This is the biggest corpus of compressed sentences we know, whatever the language. The paper provides some figures in order to compare the different annotators: compression rates, number of tokens per sentence, percentage of tokens kept according to their POS, position of dropped tokens in the sentence compression phase, etc. These figures show important differences from an annotator to the other. Another point is the different strategies of compression used according to the length of the sentence.
Availability of labeled language resources, such as annotated corpora and domain dependent labeled language resources is crucial for experiments in the field of Natural Language Processing. Most often, due to lack of resources, manual verification and annotation of electronic text material is a prerequisite for the development of NLP tools. In the context of under-resourced language, the lack of copora becomes a crucial problem because most of the research efforts are supported by organizations with limited funds. Using free, multilingual and highly structured corpora like Wikipedia to produce automatically labeled language resources can be an answer to those needs. This paper introduces NLGbAse, a multilingual linguistic resource built from the Wikipedia encyclopedic content. This system produces structured metadata which make possible the automatic annotation of corpora with syntactical and semantical labels. A metadata contains semantical and statistical informations related to an encyclopedic document. To validate our approach, we built and evaluated a Named Entity Recognition tool, trained with Wikipedia corpora annotated by our system.
This paper presents a new algorithm for automatic summarization of specialized texts combining terminological and semantic resources: a term extractor and an ontology. The term extractor provides the list of the terms that are present in the text together their corresponding termhood. The ontology is used to calculate the semantic similarity among the terms found in the main body and those present in the document title. The general idea is to obtain a relevance score for each sentence taking into account both the ”termhood” of the terms found in such sentence and the similarity among such terms and those terms present in the title of the document. The phrases with the highest score are chosen to take part of the final summary. We evaluate the algorithm with Rouge, comparing the resulting summaries with the summaries of other summarizers. The sentence selection algorithm was also tested as part of a standalone summarizer. In both cases it obtains quite good results although the perception is that there is a space for improvement.
Nous Ă©tudions diffĂ©rentes mĂ©thodes d’évaluation de rĂ©sumĂ© de documents basĂ©es sur le contenu. Nous nous intĂ©ressons en particulier Ă  la corrĂ©lation entre les mesures d’évaluation avec et sans rĂ©fĂ©rence humaine. Nous avons dĂ©veloppĂ© FRESA, un nouveau systĂšme d’évaluation fondĂ© sur le contenu qui calcule les divergences entre les distributions de probabilitĂ©. Nous appliquons notre systĂšme de comparaison aux diverses mesures d’évaluation bien connues en rĂ©sumĂ© de texte telles que la Couverture, Responsiveness, Pyramids et Rouge en Ă©tudiant leurs associations dans les tĂąches du rĂ©sumĂ© multi-document gĂ©nĂ©rique (francais/anglais), focalisĂ© (anglais) et rĂ©sumĂ© mono-document gĂ©nĂ©rique (français/espagnol).

2009

On utilise souvent des ressources lexicales externes pour amĂ©liorer les performances des systĂšmes d’étiquetage d’entitĂ©s nommĂ©es. Les contenus de ces ressources lexicales peuvent ĂȘtre variĂ©s : liste de noms propres, de lieux, de marques. On note cependant que la disponibilitĂ© de corpus encyclopĂ©diques exhaustifs et ouverts de grande taille tels que Worldnet ou Wikipedia, a fait Ă©merger de nombreuses propositions spĂ©cifiques d’exploitation de ces contenus par des systĂšmes d’étiquetage. Un problĂšme demeure nĂ©anmoins ouvert avec ces ressources : celui de l’adaptation de leur taxonomie interne, complexe et composĂ©e de dizaines de milliers catĂ©gories, aux exigences particuliĂšres de l’étiquetage des entitĂ©s nommĂ©es. Pour ces derniĂšres, au plus de quelques centaines de classes sĂ©mantiques sont requises. Dans cet article nous explorons cette difficultĂ© et proposons un systĂšme complet de transformation d’un arbre taxonomique encyclopĂ©dique en une systĂšme Ă  classe sĂ©mantiques adaptĂ© Ă  l’étiquetage d’entitĂ©s nommĂ©es.
Nous prĂ©sentons une approche exploratoire basĂ©e sur des notions thermodynamiques de la Physique statistique pour la compression de phrases. Nous dĂ©crivons le modĂšle magnĂ©tique des verres de spins, adaptĂ© Ă  notre conception de la problĂ©matique. Des simulations MĂ©tropolis Monte-Carlo permettent d’introduire des fluctuations thermiques pour piloter la compression. Des comparaisons intĂ©ressantes de notre mĂ©thode ont Ă©tĂ© rĂ©alisĂ©es sur un corpus en français.
Le rĂ©sumĂ© automatique de texte est une problĂ©matique difficile, fortement dĂ©pendante de la langue et qui peut nĂ©cessiter un ensemble de donnĂ©es d’apprentissage consĂ©quent. L’approche par extraction peut aider Ă  surmonter ces difficultĂ©s. (Mihalcea, 2004) a dĂ©montrĂ© l’intĂ©rĂȘt des approches Ă  base de graphes pour l’extraction de segments de texte importants. Dans cette Ă©tude, nous dĂ©crivons une approche indĂ©pendante de la langue pour la problĂ©matique du rĂ©sumĂ© automatique multi-documents. L’originalitĂ© de notre mĂ©thode repose sur l’utilisation d’une mesure de similaritĂ© permettant le rapprochement de segments morphologiquement proches. De plus, c’est Ă  notre connaissance la premiĂšre fois que l’évaluation d’une approche de rĂ©sumĂ© automatique multi-document est conduite sur des textes en français.
Le marchĂ© d’offres d’emploi et des candidatures sur Internet connaĂźt une croissance exponentielle. Ceci implique des volumes d’information (majoritairement sous la forme de texte libre) qu’il n’est plus possible de traiter manuellement. Une analyse et catĂ©gorisation assistĂ©es nous semble pertinente en rĂ©ponse Ă  cette problĂ©matique. Nous proposons E-Gen, systĂšme qui a pour but l’analyse et catĂ©gorisation assistĂ©s d’offres d’emploi et des rĂ©ponses des candidats. Dans cet article nous prĂ©sentons plusieurs stratĂ©gies, reposant sur les modĂšles vectoriel et probabiliste, afin de rĂ©soudre la problĂ©matique du profilage des candidatures en fonction d’une offre prĂ©cise. Nous avons Ă©valuĂ© une palette de mesures de similaritĂ© afin d’effectuer un classement pertinent des candidatures au moyen des courbes ROC. L’utilisation d’une forme de relevance feedback a permis de surpasser nos rĂ©sultats sur ce problĂšme difficile et sujet Ă  une grande subjectivitĂ©.

2008

Nous prĂ©sentons dans cet article une mĂ©thode d’extraction automatique d’informations sur des textes de trĂšs petite taille, faiblement structurĂ©s. Nous travaillons sur des textes dont la rĂ©daction n’est pas normalisĂ©e, avec trĂšs peu de mots pour caractĂ©riser chaque information. Les textes ne contiennent pas ou trĂšs peu de phrases. Il s’agit le plus souvent de morceaux de phrases ou d’expressions composĂ©es de quelques mots. Nous comparons plusieurs mĂ©thodes d’extraction, dont certaines sont entiĂšrement automatiques. D’autres utilisent en partie une connaissance du domaine que nous voulons rĂ©duite au minimum, de façon Ă  minimiser le travail manuel en amont. Enfin, nous prĂ©sentons nos rĂ©sultats qui dĂ©passent ce dont il est fait Ă©tat dans la littĂ©rature, avec une prĂ©cision Ă©quivalente et un rappel supĂ©rieur.
Dans cet article, nous prĂ©sentons des applications du systĂšme Enertex au Traitement Automatique de la Langue Naturelle. Enertex est basĂ© sur l’énergie textuelle, une approche par rĂ©seaux de neurones inspirĂ©e de la physique statistique des systĂšmes magnĂ©tiques. Nous avons appliquĂ© cette approche aux problĂšmes du rĂ©sumĂ© automatique multi-documents et de la dĂ©tection de frontiĂšres thĂ©matiques. Les rĂ©sultats, en trois langues : anglais, espagnol et français, sont trĂšs encourageants.
La croissance exponentielle de l’Internet a permis le dĂ©veloppement de sites d’offres d’emploi en ligne. Le systĂšme E-Gen (Traitement automatique d’offres d’emploi) a pour but de permettre l’analyse et la catĂ©gorisation d’offres d’emploi ainsi qu’une analyse et classification des rĂ©ponses des candidats (Lettre de motivation et CV). Nous prĂ©sentons les travaux rĂ©alisĂ©s afin de rĂ©soudre la seconde partie : on utilise une reprĂ©sentation vectorielle de texte pour effectuer une classification des piĂšces jointes contenus dans le mail Ă  l’aide de SVM. Par la suite, une Ă©valuation de la candidature est effectuĂ©e Ă  l’aide de diffĂ©rents classifieurs (SVM et n-grammes de mots).

2007

Dans cet article, nous prĂ©sentons une approche de rĂ©seaux de neurones inspirĂ©e de la physique statistique de systĂšmes magnĂ©tiques pour Ă©tudier des problĂšmes fondamentaux du Traitement Automatique de la Langue Naturelle. L’algorithme modĂ©lise un document comme un systĂšme de neurones oĂč l’on dĂ©duit l’énergie textuelle. Nous avons appliquĂ© cette approche aux problĂšmes de rĂ©sumĂ© automatique et de dĂ©tection de frontiĂšres thĂ©matiques. Les rĂ©sultats sont trĂšs encourageants.
Search
Fix author