Horacio Saggion

dblp:36/2688 · DBLP profile ↗
← Back
88ranked-venue papers
26as first author
13since 2021 · last 2026
0000-0003-0016-7807ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 74 · 25 first-author · 12 since 2021Databases, data management, data science and information retrieval · 17 · 5 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 since 2021Human-computer interaction and ubiquitous computing · 3
YearPublicationVenuePosition
2026 A Multilingual Human Annotated Corpus of Original and Easy-to-Read Texts to Support Access to Democratic Participatory Processes
Stefan Bott, Verena Riegler, Horacio Saggion, Almudena Rascón, Nouran Khallaf
LREC3
2025 UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency Assessment
abstract
Joseph Marvin Imperial, Abdullah Barayan, Regina Stodden, Rodrigo Wilkens, Ricardo Muñoz Sánchez, Lingyun Gao, Melissa Torgbi, Dawn Knight, Gail Forey, Reka R. Jablonkai, Ekaterina Kochmar, Robert Joshua Reynolds, Eugénio Ribeiro, Horacio Saggion, Elena Volodina, Sowmya Vajjala, Thomas François, Fernando Alva-Manchego, Harish Tayyar Madabushi. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Joseph Marvin Imperial, Abdullah Barayan, Regina Stodden, Rodrigo Wilkens, Ricardo Muñoz Sánchez, Lingyun Gao, Melissa Torgbi, Dawn Knight, Gail Forey, Reka R. Jablonkai, Ekaterina Kochmar, Robert Reynolds 0001, Eugénio Ribeiro, Horacio Saggion, Elena Volodina, Sowmya Vajjala, Thomas François, Fernando Alva-Manchego, Harish Tayyar Madabushi
EMNLP14
2025 Attacking Misinformation Detection Using Adversarial Examples Generated by Language Models
abstract
Large language models have many beneficial applications, but can they also be used to attack content-filtering algorithms in social media platforms?We investigate the challenge of generating adversarial examples to test the robustness of text classification algorithms detecting low-credibility content, including propaganda, false claims, rumours and hyperpartisan news.We focus on simulation of content moderation by setting realistic limits on the number of queries an attacker is allowed to attempt.Within our solution (TREPAT), initial rephrasings are generated by large language models with prompts inspired by meaning-preserving NLP tasks, such as text simplification and style transfer.Subsequently, these modifications are decomposed into small changes, applied through beam search procedure, until the victim classifier changes its decision.We perform (1) quantitative evaluation using various prompts, models and query limits, (2) targeted manual assessment of the generated text and (3) qualitative linguistic analysis.The results confirm the superiority of our approach in the constrained scenario, especially in case of long input text (news articles), where exhaustive search is not feasible.
Piotr Przybyla, Euan McGill, Horacio Saggion
EMNLP3
2025 Towards Trustworthy Lexical Simplification: Exploring Safety and Efficiency with Small LLMs
abstract
Despite their strong performance, large language models (LLMs) face challenges in real-world application of lexical simplification (LS), particularly in privacy-sensitive and resource-constrained environments. Moreover, since vulnerable user groups (e.g., people with disabilities) are one of the key target groups of this technology, it is crucial to ensure the safety and correctness of the output of LS systems. To address these issues, we propose an efficient framework for LS systems that utilizes small LLMs deployable in local environments. Within this framework, we explore knowledge distillation with synthesized data and in-context learning as baselines. Our experiments in five languages evaluate model outputs both automatically and manually. Our manual analysis reveals that while knowledge distillation boosts automatic metric scores, it also introduces a safety trade-off by increasing harmful simplifications. Importantly, we find that the model’s output probability is a useful signal for detecting harmful simplifications. Leveraging this, we propose a filtering strategy that suppresses harmful simplifications while largely preserving beneficial ones. This work establishes a benchmark for efficient and safe LS with small LLMs. It highlights the key trade-offs between performance, efficiency, and safety, and demonstrates a promising approach for safe real-world deployment.
Akio Hayakawa, Stefan Bott, Horacio Saggion
INLG3
2024 Bootstrapping Pre-trained Word Embedding Models for Sign Language Gloss Translation
abstract
This paper explores a novel method to modify existing pre-trained word embedding models of spoken languages for Sign Language glosses. These newly-generated embeddings are described, visualised, and then used in the encoder and/or decoder of models for the Text2Gloss and Gloss2Text task of machine translation. In two translation settings (one including data augmentation-based pre-training and a baseline), we find that bootstrapped word embeddings for glosses improve translation across four Signed/spoken language pairs. Many improvements are statistically significant, including those where the bootstrapped gloss embedding models are used.Languages included: American Sign Language, Finnish Sign Language, Spanish Sign Language, Sign Language of The Netherlands.
Euan McGill, Luis Chiruzzo, Horacio Saggion
EAMT (1)3
2024 SignON - a Co-creative Machine Translation for Sign and Spoken Languages (end-of-project results, contributions and lessons learned)
abstract
SignON, a 3-year Horizon 20202 project addressing the lack of technology and services for MT between sign languages (SLs) and spoken languages (SpLs) ended in December 2023. SignON was unprecedented. Not only it addressed the wider complexity of the aforementioned problem – from research and development of recognition, translation and synthesis, through development of easy-to-use mobile applications and a cloud-based framework to do the “heavy lifting” as well as to establishing ethical, privacy and inclusivenesspolicies and operation guidelines – but also engaged with the deaf and hard of hearing communities in an effective co-creation approach where these main stakeholders drove the development in the right direction and had the final say.Currently we are witnessing advances in natural language processing for SLs, including MT. SignON was one of the largest projects that contributed to this surge with 17 partners and more than 60 consortium members, working in parallel with other international and European initiatives, such as project EASIER and others.
Dimitar Sht. Shterionov, Vincent Vandeghinste, Mirella De Sisto, Aoife Brady, Mathieu De Coster, Lorraine Leeson, Andy Way, Josep Blat, Frankie Picron, Davy Van Landuyt, Marcello Paolo Scipioni, Aditya Parikh, Louis ten Bosch, John J. O'Flaherty, Joni Dambre, Caro Brosens, Jorn Rijckaert, Víctor Ubieto Nogales, Bram Vanroy, Santiago Egea Gómez, Ineke Schuurman, Gorka Labaka, Adrián Núñez-Marcos, Irene Murtagh, Euan McGill, Horacio Saggion
EAMT (2)26
2023 SignON: Sign Language Translation. Progress and challenges
abstract
SignON (https://signon-project.eu/) is a Horizon 2020 project, running from 2021 until the end of 2023, which addresses the lack of technology and services for the automatic translation between sign languages (SLs) and spoken languages, through an inclusive, human-centric solution, hence contributing to the repertoire of communication media for deaf, hard of hearing (DHH) and hearing individuals. In this paper, we present an update of the status of the project, describing the approaches developed to address the challenges and peculiarities of SL machine translation (SLMT).
Vincent Vandeghinste, Dimitar Sht. Shterionov, Mirella De Sisto, Aoife Brady, Mathieu De Coster, Lorraine Leeson, Josep Blat, Frankie Picron, Marcello Paolo Scipioni, Aditya Parikh, Louis ten Bosch, John J. O'Flaherty, Joni Dambre, Jorn Rijckaert, Bram Vanroy, Víctor Ubieto Nogales, Santiago Egea Gómez, Ineke Schuurman, Gorka Labaka, Adrián Núñez-Marcos, Irene Murtagh, Euan McGill, Horacio Saggion
EAMT23
2023 Creating a Silver Standard for Patent Simplification
abstract
Patents are legal documents that aim at protecting inventions on the one hand and at making technical knowledge circulate on the other. Their complex style -- a mix of legal, technical, and extremely vague language -- makes their content hard to access for humans and machines and poses substantial challenges to the information retrieval community. This paper proposes an approach to automatically simplify patent text through rephrasing. Since no in-domain parallel simplification data exist, we propose a method to automatically generate a large-scale silver standard for patent sentences. To obtain candidates, we use a general-domain paraphrasing system; however, the process is error-prone and difficult to control. Thus, we pair it with proper filters and construct a cleaner corpus that can successfully be used to train a simplification system. Human evaluation of the synthetic silver corpus shows that it is considered grammatical, adequate, and contains simple sentences.
Silvia Casola, Alberto Lavelli, Horacio Saggion
SIGIR3
2022 Sentence Simplification Capabilities of Transfer-Based Models
abstract
According to the official adult literacy report conducted in 24 highly-developed countries, more than 50% adults, on average, can only understand basic vocabulary, short sentences, and basic syntactic constructions. Everyday information found in news articles is thus inaccessible to many people, impeding their social inclusion and informed decision-making. Systems for automatic sentence simplification aim to provide scalable solution to this problem. In this paper, we propose new state-of-the-art sentence simplification systems for English and Spanish, and specifications for expert evaluation that are in accordance with well-established easy-to-read guidelines. We conduct expert evaluation of our new systems and the previous state-of-the-art systems for English and Spanish, and discuss strengths and weaknesses of each of them. Finally, we draw conclusions about the capabilities of the state-of-the-art sentence simplification systems and give some directions for future research.
Sanja Stajner, Kim Cheng Sheang, Horacio Saggion
AAAI3
2022 ALEXSIS: A Dataset for Lexical Simplification in Spanish
abstract
Lexical Simplification is the process of reducing the lexical complexity of a text by replacing difficult words with easier to read (or understand) expressions while preserving the original information and meaning. In this paper we introduce ALEXSIS, a new dataset for this task, and we use ALEXSIS to benchmark Lexical Simplification systems in Spanish. The paper describes the evaluation of three kind of approaches to Lexical Simplification, a thesaurus-based approach, a single transformers-based approach, and a combination of transformers. We also report state of the art results on a previous Lexical Simplification dataset for Spanish.
Daniel Ferrés, Horacio Saggion
LREC2
2022 Challenges with Sign Language Datasets for Sign Language Recognition and Translation
abstract
Sign Languages (SLs) are the primary means of communication for at least half a million people in Europe alone. However, the development of SL recognition and translation tools is slowed down by a series of obstacles concerning resource scarcity and standardization issues in the available data. The former challenge relates to the volume of data available for machine learning as well as the time required to collect and process new data. The latter obstacle is linked to the variety of the data, i.e., annotation formats are not unified and vary amongst different resources. The available data formats are often not suitable for machine learning, obstructing the provision of automatic tools based on neural models. In the present paper, we give an overview of these challenges by comparing various SL corpora and SL machine learning datasets. Furthermore, we propose a framework to address the lack of standardization at format level, unify the available resources and facilitate SL research for different languages. Our framework takes ELAN files as inputs and returns textual and visual data ready to train SL recognition and translation models. We present a proof of concept, training neural translation models on the data produced by the proposed framework.
Mirella De Sisto, Vincent Vandeghinste, Santiago Egea Gómez, Mathieu De Coster, Dimitar Sht. Shterionov, Horacio Saggion
LREC6
2022 Linguistically Enhanced Text to Sign Gloss Machine Translation
Santiago Egea Gómez, Luis Chiruzzo, Euan McGill, Horacio Saggion
NLDB4
2021 Controllable Sentence Simplification with a Unified Text-to-Text Transfer Transformer
abstract
Recently, a large pre-trained language model called T5 (A Unified Text-to-Text Transfer Transformer) has achieved state-of-the-art performance in many NLP tasks.However, no study has been found using this pre-trained model on Text Simplification.Therefore in this paper, we explore the use of T5 fine-tuning on Text Simplification combining with a controllable mechanism to regulate the system outputs that can help generate adapted text for different target audiences.Our experiments show that our model achieves remarkable results with gains of between +0.69 and +1.41 over the current state-of-the-art (BART+ACCESS).We argue that using a pre-trained model such as T5, trained on several tasks with large amounts of data, can help improve Text Simplification. 1
Kim Cheng Sheang, Horacio Saggion
INLG2
2020 A Multi-level Annotated Corpus of Scientific Papers for Scientific Document Summarization and Cross-document Relation Discovery
abstract
Related work sections or literature reviews are an essential part of every scientific article being crucial for paper reviewing and assessment. The automatic generation of related work sections can be considered an instance of the multi-document summarization problem. In order to allow the study of this specific problem, we have developed a manually annotated, machine readable data-set of related work sections, cited papers (e.g. references) and sentences, together with an additional layer of papers citing the references. We additionally present experiments on the identification of cited sentences, using as input citation contexts. The corpus alongside the gold standard are made available for use by the scientific community.
Ahmed AbuRa'ed, Horacio Saggion, Luis Chiruzzo
LREC2
2020 Cross-lingual semantic annotation of biomedical literature: experiments in Spanish and English
abstract
MOTIVATION: Biomedical literature is one of the most relevant sources of information for knowledge mining in the field of Bioinformatics. In spite of English being the most widely addressed language in the field; in recent years, there has been a growing interest from the natural language processing community in dealing with languages other than English. However, the availability of language resources and tools for appropriate treatment of non-English texts is lacking behind. Our research is concerned with the semantic annotation of biomedical texts in the Spanish language, which can be considered an under-resourced language where biomedical text processing is concerned. RESULTS: We have carried out experiments to assess the effectiveness of several methods for the automatic annotation of biomedical texts in Spanish. One approach is based on the linguistic analysis of Spanish texts and their annotation using an information retrieval and concept disambiguation approach. A second method takes advantage of a Spanish-English machine translation process to annotate English documents and transfer annotations back to Spanish. A third method takes advantage of the combination of both procedures. Our evaluation shows that a combined system has competitive advantages over the two individual procedures. AVAILABILITY AND IMPLEMENTATION: UMLSMapper (https://snlt.vicomtech.org/umlsmapper) and the annotation transfer tool (http://scientmin.taln.upf.edu/anntransfer/) are freely available for research purposes as web services and/or demos. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Naiara Pérez, Pablo Accuosto, Àlex Bravo, Montse Cuadros, Eva Martínez Garcia, Horacio Saggion, German Rigau
Bioinform.6
2020 Mining arguments in scientific abstracts with discourse-level embeddings
Pablo Accuosto, Horacio Saggion
Data Knowl. Eng.2
2020 MSC+: Language pattern learning for word sense induction and disambiguation
Fábio Bif Goularte, Danielly Sorato, Silvia M. Nassar, Renato Fileto, Horacio Saggion
Knowl. Based Syst.5
2019 Discourse-Driven Argument Mining in Scientific Abstracts
Pablo Accuosto, Horacio Saggion
NLDB2
2019 A text summarization method based on fuzzy rules and applicable to automated assessment
Fábio Bif Goularte, Silvia M. Nassar, Renato Fileto, Horacio Saggion
Expert Syst. Appl.4
2019 Improving lexical coverage of text simplification systems for Spanish
Sanja Stajner, Horacio Saggion, Simone Paolo Ponzetto
Expert Syst. Appl.2
2018 Interpretable Emoji Prediction via Label-Wise Attention LSTMs
abstract
Human language has evolved towards newer forms of communication such as social media, where emojis (i.e., ideograms bearing a visual meaning) play a key role.While there is an increasing body of work aimed at the computational modeling of emoji semantics, there is currently little understanding about what makes a computational model represent or predict a given emoji in a certain way.In this paper we propose a label-wise attention mechanism with which we attempt to better understand the nuances underlying emoji prediction.In addition to advantages in terms of interpretability, we show that our proposed architecture improves over standard baselines in emoji prediction, and does particularly well when predicting infrequent emojis.
Francesco Barbieri, Luis Espinosa Anke, José Camacho-Collados, Steven Schockaert, Horacio Saggion
EMNLP5
2018 PDFdigest: an Adaptable Layout-Aware PDF-to-XML Textual Content Extractor for Scientific Articles
Daniel Ferrés, Horacio Saggion, Francesco Ronzano, Àlex Bravo
LREC2
2017 Characterizing mention mismatching problems for improving recognition results
abstract
Mentions to real world things which are recognized by software tools in text often mismatch the ground truth. This paper proposes a formal classification of mention mismatching problems, including partial matching. Then, it depicts evidence that some longer mentions are associated with higher precision and more specific things than shorter mentions that overlap them. Based on this, some algorithms are proposed to automatically improve mentions by increasing their sizes whenever and as much as possible. Experimental results applying a variety of state-of-the-art annotation tools against several datasets made from real world texts show that over-segmentation (returned mention contained in the corresponding one of the ground truth) is the most prevalent partial matching problem among those of the proposed classification. In addition, some of the proposed algorithms for mention enhancing were able to correct most over-segmented mentions returned by tools used in the experiments with prominent benchmarks, leading to gains in precision and recall.
Jean Carlos Oliveira de Abreu, Renato Fileto, Axel-Cyrille Ngonga Ngomo, Michael Röder, Matthias Wittwer, Horacio Saggion
iiWAS6
2017 Using genre-specific features for patent summaries
Joan Codina, Nadjet Bouayad-Agha, Alicia Burga, Gerard Casamayor, Simon Mille, Andreas Müller 0012, Horacio Saggion, Leo Wanner
Inf. Process. Manag.7
2016 ExTaSem! Extending, Taxonomizing and Semantifying Domain Terminologies
abstract
We introduce ExTaSem!, a novel approach for the automatic learning of lexical taxonomies from domain terminologies. First, we exploit a very large semantic network to collect housands of in-domain textual definitions. Second, we extract (hyponym, hypernym) pairs from each definition with a CRF-based algorithm trained on manually-validated data. Finally, we introduce a graph induction procedure which constructs a full-fledged taxonomy where each edge is weighted according to its domain pertinence. ExTaSem! achieves state-of-the-art results in the following taxonomy evaluation experiments: (1) Hypernym discovery, (2) Reconstructing gold standard taxonomies, and (3) Taxonomy quality according to structural measures. We release weighted taxonomies for six domains for the use and scrutiny of the community.
Luis Espinosa Anke, Horacio Saggion, Francesco Ronzano, Roberto Navigli
AAAI2
2016 Extending WordNet with Fine-Grained Collocational Information via Supervised Distributional Learning
abstract
WordNet is probably the best known lexical resource in Natural Language Processing. While it is widely regarded as a high quality repository of concepts and semantic relations, updating and extending it manually is costly. One important type of relation which could potentially add enormous value to WordNet is the inclusion of collocational information, which is paramount in tasks such as Machine Translation, Natural Language Generation and Second Language Learning. In this paper, we present ColWordNet (CWN), an extended WordNet version with fine-grained collocational information, automatically introduced thanks to a method exploiting linear relations between analogous sense-level embeddings spaces. We perform both intrinsic and extrinsic evaluations, and release CWN for the use and scrutiny of the community.
Luis Espinosa Anke, José Camacho-Collados, Sara Rodríguez-Fernández, Horacio Saggion, Leo Wanner
COLING4
2016 Supervised Distributional Hypernym Discovery via Domain Adaptation
abstract
Comunicació presentada a la Conference on Empirical Methods in Natural Language Processing celebrada els dies 1 a 5 de novembre de 2016 a Austin, Texas.
Luis Espinosa Anke, José Camacho-Collados, Claudio Delli Bovi, Horacio Saggion
EMNLP4
2016 What does this Emoji Mean? A Vector Space Skip-Gram Model for Twitter Emojis
Francesco Barbieri, Francesco Ronzano, Horacio Saggion
LREC3
2016 A Multi-Layered Annotated Corpus of Scientific Papers
Beatríz Fisas, Francesco Ronzano, Horacio Saggion
LREC3
2016 ELMD: An Automatically Generated Entity Linking Gold Standard Dataset in the Music Domain
Sergio Oramas, Luis Espinosa Anke, Mohamed Sordo, Horacio Saggion, Xavier Serra
LREC4
2016 How Cosmopolitan Are Emojis?: Exploring Emojis Usage and Meaning over Different Languages with Distributional Semantics
abstract
Choosing the right emoji to visually complement or condense the meaning of a message has become part of our daily life. Emojis are pictures, which are naturally combined with plain text, thus creating a new form of language. These pictures are the same independently of where we live, but they can be interpreted and used in different ways. In this paper we compare the meaning and the usage of emojis across different languages. Our results suggest that the overall semantics of the subset of the emojis we studied is preserved across all the languages we analysed. However, some emojis are interpreted in a different way from language to language, and this could be related to socio-geographical differences.
Francesco Barbieri, Germán Kruszewski, Francesco Ronzano, Horacio Saggion
ACM Multimedia4
2016 YATS: Yet Another Text Simplifier
Daniel Ferrés, Montserrat Marimon, Horacio Saggion, Ahmed AbuRa'ed
NLDB3
2016 An Empirical Assessment of Citation Information in Scientific Summarization
Francesco Ronzano, Horacio Saggion
NLDB2
2016 Simplifying words in context. Experiments with two lexical resources in Spanish
Horacio Saggion, Stefan Bott, Luz Rello
Comput. Speech Lang.1
2016 Information extraction for knowledge base construction in the music domain
Sergio Oramas, Luis Espinosa Anke, Mohamed Sordo, Horacio Saggion, Xavier Serra
Data Knowl. Eng.4
2015 Hypernym Extraction: Combining Machine-Learning and Dependency Grammar
Luis Espinosa Anke, Francesco Ronzano, Horacio Saggion
CICLing (1)3
2015 Dr. Inventor Framework: Extracting Structured Information from Scientific Publications
Francesco Ronzano, Horacio Saggion
Discovery Science2
2015 Stimulating and Simulating Creativity with Dr Inventor
Diarmuid P. O'Donoghue, Yalemisew M. Abgaz, Donny Hurley, Francesco Ronzano, Horacio Saggion
ICCC5
2015 Do We Criticise (and Laugh) in the Same Way? Automatic Detection of Multi-Lingual Satirical News in Twitter
Francesco Barbieri, Francesco Ronzano, Horacio Saggion
IJCAI3
2014 Modelling Irony in Twitter
abstract
Computational creativity is one of the central research topics of Artificial Intelligence and Natural Language Processing today.Irony, a creative use of language, has received very little attention from the computational linguistics research point of view.In this study we investigate the automatic detection of irony casting it as a classification problem.We propose a model capable of detecting irony in the social network Twitter.In cross-domain classification experiments our model based on lexical features outperforms a word-based baseline previously used in opinion mining and achieves state-of-the-art performance.Our features are simple to implement making the approach easily replicable.
Francesco Barbieri, Horacio Saggion
EACL2
2014 Automatic Detection of Irony and Humour in Twitter
Francesco Barbieri, Horacio Saggion
ICCC2
2014 Towards Dr Inventor: A Tool for Promoting Scientific Creativity
Diarmuid P. O'Donoghue, Horacio Saggion, Donny Hurley, Yalemisew M. Abgaz, Óscar Corcho, Jian J. Zhang 0001, J. M. Careil, Babak Mahdian
ICCC2
2014 Modelling Irony in Twitter: Feature Analysis and Evaluation
Francesco Barbieri, Horacio Saggion
LREC2
2014 Can Numerical Expressions Be Simpler? Implementation and Demostration of a Numerical Simplification System for Spanish
Susana Bautista, Horacio Saggion
LREC2
2014 Creating Summarization Systems with SUMMA
Horacio Saggion
LREC1
2014 Applying Dependency Relations to Definition Extraction
Luis Espinosa Anke, Horacio Saggion
NLDB2
2013 An iOS reader for people with dyslexia
abstract
We present DysWebxia, an eBook reader for iOS which modifies the form and the content of the text. This tool is specifically designed for people with dyslexia according to previous research with this target group. The settings are customizable depending on the reading preferences.
Luz Rello, Ricardo Baeza-Yates, Horacio Saggion, Clara Bayarri, Simone D. J. Barbosa
ASSETS3
2013 Automatic Text Simplification in Spanish: A Comparative Evaluation of Complementing Modules
Biljana Drndarevic, Sanja Stajner, Stefan Bott, Susana Bautista, Horacio Saggion
CICLing (2)5
2013 The Impact of Lexical Simplification by Verbal Paraphrases for People with and without Dyslexia
Luz Rello, Ricardo Baeza-Yates, Horacio Saggion
CICLing (2)3
2013 Readability Indices for Automatic Evaluation of Text Simplification Systems: A Feasibility Study for Spanish
Sanja Stajner, Horacio Saggion
IJCNLP2
2013 One Half or 50%? An Eye-Tracking Study of Number Representation Readability
Luz Rello, Susana Bautista, Ricardo Baeza-Yates, Pablo Gervás, Raquel Hervás, Horacio Saggion
INTERACT (4)6
2013 Frequent Words Improve Readability and Short Words Improve Understandability for People with Dyslexia
Luz Rello, Ricardo Baeza-Yates, Laura Dempere-Marco, Horacio Saggion
INTERACT (4)4
2013 Unsupervised Learning Summarization Templates from Concise Summaries
Horacio Saggion
HLT-NAACL1
2012 Can Spanish Be Simpler? LexSiS: Lexical Simplification for Spanish
Stefan Bott, Luz Rello, Biljana Drndarevic, Horacio Saggion
COLING4
2012 Automatic Simplification of Spanish Text for e-Accessibility
Stefan Bott, Horacio Saggion
ICCHP (1)2
2012 Text Simplification Tools for Spanish
Stefan Bott, Horacio Saggion, Simon Mille
LREC2
2012 The CONCISUS Corpus of Event Summaries
Horacio Saggion, Sandra Szasz
LREC1
2012 From Ontology to NL: Generation of Multilingual User-Oriented Environmental Reports
Nadjet Bouayad-Agha, Gerard Casamayor, Simon Mille, Marco Rospocher, Horacio Saggion, Luciano Serafini, Leo Wanner
NLDB5
2012 Can Text Summaries Help Predict Ratings? A Case Study of Movie Reviews
Horacio Saggion, Elena Lloret, Manuel Palomar
NLDB1
2011 Learning Predicate Insertion Rules for Document Abstracting
Horacio Saggion
CICLing (2)1
2011 Invited Talks
Horacio Saggion
NLDB1
2010 Interpreting SentiWordNet for Opinion Classification
Horacio Saggion, Adam Funk
LREC1
2010 NLP Resources for the Analysis of Patient/Therapist Interviews
Horacio Saggion, Elena Stein-Sparvieri, David Maldavsky, Sandra Szasz
LREC1
2008 Ontology-Driven Human Language Technology for Semantic-Based Business Intelligence
abstract
In this poster submission, we describe the actual state of development of textual analysis and ontology-based information extraction in real world applications, as they are defined in the context of the European R&D project “MUSING” dealing with Business Intelligence. We present in some details the actual state of ontology development, including a time and domain ontologies, which are guiding information extraction onto an ontology population task.
Thierry Declerck, Hans-Ulrich Krieger, Horacio Saggion, Marcus Spies
ECAI3
2008 Experiments on Semantic-based Clustering for Cross-document Coreference
Horacio Saggion
IJCNLP1
2008 Introduction to Text Summarization and Other Information Access Technologies
Horacio Saggion
IJCNLP1
2008 A Framework for Identity Resolution and Merging for Multi-source Information Extraction
Milena Yankova, Horacio Saggion, Hamish Cunningham
LREC2
2006 Multilingual Multidocument Summarization Tools and Evaluation
Horacio Saggion
LREC1
2006 Language Resources for Background Gathering
Horacio Saggion, Robert J. Gaizauskas
LREC1
2006 Indexing and abstracting in theory and practice, third edition
Horacio Saggion
J. Assoc. Inf. Sci. Technol.1
2005 Context-based generic cross-lingual retrieval of documents and automated summaries
abstract
Abstract We develop a context‐based generic cross‐lingual retrieval model that can deal with different language pairs. Our model considers contexts in the query translation process. Contexts in the query as well as in the documents based on co‐occurrence statistics from different granularity of passages are exploited. We also investigate cross‐lingual retrieval of automatic generic summaries. We have implemented our model for two different cross‐lingual settings, namely, retrieving Chinese documents from English queries as well as retrieving English documents from Chinese queries. Extensive experiments have been conducted on a large‐scale parallel corpus enabling studies on retrieval performance for two different cross‐lingual settings of full‐length documents as well as automated summaries.
Wai Lam, Ki Chan, Dragomir R. Radev, Horacio Saggion, Simone Teufel
J. Assoc. Inf. Sci. Technol.4
2004 MEAD - A Platform for Multidocument Multilingual Text Summarization
Dragomir R. Radev, Timothy Allison, Sasha Blair-Goldensohn, John Blitzer, Arda Çelebi, Stanko Dimitrov, Elliott Drábek, Ali Hakim, Wai Lam, Danyu Liu, Jahna Otterbacher, Horacio Saggion, Simone Teufel, Michael Topper, Adam Winkel
LREC13
2004 Identifying Definitions in Text Collections for Question Answering
Horacio Saggion
LREC1
2004 Multimedia indexing through multi-source and multi-language information extraction: the MUMIS project
Horacio Saggion, Hamish Cunningham, Kalina Bontcheva, Diana Maynard, Oana Hamza, Yorick Wilks
Data Knowl. Eng.1
2003 Evaluation Challenges in Large-Scale Document Summarization
abstract
We present a large-scale meta evaluation of eight evaluation measures for both single-document and multi-document summarizers. To this end we built a corpus consisting of (a) 100 Million automatic summaries using six summarizers and baselines at ten summary lengths in both English and Chinese, (b) more than 10,000 manual abstracts and extracts, and (c) 200 Million automatic document and summary retrievals using 20 queries. We present both qualitative and quantitative results showing the strengths and draw-backs of all evaluation methods and how they rank the different summarizers.
Dragomir R. Radev, Simone Teufel, Horacio Saggion, Wai Lam, John Blitzer, Arda Çelebi, Danyu Liu, Elliott Drábek
ACL3
2003 Using Natural Language Processing for Semantic Indexing of Scene-of-Crime Photographs
Horacio Saggion, Katerina Pastra, Yorick Wilks
CICLing1
2003 Robust Generic and Query-based Summarization
Horacio Saggion, Kalina Bontcheva, Hamish Cunningham
EACL1
2003 Event-Coreference across Multiple, Multi-lingual Sources in the Mumis Project
Horacio Saggion, Hamish Cunningham, Jan Kuper, Thierry Declerck, Peter Wittenburg
EACL1
2003 NLP for Indexing and Retrieval of Captioned Photographs
Horacio Saggion, Katerina Pastra, Yorick Wilks
EACL1
2003 Intelligent Multimedia Indexing and Retrieval through Multi-source Information Extraction and Merging
Jan Kuper, Horacio Saggion, Hamish Cunningham, Thierry Declerck, Franciska de Jong, Dennis Reidsma, Yorick Wilks, Peter Wittenburg
IJCAI2
2003 Extracting relational facts for indexing and retrieval of crime-scene photographs
Katerina Pastra, Horacio Saggion, Yorick Wilks
Knowl. Based Syst.2
2002 Meta-evaluation of Summaries in a Cross-lingual Environment using Content-based Metrics
Horacio Saggion, Dragomir R. Radev, Simone Teufel, Wai Lam
COLING1
2002 Extracting Information for Automatic Indexing of Multimedia Material
Horacio Saggion, Hamish Cunningham, Diana Maynard, Kalina Bontcheva, Oana Hamza, Cristian Ursu, Yorick Wilks
LREC1
2002 Developing Infrastructure for the Evaluation of Single and Multi-document Summarization Systems in a Cross-lingual Environment
Horacio Saggion, Dragomir R. Radev, Simone Teufel, Wai Lam, Stephanie M. Strassel
LREC1
2002 Access to Multimedia Information through Multisource and Multilanguage Information Extraction
Horacio Saggion, Hamish Cunningham, Kalina Bontcheva, Diana Maynard, Cristian Ursu, Oana Hamza, Yorick Wilks
NLDB1
2002 Generating Indicative-Informative Summaries with SumUM
abstract
We present and evaluate SumUM, a text summarization system that takes a raw technical text as input and produces an indicative informative summary. The indicative part of the summary identifies the topics of the document, and the informative part elaborates on some of these topics according to the reader's interest. SumUM motivates the topics, describes entities, and defines concepts. It is a first step for exploring the issue of dynamic summarization. This is accomplished through a process of shallow syntactic and semantic analysis, concept identification, and text regeneration. Our method was developed through the study of a corpus of abstracts written by professional abstractors. Relying on human judgment, we have evaluated indicativeness, informativeness, and text acceptability of the automatic summaries. The results thus far indicate good performance when compared with other summarization technologies.
Horacio Saggion, Guy Lapalme
Comput. Linguistics1
2002 Architectural elements of language engineering robustness
abstract
We discuss robustness in LE systems from the perspective of engineering, and the predictability of both outputs and construction process that this entails. We present an architectural system that contributes to engineering robustness and low-overhead systems development (GATE, a General Architecture for Text Engineering). To verify our ideas we present results from the development of a multi-purpose cross-genre Named Entity recognition system. This system aims be robust across diverse input types, and to reduce the need for costly and timeconsuming adaptation of systems to new applications, with its capability to process texts from widely differing domains and genres.
Diana Maynard, Valentin Tablan, Hamish Cunningham, Cristian Ursu, Horacio Saggion, Kalina Bontcheva, Yorick Wilks
Nat. Lang. Eng.5
1999 Using Linguistic Knowledge in Automatic Abstracting
abstract
We present work on the automatic generation of short indicative-informative abstracts of scientific and technical articles. The indicative part of the abstract identifies the topics of the document while the informative part of the abstract elaborate some topics according to the reader's interest by motivating the topics, describing entities and defining concepts. We have defined our method of automatic abstracting by studying a corpus professional abstracts. The method also considers the reader's interest as essential in the process of abstracting.
Horacio Saggion
ACL1