VLDB 2026 Research / reviewers in the wild / expert
Leo Wanner
dblp:63/3922
· DBLP profile ↗
82ranked-venue papers
14as first author
11since 2021 · last 2025
0000-0002-9446-3748ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 72 · 14 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 6 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Exploring morphology-aware tokenization: A case study on Spanish language modelingabstractThis paper investigates to what extent the integration of morphological information can improve subword tokenization and thus also language modeling performance.We focus on Spanish, a language with fusional morphology, where subword segmentation can benefit from linguistic structure.Instead of relying on purely data-driven strategies like Byte Pair Encoding (BPE), we explore a linguistically grounded approach: training a tokenizer on morphologically segmented data.To do so, we develop a semi-supervised segmentation model for Spanish, building gold-standard datasets to guide and evaluate it.We then use this tokenizer to pre-train a masked language model and assess its performance on several downstream tasks.Our results show improvements over a baseline with a standard tokenizer, supporting our hypothesis that morphology-aware tokenization offers a viable and principled alternative for improving language modeling. Alba Táboas García, Piotr Przybyla, Leo Wanner |
EMNLP | 3 |
| 2025 | What the #?*!: Disentangling Hate Across Target IdentitiesabstractYiping Jin, Leo Wanner, Aneesh Moideen Koya. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yiping Jin, Leo Wanner, Aneesh Moideen Koya |
NAACL (Long Papers) | 2 |
| 2024 | GPT-HateCheck: Can LLMs Write Better Functional Tests for Hate Speech Detection?abstractOnline hate detection suffers from biases incurred in data sampling, annotation, and model pre-training. Therefore, measuring the averaged performance over all examples in held-out test data is inadequate. Instead, we must identify specific model weaknesses and be informed when it is more likely to fail. A recent proposal in this direction is HateCheck, a suite for testing fine-grained model functionalities on synthesized data generated using templates of the kind “You are just a [slur] to me.” However, despite enabling more detailed diagnostic insights, the HateCheck test cases are often generic and have simplistic sentence structures that do not match the real-world data. To address this limitation, we propose GPT-HateCheck, a framework to generate more diverse and realistic functional tests from scratch by instructing large language models (LLMs). We employ an additional natural language inference (NLI) model to verify the generations. Crowd-sourced annotation demonstrates that the generated test cases are of high quality. Using the new functional tests, we can uncover model weaknesses that would be overlooked using the original HateCheck dataset. Yiping Jin, Leo Wanner, Alexander V. Shvets |
LREC/COLING | 2 |
| 2024 | Using Large Language Models and Recruiter Expertise for Optimized Multilingual Job Offer - Applicant CV Matching
Hamit Kavas, Marc Serra-Vidal, Leo Wanner |
IJCAI | 3 |
| 2022 | Directions for NLP Practices Applied to Online Hate Speech DetectionabstractAddressing hate speech in online spaces has been conceptualized as a classification task that uses Natural Language Processing (NLP) techniques.Through this conceptualization, the hate speech detection task has relied on common conventions and practices from NLP.For instance, inter-annotator agreement is conceptualized as a way to measure dataset quality and certain metrics and benchmarks are used to assure model generalization.However, hate speech is a deeply complex and situated concept that eludes such static and disembodied practices.In this position paper, we critically reflect on these methodologies for hate speech detection, we argue that many conventions in NLP are poorly suited for the problem and encourage researchers to develop methods that are more appropriate for the task. Paula Fortuna, Mónica Domínguez, Leo Wanner, Zeerak Talat |
EMNLP | 3 |
| 2022 | Social Media and Web Sensing on Interior and Urban DesignabstractSocial media and web sites provide an access to pub-lic opinions on certain aspects and therefore play an important role in getting insights on targeted audiences. Designers have been investigating how to use them for grasping social feelings and needs associated with the arrangement of spaces that surround people in everyday life to find inspiration and come up with ideas for adaptive designs. Following this, we propose a novel design-oriented tool-set that retrieves and analyses online public information from Twitter and focused content from relevant web sites. We present the data collection pipeline and multilingual analysis algorithms like concept extraction and sentiment analysis on interior and urban design. Finally, we showcase an application of the proposed tool-set within two case studies. Evangelos A. Stathopoulos, Alexander V. Shvets, Roberto Carlini, Sotiris Diplaris, Stefanos Vrochidis, Leo Wanner, Ioannis Kompatsiaris |
ISCC | 6 |
| 2021 | Evaluating language models for the retrieval and categorization of lexical collocationsabstractComunicació presentada a: EACL 2021 celebrat del 19 a 23 d'abril de 2021 en línia. Luis Espinosa Anke, Joan Codina, Leo Wanner |
EACL | 3 |
| 2021 | On the evolution of syntactic information encoded by BERT's contextualized representationsabstractThe adaptation of pretrained language models to solve supervised tasks has become a baseline in NLP, and many recent works have focused on studying how linguistic information is encoded in the pretrained sentence representations.Among other information, it has been shown that entire syntax trees are implicitly embedded in the geometry of such models.As these models are often fine-tuned, it becomes increasingly important to understand how the encoded knowledge evolves along the fine-tuning.In this paper, we analyze the evolution of the embedded syntax trees along the fine-tuning process of BERT for six different tasks, covering all levels of the linguistic structure.Experimental results show that the encoded syntactic information is forgotten (PoS tagging), reinforced (dependency and constituency parsing) or preserved (semanticsrelated tasks) in different ways along the finetuning process depending on the task. Laura Pérez-Mayos, Roberto Carlini, Miguel Ballesteros, Leo Wanner |
EACL | 4 |
| 2021 | How much pretraining data do language models need to learn syntax?abstractTransformers-based pretrained language models achieve outstanding results in many wellknown NLU benchmarks.However, while pretraining methods are very convenient, they are expensive in terms of time and resources.This calls for a study of the impact of pretraining data size on the knowledge of the models.We explore this impact on the syntactic capabilities of RoBERTa, using models trained on incremental sizes of raw text data.First, we use syntactic structural probes to determine whether models pretrained on more data encode a higher amount of syntactic information.Second, we perform a targeted syntactic evaluation to analyze the impact of pretraining data size on the syntactic generalization performance of the models.Third, we compare the performance of the different models on three downstream applications: part-of-speech tagging, dependency parsing and paraphrase identification.We complement our study with an analysis of the cost-benefit trade-off of training such models.Our experiments show that while models pretrained on more data encode more syntactic knowledge and perform better on downstream applications, they do not always offer a better performance across the different syntactic phenomena and come at a higher financial and environmental cost. Laura Pérez-Mayos, Miguel Ballesteros, Leo Wanner |
EMNLP (1) | 3 |
| 2021 | ThemePro 2.0: Showcasing the Role of Thematic Progression in Engaging Human-Computer Interaction
Mónica Domínguez, Juan Soler Company, Leo Wanner |
Interspeech | 3 |
| 2021 | How well do hate speech, toxicity, abusive and offensive language classification models generalize across datasets?abstractA considerable body of research deals with the automatic identification of hate speech and related phenomena. However, cross-dataset model generalization remains a challenge. In this context, we address two still open central questions: (i) to what extent does the generalization depend on the model and the composition and annotation of the training data in terms of different categories?, and (ii) do specific features of the datasets or models influence the generalization potential? To answer (i), we experiment with BERT, ALBERT, fastText, and SVM models trained on nine common public English datasets, whose class (or category) labels are standardized (and thus made comparable), in intra- and cross-dataset setups. The experiments show that indeed the generalization varies from model to model and that some of the categories (e.g., ‘toxic’, ‘abusive’, or ‘offensive’) serve better as cross-dataset training categories than others (e.g., ‘hate speech’). To answer (ii), we use a Random Forest model for assessing the relevance of different model and dataset features during the prediction of the performance of 450 BERT, 450 ALBERT, 450 fastText, and 348 SVM binary abusive language classifiers (1698 in total). We find that in order to generalize well, a model already needs to perform well in an intra-dataset scenario. Furthermore, we find that some other parameters are equally decisive for the success of the generalization, including, e.g., the training and target categories and the percentage of the out-of-domain vocabulary. Paula Fortuna, Juan Soler Company, Leo Wanner |
Inf. Process. Manag. | 3 |
| 2020 | Concept Extraction Using Pointer-Generator Networks and Distant Supervision for Data Augmentation
Alexander V. Shvets, Leo Wanner |
EKAW | 2 |
| 2020 | ThemePro: A Toolkit for the Analysis of Thematic ProgressionabstractThis paper introduces ThemePro, a toolkit for the automatic analysis of thematic progression. Thematic progression is relevant to natural language processing (NLP) applications dealing, among others, with discourse structure, argumentation structure, natural language generation, summarization and topic detection. A web platform demonstrates the potential of this toolkit and provides a visualization of the results including syntactic trees, hierarchical thematicity over propositions and thematic progression over whole texts. Mónica Domínguez, Juan Soler Company, Leo Wanner |
LREC | 3 |
| 2020 | Toxic, Hateful, Offensive or Abusive? What Are We Really Classifying? An Empirical Analysis of Hate Speech DatasetsabstractThe field of the automatic detection of hate speech and related concepts has raised a lot of interest in the last years. Different datasets were annotated and classified by means of applying different machine learning algorithms. However, few efforts were done in order to clarify the applied categories and homogenize different datasets. Our study takes up this demand. We analyze six different publicly available datasets in this field with respect to their similarity and compatibility. We conduct two different experiments. First, we try to make the datasets compatible and represent the dataset classes as Fast Text word vectors analyzing the similarity between different classes in a intra and inter dataset manner. Second, we submit the chosen datasets to the Perspective API Toxicity classifier, achieving different performances depending on the categories and datasets. One of the main conclusions of these experiments is that many different definitions are being used for equivalent concepts, which makes most of the publicly available datasets incompatible. Grounded in our analysis, we provide guidelines for future dataset collection and annotation. Paula Fortuna, Juan Soler Company, Leo Wanner |
LREC | 3 |
| 2019 | Collocation Classification with Unsupervised Relation VectorsabstractLexical relation classification is the task of predicting whether a certain relation holds between a given pair of words.In this paper, we explore to which extent the current distributional landscape based on word embeddings provides a suitable basis for classification of collocations, i.e., pairs of words between which idiosyncratic lexical relations hold.First, we introduce a novel dataset with collocations categorized according to lexical functions.Second, we conduct experiments on a subset of this benchmark, comparing it in particular to the well known DiffVec dataset.In these experiments, in addition to simple word vector arithmetic operations, we also investigate the role of unsupervised relation vectors as a complementary input.While these relation vectors indeed help, we also show that lexical function classification poses a greater challenge than the syntactic and semantic relations that are typically used for benchmarks in the literature. Luis Espinosa Anke, Steven Schockaert, Leo Wanner |
ACL (1) | 3 |
| 2019 | Teaching FORGe to Verbalize DBpedia Properties in SpanishabstractStatistical generators increasingly dominate the research in NLG.However, grammarbased generators that are grounded in a solid linguistic framework remain very competitive, especially for generation from deep knowledge structures.Furthermore, if built modularly, they can be ported to other genres and languages with a limited amount of work, without the need of the annotation of a considerable amount of training data.One of these generators is FORGe, which is based on the Meaning-Text Model.In the recent WebNLG challenge (the first comprehensive task addressing the mapping of RDF triples to text) FORGe ranked first with respect to the overall quality in human evaluation.We extend the coverage of FORGE's open source grammatical and lexical resources for English, so as to further improve the English outcome, and port them to Spanish, to achieve a comparable quality.This confirms that, as already observed in the case of SimpleNLG, a robust universal grammar-driven framework and a systematic organization of the linguistic resources can be an adequate choice for NLG applications. Simon Mille, Stamatia Dasiopoulou, Beatríz Fisas, Leo Wanner |
INLG | 4 |
| 2019 | Automatic Classification and Linguistic Analysis of Extremist Online Material
Juan Soler Company, Leo Wanner |
MMM (2) | 2 |
| 2018 | Automatic Identification of Texts Written by Authors with Alzheimer's Disease
Juan Soler Company, Leo Wanner |
CogSci | 2 |
| 2018 | Underspecified Universal Dependency Structures as Inputs for Multilingual Surface RealisationabstractIn this paper, we present the datasets used in the Shallow and Deep Tracks of the First Multilingual Surface Realisation Shared Task (SR'18).For the Shallow Track, data in ten languages has been released: Arabic, Czech, Dutch, English, Finnish, French, Italian, Portuguese, Russian and Spanish.For the Deep Track, data in three languages is made available: English, French and Spanish.We describe in detail how the datasets were derived from the Universal Dependencies V2.0, and report on an evaluation of the Deep Track input quality.In addition, we examine the motivation for, and likely usefulness of, deriving NLG inputs from annotations in resources originally developed for Natural Language Understanding (NLU), and assess whether the resulting inputs supply enough information of the right kind for the final stage in the NLG process. Simon Mille, Anya Belz, Bernd Bohnet, Leo Wanner |
INLG | 4 |
| 2018 | Sentence Packaging in Text Generation from Semantic Graphs as a Community Detection ProblemabstractAn increasing amount of research tackles the challenge of text generation from abstract ontological or semantic structures, which are in their very nature potentially large connected graphs.These graphs must be "packaged" into sentence-wise subgraphs.We interpret the problem of sentence packaging as a community detection problem with post optimization.Experiments on the texts of the Verb-Net/FrameNet structure annotated-Penn Treebank, which have been converted into graphs by a coreference merge using Stanford CoreNLP, show a high F 1 -score of 0.738. Alexander V. Shvets, Simon Mille, Leo Wanner |
INLG | 3 |
| 2018 | Compilation of Corpora for the Study of the Information Structure-Prosody Interface
Alicia Burga, Mónica Domínguez, Mireia Farrús, Leo Wanner |
LREC | 4 |
| 2018 | Generation of a Spanish Artificial Collocation Error Corpus
Sara Rodríguez-Fernández, Roberto Carlini, Leo Wanner |
LREC | 3 |
| 2018 | Improving the Quality of Video-to-Language Models by Optimizing Annotation of the Training Material
Laura Pérez-Mayos, Federico Sukno, Leo Wanner |
MMM (1) | 3 |
| 2018 | On the role of syntactic dependencies and discourse relations for author and gender identification
Juan Soler Company, Leo Wanner |
Pattern Recognit. Lett. | 2 |
| 2017 | Shared Task Proposal: Multilingual Surface Realization Using Universal Dependency TreesabstractWe propose a shared task on multilingual Surface Realization, i.e., on mapping unordered and uninflected universal dependency trees to correctly ordered and inflected sentences in a number of languages.A second deeper input will be available in which, in addition, functional words, fine-grained PoS and morphological information will be removed from the input trees.The first shared task on Surface Realization was carried out in 2011 with a similar setup, with a focus on English.We think that it is time for relaunching such a shared task effort in view of the arrival of Universal Dependencies annotated treebanks for a large number of languages on the one hand, and the increasing dominance of Deep Learning, which proved to be a game changer for NLP, on the other hand. Simon Mille, Bernd Bohnet, Leo Wanner, Anya Belz |
INLG | 3 |
| 2017 | A demo of FORGe: the Pompeu Fabra Open Rule-based GeneratorabstractThis demo paper presents the multilingual deep sentence generator developed by the TALN group at Universitat Pompeu Fabra, implemented as a series of rule-based graphtransducers. Simon Mille, Leo Wanner |
INLG | 2 |
| 2017 | A Thematicity-Based Prosody Enrichment Tool for CTS
Mónica Domínguez, Mireia Farrús, Leo Wanner |
INTERSPEECH | 3 |
| 2017 | Using Prosody to Classify Discourse RelationsabstractComunicació presentada a: The 18th Annual Conference of the International Speech Communication Association (INTERSPEECH 2017), celebrada a Estocolm, Suència, del 20 al 24 d'agost de 2017. Janine Kleinhans, Mireia Farrús, Agustín Gravano, Juan Manuel Pérez, Catherine Lai, Leo Wanner |
INTERSPEECH | 6 |
| 2017 | Prosograph: A Tool for Prosody Visualisation of Large Speech Corpora
Alp Öktem, Mireia Farrús, Leo Wanner |
INTERSPEECH | 3 |
| 2017 | Towards Reasoned Modality Selection in an Embodied Conversation Agent
Carla Ten-Ventura, Roberto Carlini, Stamatia Dasiopoulou, Gerard Llorach Tó, Leo Wanner |
IVA | 5 |
| 2017 | Using genre-specific features for patent summaries
Joan Codina, Nadjet Bouayad-Agha, Alicia Burga, Gerard Casamayor, Simon Mille, Andreas Müller 0012, Horacio Saggion, Leo Wanner |
Inf. Process. Manag. | 8 |
| 2016 | Extending WordNet with Fine-Grained Collocational Information via Supervised Distributional LearningabstractWordNet is probably the best known lexical resource in Natural Language Processing. While it is widely regarded as a high quality repository of concepts and semantic relations, updating and extending it manually is costly. One important type of relation which could potentially add enormous value to WordNet is the inclusion of collocational information, which is paramount in tasks such as Machine Translation, Natural Language Generation and Second Language Learning. In this paper, we present ColWordNet (CWN), an extended WordNet version with fine-grained collocational information, automatically introduced thanks to a method exploiting linear relations between analogous sense-level embeddings spaces. We perform both intrinsic and extrinsic evaluations, and release CWN for the use and scrutiny of the community. Luis Espinosa Anke, José Camacho-Collados, Sara Rodríguez-Fernández, Horacio Saggion, Leo Wanner |
COLING | 5 |
| 2016 | An Automatic Prosody Tagger for Spontaneous SpeechabstractSpeech prosody is known to be central in advanced communication technologies. However, despite the advances of theoretical studies in speech prosody, so far, no large scale prosody annotated resources that would facilitate empirical research and the development of empirical computational approaches are available. This is to a large extent due to the fact that current common prosody annotation conventions offer a descriptive framework of intonation contours and phrasing based on labels. This makes it difficult to reach a satisfactory inter-annotator agreement during the annotation of gold standard annotations and, subsequently, to create consistent large scale annotations. To address this problem, we present an annotation schema for prominence and boundary labeling of prosodic phrases based upon acoustic parameters and a tagger for prosody annotation at the prosodic phrase level. Evaluation proves that inter-annotator agreement reaches satisfactory values, from 0.60 to 0.80 Cohen’s kappa, while the prosody tagger achieves acceptable recall and f-measure figures for five spontaneous samples used in the evaluation of monologue and dialogue formats in English and Spanish. The work presented in this paper is a first step towards a semi-automatic acquisition of large corpora for empirical prosodic analysis. Mónica Domínguez, Mireia Farrús, Leo Wanner |
COLING | 3 |
| 2016 | A Neural Network Architecture for Multilingual Punctuation GenerationabstractEven syntactically correct sentences are perceived as awkward if they do not contain correct punctuation.Still, the problem of automatic generation of punctuation marks has been largely neglected for a long time.We present a novel model that introduces punctuation marks into raw text material with transition-based algorithm using LSTMs.Unlike the state-of-the-art approaches, our model is language-independent and also neutral with respect to the intended use of the punctuation.Multilingual experiments show that it achieves high accuracy on the full range of punctuation marks across languages. Miguel Ballesteros, Leo Wanner |
EMNLP | 2 |
| 2016 | Towards Multiple Antecedent Coreference Resolution in Specialized Discourse
Alicia Burga, Sergio Cajal, Joan Codina, Leo Wanner |
LREC | 4 |
| 2016 | Example-based Acquisition of Fine-grained Collocation Resources
Sara Rodríguez-Fernández, Roberto Carlini, Luis Espinosa Anke, Leo Wanner |
LREC | 4 |
| 2016 | A Semi-Supervised Approach for Gender Identification
Juan Soler Company, Leo Wanner |
LREC | 2 |
| 2016 | Environmental data extraction from heatmaps using the AirMerge system
Victor Epitropou, Anastasios Bassoukos, Kostas D. Karatzas, Ari Karppinen, Leo Wanner, Stefanos Vrochidis, Ioannis Kompatsiaris, Jaakko Kukkonen |
Multim. Tools Appl. | 5 |
| 2016 | Data-driven deep-syntactic dependency parsingabstractAbstract ‘Deep-syntactic’ dependency structures that capture the argumentative, attributive and coordinative relations between full words of a sentence have a great potential for a number of NLP-applications. The abstraction degree of these structures is in between the output of a syntactic dependency parser (connected trees defined over all words of a sentence and language-specific grammatical functions) and the output of a semantic parser (forests of trees defined over individual lexemes or phrasal chunks and abstract semantic role labels which capture the frame structures of predicative elements and drop all attributive and coordinative dependencies). We propose a parser that provides deep-syntactic structures. The parser has been tested on Spanish, English and Chinese. Miguel Ballesteros, Bernd Bohnet, Simon Mille, Leo Wanner |
Nat. Lang. Eng. | 4 |
| 2015 | Multiple Language Gender Identification for Blog Posts
Juan Soler Company, Leo Wanner |
CogSci | 2 |
| 2015 | Data-driven sentence generation with non-isomorphic treesabstractMiguel Ballesteros, Bernd Bohnet, Simon Mille, Leo Wanner. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Miguel Ballesteros, Bernd Bohnet, Simon Mille, Leo Wanner |
HLT-NAACL | 4 |
| 2015 | Visualizing Deep-Syntactic Parser OutputabstractJuan Soler-Company, Miguel Ballesteros, Bernd Bohnet, Simon Mille, Leo Wanner. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations. 2015. Juan Soler Company, Miguel Ballesteros, Bernd Bohnet, Simon Mille, Leo Wanner |
HLT-NAACL | 5 |
| 2015 | Getting the environmental information across: from the Web to the userabstractAbstract Environmental and meteorological conditions are of utmost importance for the population, as they are strongly related to the quality of life. Citizens are increasingly aware of this importance. This awareness results in an increasing demand for environmental information tailored to their specific needs and background. We present an environmental information platform that supports submission of user queries related to environmental conditions and orchestrates results from complementary services to generate personalized suggestions. The system discovers and processes reliable data in the Web in order to convert them into knowledge. At runtime, this information is transferred into an ontology‐structured knowledge base, from which then information relevant to the specific user is deduced and communicated in the language of their preference. The platform is demonstrated with real world use cases in the south area of Finland, showing the impact it can have on the quality of everyday life. Leo Wanner, Harald Bosch, Nadjet Bouayad-Agha, Gerard Casamayor, Thomas Ertl, Désirée Hilbring, Lasse Johansson, Kostas D. Karatzas, Ari Karppinen, Ioannis Kompatsiaris, Tarja Koskentalo, Simon Mille, Jürgen Moßgraber, Anastasia Moumtzidou, Maria Myllynen, Emanuele Pianta, Marco Rospocher, Luciano Serafini, Virpi Tarvainen, Sara Tonelli, Stefanos Vrochidis |
Expert Syst. J. Knowl. Eng. | 1 |
| 2015 | Ontology-centered environmental information delivery for personalized decision support
Leo Wanner, Marco Rospocher, Stefanos Vrochidis, Lasse Johansson, Nadjet Bouayad-Agha, Gerard Casamayor, Ari Karppinen, Ioannis Kompatsiaris, Simon Mille, Anastasia Moumtzidou, Luciano Serafini |
Expert Syst. Appl. | 1 |
| 2014 | Deep-Syntactic Parsing
Miguel Ballesteros, Bernd Bohnet, Simon Mille, Leo Wanner |
COLING | 4 |
| 2014 | Classifiers for data-driven deep sentence generationabstractState-of-the-art statistical sentence gener-ators deal with isomorphic structures only. Therefore, given that semantic and syntac-tic structures tend to differ in their topol-ogy and number of nodes, i.e., are not iso-morphic, statistical generation saw so far itself confined to shallow, syntactic gener-ation. In this paper, we present a series of fine-grained classifiers that are essen-tial for data-driven deep sentence genera-tion in that they handle the problem of the projection of non-isomorphic structures. 1 Miguel Ballesteros, Simon Mille, Leo Wanner |
INLG | 3 |
| 2014 | An Exercise in Reuse of Resources: Adapting General Discourse Coreference Resolution for Detecting Lexical Chains in Patent Documentation
Nadjet Bouayad-Agha, Alicia Burga, Gerard Casamayor, Joan Codina, Rogelio Nazar, Leo Wanner |
LREC | 6 |
| 2014 | How to Use less Features and Reach Better Performance in Author Gender Identification
Juan Soler Company, Leo Wanner |
LREC | 2 |
| 2013 | The fusion of meteorological- and air quality information for orchestrated services using environmental profiling
Lasse Johansson, Ari Karppinen, Leo Wanner |
FUSION | 3 |
| 2013 | Towards the Annotation of Penn TreeBank with Information Structure
Bernd Bohnet, Alicia Burga, Leo Wanner |
IJCNLP | 3 |
| 2012 | The Surface Realisation Task: Recent Developments and Future Plans
Anya Belz, Bernd Bohnet, Simon Mille, Leo Wanner, Michael White 0001 |
INLG | 4 |
| 2012 | Content Selection From Semantic Web Data
Nadjet Bouayad-Agha, Gerard Casamayor, Leo Wanner, Chris Mellish |
INLG | 3 |
| 2012 | Towards a Surface Realization-Oriented Corpus Annotation
Leo Wanner, Simon Mille, Bernd Bohnet |
INLG | 1 |
| 2012 | From Ontology to NL: Generation of Multilingual User-Oriented Environmental Reports
Nadjet Bouayad-Agha, Gerard Casamayor, Simon Mille, Marco Rospocher, Horacio Saggion, Luciano Serafini, Leo Wanner |
NLDB | 7 |
| 2011 | FootbOWL: Using a Generic Ontology of Football Competition for Planning Match Summaries
Nadjet Bouayad-Agha, Gerard Casamayor, Leo Wanner, Fernando Díez, Sergio López Hernández |
ESWC (1) | 3 |
| 2011 | Towards the derivation of verbal content relations from patent claims using deep syntactic structures
Gabriela Ferraro, Leo Wanner |
Knowl. Based Syst. | 2 |
| 2010 | Broad Coverage Multilingual Deep Sentence Generation with a Stochastic Multi-Level Realizer
Bernd Bohnet, Leo Wanner, Simon Mille, Alicia Burga |
COLING | 2 |
| 2010 | Open Soucre Graph Transducer Interpreter and Grammar Development Environment
Bernd Bohnet, Leo Wanner |
LREC | 2 |
| 2010 | Syntactic Dependencies for Multilingual and Multilevel Corpus Annotation
Simon Mille, Leo Wanner |
LREC | 2 |
| 2010 | Towards a Motivated Annotation Schema of Collocation Errors in Learner Corpora
Margarita Alonso Ramos, Leo Wanner, Orsolya Vincze, Gerard Casamayor, Nancy Vázquez Veiga, Estela Mosqueira Suárez, Sabela Prieto González |
LREC | 2 |
| 2009 | Improving the comprehension of legal documentation: the case of patent claimsabstractWith their abstract vocabulary and overly long sentences, patent claims, like several other genres of legal discourse, are notoriously difficult to read and comprehend. The enormous number of both native and non-native users reading patent claims on a daily basis raises the demand for means that make them easier and faster to understand. An obvious way to satisfy this demand is to paraphrase the original material, i.e., to rewrite it in a more appropriate style, or---even better---to summarize it in the language of preference of the reader such that the reader can rapidly grasp its essence. PATExpert is a patent processing service which incorporates, among other technologies, paraphrasing and multilingual summarization of patent claims. With the goal to offer the user the most suitable options and to evaluate alternative techniques that are based on different contextual and linguistic criteria, both paraphrasing and summarization implement "surface-oriented" strategies and "deep" strategies. The surface strategies make use of shallow linguistic criteria such as punctuation and syntactic and lexical markers. The deep strategies operate on deep-syntactic structures of the claims, using a full fledged text generator for synthesis of the paraphrase or summary, respectively. Nadjet Bouayad-Agha, Gerard Casamayor, Gabriela Ferraro, Simon Mille, Vanesa Vidal, Leo Wanner |
ICAIL | 6 |
| 2008 | Multilingual summarization in practice: the case of patent claims
Simon Mille, Leo Wanner |
EAMT | 2 |
| 2008 | Two-step flow in bilingual lexicon extraction from unrelated corpora
Rogelio Nazar, Leo Wanner, Jorge Vivaldi |
EAMT | 2 |
| 2008 | Making Text Resources Accessible to the Reader: the Case of Patent Claims
Simon Mille, Leo Wanner |
LREC | 2 |
| 2008 | Using Semantically Annotated Corpora to Build Collocation Resources
Margarita Alonso Ramos, Owen Rambow, Leo Wanner |
LREC | 3 |
| 2008 | Morphological mismatches in machine translation
Igor Mel'cuk, Leo Wanner |
Mach. Transl. | 2 |
| 2007 | A Modular Framework for Ontology-based Representation of Patent Information
Mark Giereth, Steffen Koch 0001, Ioannis Kompatsiaris, Symeon Papadopoulos, Emanuele Pianta, Luciano Serafini, Leo Wanner |
JURIX | 7 |
| 2006 | Local Document Relevance Clustering in IR Using Collocation Information
Leo Wanner, Margarita Alonso Ramos |
LREC | 1 |
| 2006 | Making sense of collocations
Leo Wanner, Bernd Bohnet, Mark Giereth |
Comput. Speech Lang. | 1 |
| 2006 | Syntactic mismatches in machine translation
Igor Mel'cuk, Leo Wanner |
Mach. Transl. | 2 |
| 2004 | Enriching the Spanish EuroWordNet by Collocations
Leo Wanner, Margarita Alonso Ramos, Maria Antònia Martí |
LREC | 1 |
| 2004 | Towards automatic fine-grained semantic classification of verb-noun collocationsabstractPlain lists of collocations as provided to date by most approaches to automatic acquisition of collocations from corpora are useful as a resource for dictionary construction. However, their use is rather limited in the case of NLP-applications such as Text Generation, Machine Translation and Text Summarization if not enriched by information on the grammatical function of the collocation elements and by information on the semantics of the collocations as multiword units. In this article, we describe an approach to a fine-grained classification of verb-noun bigrams according to a semantically motivated typology of collocations and illustrate this with Spanish material. The typology of collocations that underlies our classification is based on verb-noun Lexical Functions (LFs) from the Explanatory Combinatorial Lexicology. In the first stage of the approach, the program learns the semantic features of each LF from training data. In the second stage, it examines the semantic features of verb-noun candidate bigrams and compares them with the features of all the LFs taken into account. A candidate whose features are sufficiently similar to those of a specific LF is considered to be an instance of this LF. The semantic features of both the training material and the candidate bigrams are derived from the hyperonymy hierarchies provided by the EuroWordNet. In the experiments carried out to validate the approach, we achieved an average $f$ -score of about 70%. Leo Wanner |
Nat. Lang. Eng. | 1 |
| 2001 | Towards a Lexicographic Approach to Lexical Transfer in Machine Translation (Illustrated by the German-Russian Language Pair)
Igor Mel'cuk, Leo Wanner |
Mach. Transl. | 2 |
| 2000 | A development Environment for an MTT-Based Sentence GeneratorabstractWith the rising standard of the state of the art in text generation and the increase of the number of practical generation applications, it becomes more and more important to provide means for the maintenance of the generator, i.e. its extension, modification, and monitoring by grammarians who are not familiar with its internals. However, only a few sentence and text generators developed to date actually provide these means. One of these generators is KPML (Bateman, 1997). KPML comes with a Development Environment and there is no doubt about the contribution of this environment to the popularity of the systemic approach in generation. Bernd Bohnet, Andreas Langjahr, Leo Wanner |
INLG | 3 |
| 1998 | De-Constraining Text Generation
Stephen Beale, Sergei Nirenburg, Evelyne Viegas, Leo Wanner |
INLG | 4 |
| 1996 | The HealthDoc Sentence PlannerabstractThis paper describes the Sentence Planner (sP) in the HealthDoc project, which is concerned with the production of customized patienteducation material from a source encoded in terms of plans.The task of the sP is to transform selected, not necessarily consecutive, plans (which may vary in detail, from text plans specifying only content and discourse organization to fine-grained but incohesive, sentence plans) into completely specified specifications for the surface generator.The paper identifies the sentence planning tasks, which are highly interdependent and partially parallel, and argues, in accordance with [Nirenburg et al., 1989], that' a blackboard architecture with several independent modules is most suitable to deal with them.The architecture is presented, and the interaction of the sentence planning modules within this architecture is shown.The first implementation of the sP is discussed; examples illustrate the planning process in action. Leo Wanner, Eduard H. Hovy |
INLG (1) | 1 |
| 1996 | Editor's note
Leo Wanner |
Mach. Transl. | 1 |
| 1996 | Lexical choice in text generation and machine translation
Leo Wanner |
Mach. Transl. | 1 |
| 1994 | On Lexically Biased Discourse Organization In Text Generation
Leo Wanner |
COLING | 1 |
| 1994 | Building Another Bridge over the Generation Gap
Leo Wanner |
INLG | 1 |
| 1992 | Lexical Choice and the Organization of Lexical Resources in Text Generation
Leo Wanner |
ECAI | 1 |
| 1990 | A collocational based approach to salience-sensitive lexical selection
Leo Wanner, John A. Bateman |
INLG | 1 |