Aline Villavicencio

dblp:v/AlineVillavicencio · DBLP profile ↗
← Back
47ranked-venue papers
10as first author
17since 2021 · last 2026
0000-0002-3731-9168ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 45 · 10 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Rethinking the Idiomaticity Decomposability Hypothesis: Evidence from Distributional Learning
abstract
Maggie Mi, Golzar Atefi, Atsuki Yamaguchi, Felix Gers, Aline Villavicencio, Nafise Sadat Moosavi. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Maggie Mi, Golzar Atefi, Atsuki Yamaguchi, Felix A. Gers, Aline Villavicencio, Nafise Sadat Moosavi
ACL (1)5
2026 Mitigating Catastrophic Forgetting in Target Language Adaptation of LLMs via Source-Shielded Updates
abstract
Expanding the linguistic diversity of instruct large language models (LLMs) is crucial for global accessibility but is often hindered by the reliance on costly specialized target language labeled data and catastrophic forgetting during adaptation.We tackle this challenge under a realistic, low-resource constraint: adapting instruct LLMs using only unlabeled target language data.We introduce Source-Shielded Updates (SSU), a selective parameter update strategy that proactively preserves source knowledge.Using a small set of source data and a parameter importance scoring method, SSU identifies parameters critical to maintaining source abilities.It then applies a column-wise freezing strategy to protect these parameters before adaptation.Experiments across five typologically diverse languages and 7B and 13B models demonstrate that SSU successfully mitigates catastrophic forgetting.It reduces performance degradation on monolingual source tasks to just 3.4% (7B) and 2.8% (13B) on average, a stark contrast to the 20.3% and 22.3% from full fine-tuning.SSU also achieves target-language performance highly competitive with full fine-tuning, outperforming it on all benchmarks for 7B models and the majority for 13B models. 1
Atsuki Yamaguchi, Terufumi Morishita, Aline Villavicencio, Nikolaos Aletras
ACL (1)3
2026 Figurative Language in Alzheimer's Discourse: Linguistic and Neural Alignment in Clinical Narratives
Diana Kylymnyk, Vitória Hilgert Tomasel, Helena de Medeiros Caseli, Edward Watkins, Aline Villavicencio, Rodrigo Wilkens
LREC5
2026 Stands to Reason: Investigating the Effect of Reasoning on Idiomaticity Detection
abstract
The recent trend towards utilisation of reasoning models has improved the performance of Large Language Models (LLMs) across many tasks which involve logical steps. One linguistic task that could benefit from this framing is idiomaticity detection, as a potentially idiomatic expression must first be understood in relation to the context before it can be disambiguated. In this paper, we explore how reasoning capabilities in LLMs affect idiomaticity detection performance and examine the effect of model size. We evaluate, as open source representative models, the suite of DeepSeek-R1 distillation models ranging from 1.5B to 70B parameters across four idiomaticity detection datasets. We find the effect of reasoning to be smaller and more varied than expected. For smaller models, producing chain-of-thought (CoT) reasoning increases performance from Math-tuned intermediate models, but not to the levels of the base models, whereas larger models (14B, 32B, and 70B) show modest improvements. Our in-depth analyses reveal that larger models demonstrate good understanding of idiomaticity, successfully producing accurate definitions of expressions, while smaller models often fail to output the actual meaning. For this reason, we also experiment with providing definitions in the prompts of smaller models, which we show can improve performance in some cases.
Dylan Phelps, Rodrigo Wilkens, Edward Gow-Smith, Thomas Pickard, Maggie Mi, Marco Idiart, Aline Villavicencio
LREC7
2026 A Parallel Cross-Lingual Benchmark for Multimodal Idiomaticity Understanding
abstract
Potentially idiomatic expressions (PIEs) carry meanings inherently tied to the everyday experience of a given language community. As such, they constitute an interesting challenge for assessing the linguistic (and to some extent cultural) capabilities of NLP systems. In this paper, we present XMPIE, a parallel multilingual and multimodal dataset of potentially idiomatic expressions. The dataset, containing 34 languages and over ten thousand items, allows comparative analyses of idiomatic patterns among language-specific realisations and preferences in order to gather insights about shared cultural aspects. This parallel dataset allows evaluation of language model performance for a given PIE in different languages and whether idiomatic understanding in one language can be transferred to another. Moreover, the dataset supports the study of PIEs across textual and visual modalities, to measure to what extent PIE understanding in one modality transfers or implies in understanding in another modality (text vs. image). The data was created by language experts, with both textual and visual components crafted under multilingual guidelines, and each PIE is accompanied by five images representing a spectrum from idiomatic to literal meanings, including semantically related and random distractors. The result is a high-quality benchmark for evaluating multilingual and multimodal idiomatic language understanding.
Dilara Torunoglu-Selamet, Dogukan Arslan, Rodrigo Wilkens, Wei He 0017, Doruk Eryigit, Thomas Pickard, Adriana S. Pagano, Aline Villavicencio, Gülsen Eryigit, Ágnes Abuczki, Aida Cardoso, Alesia Lazarenka, Dina Almassova, Amália Mendes, Anna Kanellopoulou, Antoni Brosa-Rodríguez, Baiba Valkovska, Beata Wojtowicz, Bolette Pedersen, Carlos Manuel Hidalgo-Ternero, Chaya Liebeskind, Danka Jokic, Diego Alves, Eleni Triantafyllidi, Erik Velldal, Fred Philippy, Giedre Valunaite Oleskeviciene, Ieva Rizgeliene, Inguna Skadina, Irina Lobzhanidze, Isabell Stinessen Haugen, Jauza Akbar Krito, Jelena M. Markovic, Johanna Monti, Josue Alejandro Sauca, Kaja Dobrovoljc, Kingsley O. Ugwuanyi, Laura Rituma, Lilja Øvrelid, Maha Tufail Agro, Manzura Abjalova, Maria Chatzigrigoriou, María del Mar Sánchez Ramos, Marija Pendevska, Masoumeh Seyyedrezaei, Mehrnoush Shamsfard, Momina Ahsan, Muhammad Ahsan Riaz Khan, Nathalie Carmen Hau Norman, Nilay Erdem Ayyildiz, Nina Hosseini-Kivanani, Noémi Ligeti-Nagy, Numaan Naeem, Olha Kanishcheva, Olha Yatsyshyna, Daniil Orel, Petra Giommarelli, Petya Osenova, Radovan Garabík, Regina E. Semou, Rozane Rebechi, Salsabila Zahirah Pranida, Samia Touileb, Sanni Nimb, Sarvinoz Sharipova, Shahar Golan, Shaoxiong Ji, Sopuruchi Christian Aboh, Srdjan Sucur, Stella Markantonatou, Sussi Olsen, Vahideh Tajalli, Veronika Lipp, Voula Giouli, Yelda Yesildal Eraydin, Zahra Saaberi, Zhuohan Xie
LREC8
2026 Automated Machine Learning in medical research: A systematic literature mapping study
Giovanna A. Castro, Luiza G. Barioto, Yu H. Cao, Renato Moraes Silva, Helena de Medeiros Caseli, João A. Machado-Neto, Ricardo Cerri, Aline Villavicencio, Tiago A. Almeida 0001
Artif. Intell. Medicine8
2026 StealthMark: Harmless and Stealthy Ownership Verification for Medical Segmentation via Uncertainty-Guided Backdoors
abstract
Annotating medical data for training AI models is often costly and limited due to the shortage of specialists with relevant clinical expertise. This challenge is further compounded by privacy and ethical concerns associated with sensitive patient information. As a result, well-trained medical segmentation models on private datasets constitute valuable intellectual property requiring robust protection mechanisms. Existing model protection techniques primarily focus on classification and generative tasks, while segmentation models-crucial to medical image analysis-remain largely underexplored. In this paper, we propose a novel, stealthy, and harmless method, StealthMark, for verifying the ownership of medical segmentation models under closed-box conditions. Our approach subtly modulates model uncertainty without altering the final segmentation outputs, thereby preserving the model's performance. To enable ownership verification, we incorporate model-agnostic explanation methods, e.g. LIME, to extract feature attributions from the model outputs. Under specific triggering conditions, these explanations reveal a distinct and verifiable watermark. We further design the watermark as a QR code to facilitate robust and recognizable ownership claims. We conducted extensive experiments across four medical imaging datasets (CMR dataset from UK Biobank, the SEG fundus dataset, the EchoNet echocardiography dataset, and the PraNet colonoscopy dataset) and five mainstream segmentation models. The results demonstrate the effectiveness, stealthiness, and harmlessness of our method on the original model's segmentation performance. For example, when applied to the SAM model, StealthMark consistently achieved attack success rates (ASR) above 95% across various datasets while maintaining less than a 1% drop in Dice and AUC scores-significantly outperforming backdoor-based watermarking methods and highlighting its strong potential for practical deployment. Our implementation code is made available at https://github.com/Qinkaiyu/StealthMark.
Qinkai Yu, Chong Zhang 0006, Gaojie Jin, Tianjin Huang, Wei Zhou 0021, Xiao-Bo Jin, Bo Huang 0012, Yitian Zhao, Gregory Yoke Hong Lip, Yalin Zheng, Aline Villavicencio, Yanda Meng
IEEE Trans. Image Process.13
2026 Introduction to the Special Issue on Transformers
Feng Xia 0001, Tyler Derr, Anh Tuan Luu, Richa Singh 0001, Aline Villavicencio
ACM Trans. Intell. Syst. Technol.5
2025 Rolling the DICE on Idiomaticity: How LLMs Fail to Grasp Context
abstract
Human processing of idioms heavily depends on interpreting the surrounding context in which they appear.While large language models (LLMs) have achieved impressive performance on idiomaticity detection benchmarks, this success may be driven by reasoning shortcuts present in existing datasets.To address this, we introduce a novel, controlled contrastive dataset (DICE) specifically designed to assess whether LLMs can effectively leverage context to disambiguate idiomatic meanings.Furthermore, we investigate the influence of collocational frequency and sentence probability-proxies for human processing known to affect idiom resolution-on model performance.Our results show that LLMs frequently fail to resolve idiomaticity when it depends on contextual understanding, and they perform better on sentences deemed more likely by the model.Additionally, idiom frequency influences performance but does not guarantee accurate interpretation.Our findings emphasize the limitations of current models in grasping contextual meaning and highlight the need for more context-sensitive evaluation.
Maggie Mi, Aline Villavicencio, Nafise Sadat Moosavi
ACL (1)2
2025 From Input Perception to Predictive Insight: Modeling Model Blind Spots Before They Become Errors
abstract
Language models often struggle with idiomatic, figurative, or context-sensitive inputs, not because they produce flawed outputs, but because they misinterpret the input from the outset.We propose an input-only method for anticipating such failures using token-level likelihood features inspired by surprisal and the Uniform Information Density hypothesis.These features capture localized uncertainty in input comprehension and outperform standard baselines across five linguistically challenging datasets.We show that span-localized features improve error detection for larger models, while smaller models benefit from global patterns.Our method requires no access to outputs or hidden activations, offering a lightweight and generalizable approach to pre-generation error prediction. https://github.com/mi-m1/input_ perceptionHow is the expression "caught between a rock and a hard place" used in the following sentence?Literally or figuratively? LLM AnswerOutput-level Error Detection External Knowledge BaseLog Probabilities LLM as a Judge Suddenly she was caught between a rock and a hard place.'Suddenly', 'Ġshe', 'Ġwas', 'Ġcaught', 'Ġbetween', 'Ġa', 'Ġrock', 'Ġand', 'Ġa', 'Ġhard', 'Ġplace' Error: 81% Correct:19% Prompt: How is the expression "caught between a rock and a hard place" used in the following sentence?Literally or figuratively?Sentence: Suddenly she was caught between a rock and a hard place.
Maggie Mi, Aline Villavicencio, Nafise Sadat Moosavi
EMNLP2
2024 Representation transfer and data cleaning in multi-views for text simplification
abstract
Representation transfer is a widely used technique in natural language processing. We propose methods of cleaning the dominant dataset of text simplification (TS) WikiLarge in multi-views to remove errors that impact model training and fine-tuning. The results show that our method can effectively refine the dataset. We propose to take the pre-trained text representations from a similar task (e.g., text summarization) to text simplification to conduct a continue-fine-tuning strategy to improve the performance of pre-trained models on TS. This approach will speed up the training and make the model convergence easier. Besides, we also propose a new decoding strategy for simple text generation. It is able to generate simpler and more comprehensible text with controllable lexical simplicity. The experimental results show that our method can achieve good performance on many evaluation metrics.
Wei He 0017, Katayoun Farrahi, Bohua Peng, Aline Villavicencio
Pattern Recognit. Lett.5
2024 Multi-perspective thought navigation for source-free entity linking
Bohua Peng, Wei He 0017, Aline Villavicencio, Chengfu Wu
Pattern Recognit. Lett.4
2023 Evaluating Open-Domain Dialogues in Latent Space with Next Sentence Prediction and Mutual Information
abstract
The long-standing one-to-many issue of the open-domain dialogues poses significant challenges for automatic evaluation methods, i.e., there may be multiple suitable responses which differ in semantics for a given conversational context.To tackle this challenge, we propose a novel learning-based automatic evaluation metric (CMN), which can robustly evaluate open-domain dialogues by augmenting Conditional Variational Autoencoders (CVAEs) with a Next Sentence Prediction (NSP) objective and employing Mutual Information (MI) to model the semantic similarity of text in the latent space.Experimental results on two opendomain dialogue datasets demonstrate the superiority of our method compared with a wide range of baselines, especially in handling responses which are distant to the golden reference responses in semantics.
Kun Zhao 0007, Bohao Yang, Chenghua Lin 0002, Wenge Rong, Aline Villavicencio, Xiaohui Cui
ACL (1)5
2023 Understanding the effects of negative (and positive) pointwise mutual information on word vectors
abstract
Despite the recent popularity of contextual word embeddings, static word embeddings still dominate lexical semantic tasks, making their study of continued relevance. A widely adopted family of such static word embeddings is derived by explicitly factorising the Pointwise Mutual Information (PMI) weighting of the co-occurrence matrix. As unobserved co-occurrences lead PMI to negative infinity, a common workaround is to clip negative PMI at 0. However, it is unclear what information is lost by collapsing negative PMI values to 0. To answer this question, we isolate and study the effects of negative (and positive) PMI on the semantics and geometry of models adopting factorisation of different PMI matrices. Word and sentence-level evaluations show that only accounting for positive PMI in the factorisation strongly captures both semantics and syntax, whereas using only negative PMI captures little of semantics but a surprising amount of syntactic information. Results also reveal that incorporating negative PMI induces stronger rank invariance of vector norms and directions, as well as improved rare word representations.
Alexandre Salle, Aline Villavicencio
J. Exp. Theor. Artif. Intell.2
2022 Improving Tokenisation by Alternative Treatment of Spaces
abstract
Tokenisation is the first step in almost all NLP tasks, and state-of-the-art transformer-based language models all use subword tokenisation algorithms to process input text.Existing algorithms have problems, often producing tokenisations of limited linguistic validity and representing equivalent strings differently depending on their position within a word.We hypothesise that these problems hinder the ability of transformer-based models to handle complex words, and suggest that these problems are a result of allowing tokens to include spaces.We thus experiment with an alternative tokenisation approach where spaces are always treated as individual tokens.Specifically, we apply this modification to the BPE and Unigram algorithms.We find that our modified algorithms lead to improved performance on downstream NLP tasks that involve handling complex words, whilst having no detrimental effect on performance in general natural language understanding tasks.Intrinsically, we find that our modified algorithms give more morphologically correct tokenisations, in particular when handling prefixes.Given the results of our experiments, we advocate for always treating spaces as individual tokens as an improved tokenisation method.
Edward Gow-Smith, Harish Tayyar Madabushi, Carolina Scarton, Aline Villavicencio
EMNLP4
2021 Assessing the Representations of Idiomaticity in Vector Models with a Noun Compound Dataset Labeled at Type and Token Levels
abstract
Marcos Garcia, Tiago Kramer Vieira, Carolina Scarton, Marco Idiart, Aline Villavicencio. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Marcos García 0001, Tiago Kramer Vieira, Carolina Scarton, Marco Idiart, Aline Villavicencio
ACL/IJCNLP (1)5
2021 Probing for idiomaticity in vector space models
abstract
Marcos Garcia, Tiago Kramer Vieira, Carolina Scarton, Marco Idiart, Aline Villavicencio. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Marcos García 0001, Tiago Kramer Vieira, Carolina Scarton, Marco Idiart, Aline Villavicencio
EACL5
2020 Investigating alignment interpretability for low-resource NMT
Marcely Zanon Boito, Aline Villavicencio, Laurent Besacier
Mach. Transl.2
2019 Empirical Evaluation of Sequence-to-Sequence Models for Word Discovery in Low-Resource Settings
abstract
International audience
Marcely Zanon Boito, Aline Villavicencio, Laurent Besacier
INTERSPEECH2
2019 Unsupervised Compositionality Prediction of Nominal Compounds
abstract
Nominal compounds such as red wine and nut case display a continuum of compositionality, with varying contributions from the components of the compound to its semantics. This article proposes a framework for compound compositionality prediction using distributional semantic models, evaluating to what extent they capture idiomaticity compared to human judgments. For evaluation, we introduce data sets containing human judgments in three languages: English, French, and Portuguese. The results obtained reveal a high agreement between the models and human predictions, suggesting that they are able to incorporate information about idiomaticity. We also present an in-depth evaluation of various factors that can affect prediction, such as model and corpus parameters and compositionality operations. General crosslingual analyses reveal the impact of morphological variation and corpus size in the ability of the model to predict compositionality, and of a uniform combination of the components for best results.
Silvio Cordeiro, Aline Villavicencio, Marco Idiart, Carlos Ramisch
Comput. Linguistics2
2019 Discovering multiword expressions
abstract
Abstract In this paper, we provide an overview of research on multiword expressions (MWEs), from a natural language processing perspective. We examine methods developed for modelling MWEs that capture some of their linguistic properties, discussing their use for MWE discovery and for idiomaticity detection. We concentrate on their collocational and contextual preferences, along with their fixedness in terms of canonical forms and their lack of word-for-word translatatibility. We also discuss a sample of the MWE resources that have been used in intrinsic evaluation setups for these methods.
Aline Villavicencio, Marco Idiart
Nat. Lang. Eng.1
2018 Unsupervised Word Segmentation from Speech with Attention
abstract
International audience
Pierre Godard, Marcely Zanon Boito, Lucas Ondel Yang, Alexandre Berard, François Yvon, Aline Villavicencio, Laurent Besacier
INTERSPEECH6
2018 The brWaC Corpus: A New Open Resource for Brazilian Portuguese
Jorge A. Wagner Filho, Rodrigo Wilkens, Marco Idiart, Aline Villavicencio
LREC4
2017 Unwritten languages demand attention too! Word discovery with encoder-decoder models
abstract
Word discovery is the task of extracting words from un-segmented text. In this paper we examine to what extent neural networks can be applied to this task in a realistic unwritten language scenario, where only small corpora and limited annotations are available. We investigate two scenarios: one with no supervision and another with limited supervision with access to the most frequent words. Obtained results show that it is possible to retrieve at least 27% of the gold standard vocabulary by training an encoder-decoder neural machine translation system with only 5,157 sentences. This result is close to those obtained with a task-specific Bayesian nonparametric model. Moreover, our approach has the advantage of generating translation alignments, which could be used to create a bilingual lexicon. As a future perspective, this approach is also well suited to work directly from speech.
Marcely Zanon Boito, Alexandre Berard, Aline Villavicencio, Laurent Besacier
ASRU3
2016 Predicting the Compositionality of Nominal Compounds: Giving Word Embeddings a Hard Time
abstract
Distributional semantic models (DSMs) are often evaluated on artificial similarity datasets containing single words or fully compositional phrases. We present a large-scale multilingual evaluation of DSMs for predicting the degree of semantic compositionality of nominal compounds on 4 datasets for English and French. We build a total of 816 DSMs and perform 2,856 evaluations using word2vec, GloVe, and PPMI-based models. In addition to the DSMs, we compare the impact of different parameters, such as level of corpus preprocessing, context window size and number of dimensions. The results obtained have a high correlation with human judgments, being comparable to or outperforming the state of the art for some datasets (Spearman's ρ=.82 for the Reddy dataset).
Silvio Cordeiro, Carlos Ramisch, Marco Idiart, Aline Villavicencio
ACL (1)4
2016 mwetoolkit+sem: Integrating Word Embeddings in the mwetoolkit for Semantic MWE Processing
Silvio Cordeiro, Carlos Ramisch, Aline Villavicencio
LREC3
2016 Multiword Expressions in Child Language
Rodrigo Wilkens, Marco Idiart, Aline Villavicencio
LREC3
2016 B2SG: a TOEFL-like Task for Portuguese
Rodrigo Wilkens, Leonardo Zilio, Aline Villavicencio
LREC4
2016 VerbLexPor: a lexical resource with semantic roles for Portuguese
Leonardo Zilio, Maria José Bocorny Finatto, Aline Villavicencio
LREC3
2014 Nothing like Good Old Frequency: Studying Context Filters for Distributional Thesauri
abstract
Much attention has been given to the impact of informativeness and similarity measures on distributional thesauri.We investigate the effects of context filters on thesaurus quality and propose the use of cooccurrence frequency as a simple and inexpensive criterion.For evaluation, we measure thesaurus agreement with WordNet and performance in answering TOEFL-like questions.Results illustrate the sensitivity of distributional thesauri to filters.
Muntsa Padró, Marco Idiart, Aline Villavicencio, Carlos Ramisch
EMNLP3
2014 Identification of Multiword Expressions in the brWaC
Rodrigo Boos, Kassius Prestes, Aline Villavicencio
LREC3
2014 Comparing the Quality of Focused Crawlers and of the Translation Resources Obtained from them
Bruno Laranjeira, Viviane Pereira Moreira, Aline Villavicencio, Carlos Ramisch, Maria José Bocorny Finatto
LREC3
2014 Comparing Similarity Measures for Distributional Thesauri
Muntsa Padró, Marco Idiart, Aline Villavicencio, Carlos Ramisch
LREC3
2013 Language Acquisition and Probabilistic Models: keeping it simple
Aline Villavicencio, Marco Idiart, Robert C. Berwick, Igor Malioutov
ACL (1)1
2012 A large scale annotated child language construction database
Aline Villavicencio, Beracah Yankama, Marco Idiart, Robert C. Berwick
LREC1
2012 Syntax-Based Collocation Extraction, by Violeta Seretan. Berlin: Springer, 2011. ISBN-10 9400701330, ISBN-13 978-9400701335. $139.00/£90.00 (Hardcover) xi + 220 pages
Aline Villavicencio
Nat. Lang. Eng.1
2010 mwetoolkit: a Framework for Multiword Expression Identification
Carlos Ramisch, Aline Villavicencio, Christian Boitet
LREC2
2009 Prepositions in Applications: A Survey and Introduction to the Special Issue
abstract
C1 - Journal Articles Refereed
Timothy Baldwin, Valia Kordoni, Aline Villavicencio
Comput. Linguistics3
2008 Picking them up and Figuring them out: Verb-Particle Constructions, Noise and Idiomaticity
Carlos Ramisch, Aline Villavicencio, Leonardo Moura, Marco Idiart
CoNLL2
2007 Validation and Evaluation of Automatically Acquired Multiword Expressions for Grammar Engineering
Aline Villavicencio, Valia Kordoni, Yi Zhang 0003, Marco Idiart, Carlos Ramisch
EMNLP-CoNLL1
2005 The availability of verb-particle constructions in lexical resources: How much is enough?
Aline Villavicencio
Comput. Speech Lang.1
2005 Introduction to the special issue on multiword expressions: Having a crack at a hard nut
Aline Villavicencio, Francis Bond, Anna Korhonen, Diana McCarthy
Comput. Speech Lang.1
2004 A Multilingual Database of Idioms
Aline Villavicencio, Timothy Baldwin, Benjamin Waldron
LREC1
2002 Extracting the Unextractable: A Case Study on Verb-particles
Timothy Baldwin, Aline Villavicencio
CoNLL2
2002 Learning to Distinguish PP Arguments from Adjuncts
Aline Villavicencio
CoNLL1
2002 Multiword expressions: linguistic precision and reusability
Ann A. Copestake, Fabre Lambeau, Aline Villavicencio, Francis Bond, Timothy Baldwin, Ivan A. Sag, Dan Flickinger
LREC3
1999 Representing a System of Lexical Types Using Default Unification
Aline Villavicencio
EACL1