VLDB 2026 Research / reviewers in the wild / expert
Alicia Pérez
dblp:68/706
· DBLP profile ↗
27ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0003-2638-9598ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Early Mortality Prediction With Clinical Notes as Time-Series ScoresabstractThe optimization of early mortality prediction models to accurately identify high-risk patients during a hospital admission is key in medical research to assist in healthcare decision-making. While these models use structured (physiological features) or unstructured (free-text clinical notes) data, integrating them synergistically, especially for early prediction, remains a challenge. We introduce a novel, model-agnostic approach to represent clinical notes as quantitative, task-specific time-series, structurally aligning them with physiological features. Instead of relying on complex multimodal architectures, we shift fusion from the architecture level to the data level by transforming each note into a numerical mortality risk score, creating a new time-series feature that can be seamlessly integrated into any standard time-series prediction model. We evaluate our approach on the following: first, early prediction, second, data scalability, and third, versatility across centralized, federated, and local learning frameworks, in an attempt to assess the stability of the approach in scenarios with different amounts of data to learn from. Using only readily accessible data from the first 8–24 h postadmission, our method significantly boosts baseline model performance, with relative improvements of up to 7 % in area under the receiver operating characteristic curve and 28 % in area under the precision–recall curve over prior work on early prediction that relies solely on physiological data. Moreover, we find that incorporating our text-derived scores significantly achieves higher performance gains than nearly doubling the training data size in scenarios where no textual information is considered. Rather than proposing a predictive model, we present a new technique to structure clinical narratives, enabling a more effective fusion with temporal, quantitative health data for improved early prediction of mortality. Gabriel Vázquez-Serrano, Maite Oronoz, Alicia Pérez |
IEEE Trans. Hum. Mach. Syst. | 3 |
| 2025 | Medical Prognosis from Electronic Health Records in Spanish
Nuria Lebeña, Claudia Borg, Arantza Casillas, Alicia Pérez |
AIME (1) | 4 |
| 2023 | Cause of Death estimation from Verbal Autopsies: Is the Open Response redundant or synergistic?abstractCivil registration and vital statistics systems capture birth and death events to compile vital statistics and to provide legal rights to citizens. Vital statistics are a key factor in promoting public health policies and the health of the population. Medical certification of cause of death is the preferred source of cause of death information. However, two thirds of all deaths worldwide are not captured in routine mortality information systems and their cause of death is unknown. Verbal autopsy is an interim solution for estimating the cause of death distribution at the population level in the absence of medical certification. A Verbal Autopsy (VA) consists of an interview with the relative or the caregiver of the deceased. The VA includes both Closed Questions (CQs) with structured answer options, and an Open Response (OR) consisting of a free narrative of the events expressed in natural language and without any pre-determined structure. There are a number of automated systems to analyze the CQs to obtain cause specific mortality fractions with limited performance. We hypothesize that the incorporation of the text provided by the OR might convey relevant information to discern the CoD. The experimental layout compares existing Computer Coding Verbal Autopsy methods such as Tariff 2.0 with other approaches well suited to the processing of structured inputs as is the case of the CQs. Next, alternative approaches based on language models are employed to analyze the OR. Finally, we propose a new method with a bi-modal input that combines the CQs and the OR. Empirical results corroborated that the CoD prediction capability of the Tariff 2.0 algorithm is outperformed by our method taking into account the valuable information conveyed by the OR. As an added value, with this work we made available the software to enable the reproducibility of the results attained with a version implemented in R to make the comparison with Tariff 2.0 evident. Ander Cejudo, Arantza Casillas, Alicia Pérez, Maite Oronoz, Daniel Cobos |
Artif. Intell. Medicine | 3 |
| 2022 | Preliminary exploration of topic modelling representations for Electronic Health Records coding according to the International Classification of Diseases in Spanish
Nuria Lebeña, Alberto Blanco 0001, Alicia Pérez, Arantza Casillas |
Expert Syst. Appl. | 3 |
| 2022 | Implementation of specialised attention mechanisms: ICD-10 classification of Gastrointestinal discharge summaries in English, Spanish and Swedish
Alberto Blanco 0001, Sonja Remmer, Alicia Pérez, Hercules Dalianis, Arantza Casillas |
J. Biomed. Informatics | 3 |
| 2022 | Exploiting ICD Hierarchy for Classification of EHRs in Spanish Through Multi-Task TransformersabstractElectronic Health Records (EHRs) convey valuable information. Experts in clinical documentation read the report, understand the prior work, procedures, tests carried out, and encode the EHRs according to the International Classification of Diseases (ICD). Assigning these codes to the EHRs helps to share information, and extract statistics. In this paper, we explore computer-aided multi-label classification approaches. While Natural Language Understanding has evolved for clinical text mining, there is still a gap for languages other than English. Language-modeling aware Transformers has demonstrated state of the art approaches through exploiting contextual dependencies. Here we focus on EHRs written in Spanish, and try to benefit from the Language Model itself, with unannotated corpus with less data but in-house, in-domain and closely-related EHRs to that of the downstream task. The International Classification of Diseases coding scheme is hierarchical, but its synergies among hierarchical levels are rarely exploited. In this work, we implement and release a hierarchical head for multi-label classification, which benefits from the hierarchy of the ICD via multi-task classification. Alberto Blanco 0001, Alicia Pérez, Arantza Casillas |
IEEE J. Biomed. Health Informatics | 2 |
| 2021 | Extracting Cause of Death From Verbal Autopsy With Deep Learning Interpretable MethodsabstractThe international standard to ascertain the cause of death is medical certification. However, in many low and middle-income countries, the majority of deaths occur outside of health facilities. In these cases, Verbal Autopsy (VA), the narrative provided by a family member or friend together with a questionnaire is designed by the World Health Organization as the main information source. Until now technology allowed us to automatically analyze the responses of the VA questionnaire with the narrative captured by the interviewer excluded. Our work addresses this gap by developing a set of models for automatic Cause of Death (CoD) ascertainment in VAs with a focus on the textual information. Empirical results show that the open response conveys valuable information towards the ascertainment of the Cause of Death, and the combination of the closed-ended questions and the open response lead to the best results. Model interpretation capabilities position the Deep Learning models as the most encouraging choice. Alberto Blanco 0001, Alicia Pérez, Arantza Casillas, Daniel Cobos |
IEEE J. Biomed. Health Informatics | 2 |
| 2020 | Neural negated entity recognition in Spanish electronic health records
Sara Santiso Gonzáles, Alicia Pérez, Arantza Casillas, Maite Oronoz |
J. Biomed. Informatics | 2 |
| 2019 | Multi-label clinical document classification: Impact of label-density
Alberto Blanco 0001, Arantza Casillas, Alicia Pérez, Arantza Díaz de Ilarraza |
Expert Syst. Appl. | 3 |
| 2019 | Word embeddings for negation detection in health records written in Spanish
Sara Santiso Gonzáles, Arantza Casillas, Alicia Pérez, Maite Oronoz |
Soft Comput. | 3 |
| 2019 | Exploring Joint AB-LSTM With Embedded Lemmas for Adverse Drug Reaction DiscoveryabstractThis work focuses on the detection of adverse drug reactions (ADRs) in electronic health records (EHRs) written in Spanish. The World Health Organization underlines the importance of reporting ADRs for patients' safety. The fact is that ADRs tend to be under-reported in daily hospital praxis. In this context, automatic solutions based on text mining can help to alleviate the workload of experts. Nevertheless, these solutions pose two challenges: 1) EHRs show high lexical variability, the characterization of the events must be able to deal with unseen words or contexts and 2) ADRs are rare events, hence, the system should be robust against skewed class distribution. To tackle these challenges, deep neural networks seem appropriate because they allow a high-level representation. Specifically, we opted for a joint AB-LSTM network, a sub-class of the bidirectional long short-term memory network. Besides, in an attempt to reinforce lexical variability, we proposed the use of embeddings created using lemmas. We compared this approach with supervised event extraction approaches based on either symbolic or dense representations. Experimental results showed that the joint AB-LSTM approach outperformed previous approaches, achieving an f-measure of 73.3. Sara Santiso Gonzáles, Alicia Pérez, Arantza Casillas |
IEEE J. Biomed. Health Informatics | 2 |
| 2018 | Can I find information about rare diseases in some other language?
Mikel Laburu, Alicia Pérez, Arantza Casillas, Iakes Goenaga, Maite Oronoz |
BIBM | 2 |
| 2018 | Deep Medical Entity Recognition for Swedish and Spanish
Rebecka Weegar, Alicia Pérez, Arantza Casillas, Maite Oronoz |
BIBM | 2 |
| 2018 | Machine Learning Approaches on Diagnostic Term Encoding With the ICD for Clinical DocumentationabstractThis work focuses on data mining applied to the clinical documentation domain. Diagnostic terms (DTs) are used as keywords to retrieve valuable information from electronic health records. Indeed, they are encoded manually by experts following the International Classification of Diseases (ICD). The goal of this work is to explore the aid of text mining on DT encoding. From the machine learning (ML) perspective, this is a high-dimensional classification task, as it comprises thousands of codes. This work delves into a robust representation of the instances to improve ML results. The proposed system is able to find the right ICD code among more than 1500 possible ICD codes with 92% precision for the main disease (primary class) and 88% for the main disease together with the nonessential modifiers (fully specified class). The methodology employed is simple and portable. According to the experts from public hospitals, the system is very useful in particular for documentation and pharmacosurveillance services. In fact, they reported an accuracy of 91.2% on a small randomly extracted test. Hence, together with this paper, we made the software publicly available in order to help the clinical and research community. Aitziber Atutxa, Alicia Pérez, Arantza Casillas |
IEEE J. Biomed. Health Informatics | 2 |
| 2017 | Semi-supervised medical entity recognition: A study on Spanish and Swedish clinical corpora
Alicia Pérez, Rebecka Weegar, Arantza Casillas, Koldo Gojenola, Maite Oronoz, Hercules Dalianis |
J. Biomed. Informatics | 1 |
| 2016 | Clinical text mining for efficient extraction of drug-allergy reactionsabstractThis work focuses on the extraction of allergic drug reactions in electronic health records. The goal is to annotate a sub-class of cause-effect events, those in which drugs are causing allergies. Little work has carried out in this field, seldom for Spanish clinical text mining, which is, indeed, the aim of this work. We present two approaches: a rule-based method and another one based on machine learning. Both approaches incorporate semantic knowledge derived from FreeLing-Med, a software explicitly developed to parse texts in the medical domain. Having recognised the medical entities for a given record, the challenge stands on triggering the underlying allergies. To this end, the knowledge is expressed as a set of semantic, syntactic and structural features. Our best system, based on machine learning, obtained a precision of 0.90 with a recall of 0.87, outperforming a rule-based approach. Arantza Casillas, Koldo Gojenola, Alicia Pérez, Maite Oronoz |
BIBM | 3 |
| 2016 | IXAmed-IE: On-line medical entity identification and ADR event extraction in SpanishabstractThis work presents an on-line system developed for medical information extraction. The goal is to provide a web-based service addressed to the medical community for the efficient processing of electronic health records in Spanish and support the clinical decision making process. This tool assists in the identification of medical entities as well as cause-effect reactions in order to help professionals both to prevent and to document adverse drug events. So far, the prototype is in its early stage of testing and validation by experts from the Galdakao-Usansolo and Basurto hospitals from the Basque Sanitary System (Osakidetza). Arantza Casillas, Arantza Díaz de Ilarraza, K. Fernandez, Koldo Gojenola, Maite Oronoz, Alicia Pérez, Sara Santiso Gonzáles |
BIBM | 6 |
| 2016 | Learning to extract adverse drug reaction events from electronic health records in Spanish
Arantza Casillas, Alicia Pérez, Maite Oronoz, Koldo Gojenola, Sara Santiso Gonzáles |
Expert Syst. Appl. | 2 |
| 2015 | Computer aided classification of diagnostic terms in spanish
Alicia Pérez, Koldo Gojenola, Arantza Casillas, Maite Oronoz, Arantza Díaz de Ilarraza |
Expert Syst. Appl. | 1 |
| 2015 | On the creation of a clinical gold standard corpus in Spanish: Mining adverse drug reactions
Maite Oronoz, Koldo Gojenola, Alicia Pérez, Arantza Díaz de Ilarraza, Arantza Casillas |
J. Biomed. Informatics | 3 |
| 2013 | Automatic Annotation of Medical Records in Spanish with Disease, Drug and Substance Names
Maite Oronoz, Arantza Casillas, Koldo Gojenola, Alicia Pérez |
CIARP (2) | 4 |
| 2012 | EuskoParl: a speech and text Spanish-Basque parallel corpus
Alicia Pérez, José M. Alcaide, M. Inés Torres |
INTERSPEECH | 1 |
| 2010 | Hierarchical Finite-State Models for Speech Translation Using Categorization of Phrases
Raquel Justo, Alicia Pérez, M. Inés Torres, Francisco Casacuberta |
CICLing | 2 |
| 2010 | Potential scope of a fully-integrated architecture for speech translation
Alicia Pérez, M. Inés Torres, Francisco Casacuberta |
EAMT | 1 |
| 2008 | Joining linguistic and statistical methods for Spanish-to-Basque speech translation
Alicia Pérez, M. Inés Torres, Francisco Casacuberta |
Speech Commun. | 1 |
| 2007 | Speech Translation with Phrase Based Stochastic Finite-State TransducersabstractStochastic finite-state transducers constitute a type of word-based models that allow an easy integration with acoustic model for speech translation. The aim of this work is to develop a novel approach to phrase-based statistical finite-state transducers. In this work, we explore the use of linguistically motivated phrases to build phrase-based models. The proposed phrase-based transducer has been tested and compared to a word-based equivalent machine, yielding promising results in the reported preliminary text and speech translation experiments. Alicia Pérez, M. Inés Torres, Francisco Casacuberta |
ICASSP (4) | 1 |
| 2007 | Evaluation of alternatives on speech to sign language translationabstractThis paper evaluates different approaches on speech to sign language machine translation. The framework of the application focuses on assisting deaf people to apply for the passport or related information. In this context, the main aim is to automatically translate the spontaneous speech, uttered by an officer, into Spanish Sign Language (SSL). In order to get the best translation quality, three alternative techniques have been evaluated: a rule-based approach, a phrase-based statistical approach, and a approach that makes use of stochastic finite state transducers. The best speech translation experiments have reported a 32.0 % SER (Sign Error Rate) and a 7.1 BLEU (BiLingual Evaluation Understudy) including speech recognition errors. Rubén San-Segundo-Hernández, Alicia Pérez, Daniel Ortiz-Martínez, Luis Fernando D'Haro, M. Inés Torres, Francisco Casacuberta |
INTERSPEECH | 2 |