Arantza Casillas

dblp:42/6563 · also Arantza Casillas Rubio · DBLP profile ↗
← Back
31ranked-venue papers
11as first author
6since 2021 · last 2025
0000-0003-4248-8182ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 8 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 Medical Prognosis from Electronic Health Records in Spanish
Nuria Lebeña, Claudia Borg, Arantza Casillas, Alicia Pérez
AIME (1)3
2023 Cause of Death estimation from Verbal Autopsies: Is the Open Response redundant or synergistic?
abstract
Civil registration and vital statistics systems capture birth and death events to compile vital statistics and to provide legal rights to citizens. Vital statistics are a key factor in promoting public health policies and the health of the population. Medical certification of cause of death is the preferred source of cause of death information. However, two thirds of all deaths worldwide are not captured in routine mortality information systems and their cause of death is unknown. Verbal autopsy is an interim solution for estimating the cause of death distribution at the population level in the absence of medical certification. A Verbal Autopsy (VA) consists of an interview with the relative or the caregiver of the deceased. The VA includes both Closed Questions (CQs) with structured answer options, and an Open Response (OR) consisting of a free narrative of the events expressed in natural language and without any pre-determined structure. There are a number of automated systems to analyze the CQs to obtain cause specific mortality fractions with limited performance. We hypothesize that the incorporation of the text provided by the OR might convey relevant information to discern the CoD. The experimental layout compares existing Computer Coding Verbal Autopsy methods such as Tariff 2.0 with other approaches well suited to the processing of structured inputs as is the case of the CQs. Next, alternative approaches based on language models are employed to analyze the OR. Finally, we propose a new method with a bi-modal input that combines the CQs and the OR. Empirical results corroborated that the CoD prediction capability of the Tariff 2.0 algorithm is outperformed by our method taking into account the valuable information conveyed by the OR. As an added value, with this work we made available the software to enable the reproducibility of the results attained with a version implemented in R to make the comparison with Tariff 2.0 evident.
Ander Cejudo, Arantza Casillas, Alicia Pérez, Maite Oronoz, Daniel Cobos
Artif. Intell. Medicine2
2022 Preliminary exploration of topic modelling representations for Electronic Health Records coding according to the International Classification of Diseases in Spanish
Nuria Lebeña, Alberto Blanco 0001, Alicia Pérez, Arantza Casillas
Expert Syst. Appl.4
2022 Implementation of specialised attention mechanisms: ICD-10 classification of Gastrointestinal discharge summaries in English, Spanish and Swedish
Alberto Blanco 0001, Sonja Remmer, Alicia Pérez, Hercules Dalianis, Arantza Casillas
J. Biomed. Informatics5
2022 Exploiting ICD Hierarchy for Classification of EHRs in Spanish Through Multi-Task Transformers
abstract
Electronic Health Records (EHRs) convey valuable information. Experts in clinical documentation read the report, understand the prior work, procedures, tests carried out, and encode the EHRs according to the International Classification of Diseases (ICD). Assigning these codes to the EHRs helps to share information, and extract statistics. In this paper, we explore computer-aided multi-label classification approaches. While Natural Language Understanding has evolved for clinical text mining, there is still a gap for languages other than English. Language-modeling aware Transformers has demonstrated state of the art approaches through exploiting contextual dependencies. Here we focus on EHRs written in Spanish, and try to benefit from the Language Model itself, with unannotated corpus with less data but in-house, in-domain and closely-related EHRs to that of the downstream task. The International Classification of Diseases coding scheme is hierarchical, but its synergies among hierarchical levels are rarely exploited. In this work, we implement and release a hierarchical head for multi-label classification, which benefits from the hierarchy of the ICD via multi-task classification.
Alberto Blanco 0001, Alicia Pérez, Arantza Casillas
IEEE J. Biomed. Health Informatics3
2021 Extracting Cause of Death From Verbal Autopsy With Deep Learning Interpretable Methods
abstract
The international standard to ascertain the cause of death is medical certification. However, in many low and middle-income countries, the majority of deaths occur outside of health facilities. In these cases, Verbal Autopsy (VA), the narrative provided by a family member or friend together with a questionnaire is designed by the World Health Organization as the main information source. Until now technology allowed us to automatically analyze the responses of the VA questionnaire with the narrative captured by the interviewer excluded. Our work addresses this gap by developing a set of models for automatic Cause of Death (CoD) ascertainment in VAs with a focus on the textual information. Empirical results show that the open response conveys valuable information towards the ascertainment of the Cause of Death, and the combination of the closed-ended questions and the open response lead to the best results. Model interpretation capabilities position the Deep Learning models as the most encouraging choice.
Alberto Blanco 0001, Alicia Pérez, Arantza Casillas, Daniel Cobos
IEEE J. Biomed. Health Informatics3
2020 Neural negated entity recognition in Spanish electronic health records
Sara Santiso Gonzáles, Alicia Pérez, Arantza Casillas, Maite Oronoz
J. Biomed. Informatics3
2019 Multi-label clinical document classification: Impact of label-density
Alberto Blanco 0001, Arantza Casillas, Alicia Pérez, Arantza Díaz de Ilarraza
Expert Syst. Appl.2
2019 Word embeddings for negation detection in health records written in Spanish
Sara Santiso Gonzáles, Arantza Casillas, Alicia Pérez, Maite Oronoz
Soft Comput.2
2019 Exploring Joint AB-LSTM With Embedded Lemmas for Adverse Drug Reaction Discovery
abstract
This work focuses on the detection of adverse drug reactions (ADRs) in electronic health records (EHRs) written in Spanish. The World Health Organization underlines the importance of reporting ADRs for patients' safety. The fact is that ADRs tend to be under-reported in daily hospital praxis. In this context, automatic solutions based on text mining can help to alleviate the workload of experts. Nevertheless, these solutions pose two challenges: 1) EHRs show high lexical variability, the characterization of the events must be able to deal with unseen words or contexts and 2) ADRs are rare events, hence, the system should be robust against skewed class distribution. To tackle these challenges, deep neural networks seem appropriate because they allow a high-level representation. Specifically, we opted for a joint AB-LSTM network, a sub-class of the bidirectional long short-term memory network. Besides, in an attempt to reinforce lexical variability, we proposed the use of embeddings created using lemmas. We compared this approach with supervised event extraction approaches based on either symbolic or dense representations. Experimental results showed that the joint AB-LSTM approach outperformed previous approaches, achieving an f-measure of 73.3.
Sara Santiso Gonzáles, Alicia Pérez, Arantza Casillas
IEEE J. Biomed. Health Informatics3
2018 Can I find information about rare diseases in some other language?
Mikel Laburu, Alicia Pérez, Arantza Casillas, Iakes Goenaga, Maite Oronoz
BIBM3
2018 Deep Medical Entity Recognition for Swedish and Spanish
Rebecka Weegar, Alicia Pérez, Arantza Casillas, Maite Oronoz
BIBM3
2018 Machine Learning Approaches on Diagnostic Term Encoding With the ICD for Clinical Documentation
abstract
This work focuses on data mining applied to the clinical documentation domain. Diagnostic terms (DTs) are used as keywords to retrieve valuable information from electronic health records. Indeed, they are encoded manually by experts following the International Classification of Diseases (ICD). The goal of this work is to explore the aid of text mining on DT encoding. From the machine learning (ML) perspective, this is a high-dimensional classification task, as it comprises thousands of codes. This work delves into a robust representation of the instances to improve ML results. The proposed system is able to find the right ICD code among more than 1500 possible ICD codes with 92% precision for the main disease (primary class) and 88% for the main disease together with the nonessential modifiers (fully specified class). The methodology employed is simple and portable. According to the experts from public hospitals, the system is very useful in particular for documentation and pharmacosurveillance services. In fact, they reported an accuracy of 91.2% on a small randomly extracted test. Hence, together with this paper, we made the software publicly available in order to help the clinical and research community.
Aitziber Atutxa, Alicia Pérez, Arantza Casillas
IEEE J. Biomed. Health Informatics3
2017 Semi-supervised medical entity recognition: A study on Spanish and Swedish clinical corpora
Alicia Pérez, Rebecka Weegar, Arantza Casillas, Koldo Gojenola, Maite Oronoz, Hercules Dalianis
J. Biomed. Informatics3
2016 Clinical text mining for efficient extraction of drug-allergy reactions
abstract
This work focuses on the extraction of allergic drug reactions in electronic health records. The goal is to annotate a sub-class of cause-effect events, those in which drugs are causing allergies. Little work has carried out in this field, seldom for Spanish clinical text mining, which is, indeed, the aim of this work. We present two approaches: a rule-based method and another one based on machine learning. Both approaches incorporate semantic knowledge derived from FreeLing-Med, a software explicitly developed to parse texts in the medical domain. Having recognised the medical entities for a given record, the challenge stands on triggering the underlying allergies. To this end, the knowledge is expressed as a set of semantic, syntactic and structural features. Our best system, based on machine learning, obtained a precision of 0.90 with a recall of 0.87, outperforming a rule-based approach.
Arantza Casillas, Koldo Gojenola, Alicia Pérez, Maite Oronoz
BIBM1
2016 IXAmed-IE: On-line medical entity identification and ADR event extraction in Spanish
abstract
This work presents an on-line system developed for medical information extraction. The goal is to provide a web-based service addressed to the medical community for the efficient processing of electronic health records in Spanish and support the clinical decision making process. This tool assists in the identification of medical entities as well as cause-effect reactions in order to help professionals both to prevent and to document adverse drug events. So far, the prototype is in its early stage of testing and validation by experts from the Galdakao-Usansolo and Basurto hospitals from the Basque Sanitary System (Osakidetza).
Arantza Casillas, Arantza Díaz de Ilarraza, K. Fernandez, Koldo Gojenola, Maite Oronoz, Alicia Pérez, Sara Santiso Gonzáles
BIBM1
2016 Learning to extract adverse drug reaction events from electronic health records in Spanish
Arantza Casillas, Alicia Pérez, Maite Oronoz, Koldo Gojenola, Sara Santiso Gonzáles
Expert Syst. Appl.1
2015 Computer aided classification of diagnostic terms in spanish
Alicia Pérez, Koldo Gojenola, Arantza Casillas, Maite Oronoz, Arantza Díaz de Ilarraza
Expert Syst. Appl.3
2015 On the creation of a clinical gold standard corpus in Spanish: Mining adverse drug reactions
Maite Oronoz, Koldo Gojenola, Alicia Pérez, Arantza Díaz de Ilarraza, Arantza Casillas
J. Biomed. Informatics5
2013 Automatic Annotation of Medical Records in Spanish with Disease, Drug and Substance Names
Maite Oronoz, Arantza Casillas, Koldo Gojenola, Alicia Pérez
CIARP (2)2
2007 Multilingual news clustering: Feature translation vs. identification of cognate named entities
Soto Montalvo, Raquel Martínez-Unanue 0001, Arantza Casillas, Víctor Fresno-Fernández
Pattern Recognit. Lett.3
2006 Multilingual Document Clustering: An Heuristic Approach Based on Cognate Named Entities
abstract
This paper presents an approach for Multilingual Document Clustering in comparable corpora. The algorithm is of heuristic nature and it uses as unique evidence for clustering the identification of cognate named entities between both sides of the comparable corpora. One of the main advantages of this approach is that it does not depend on bilingual or multilingual resources. However, it depends on the possibility of identifying cognate named entities between the languages used in the corpus. An additional advantage of the approach is that it does not need any information about the right number of clusters; the algorithm calculates it. We have tested this approach with a comparable corpus of news written in English and Spanish. In addition, we have compared the results with a system which translates selected document features. The obtained results are encouraging.
Soto Montalvo, Raquel Martínez-Unanue 0001, Arantza Casillas, Víctor Fresno-Fernández
ACL3
2004 Sentence Alignment for Spanish-Basque Bitexts: Word Correspondences vs. Markup Similarity
Arantza Casillas, Idoia Fernández, Raquel Martínez-Unanue 0001
CICLing1
2004 Sampling and Feature Selection in a Genetic Algorithm for Document Clustering
Arantza Casillas, María Teresa González de Lena, Raquel Martínez-Unanue 0001
CICLing1
2004 Evaluation of Web Page Representations by Content Through Clustering
Arantza Casillas, Víctor Fresno-Fernández, María Teresa González de Lena, Raquel Martínez-Unanue 0001
SPIRE1
2003 Partitional Clustering Experiments with News Documents
Arantza Casillas, María Teresa González de Lena, Raquel Martínez-Unanue 0001
CICLing1
2003 Experiments with Linguistic Categories for Language Model Optimization
Arantza Casillas, Amparo Varona, M. Inés Torres
CICLing1
2002 Aligning Multiword Terms Using a Hybrid Approach
Arantza Casillas, Raquel Martínez-Unanue 0001
CICLing1
2002 Experiments with a Bilingual Document Generation Environment
Arantza Casillas, Raquel Martínez-Unanue 0001
CICLing1
2000 DTD-driven bilingual document generation
abstract
Extensively annotated bilingual parallel corpora can be exploited to feed editing tools that integrate the processes of document composition and translation. Here we discuss the architecture of an interactive editing tool that, on top of techniques common to most Translation Memory-based systems, applies the potential of SGML's DTDs to guide the process of bilingual document generation. Rather than employing just simple task-oriented mark-up, we selected a set of TEI's highly complex and versatile collection of tags to help disclose the underlying logical structure of documents in the test-corpus. DTDs were automatically induced and later integrated in the editing tool to provide the basic scheme for new documents.
Arantza Casillas, Joseba Abaitua, Raquel Martínez-Unanue 0001
INLG1
1998 Value added tagging for multilingual resources management
Joseba Abaitua, Arantza Casillas, Raquel Martínez-Unanue 0001
LREC2