Roselyne Tchoua

dblp:10/1921 · also Roselyne B. Tchoua, Roselyne Barreto · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0002-7195-4456ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 9 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Enhancing Pulmonary Nodule Localization Based on Latent Representations
abstract
Accurately localizing pulmonary nodules relative to other anatomical structures is crucial for disease management, guiding biopsies, and formulating effective treatment strategies. This study introduces a fully automated approach for classifying nodules detected in computed tomography (CT) images as pleural (near the pleura as$N_{p}$) or non-pleural (distant from the pleura as$D_{p}$). We propose a combination of Principal Component Analysis (PCA) and deep learning approach to determine a threshold for the optimal correlation between the original image and the latent PCA-based image reconstruction used to classify lung nodule as$N_{p}$or$D_{p}$. Applying our methodology to the Lung Image Database Consortium and Image Database Resource Initiative (LIDC-IDRI) dataset, we found that the PCA-based approach demonstrated substantial agreement with human evaluations of the proximity of the lung nodule to the pleura, achieving a Cohen's Kappa value of 0.76, outperforming an intensitybased baseline model with a Cohen's Kappa value of 0.55. Further analysis revealed significant variability in radiologists' semantic ratings for$N_{p}$versus$D_{p}$, with the highest variability observed in the texture feature. These findings demonstrate that our approach has the potential to enhance the classification of lung nodule localization when integrated into computer-aided diagnosis (CAD) systems.
Charmi Patel, Yiyang Wang 0003, Roselyne Tchoua, Thiruvarangan Ramaraj, Jacob D. Furst, Daniela Raicu
CBMS3
2025 Novel Perspective on Ensemble Clustering for Persistent Cluster Patterns: A Case Study in Disease Cluster Discovery
abstract
Clustering is the process of finding natural groups within a data set such that patterns within a group are more similar than patterns belonging to different groups. It has been used in a wide range of scientific and engineering disciplines. Yet, clustering is also a difficult unsupervised problem without an absolute ground truth. In practice, the "best" clustering method is the one that produces the most interpretable results. There is no universal optimum way to select the number of clusters. Our ultimate goal is to enable domain experts in high-stakes fields to make informed decisions when choosing a final, insightful, and actionable clustering solution for complex problems. In such challenging scenarios, ensemble clustering is a technique where multiple clustering results are combined to produce a more robust and stable final clustering. Here, inspired by a medical application, we take a novel approach to ensemble clustering; rather than focusing on an optimum global clustering, we identify persistent cluster(s) by multiple techniques across multiple clustering experiments among patients from Rush University Medical Center in Chicago. The key novelty resides in reducing the uncertainty in clustering, in the absence of ground truth, by looking at persistent clusters across the ensemble clustering. In healthcare, physicians aim to treat a patient based on a series of symptoms that contribute to a single disease. Because most clinical guidelines focus on individual diseases, managing patients with multimorbidity—when a person has multiple (chronic) conditions—can be challenging. We aim to discover clusters of diseases with some level of confidence to assist physicians in diagnosing and treating patients with multimorbidity. In this case study, we identify and thoroughly characterize a cluster of diseases that affects primarily women and persists through clustering the data from k=16 to k=20 clusters using three clustering techniques. Additionally, our first proposed multimorbidity cluster identified a network connecting asthma, breast cancer, hip/pelvic fracture, endometrial cancer, and non-Alzheimer’s dementia. We use multiple clustering techniques, consensus metrics, as well as graph analysis and visualization to provide confidence and foster trust from domain experts in our results. We are working with physicians to provide a medical explanation for this data-driven disease multimorbidity discovery.
Fernanda Cabral, Maria Alexandra Hubbard, Rudhvish Patel, Adeola Badmos, Daniela Raicu, Raj Shah, Roselyne Tchoua
eScience7
2025 Leveraging Hidden Patterns in Open-Ended Community Health Workers' Notes to Improve Prediction of Patient Readmission
abstract
Patient readmissions to Emergency Departments (EDs) pose significant challenges to healthcare systems, often indicating suboptimal care transitions, inadequate patient education, or insufficient post-discharge support. These unplanned readmissions not only compromise patient outcomes but also contribute to escalating healthcare costs. In response, healthcare providers are increasingly seeking predictive models to identify at-risk patients and implement preventive strategies. While advanced deep learning models have shown promise in predicting readmission risks, their "black-box" nature and substantial data requirements often limit clinical applicability due to a lack of interpretability. This study investigates the integration of Machine Learning (ML) and natural language processing (NLP) techniques to enhance the accuracy and explainability of patient readmission risk predictions. Utilizing data from the Sinai Urban Health Institute (SUHI), we analysed both structured data including demographics, interaction logs, social determinants of health (SDoH) survey answers and unstructured data, specifically notes capturing conversations between patients and Community Health Workers (CHWs). Our findings indicate that incorporating unstructured textual data improved model performance, with the area under the receiver operating characteristic curve (AUC) increasing from 0.68 to 0.74. This enhancement suggests that patient-CHW conversations capture critical, non-medical factors influencing readmissions, such as personal needs and social support deficits, which are not typically recorded in standard medical records. The study underscores the value of integrating patient-CHW contact notes into predictive modelling, helping to highlight the important role of community health in informing targeted interventions to reduce preventable readmissions.
Naveen Kumar Reddy Veeramreddy, Ankita Mishra, Navika Maglani, Sameer Shaik, Kelly McCabe, Jacob D. Furst, Daniela Raicu, Roselyne Tchoua, Jamshid Sourati
eScience8
2023 Curriculum gDRO: Improving Lung Malignancy Classification through Robust Curriculum Task Learning
abstract
Deep learning models used in Computer-Aided Diagnosis (CAD) systems are often trained with Empirical Risk Minimization (ERM) loss. These models often achieve high overall classification accuracy but with lower classification accuracy on certain subgroups. In the context of lung nodule malignancy classification task, these atypical subgroups exist due to the lung cancer heterogeneity. In this study, we characterize lung nodule malignancy subgroups using the malignancy likelihood ratings given by radiologists and improve the worst subgroup performance by utilizing group Distributionally Robust Optimization (gDRO). However, we noticed that gDRO improves on worst subgroup performance from the benign category, which has less clinical importance than improving classification accuracy for a malignant subgroup. Therefore, we propose a novel curriculum gDRO training scheme that trains for an “easy” task (nodule malignancy is determinate or indeterminate for radiologists) first, then for a “hard” task (malignant, benign, or indeterminate nodule). Our results indicate that our approach boosts the worst group subclass accuracy from the malignant category, by up to 6 percentage points compared to standard methods that address and improve worst group classification performance.
Arun Sivakumar, Yiyang Wang 0003, Roselyne Tchoua, Thiruvarangan Ramaraj, Jacob D. Furst, Daniela Raicu
CBMS3
2022 Failure Sources in Machine Learning for Medicine - A Study
abstract
Machine learning (ML) inherently suffers from at least a small amount of inaccuracy. Typically, these errors are acceptable in trade for either speed to an answer or the ability to find an answer at all. For high consequence domains, such as medicine where a wrong diagnosis can mean the difference between catching a disease early or not or prescribing debilitating treatment when it may not be needed, certain kinds and types errors are less acceptable. In a study attempting to reproduce ML for medicine research, many difficulties are encountered. These difficulties highlight both the need for higher standards to achieve reproducible ML in general and especially when it comes to high-stakes domains. This paper explores some of those difficulties with a focus on the error sources and discussions about how they may be addressed.
Hana Ahmed, Roselyne Tchoua, Jay F. Lofstead
e-Science2
2022 Text Summarization towards Scientific Information Extraction
abstract
Despite the exponential growth in scientific textual content, publications remain the primary means of disseminating vital research to experts within their respective fields. These texts are predominantly written for human consumption, resulting in two fundamental challenges; experts cannot efficiently remain well-informed to leverage the latest discoveries, and applications which rely on valuable insights buried in these texts cannot effectively build upon published results. Consequently, scientific progress stalls. Automatic Text Summarization (ATS) and Information Extraction (IE) are two essential fields which address this problem. While the two research topics are often studied independently, this work proposes to look at ATS in the context of IE, specifically as it relates to Scientific IE. However, Scientific Information Extraction faces several challenges; chiefly, the scarcity of relevant entities and insufficient training data. In this paper, we focus on extractive ATS, which identifies the most valuable sentences from textual content for the purpose of ultimately extracting scientific relations. We account for the associated challenges by means of an ensemble method through the integration of three weakly supervised learning models, one for each entity of the target relation. Notably, while the relation is well defined, we do not require previously annotated data for the entities composing the relation. The central objective is to generate balanced training data, which many advanced natural language processing models require. We apply this idea in the domain of materials science, extracting the polymer-glass transition temperature relation and achieve 94.7% recall (i.e., sentences which contain relations annotated by humans), while reducing the text by 99.3% of the original document.
Abigail Keller, Jacob D. Furst, Daniela Raicu, Peter M. Hastings, Roselyne Tchoua
e-Science5
2021 Ensemble Labeling towards Scientific Information Extraction (ELSIE) - Blob Extraction
abstract
Scientific publications constitute an extremely valuable repository of knowledge and collection of facts crucial to the advancement of science and development of applications, which grows as researchers learn from previous works and scientists use results in the literature to design and create. With the exponential growth of available publications, reading and extracting this wealth of information has become impractical for humans. Despite great progress in natural language processing, machine-learned solutions require large amounts of carefully annotated data for good performance. This is especially true in the context of accurately labeling and extracting complex scientific data. Towards our ultimate goal of extracting scientific facts from the literature, we first aim to identify blobs of text that contain all of the facts in a publication to be later automatically extracted or scrutinized by experts. Our previous work identified some facts missed by experts yet missed others due to the assumption that the target relation—here, a polymer and its glass transition temperature—would be contained within the same sentence. We set out to enhance our approximate labeling system to look back and ahead for missing information and successfully achieved 100% recall of scientific facts while reducing the full-text publication to 6% of its original size. Moreover, we assign confidence scores to sentences to further assist expert curators in identifying important sentences and facts locked in unstructured text.
Erin Murphy, Alexander Rasin, Jacob D. Furst, Daniela Raicu, Roselyne Tchoua
e-Science5
2020 Enhancing Recall Using Data Cleaning for Biomedical Big Data
abstract
In clinical practice, large amounts of heterogeneous medical data are generated on a daily basis. This data has the potential to be used for biomedical research and as a diagnostic reference for physicians. However, leveraging heterogeneous data for analysis requires integrating it first. Integration process includes a pre-processing data cleaning phase that eliminates inconsistencies and errors originating from each data source. In this paper, we describe a workflow for cleaning heterogeneous biomedical data sources. Our novel data cleaning approach can be applied for replacement of missing text and to improve the number of relevant cases retrieved by search queries. When the threshold for missing category replacement is met, our results show that our method achieves a missing content replacement precision of 85%, which represents an improvement of 18% over the baseline state of our datasets.
Priya Deshpande, Alexander Rasin, Roselyne Tchoua, Jacob D. Furst, Daniela Raicu, Sameer K. Antani
CBMS3
2019 Active Learning Yields Better Training Data for Scientific Named Entity Recognition
abstract
Despite significant progress in natural language processing, machine learning models require substantial expertannotated training data to perform well in tasks such as named entity recognition (NER) and entity relations extraction. Furthermore, NER is often more complicated when working with scientific text. For example, in polymer science, chemical structure may be encoded using nonstandard naming conventions, the same concept can be expressed using many different terms (synonymy), and authors may refer to polymers with ad-hoc labels. These challenges, which are not unique to polymer science, make it difficult to generate training data, as specialized skills are needed to label text correctly. We have previously designed polyNER, a semi-automated system for efficient identification of scientific entities in text. PolyNER applies word embedding models to generate entity-rich corpora for productive expert labeling, and then uses the resulting labeled data to bootstrap a context-based classifier. PolyNER facilitates a labeling process that is otherwise tedious and expensive. Here, we use active learning to efficiently obtain more annotations from experts and improve performance. Our approach requires just five hours of expert time to achieve discrimination capacity comparable to that of a state-of-the-art chemical NER toolkit.
Roselyne Tchoua, Aswathy Ajith, Zhi Hong, Logan T. Ward, Kyle Chard, Debra Audus, Shrayesh Patel, Juan de Pablo, Ian T. Foster
eScience1
2017 Towards a Hybrid Human-Computer Scientific Information Extraction Pipeline
abstract
The emerging field of materials informatics has the potential to greatly reduce time-to-market and development costs for new materials. The success of such efforts hinges on access to large, high-quality databases of material properties. However, many such data are only to be found encoded in text within esoteric scientific articles, a situation that makes automated extraction difficult and manual extraction time-consuming and error-prone. To address this challenge, we present a hybrid Information Extraction (IE) pipeline to improve the machine-human partnership with respect to extraction quality and person-hours, through a combination of rule-based, machine learning, and crowdsourcing approaches. Our goal is to leverage computer and human strengths to alleviate the burden on human curators by automating initial extraction tasks before prioritizing and assigning specialized curation tasks to humans with different levels of training: using non-experts for straightforward tasks such as validation of higher accuracy results (e.g., completing partial facts) and domain experts for low-certainty results (e.g., reviewing specialized compound labels). To validate our approaches, we focus on the task of extracting the glass transition temperature of polymers from published articles. Applying our approaches to 6 090 articles, we have so far extracted 259 refined data values. We project that this number will grow considerably as we tune our methods and process more articles, to exceed that found in standard, expert-curated polymer data handbooks while also being easier to keep up-to-date. The freely available data can be found on our Polymer Properties Predictor and Database website at http://pppdb.uchicago.edu.
Roselyne Tchoua, Kyle Chard, Debra Audus, Logan T. Ward, Joshua Lequieu, Juan de Pablo, Ian T. Foster
eScience1
2014 Hello ADIOS: the challenges and lessons of developing leadership class I/O frameworks
abstract
SUMMARY Applications running on leadership platforms are more and more bottlenecked by storage input/output (I/O). In an effort to combat the increasing disparity between I/O throughput and compute capability, we created Adaptable IO System (ADIOS) in 2005. Focusing on putting users first with a service oriented architecture, we combined cutting edge research into new I/O techniques with a design effort to create near optimal I/O methods. As a result, ADIOS provides the highest level of synchronous I/O performance for a number of mission critical applications at various Department of Energy Leadership Computing Facilities. Meanwhile ADIOS is leading the push for next generation techniques including staging and data processing pipelines. In this paper, we describe the startling observations we have made in the last half decade of I/O research and development, and elaborate the lessons we have learned along this journey. We also detail some of the challenges that remain as we look toward the coming Exascale era. Copyright © 2013 John Wiley & Sons, Ltd.
Qing Liu 0002, Jeremy Logan, Yuan Tian 0004, Hasan Abbasi, Norbert Podhorszki, Jong Choi 0001, Scott Klasky, Roselyne Tchoua, Jay F. Lofstead, Ron A. Oldfield, Manish Parashar, Nagiza F. Samatova, Karsten Schwan, Arie Shoshani, Matthew Wolf, Kesheng Wu, Weikuan Yu
Concurr. Comput. Pract. Exp.8
2013 ADIOS Visualization Schema: A First Step Towards Improving Interdisciplinary Collaboration in High Performance Computing
abstract
Scientific communities have benefitted from a significant increase of available computing and storage resources in the last few decades. For science projects that have access to leadership scale computing resources, the capacity to produce data has been growing exponentially. Teams working on such projects must now include, in addition to the traditional application scientists, experts in various disciplines including applied mathematicians for development of algorithms, visualization specialists for large data, and I/O specialists. Sharing of knowledge and data is becoming a requirement for scientific discovery, providing useful mechanisms to facilitate this sharing is a key challenge for e-Science. Our hypothesis is that in order to decrease the time to solution for application scientists we need to lower the barrier of entry into related computing fields. We aim at improving users' experience when interacting with a vast software ecosystem and/or huge amount of data, while maintaining focus on their primary research field. In this context we present our approach to bridge the gap between the application scientists and the visualization experts through a visualization schema as a first step and proof of concept for a new way to look at interdisciplinary collaboration among scientists dealing with big data. The key to our approach is recognizing that our users are scientists who mostly work as islands. They tend to work in very specialized environment but occasionally have to collaborate with other researchers in order to take full advantage of computing innovations and get insight from big data. We present an example of identifying the connecting elements between one of such relationships and offer a liaison schema to facilitate their collaboration.
Roselyne Tchoua, Jong Choi 0001, Scott Klasky, Qing Liu 0002, Jeremy Logan, Kenneth Moreland, Jingqing Mu, Manish Parashar, Norbert Podhorszki, David Pugmire, Matthew Wolf
e-Science1
2010 EFFIS: An End-to-end Framework for Fusion Integrated Simulation
abstract
The purpose of the Fusion Simulation Project is to develop a predictive capability for integrated modeling of magnetically confined burning plasmas. In support of this mission, the Center for Plasma Edge Simulation has developed an End-to-end Framework for Fusion Integrated Simulation (EFFIS) that combines critical computer science technologies in an effective manner to support leadership class computing and the coupling of complex plasma physics models. We describe here the main components of EFFIS and how they are being utilized to address our goal of integrated predictive plasma edge simulation.
Julian C. Cummings, Jay F. Lofstead, Karsten Schwan, Alex Sim, Arie Shoshani, Ciprian Docan, Manish Parashar, Scott Klasky, Norbert Podhorszki, Roselyne Tchoua
PDP10
2009 Enabling Advanced Visualization Tools in a Web-Based Simulation Monitoring System
abstract
Simulations that require massive amounts of computing power and generate tens of terabytes of data are now part of the daily lives of scientists. Analyzing and visualizing the results of these simulations as they are computed can lead not only to early insights but also to useful knowledge that can be provided as feedback to the simulation, avoiding unnecessary use of computing power. Our work is aimed at making advanced visualization tools available to scientists in a user-friendly, Web-based environment where they can be accessed anytime from anywhere. In the context of turbulent combustion for example, visualization is used to understand the coupling between turbulence and the turbulent mixing of scalars. Although isosurface generation is a useful technique in this scenario, computing and rendering isosurfaces one at a time is expensive and not particularly well-suited for such a Web-based framework. In this paper we propose the use of a summary structure, called contour tree, that captures the topological structure of a scalar field and guides the user in identifying useful isosurfaces. We have also designed an interface which has been integrated with a Web-based simulation monitoring system, that allows users to interact with and explore multiple isosurfaces.
Emanuele Santos, Julien Tierny, Ayla Khan, Brad Grimm, Lauro Didier Lins, Juliana Freire, Valerio Pascucci, Cláudio T. Silva, Scott Klasky, Roselyne Tchoua, Norbert Podhorszki
eScience10
2009 Tracking Files in the Kepler Provenance Framework
Pierre Mouallem, Roselyne Tchoua, Scott Klasky, Norbert Podhorszki, Mladen A. Vouk
SSDBM2