VLDB 2026 Research / reviewers in the wild / expert
Paea LePendu
dblp:40/1755
· DBLP profile ↗
23ranked-venue papers
5as first author
5since 2021 · last 2026
0000-0001-7358-931XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 1 first-authorDatabases, data management, data science and information retrieval · 7 · 3 first-authorHuman-computer interaction and ubiquitous computing · 5 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Late Test Takers Do Worse On Exams
Kelly Downey, Paea LePendu |
SIGCSE (2) | 2 |
| 2026 | What About Cheatsheets are Useful?
Kevin Loritsch, Kelly Downey, Paea LePendu |
SIGCSE (2) | 3 |
| 2026 | Creating a Second Pathway to the Computing MajorabstractIn 2021, the Computer Science and Engineering department at UC Riverside added a second pathway to the major with a new CS1 and CS2 course sequence. Our goal was to address disparities in course outcomes between different populations, particularly between those with and without prior coding experience and majors versus non-majors. The hope was that students new to computing would have a viable path to discover whether they had interest and aptitude in computing without the added stress of being in a classroom with students who have prior coding experience. In this paper, we report the data analysis that led to our decision to add a second pathway, design choices, and challenges (particularly in university politics) in launching and adding a second pathway. We present the results over the last three years, which illustrate that the new series has leveled the playing field. Most importantly, the course outcome data and changes to computing major enrollment illustrate that we achieved our goal of attracting students new to computing and thus are broadening participation in computing. Ashley Pang, Paea LePendu, Mariam Salloum, Neftali Watkinson Medina, Carla E. Brodley |
SIGCSE (1) | 2 |
| 2024 | KB2Bench: Toward a Benchmark Framework for Large Language Models on Medical KnowledgeabstractWhile Large Language Models (LLMs) have trans-formed question answering tasks, their propensity for hallucinations continues to drive an area of active research. Efforts toward creating benchmarks to test LLMs' performance on queries, in particular for the field of medicine, have led to a few reputable benchmarks, but these are limited in scope because of the amount of human annotation required. Our framework addresses this issue by leveraging existing, large knowledge bases for medicine to generate vast query and answer sets dynamically, which are less likely to be memorized by LLMs. The framework rests on designing a few key knowledge patterns, which can then generate millions (potentially billions) of queries. This offers a more efficient, cost-effective, and scalable alternative to human-curated annotations used in medical question-and-answer benchmarks. Applying our framework to a small sample of five drug related ontologies, we are already capable of more than 100,000 unique drug related queries, which is 10 to 1000 times larger than existing various human annotation efforts. This paper introduces the KB2Bench framework. Douglas Adjei-Frempah, Lisa Chen, Paea LePendu |
ICTAI | 3 |
| 2023 | Broadening Participation: The Data Science Academy for K-12abstractThe Data Science Academy (DSA) is an extra-curricular program for 7-12th graders that has evolved over the last four years with valuable lessons learned along the way. DSA was created in part because the K-12 curriculum is already packed, so deploying this informal learning environment and outreach program was one way to meet the challenge of the CS4ALL initiative and broaden participation in computing. DSA serves multiple purposes in that regard: it teaches teachers, it gives undergraduates mentoring experience, and it provides a platform for educational research and development. Currently, the DSA comprises five teaching modules, which have been repackaged and delivered as quarter or semester long weekend sessions, or shorter intensive summer programs. That is to say, the DSA accommodates flexible formats, including virtual or in-person ones and soon asynchronous options as well. The DSA provides an opportunity to develop and test lesson modules, including those derived from research projects in partnership with DS-PATH participants (an NSF project to create DS Pathways). Ultimately, the goal is to apply our experience with the DSA in order to expand the K-12 curriculum with new Data Science courses (grades 9-12), modules and pallets (grades 6-8), which is the next phase of our project. Taneesha Sharma, Paea LePendu |
SIGCSE (2) | 2 |
| 2020 | Summer Coding Camp as a Gateway to STEMabstractJust about everyone in the U.S., from the National Science Foundation down to local districts, has been pushing to introduce computer science concepts into K-12. Nevertheless, many students complete high school never having the chance to learn CS. We have created a summer coding camp for high-school students (including 8th graders entering 9th grade) and designed a multi-year study to assess its effectiveness as an informal learning environment, based on theories of human motivation such as Self-Determination Theory. The camp is a 1-week immersion experience, 9am to 5pm with food and activities, that introduces basic programming via MIT APP Inventor. Lecture material and in-class exercises draw upon meaningful applications, ones appealing to "social good." One unique aspect is the inclusion of professional and career development activities that engage students and broaden perspectives on CS and its applications. For example, the camp includes a college information session, alumni Skype and in-person talks, off-site visits to nearby companies, and research talks and demos by faculty. Using a pre-and-post survey design, the current study examines the effects of the camp on student self-efficacy and interest in computing, as well as general school engagement and motivation. Results confirm that participation in the summer camp increased students' self-efficacy and interest in computing, enhanced engagement in school on topics in general, and strengthened intrinsic motivation for completing schoolwork. The effects were similar for boys and girls. Paea LePendu, Cecilia Cheung, Mariam Salloum, Pamela Sheffler, Kelly Downey |
SIGCSE | 1 |
| 2018 | The HII-C Knowledge-Based App Project: Goals and Design Approach
Austin Michne, Andrew J. Solomon, Yael Ozair, Kaan Aksoy, Paea LePendu, Robert A. Greenes |
AMIA | 5 |
| 2015 | Functional evaluation of out-of-the-box text-mining tools for data-mining tasksabstractOBJECTIVE: The trade-off between the speed and simplicity of dictionary-based term recognition and the richer linguistic information provided by more advanced natural language processing (NLP) is an area of active discussion in clinical informatics. In this paper, we quantify this trade-off among text processing systems that make different trade-offs between speed and linguistic understanding. We tested both types of systems in three clinical research tasks: phase IV safety profiling of a drug, learning adverse drug-drug interactions, and learning used-to-treat relationships between drugs and indications. MATERIALS: We first benchmarked the accuracy of the NCBO Annotator and REVEAL in a manually annotated, publically available dataset from the 2008 i2b2 Obesity Challenge. We then applied the NCBO Annotator and REVEAL to 9 million clinical notes from the Stanford Translational Research Integrated Database Environment (STRIDE) and used the resulting data for three research tasks. RESULTS: There is no significant difference between using the NCBO Annotator and REVEAL in the results of the three research tasks when using large datasets. In one subtask, REVEAL achieved higher sensitivity with smaller datasets. CONCLUSIONS: For a variety of tasks, employing simple term recognition methods instead of advanced NLP methods results in little or no impact on accuracy when using large datasets. Simpler dictionary-based methods have the advantage of scaling well to very large datasets. Promoting the use of simple, dictionary-based methods for population level analyses can advance adoption of NLP in practice. Kenneth Jung, Paea LePendu, Srinivasan Iyer 0002, Anna Bauer-Mehren, Bethany Percha, Nigam H. Shah |
J. Am. Medical Informatics Assoc. | 2 |
| 2014 | Finding progression stages in time-evolving event sequencesabstractEvent sequences, such as patients' medical histories or users' sequences of product reviews, trace how individuals progress over time. Identifying common patterns, or progression stages, in such event sequences is a challenging task because not every individual follows the same evolutionary pattern, stages may have very different lengths, and individuals may progress at different rates. In this paper, we develop a model-based method for discovering common progression stages in general event sequences. We develop a generative model in which each sequence belongs to a class, and sequences from a given class pass through a common set of stages, where each sequence evolves at its own rate. We then develop a scalable algorithm to infer classes of sequences, while also segmenting each sequence into a set of stages. We evaluate our method on event sequences, ranging from patients' medical histories to online news and navigational traces from the Web. The evaluation shows that our methodology can predict future events in a sequence, while also accurately inferring meaningful progression stages, and effectively grouping sequences based on common progression patterns. More generally, our methodology allows us to reason about how event sequences progress over time, by discovering patterns and categories of temporal evolution in large-scale datasets of events. Jaewon Yang, Julian J. McAuley, Jure Leskovec, Paea LePendu, Nigam H. Shah |
WWW | 4 |
| 2014 | Toward personalizing treatment for depression: predicting diagnosis and severityabstractOBJECTIVE: Depression is a prevalent disorder difficult to diagnose and treat. In particular, depressed patients exhibit largely unpredictable responses to treatment. Toward the goal of personalizing treatment for depression, we develop and evaluate computational models that use electronic health record (EHR) data for predicting the diagnosis and severity of depression, and response to treatment. MATERIALS AND METHODS: We develop regression-based models for predicting depression, its severity, and response to treatment from EHR data, using structured diagnosis and medication codes as well as free-text clinical reports. We used two datasets: 35,000 patients (5000 depressed) from the Palo Alto Medical Foundation and 5651 patients treated for depression from the Group Health Research Institute. RESULTS: Our models are able to predict a future diagnosis of depression up to 12 months in advance (area under the receiver operating characteristic curve (AUC) 0.70-0.80). We can differentiate patients with severe baseline depression from those with minimal or mild baseline depression (AUC 0.72). Baseline depression severity was the strongest predictor of treatment response for medication and psychotherapy. CONCLUSIONS: It is possible to use EHR data to predict a diagnosis of depression up to 12 months in advance and to differentiate between extreme baseline levels of depression. The models use commonly available data on diagnosis, medication, and clinical progress notes, making them easily portable. The ability to automatically determine severity can facilitate assembly of large patient cohorts with similar severity from multiple sites, which may enable elucidation of the moderators of treatment response in the future. Sandy Huang, Paea LePendu, Srinivasan Iyer 0002, Ming Tai-Seale, David Carrell, Nigam H. Shah |
J. Am. Medical Informatics Assoc. | 2 |
| 2014 | Mining clinical text for signals of adverse drug-drug interactionsabstractBACKGROUND AND OBJECTIVE: Electronic health records (EHRs) are increasingly being used to complement the FDA Adverse Event Reporting System (FAERS) and to enable active pharmacovigilance. Over 30% of all adverse drug reactions are caused by drug-drug interactions (DDIs) and result in significant morbidity every year, making their early identification vital. We present an approach for identifying DDI signals directly from the textual portion of EHRs. METHODS: We recognize mentions of drug and event concepts from over 50 million clinical notes from two sites to create a timeline of concept mentions for each patient. We then use adjusted disproportionality ratios to identify significant drug-drug-event associations among 1165 drugs and 14 adverse events. To validate our results, we evaluate our performance on a gold standard of 1698 DDIs curated from existing knowledge bases, as well as with signaling DDI associations directly from FAERS using established methods. RESULTS: Our method achieves good performance, as measured by our gold standard (area under the receiver operator characteristic (ROC) curve >80%), on two independent EHR datasets and the performance is comparable to that of signaling DDIs from FAERS. We demonstrate the utility of our method for early detection of DDIs and for identifying alternatives for risky drug combinations. Finally, we publish a first of its kind database of population event rates among patients on drug combinations based on an EHR corpus. CONCLUSIONS: It is feasible to identify DDI signals and estimate the rate of adverse events among patients on drug combinations, directly from clinical text; this could have utility in prioritizing drug interaction surveillance as well as in clinical decision support. Srinivasan Iyer 0002, Rave Harpaz, Paea LePendu, Anna Bauer-Mehren, Nigam H. Shah |
J. Am. Medical Informatics Assoc. | 3 |
| 2014 | Cross-domain targeted ontology subsets for annotation: The case of SNOMED CORE and RxNorm
Pablo López-García, Paea LePendu, Mark A. Musen, Arantza Illarramendi |
J. Biomed. Informatics | 2 |
| 2013 | Predictive Models in Mental Health: From Diagnosis to Treatment
Sandy Huang, Paea LePendu, Srinivasan Iyer 0002, Ming Tai-Seale, David Carrell, Nigam H. Shah |
AMIA | 2 |
| 2013 | Learning Practice-based Evidence from Unstructured Clinical Notes
Nigam H. Shah, Paea LePendu, Anna Bauer-Mehren, Srinivasan Iyer 0002, Kenneth Jung, Tyler Cole, Rave Harpaz |
AMIA | 2 |
| 2013 | Mining Biomedical Ontologies and Data Using RDF HypergraphsabstractAs researchers analyze huge amounts of data that are annotated by large biomedical ontologies, one of the major challenges for data mining and machine learning is to leverage both ontologies and data together in a systematic and scalable way. In this paper, we address two interesting and related problems for mining biomedical ontologies and data: i) how to discover semantic associations with the help of formal ontologies, ii) how to identify potential errors in the ontologies with the help of data. By representing both ontologies and data using RDF hyper graphs, and subsequently transforming the hyper graphs to corresponding bipartite forms, we provide a generalized data mining method that scales beyond what existing ontology-based approaches can provide. We show the proposed method is indeed capable of capturing semantic associations while seamlessly incorporate domain knowledge in ontologies by performing evaluations on real-world electronic health dataset and NCBO ontologies. We also show that our data mining methods can discover and suggest corrections for misinformation in biomedical ontologies. Haishan Liu, Dejing Dou, Ruoming Jin, Paea LePendu, Nigam H. Shah |
ICMLA (1) | 4 |
| 2013 | Empirical bayes model to combine signals of adverse drug reactionsabstractData mining is a crucial tool for identifying risk signals of potential adverse drug reactions (ADRs). However, mining of ADR signals is currently limited to leveraging a single data source at a time. It is widely believed that combining ADR evidence from multiple data sources will result in a more accurate risk identification system. We present a methodology based on empirical Bayes modeling to combine ADR signals mined from ~5 million adverse event reports collected by the FDA, and healthcare data corresponding to 46 million patients' the main two types of information sources currently employed for signal detection. Based on four sets of test cases (gold standard), we demonstrate that our method leads to a statistically significant and substantial improvement in signal detection accuracy, averaging 40% over the use of each source independently, and an area under the ROC curve of 0.87. We also compare the method with alternative supervised learning approaches, and argue that our approach is preferable as it does not require labeled (training) samples whose availability is currently limited. To our knowledge, this is the first effort to combine signals from these two complementary data sources, and to demonstrate the benefits of a computationally integrative strategy for drug safety surveillance. Rave Harpaz, William DuMouchel, Paea LePendu, Nigam H. Shah |
KDD | 3 |
| 2012 | Using ontology-based annotation to profile disease researchabstractBACKGROUND: Profiling the allocation and trend of research activity is of interest to funding agencies, administrators, and researchers. However, the lack of a common classification system hinders the comprehensive and systematic profiling of research activities. This study introduces ontology-based annotation as a method to overcome this difficulty. Analyzing over a decade of funding data and publication data, the trends of disease research are profiled across topics, across institutions, and over time. RESULTS: This study introduces and explores the notions of research sponsorship and allocation and shows that leaders of research activity can be identified within specific disease areas of interest, such as those with high mortality or high sponsorship. The funding profiles of disease topics readily cluster themselves in agreement with the ontology hierarchy and closely mirror the funding agency priorities. Finally, four temporal trends are identified among research topics. CONCLUSIONS: This work utilizes disease ontology (DO)-based annotation to profile effectively the landscape of biomedical research activity. By using DO in this manner a use-case driven mechanism is also proposed to evaluate the utility of classification hierarchies. Adrien Coulet, Paea LePendu, Nigam H. Shah |
J. Am. Medical Informatics Assoc. | 3 |
| 2011 | A Hypergraph-based Method for Discovering Semantically Associated ItemsetsabstractIn this paper, we address an interesting data mining problem of finding semantically associated itemsets, i.e., items connected via indirect links. We propose a novel method for discovering semantically associated itemsets based on a hypergraph representation of the database. We describe two similarity measures to compute the strength of associations between items. Specifically, we introduce the average commute time similarity, sCT, based on the random walk model on hypergraph, and the inner-product similarity, sL+, based on the Moore-Penrose pseudoinverse of the hypergraph Laplacian matrix. Given semantically associated 2-itemsets generated by these measures, we design a hypergraph expansion method with two search strategies, namely, the clique and connected component search, to generate k-itemsets (k >; 2). We show the proposed method is indeed capable of capturing semantically associated itemsets through experiments performed on three datasets ranging from low to high dimensionality. The semantically associated itemsets discovered in our experiment is promising to provide valuable insights on interrelationship between medical concepts and other domain specific concepts. Haishan Liu, Paea LePendu, Ruoming Jin, Dejing Dou |
ICDM | 2 |
| 2011 | Enabling enrichment analysis with the Human Disease Ontology
Paea LePendu, Mark A. Musen, Nigam H. Shah |
J. Biomed. Informatics | 1 |
| 2011 | Using ontology databases for scalable query answering, inconsistency detection, and data integration
Paea LePendu, Dejing Dou |
J. Intell. Inf. Syst. | 1 |
| 2011 | NCBO Resource Index: Ontology-based search and mining of biomedical resources
Clément Jonquet, Paea LePendu, Sean M. Falconer, Adrien Coulet, Natasha F. Noy, Mark A. Musen, Nigam H. Shah |
J. Web Semant. | 2 |
| 2010 | Optimize First, Buy Later: Analyzing Metrics to Ramp-Up Very Large Knowledge Bases
Paea LePendu, Natasha F. Noy, Clément Jonquet, Paul R. Alexander, Nigam H. Shah, Mark A. Musen |
ISWC (1) | 1 |
| 2008 | Ontology Database: A New Method for Semantic Modeling and an Application to Brainwave Data
Paea LePendu, Dejing Dou, Gwen A. Frishkoff, Jiawei Rong |
SSDBM | 1 |