Jake Lever

dblp:81/11416 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0001-8198-2939ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 A Systematic Mapping of Large Language Models as Feedback Provider in Higher Education
abstract
The rapid adoption of LLMs in higher education (HE) has raised critical questions about their effectiveness in feedback provision. This systematic scoping review explores key trends, research gaps, and the pedagogical alignment of LLM-generated feedback. The findings highlight the growing integration of LLMs in assessment practices, their potential to improve feedback quality, and the need for pedagogically structured implementation. This review offers valuable insights for researchers, educators, and institutions, supporting the strategic adoption of generative AI to optimize feedback systems and improve student learning outcomes.
Eyman A. Alyahyan, Mireilla Bikanga Ada, Jake Lever
ICALT3
2025 CeRTS: certainty retrieval token search in large language model clinical information extraction
abstract
OBJECTIVE: Large language models (LLMs) must effectively communicate their uncertainty to be viable in clinical settings. As such, the need for reliable uncertainty estimation grows increasingly urgent with the expanding use of LLMs for information extraction from electronic health records. Previous token-level uncertainty estimators have only used token probabilities within a single output sequence. Here, by leveraging the constraints of JSON output structure, we instead consider all likely sequences and their respective probabilities to obtain a more robust measure of model confidence. We develop Certainty Retrieval Token Search (CeRTS), a new uncertainty estimator for structured information extraction. METHODS: We evaluated CeRTS against a previous gold-standard uncertainty estimator when extracting clinical features from lung cancer discharge summaries across eight open-source LLMs. Calibration (Brier score) and discrimination (AUROC) were used to quantify performance. RESULTS: CeRTS surpassed the previous gold-standard estimator in discriminatory power across every model and achieved better calibration in most cases. CeRTS had the strongest agreement between model confidence and accuracy with Qwen-2.5. CONCLUSION: CeRTS enhances LLM-based information extraction from unstructured clinical text by assigning well-calibrated confidence scores to each extracted item, providing medical researchers with a quantitative measure of reliability at minimal additional cost. Although its performance was generally robust, CeRTS struggled with DeepSeek-R1, which we attribute to the model's Chain-of-Thought reasoning steps. Our evaluation focused on clinical data, but CeRTS can be applied to any domain requiring reliable uncertainty estimation.
Lars E. Schimmelpfennig, Kriti Bhattarai, Inez Y. Oh, Jake Lever, Obi L. Griffith, Malachi Griffith, Albert M. Lai, Zachary B. Abrams
J. Biomed. Informatics4
2023 Associating biological context with protein-protein interactions through text mining at PubMed scale
Daniel N. Sosa, Rogier Hintzen, Betty Xiong, Alex de Giorgio, Julien Fauqueur, Mark Davies, Jake Lever, Russ B. Altman
J. Biomed. Informatics7
2021 Repurposing biomedical informaticians for COVID-19
Daniel N. Sosa, Binbin Chen 0002, Amit Kaushal, Adam Lavertu, Jake Lever, Stefano E. Rensi, Russ B. Altman
J. Biomed. Informatics5
2021 Search and visualization of gene-drug-disease interactions for pharmacogenomics and precision medicine research using GeneDive
abstract
BACKGROUND: Understanding the relationships between genes, drugs, and disease states is at the core of pharmacogenomics. Two leading approaches for identifying these relationships in medical literature are: human expert led manual curation efforts, and modern data mining based automated approaches. The former generates small amounts of high-quality data, and the latter offers large volumes of mixed quality data. The algorithmically extracted relationships are often accompanied by supporting evidence, such as, confidence scores, source articles, and surrounding contexts (excerpts) from the articles, that can be used as data quality indicators. Tools that can leverage these quality indicators to help the user gain access to larger and high-quality data are needed. APPROACH: We introduce GeneDive, a web application for pharmacogenomics researchers and precision medicine practitioners that makes gene, disease, and drug interactions data easily accessible and usable. GeneDive is designed to meet three key objectives: (1) provide functionality to manage information-overload problem and facilitate easy assimilation of supporting evidence, (2) support longitudinal and exploratory research investigations, and (3) offer integration of user-provided interactions data without requiring data sharing. RESULTS: GeneDive offers multiple search modalities, visualizations, and other features that guide the user efficiently to the information of their interest. To facilitate exploratory research, GeneDive makes the supporting evidence and context for each interaction readily available and allows the data quality threshold to be controlled by the user as per their risk tolerance level. The interactive search-visualization loop enables relationship discoveries between diseases, genes, and drugs that might not be explicitly described in literature but are emergent from the source medical corpus and deductive reasoning. The ability to utilize user's data either in combination with the GeneDive native datasets or in isolation promotes richer data-driven exploration and discovery. These functionalities along with GeneDive's applicability for precision medicine, bringing the knowledge contained in biomedical literature to bear on particular clinical situations and improving patient care, are illustrated through detailed use cases. CONCLUSION: GeneDive is a comprehensive, broad-use biological interactions browser. The GeneDive application and information about its underlying system architecture are available at http://www.genedive.net. GeneDive Docker image is also available for download at this URL, allowing users to (1) import their own interaction data securely and privately; and (2) generate and test hypotheses across their own and other datasets.
Mike Wong 0001, Paul Previde, Jack Cole, Brook Thomas, Nayana Laxmeshwar, Emily K. Mallory, Jake Lever, Dragutin Petkovic, Russ B. Altman, Anagha Kulkarni 0001
J. Biomed. Informatics7
2019 LIONS: analysis suite for detecting and quantifying transposable element initiated transcription from RNA-seq
abstract
SUMMARY: Transposable elements (TEs) influence the evolution of novel transcriptional networks yet the specific and meaningful interpretation of how TE-derived transcriptional initiation contributes to the transcriptome has been marred by computational and methodological deficiencies. We developed LIONS for the analysis of RNA-seq data to specifically detect and quantify TE-initiated transcripts. AVAILABILITY AND IMPLEMENTATION: Source code, container, test data and instruction manual are freely available at www.github.com/ababaian/LIONS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Artem Babaian, I. Richard Thompson, Jake Lever, Liane Gagnier, Mohammad M. Karimi, Dixie L. Mager
Bioinform.3
2018 A collaborative filtering-based approach to biomedical knowledge discovery
abstract
Motivation: The increase in publication rates makes it challenging for an individual researcher to stay abreast of all relevant research in order to find novel research hypotheses. Literature-based discovery methods make use of knowledge graphs built using text mining and can infer future associations between biomedical concepts that will likely occur in new publications. These predictions are a valuable resource for researchers to explore a research topic. Current methods for prediction are based on the local structure of the knowledge graph. A method that uses global knowledge from across the knowledge graph needs to be developed in order to make knowledge discovery a frequently used tool by researchers. Results: We propose an approach based on the singular value decomposition (SVD) that is able to combine data from across the knowledge graph through a reduced representation. Using cooccurrence data extracted from published literature, we show that SVD performs better than the leading methods for scoring discoveries. We also show the diminishing predictive power of knowledge discovery as we compare our predictions with real associations that appear further into the future. Finally, we examine the strengths and weaknesses of the SVD approach against another well-performing system using several predicted associations. Availability and implementation: All code and results files for this analysis can be accessed at https://github.com/jakelever/knowledgediscovery. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Jake Lever, Sitanshu Gakkhar, Michael Gottlieb, Tahereh Rashnavadi, Santina Lin, Celia Siu, Maia Smith, Martin R. Jones, Martin Krzywinski, Steven J. M. Jones
Bioinform.1
2012 Real-time controllable fire using textured forces
Jake Lever, Taku Komura
Vis. Comput.1