Minji Jeon

dblp:123/7967 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 From histology to spatial transcriptomics: establishing a lightweight single-patch baseline
abstract
Predicting spatial gene expression from hematoxylin and eosin (H&E)-stained histological images is a central challenge in the field of computational pathology. Although many models leverage the spatial context from entire whole-slide images to boost performance, the predictive capability achievable from single, isolated tissue patches remains unexplored. Without this fundamental baseline, it is difficult to fairly measure the marginal benefit of incorporating spatial context. In this study, we established a lightweight, efficient, and reproducible baseline for single-patch gene expression prediction, thereby providing a standardized foundation for future research. We trained various convolution-based pretrained architectures, such as EfficientNet, ResNet, and DenseNet, using a fixed data split and seed on a curated human liver Visium dataset. To assess cross-tissue generalizability, we further validated the proposed baseline on an independent breast cancer spatial transcriptomics dataset. Through full fine-tuning and morphology-preserving augmentation, EfficientNet-B0 achieved a maximum Pearson correlation coefficient (PCC) of 0.310 for highly expressed genes on liver tissue, surpassing previous single-patch methods. Furthermore, the model predicted 50 genes with PCC ≥ 0.30 using only 5.3M parameters--2.5× more genes than a ResNet-50-based baseline despite using 5× fewer parameters-suggesting a favorable performance-efficiency tradeoff within this single-patch setting. On the independent breast cancer dataset, EfficientNet-B0 consistently achieved best or near-best performance across all gene sets under both single-donor and multi-donor evaluations, demonstrating stable generalization across tissue types. Subsequent biological analysis reveals that model prediction accuracy is primarily driven by the strength of spatial organization in H&E histology, measured by Moran's I spatial autocorrelation, rather than transcript abundance or chemical properties. Consequently, this work provides a reference for quantifying the contributions of spatial context, multiresolution integration, and architectural complexity in future spatial transcriptomics models.
Hyungyum Jang, Hyunsoo Shin, Hawon Lee, Yena Jang, Sung Hoon Jung, Minji Jeon
BMC Bioinform.6
2025 LLaVA-Docent-V2: Improving Data Quality and Pedagogical Data Generation to Train Large Multimodal Models for Art Appreciation Education
Unggi Lee, Yoorim Son, Jaeyoon Shin, Gyuri Byun, Yunseo Lee, Junbo Koh, Minji Jeon, Hyeoncheol Kim
ITS (2)7
2025 ArcDFI: Attention regularization guided by CYP450 interactions for predicting drug-food interactions
abstract
CYP450 isoenzymes are known to be deeply involved in the formation of drug-food interactions (DFI). Previously introduced computational approaches for predicting DFIs do not take drug-CYP450 interactions (DCI) into account and have limited generalizability in handling compounds unseen during model training. We introduce ArcDFI, a model that utilizes attention regularization guided by CYP450 interactions to predict drug-food interactions. Experiments on DFI prediction-evaluated under stringent cold-drug and cold-food settings-show that our model outperforms ten baseline approaches, demonstrating the effectiveness of incorporating CYP450 interactions. Analysis of its attention mechanism provides insight into its current understanding of DCI and how they are related to its DFI predictions. To the best of our knowledge, ArcDFI is the first DFI prediction model that incorporates the concept of DCI, resulting in improved predictive generalizability and model explainability. ArcDFI is available at https://github.com/KU-MedAI/ArcDFI.
Keonwoo Kim 0002, Jaewoo Kang, Donghyeon Park, Minji Jeon
PLoS Comput. Biol.4
2023 Is Elementary AI Education Possible?
abstract
As artificial intelligence (AI) technology becomes increasingly pervasive, it is critical that students recognize AI and how it can be used. There is little research exploring learning capabilities of elementary students and the pedagogical supports necessary to facilitate students' learning. PrimaryAI was created as a 3rd-5th grade AI curriculum that utilizes problem-based and immersive learning within an authentic life science context through four units that cover machine learning, computer vision, AI planning, and AI ethics. The curriculum was implemented by two upper elementary teachers during Spring 2022. Based on pre-test/post-test results, students were able to conceptualize AI concepts related to machine learning and computer vision. Results showed no significant differences based on gender. Teachers indicated the curriculum engaged students and provided teachers with sufficient scaffolding to teach the content in their classrooms. Recommendations for future implementations include greater alignment between the AI and life science concepts, alterations to the immersive problem-based learning environment, and enhanced connections to local animal populations.
Anne T. Ottenbreit-Leftwich, Krista D. Glazewski, Cindy E. Hmelo-Silver, Katie Jantaraweragul, Minji Jeon, Srijita Chakraburty, J. Adam Scribner, Seung Y. Lee, Bradford W. Mott, James C. Lester
SIGCSE (2)5
2022 PrimaryAI: Co-Designing Immersive Problem-Based Learning for Upper Elementary Student Learning of AI Concepts and Practices
abstract
There is growing awareness of the central role that artificial intelligence (AI) plays now and in children's futures. This has led to increasing interest in engaging K-12 students in AI education to promote their understanding of AI concepts and practices. Leveraging principles from problem-based pedagogies and game-based learning, our approach integrates AI education into a set of unplugged activities and a game-based learning environment. In this work, we describe outcomes from our efforts to co design problem-based AI curriculum with elementary school teachers.
Krista D. Glazewski, Anne T. Ottenbreit-Leftwich, Katie Jantaraweragul, Minji Jeon, Cindy E. Hmelo-Silver, J. Adam Scribner, Seung Y. Lee, Bradford W. Mott, James C. Lester
ITiCSE (2)4
2022 Principles for AI Education for Elementary Grades Students
abstract
AI is beginning to transform every aspect of society. With the dramatic increases in AI, K-12 students need to be prepared to understand AI. To succeed as the workers, creators, and innovators of the future, students must be introduced to core concepts of AI as early as elementary school. However, building a curriculum that introduces AI content to K-12 students present significant challenges, such as connecting to prior knowledge, and developing curricula that are meaningful for students and possible for teachers to teach. To lay the groundwork for elementary AI education, we conducted a qualitative study into the design of AI curricular approaches with elementary teachers and students. Interviews with elementary teachers and students suggests four design principles for creating an effective elementary AI curriculum to promote uptake by teachers. This example will present the co-designed curriculum with teachers (PRIMARYAI) and describe how these four elements were incorporated into real-world problem-based learning scenarios.
Anne T. Ottenbreit-Leftwich, Krista D. Glazewski, Minji Jeon, Katie Jantaraweragul, Cindy E. Hmelo-Silver, J. Adam Scribner, Seung Y. Lee, Bradford W. Mott, James C. Lester
ITiCSE (2)3
2022 Transforming L1000 profiles to RNA-seq-like profiles with deep learning
abstract
The L1000 technology, a cost-effective high-throughput transcriptomics technology, has been applied to profile a collection of human cell lines for their gene expression response to > 30,000 chemical and genetic perturbations. In total, there are currently over 3 million available L1000 profiles. Such a dataset is invaluable for the discovery of drug and target candidates and for inferring mechanisms of action for small molecules. The L1000 assay only measures the mRNA expression of 978 landmark genes while 11,350 additional genes are computationally reliably inferred. The lack of full genome coverage limits knowledge discovery for half of the human protein coding genes, and the potential for integration with other transcriptomics profiling data. Here we present a Deep Learning two-step model that transforms L1000 profiles to RNA-seq-like profiles. The input to the model are the measured 978 landmark genes while the output is a vector of 23,614 RNA-seq-like gene expression profiles. The model first transforms the landmark genes into RNA-seq-like 978 gene profiles using a modified CycleGAN model applied to unpaired data. The transformed 978 RNA-seq-like landmark genes are then extrapolated into the full genome space with a fully connected neural network model. The two-step model achieves 0.914 Pearson's correlation coefficients and 1.167 root mean square errors when tested on a published paired L1000/RNA-seq dataset produced by the LINCS and GTEx programs. The processed RNA-seq-like profiles are made available for download, signature search, and gene centric reverse search with unique case studies.
Minji Jeon, Zhuorui Xie, John Erol Evangelista, Megan L. Wojciechowicz, Daniel J. B. Clarke, Avi Ma'ayan
BMC Bioinform.1
2021 Can Language Models be Biomedical Knowledge Bases?
abstract
Pre-trained language models (LMs) have become ubiquitous in solving various natural language processing (NLP) tasks.There has been increasing interest in what knowledge these LMs contain and how we can extract that knowledge, treating LMs as knowledge bases (KBs).While there has been much work on probing LMs in the general domain, there has been little attention to whether these powerful LMs can be used as domain-specific KBs.To this end, we create the BIOLAMA benchmark, which is comprised of 49K biomedical factual knowledge triples for probing biomedical LMs.We find that biomedical LMs with recently proposed probing methods can achieve up to 18.51% Acc@5 on retrieving biomedical knowledge.Although this seems promising given the task difficulty, our detailed analyses reveal that most predictions are highly correlated with prompt templates without any subjects, hence producing similar results on each relation and hindering their capabilities to be used as domain-specific KBs.We hope that BIOLAMA can serve as a challenging benchmark for biomedical factual probing. 1
Mujeen Sung, Jinhyuk Lee, Sean S. Yi, Minji Jeon, Sungdong Kim, Jaewoo Kang
EMNLP (1)4
2021 Document Analysis of ECEP Longitudinal Data: A Case Study with Indiana
abstract
In recent years, state members of the Expanding Computing Education Pathways (ECEP) Alliance have made efforts to increase access to and broaden participation in computing at the K-12 levels. Each ECEP state's K-12 computer science (CS) education journey has been documented during their ECEP membership resulting in over 25,000 digital documents. Over the course of the project it was necessary to track key events, identify trends across states, and maintain consumable records of state progress. A systematic way to collect and track the data is critical to conduct historical and cross-state analyses. In an effort to quantify and categorize, the researchers engaged in a review process of all ECEP reports, artifacts, and other relevant data to develop a system. Relevant and important components were identified in each type of document and assigned codes using ECEP's Five Stage Model ("a five-step process toward state-level CS education reform"), the Capacity, Access, Participation, and Experience (CAPE) framework (to measure equity in CS education implementation), and specific policies initiatives (alignment with various policy initiatives - Code.org's "Nine Policy Ideas to Make CS Fundamental to K?12 Education"). Indiana was identified as a state to conduct an initial, in-depth case study using this process. Indiana's case will be used as a model to further develop the stories of other ECEP Alliance member states. Through the development of a data dashboard, we hope to organize all of this information to make it more easily accessible for review and further analysis. The ECEP data dashboard development is currently in progress.
Minji Jeon, Jacob Koressel, Anne T. Ottenbreit-Leftwich, Alan Peterfreund, Sarah Dunton, Jeffrey Xavier, Carol L. Fletcher, Rebecca Zarch, Maureen Biggers, Debra J. Richardson, Joshua Childs, Leigh Ann Sudol-DeLyser, John Goodhue
SIGCSE1
2021 How do Elementary Students Conceptualize Artificial Intelligence?
abstract
Countries around the globe have acknowledged the pervasiveness of artificial intelligence (AI) in our lives and the importance of educating our students on how AI technologies work. For example, the Chinese education ministry has integrated AI into the mandatory high school curriculum, including a pilot textbook to teach students about the fundamental AI technologies like deep learning and recognition. However, there are fewer examples of AI education at the primary level, and these are typically focused on decision-making, machine learning, and programming. We have recently started to develop our own elementary AI curriculum. To develop the curriculum, we first needed to explore what students already knew about AI. Although there are a few small studies on how primary students conceptualize AI, we conducted a study to examine how 10 nine- and ten-year-old students conceptualized artificial intelligence by focusing on two main questions: (1) How do elementary students conceptualize artificial intelligence? (2) What are elementary students' experiences with artificial intelligence? Students? definitions of AI tended to focus on programming and robotics. Students described examples of AI that included robotic vacuums, Siri/Alexa, YouTube, and search engines. They also showcased some misconceptions around AI designs and implementations.
Anne T. Ottenbreit-Leftwich, Krista D. Glazewski, Minji Jeon, Cindy E. Hmelo-Silver, Bradford W. Mott, Seung Y. Lee, James C. Lester
SIGCSE3
2021 Crowdsourced identification of multi-target kinase inhibitors for RET- and TAU- based disease: The Multi-Targeting Drug DREAM Challenge
abstract
A continuing challenge in modern medicine is the identification of safer and more efficacious drugs. Precision therapeutics, which have one molecular target, have been long promised to be safer and more effective than traditional therapies. This approach has proven to be challenging for multiple reasons including lack of efficacy, rapidly acquired drug resistance, and narrow patient eligibility criteria. An alternative approach is the development of drugs that address the overall disease network by targeting multiple biological targets ('polypharmacology'). Rational development of these molecules will require improved methods for predicting single chemical structures that target multiple drug targets. To address this need, we developed the Multi-Targeting Drug DREAM Challenge, in which we challenged participants to predict single chemical entities that target pro-targets but avoid anti-targets for two unrelated diseases: RET-based tumors and a common form of inherited Tauopathy. Here, we report the results of this DREAM Challenge and the development of two neural network-based machine learning approaches that were applied to the challenge of rational polypharmacology. Together, these platforms provide a potentially useful first step towards developing lead therapeutic compounds that address disease complexity through rational polypharmacology.
Zhaoping Xiong, Minji Jeon, Robert J. Allaway, Jaewoo Kang, Donghyeon Park, Jinhyuk Lee, Hwisang Jeon, Miyoung Ko, Hualiang Jiang, Mingyue Zheng, Aik Choon Tan, Xindi Guo, Kristen K. Dang, Alexander Tropsha, Chana Hecht, Tirtha K. Das, Heather A. Carlson, Ruben Abagyan, Justin Guinney, Avner Schlessinger, Ross L. Cagan
PLoS Comput. Biol.2
2019 ReSimNet: drug response similarity prediction using Siamese neural networks
abstract
MOTIVATION: Traditional drug discovery approaches identify a target for a disease and find a compound that binds to the target. In this approach, structures of compounds are considered as the most important features because it is assumed that similar structures will bind to the same target. Therefore, structural analogs of the drugs that bind to the target are selected as drug candidates. However, even though compounds are not structural analogs, they may achieve the desired response. A new drug discovery method based on drug response, which can complement the structure-based methods, is needed. RESULTS: We implemented Siamese neural networks called ReSimNet that take as input two chemical compounds and predicts the CMap score of the two compounds, which we use to measure the transcriptional response similarity of the two compounds. ReSimNet learns the embedding vector of a chemical compound in a transcriptional response space. ReSimNet is trained to minimize the difference between the cosine similarity of the embedding vectors of the two compounds and the CMap score of the two compounds. ReSimNet can find pairs of compounds that are similar in response even though they may have dissimilar structures. In our quantitative evaluation, ReSimNet outperformed the baseline machine learning models. The ReSimNet ensemble model achieves a Pearson correlation of 0.518 and a precision@1% of 0.989. In addition, in the qualitative analysis, we tested ReSimNet on the ZINC15 database and showed that ReSimNet successfully identifies chemical compounds that are relevant to a prototype drug whose mechanism of action is known. AVAILABILITY AND IMPLEMENTATION: The source code and the pre-trained weights of ReSimNet are available at https://github.com/dmis-lab/ReSimNet. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Minji Jeon, Donghyeon Park, Jinhyuk Lee, Hwisang Jeon, Miyoung Ko, Sunkyu Kim, Yonghwa Choi, Aik Choon Tan, Jaewoo Kang
Bioinform.1
2016 HiPub: translating PubMed and PMC texts to networks for knowledge discovery
abstract
UNLABELLED: We introduce HiPub, a seamless Chrome browser plug-in that automatically recognizes, annotates and translates biomedical entities from texts into networks for knowledge discovery. Using a combination of two different named-entity recognition resources, HiPub can recognize genes, proteins, diseases, drugs, mutations and cell lines in texts, and achieve high precision and recall. HiPub extracts biomedical entity-relationships from texts to construct context-specific networks, and integrates existing network data from external databases for knowledge discovery. It allows users to add additional entities from related articles, as well as user-defined entities for discovering new and unexpected entity-relationships. HiPub provides functional enrichment analysis on the biomedical entity network, and link-outs to external resources to assist users in learning new entities and relations. AVAILABILITY AND IMPLEMENTATION: HiPub and detailed user guide are available at http://hipub.korea.ac.kr CONTACT: [email protected], [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Kyubum Lee, Won-Ho Shin, Byounggun Kim, Sunwon Lee, Yonghwa Choi, Sunkyu Kim, Minji Jeon, Aik Choon Tan, Jaewoo Kang
Bioinform.7
2015 DSigDB: drug signatures database for gene set analysis
abstract
UNLABELLED: We report the creation of Drug Signatures Database (DSigDB), a new gene set resource that relates drugs/compounds and their target genes, for gene set enrichment analysis (GSEA). DSigDB currently holds 22 527 gene sets, consists of 17 389 unique compounds covering 19 531 genes. We also developed an online DSigDB resource that allows users to search, view and download drugs/compounds and gene sets. DSigDB gene sets provide seamless integration to GSEA software for linking gene expressions with drugs/compounds for drug repurposing and translational research. AVAILABILITY AND IMPLEMENTATION: DSigDB is freely available for non-commercial use at http://tanlab.ucdenver.edu/DSigDB. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. CONTACT: [email protected].
Minjae Yoo, Jimin Shin, Jihye Kim 0004, Karen A. Ryall, Kyubum Lee, Sunwon Lee, Minji Jeon, Jaewoo Kang, Aik Choon Tan
Bioinform.7
2014 BEReX: Biomedical Entity-Relationship eXplorer
abstract
SUMMARY: Biomedical Entity-Relationship eXplorer (BEReX) is a new biomedical knowledge integration, search and exploration tool. BEReX integrates eight popular databases (STRING, DrugBank, KEGG, PhamGKB, BioGRID, GO, HPRD and MSigDB) and delineates an integrated network by combining the information available from these databases. Users search the integrated network by entering key words, and BEReX returns a sub-network matching the key words. The resulting graph can be explored interactively. BEReX allows users to find the shortest paths between two remote nodes, find the most relevant drugs, diseases, pathways and so on related to the current network, expand the network by particular types of entities and relations and modify the network by removing or adding selected nodes. BEReX is implemented as a standalone Java application. AVAILABILITY AND IMPLEMENTATION: BEReX and a detailed user guide are available for download at our project Web site (http://infos.korea.ac.kr/berex).
Minji Jeon, Sunwon Lee, Kyubum Lee, Aik Choon Tan, Jaewoo Kang
Bioinform.1
2012 Drug-drug interaction analysis using heterogeneous biological information network
abstract
As the number of drugs increases, more prescription choices are available for physicians, and consequently the number of drugs administered together has increased. Researchers are working on finding multi-drug prescriptions that are effective and safe. An efficient method for finding DDIs plays a crucial role in this research. In order to address the problem, we construct a heterogeneous biological information network by combining multiple different databases and interaction information. Our network includes the information about genes, proteins, pathways, drugs, side effects, targets and their interactions. We propose a metric to measure the relation strength between two nodes in the network, which is based on the weighted sum of the numbers of paths containing different interaction types. We use the metric to score DDI candidates. We found that the drugs sharing a disease are more likely to have a DDI than the drugs sharing a biomolecular target, and the metric using the weighted sum of the path numbers is effective to rank the potential DDIs. We validated the result with the PharmGKB DDI dataset and the Drugs.com drug interaction checker.
Kyubum Lee, Sunwon Lee, Minji Jeon, Jaewoo Kang
BIBM3