Haoze Du

dblp:318/3909 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Generative AI-Driven Mechanism for Pan-Cancer Drug Molecule Generation
abstract
Traditional drug discovery is a time-consuming and costly endeavor. To address this challenge, this research developed a deep-learning-based conditional generative model. This model is designed to generate small molecules guided by specific protein sequences (e.g., EGFR), endowing them with targeted pharmacological activity, high novelty, uniqueness, and promising drug-likeness, thereby accelerating pan-cancer drug molecule design. The study employed a hybrid architecture combining Graph Neural Networks (GNNs) for processing molecular structures and a Transformer model for encoding protein sequences, which together provide conditional guidance for the generation process. The model's performance was evaluated on a validation set using metrics such as training loss curves, Mean Squared Error (MSE), and$\mathrm{R}^{2}$(coefficient of determination). The quality of the generated molecules was assessed based on their validity, uniqueness, novelty, and the number of violations against Lipinski's Rule of Five. The results indicate that the model training was stable and convergent, demonstrating good generalization ability. The generated molecules exhibited excellent novelty (100%) and drug-likeness (with an average of zero Lipinski's rule violations), while achieving 100% uniqueness. Although the validity rate was 22%, suggesting room for improvement, these findings collectively underscore the model's significant potential for conditional innovative drug molecule design. This work lays a solid foundation for future model optimization to enhance generation efficiency and explore its broader applications in drug discovery.
Chongyang Ma, Haoze Du, Xianfang Wang
BIBM2
2025 Objective Metrics for Evaluating Large Language Models Using External Data Sources
Haoze Du, Edward F. Gehringer
EDM1
2025 CT-Semi-net: Segmentation of Infected Areas in Lung CT Images Based on Attention Mechanism and Semi-supervised Learning
Haoze Du, Shumei Hou, Junliang Du, Qingkai Hu, Weifeng Guo, Xianfang Wang
ISBRA (2)1
2025 Enhancing Drug Synergy Combination: Integrating Graph Transformers and BiLSTM for Accurate Drug Synergy Prediction
abstract
Combination therapy of drugs showed significant potential in treating complex diseases by overcoming drug resistance and improving therapeutic efficacy. However, due to the rapid increase in the number of available drugs, the cost and time required for experimentally screening synergistic drug combinations became increasingly burdensome. In this work, we proposed a novel drug synergy prediction model called GraphTranSynergy, which utilized graph transformer and BiLSTM to capture the molecular structure of drugs and gene expression features of cell lines. GraphTranSynergy extracted graphical features of drug pairs through the graph transformer module and integrated information from the BiLSTM module to extract useful features from gene expression profiles of cell lines. The final prediction of drug synergy was made through a fully connected neural network. Our model achieved AUC and PRAUC scores of 0.94, outperforming most existing models. Independent test results demonstrated that GraphTranSynergy exhibited superior generalization ability on the AstraZeneca dataset, particularly excelling in ACC and TPR metrics. Through a series of experiments and analyses, our model not only improved prediction accuracy but also demonstrated advantages in biological interpretability.
Haoze Du, Shumei Hou, Qingkai Hu, Xiaoxiao Pang, Xianfang Wang
IEEE J. Biomed. Health Informatics2
2024 LLM-generated Feedback in Real Classes and Beyond: Perspectives from Students and Instructors
Qinjin Jia, Jialin Cui, Haoze Du, M. Parvez Rashid, Ruijie Xi, Ruochi Li, Edward F. Gehringer
EDM3
2024 Interactive Rubric Generator for Instructor's Assessment Using Prompt Engineering and Large Language Models
abstract
This research paper describes an interactive system using large language models and prompt engineering to generate rubrics. Rubrics have long been employed to ensure a grading system that is both equitable and consistent. In practice, generating rubrics could be challenging for instructors for many reasons (e.g., a tight course schedule, limited resources, and varying materials for different projects in the same course), which urges the need to generate the rubric automatically. To the best of our knowledge, little research has been performed on generating rubrics. In this work, we present a novel system based on Large Language Models (LLMs) and Prompt Engineering to help instructors generate rubric items interactively based on course materials, as well as assess the student's work using these rubrics to give timely feedback automatically. In this system, we applied several LLMs (e.g., GPT4, Llama, Falcon, and Hermes) to generate both rubric and feedback using this process: 1) a set of text chunks are initially generated from the textual materials (these textual materials may from various sources), then LDA (Latent Dirichlet Allocation) is applied to extract a set of keywords from the preprocessed text chunks for rubric generation; 2) a web page was designed to let the instructor choose if the keywords from the set are adequate as rubric words; 3) the rubric items are generated by LLMs from the rubric words. In our experiments, a total number of 1017 documents (including the syllabus, the course website, the requirement of projects, the students' works, and the instructors' feedback) were used to build the corpus to generate the rubric-related keywords. Three users (including one instructor and two teaching assistants) participated in generating the rubric interactively using the webpage. The results of experiments show that the interactively generated rubrics from the LLM-instructor system can achieve a level similar to manually created rubrics. We utilized different prompts to let the LLMs generate feedback for the student's work, based on the generated rubrics. Our study shows that generating automatic rubrics and feedback for student project reports is feasible, yet it also identifies significant challenges that future research needs to address.
Haoze Du, Parvez Rashid, Qinjin Jia, Edward F. Gehringer
FIE1