Jake Luo

dblp:157/6484 · also Zhihui Luo 0001 · DBLP profile ↗
← Back
27ranked-venue papers
3as first author
17since 2021 · last 2026
0000-0002-3900-643XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 11 · 10 since 2021Software engineering, systems software and programming languages · 6 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SymRefine: A symbolic regression approach for refining and compressing neural networks
Qiang Lu 0005, Can Huang 0009, Jake Luo
Neurocomputing5
2026 Sequential pattern transformer (SPT): a generative and interpretable framework for predicting disease trajectories
abstract
The effective integration of artificial intelligence into clinical workflows requires models that go beyond simple prediction to generate comprehensive, explainable, and actionable disease trajectories. Addressing the limitations of opaque deep learning architectures and the noise inherent in electronic health records, we introduce the sequential pattern transformer (SPT), a novel framework that synergizes sequential pattern mining with generative transformer modeling. Using four years of inpatient data from 258,460 type 2 diabetes patients, we applied the PrefixSpan algorithm to distill noisy diagnostic histories into a curated vocabulary of 95,630 statistically validated disease progression patterns. A decoder-only transformer was trained exclusively on these evidence-based sequences to learn the temporal dynamics of disease evolution. This pattern-guided approach shifts the modeling paradigm from classification to probabilistic trajectory generation. The model achieved a robust 85.78% Top-5 accuracy, significantly outperforming a standard LSTM baseline (71.47%). Beyond predictive accuracy, the framework constructs a dynamic Disease Atlas, a branching tree structure that visualizes likely future pathways, augmented by multi-level explainable AI (XAI) including learned clinical clusters, SHAP-based feature attribution, and counterfactual simulations. Crucially, this methodology is domain-agnostic and capable of efficient fine-tuning, making it a transferable solution for adapting to diverse clinical conditions and local hospital settings. SPT thus offers a transparent, robust, and scalable framework for mapping the complex temporal dynamics of disease, bridging the gap between high-performance AI and interpretable clinical application.
Mohammad Assadi Shalmani, Masoud Khani, Amirsajjad Taleban, Zihao Yi, Jennifer T. Fink, Christopher E. Weber, Qiang Lu 0005, Jake Luo
Neural Comput. Appl.8
2025 Voice-Activated Self-Monitoring Application (VoiS): User Acceptance and Satisfaction in the Field
abstract
This paper illustrates the processes and results of user acceptance and satisfaction tests for the voice-activated self-monitoring (VoiS) application conducted in the field. VoiS was designed and developed for individuals with diabetes (DM) and hypertension (HTN) to support their routine and convenient self-management using a smart speaker platform. VoiS is also accessible on users' mobile devices to visualize user-generated data. A total of nine adults with DM and HTN participated and were asked to use the VoiS system at home for a week to identify its acceptability in real-world conditions. Participants completed phone interviews to report operational errors and a structured survey to assess the acceptability of VoiS. The results showed that participants agreed or strongly agreed that VoiS was easy to use and useful, and they intended to continue using it. They were highly satisfied with VoiS. They also reported operational errors and challenges including device-related technical issues and the smart reminder feature of VoiS. Overall, participants perceived VoiS as an easy and useful tool for managing their conditions and felt motivated to monitor their biomarkers routinely.
Hyunkyoung Oh, Tala Abu Zahra, Shiyu Tian, Min Sook Park, Jake Luo, Sheikh Iqbal Ahamed, Evelyn Chan, Jeff Whittle
COMPSAC6
2025 End-to-End Multi-Modal Diffusion Mamba
abstract
Current end-to-end multi-modal models utilize different encoders and decoders to process input and output information. This separation hinders the joint representation learning of various modalities. To unify multi-modal processing, we propose a novel architecture called MDM (Multi-modal Diffusion Mamba). MDM utilizes a Mamba-based multi-step selection diffusion model to progressively generate and refine modality-specific information through a unified variational autoencoder for both encoding and decoding. This innovative approach allows MDM to achieve superior performance when processing high-dimensional data, particularly in generating high-resolution images and extended text sequences simultaneously. Our evaluations in areas such as image generation, image captioning, visual question answering, text comprehension, and reasoning tasks demonstrate that MDM significantly outperforms existing end-to-end models (MonoFormer, LlamaGen, and Chameleon etc.) and competes effectively with SOTA models like GPT-4V, Gemini Pro, and Mistral. Our results validate MDM's effectiveness in unifying multi-modal processes while maintaining computational efficiency, establishing a new direction for end-to-end multi-modal architectures.
Chunhao Lu, Qiang Lu 0005, Meichen Dong, Jake Luo
ICCV4
2025 Deep Differentiable Symbolic Regression Neural Network
Qiang Lu 0005, Yuanzhen Luo, Jake Luo, Zhiguang Wang
Neurocomputing4
2025 Discovering Acoustic Impedance Inversion Equation
abstract
Classical acoustic impedance inversion methods rely on mathematical and physical models to estimate subsurface acoustic impedance distribution. However, these methods face difficulties in accurately fitting complex impedance data. While deep learning methods have achieved higher accuracy and efficiency in impedance inversion, they remain black box models, lacking interpretability. So, it is difficult to analyze the reason why they are (or are not) effective. To address the limitations of both traditional and deep learning methods, this paper proposes a novel approach, AII-SR (Acoustic Impedance Inversion with Symbolic Regression), which discovers partial differential equations (PDEs) from impedance data to model acoustic impedance inversion. To discover these PDEs, AII-SR adopts a dual-learning framework. It employs a forward model based on the Robinson convolution principle to ensure the physical consistency and reliability of predictions. Subsequently, AII-SR creates an inversion model that combines a symbolic regression-based PDE generator with a physics-informed solving neural network (PSNN) to identify PDEs that accurately fit the impedance data. Experiments demonstrate that AII-SR outperforms deep learning methods, such as SSEI, TCN and Se-Unet, in terms of accuracy and interpretability. AII-SR generates concise, interpretable mathematical expressions in the form of PDEs, offering profound insights into the physical relationships between seismic and impedance data.
Baimou Li, Qiang Lu 0005, Jake Luo, Zhiguang Wang
IEEE Trans. Geosci. Remote. Sens.3
2024 Identifying Medical Concepts and Semantic Types in Lay Vocabularies of Health Consumers Who are Concerned with Diabetes on Social Media Using the UMLS and NLP
abstract
This study suggests a way to utilize the existing medical ontology and natural language processing techniques to extract major medical concepts from lay vocabularies of health consumers on social media and group them based on the defined semantic types in the ontology. Diabetes-related discussions on Tumblr was used to test the efficiency of SpaCy and the Markov-Viterbi algorithm to map lay medical terms to the defined medical concepts in the UMLS. The system discussed in this paper can better analyze free texts, take care of word ambiguity and extract the lifestyle indicators from the daily life discussions of diabetic people on Tumblr. The findings of this study can contribute to developing health applications that track the health behavior of those living with chronic conditions such as diabetes. This approach can also assist researchers who are interested in processing lay languages used by health consumers to foster an understanding of their health behavior.
Adib Ahmed Anik, Paramita Basak Upama, Masud Rabbani, Shiyu Tian, Min Sook Park, Sheikh Iqbal Ahamed, Jake Luo, Hyunkyoung Oh
COMPSAC7
2024 Differentiable Neural Network for Assembling Blocks
abstract
The goal of assembly blocks is to select blocks from pre-trained neural network (NN) models and combine them into a new NN for a different dataset. By reusing the weights of these blocks, training the NN with the new dataset becomes cost-effective. To achieve this goal, we propose an end-to-end differentiable neural network called PA-DNN. PA-DNN consists of two modules: a partition NN module and an assembly NN module. For the new dataset, the partition NN module divides existing pre-trained NN models into blocks. The assembly NN module then selects some of these blocks and combines them into a new NN using a stitching component. To train PA-DNN, we design a score function that evaluates the performance of each new NN generated by PA-DNN. The evaluated value is used to train the partition NN module. Additionally, two loss functions are created to train the assembly NN module and the stitching component in the new NN, respectively. After the training process, PA-DNN infers a new NN, and only the stitching component of the NN is fine-tuned with the new dataset. Experiments show that, compared to manual models, neural architecture search, and the assembly model DeRy, PA-DNN can generate a more accurate and lightweight NN with lower training costs.
Qiang Lu 0005, Yanhong Zhao, Jake Luo
ECAI5
2024 An Explainable Vision Question Answer Model via Diffusion Chain-of-Thought
Chunhao Lu, Qiang Lu 0005, Jake Luo
ECCV (67)3
2024 Alleviating Semantic Drift in Multi-Hop Question Answering on Knowledge Graphs with Bidirectional Semantics
abstract
Multi-hop question answering over knowledge graph utilizes the knowledge graph (KG) structure to infer answers. However, KG often lacks edges in the reasoning path from the question entity to the answer entity. Recent research focused on various KG embedding methods to obtain the semantics of the reasoning path (called forward semantics) to repair missing edges. However, the forward semantics method could drift as the path get longer. This paper proposes a bidirectional semantics embedding and matching method (BSEM) to alleviate the forward semantics drift problem. BSEM first leverages a backward semantics method to deduce the semantics of the opposite direction of the reasoning path. Then, BSEM constructs a two-stage learning method to merge the bidirectional (forward or backward) semantics and find the correct answer. In the two-stage learning method, joint learning is created to learn the bidirectional semantics of the reasoning path simultaneously; contrast learning is also used to improve the ability of the backward semantics to identify the correct answers that are not found by the forward semantics. Experiments on the two benchmarks, MetaQA and WebQSP, show that BSEM surpasses the five baseline methods, PullNet, EmQL, LEGO, EmbedKGQA and KGT5. Especially for the incomplete KG – WebQSP, compared with the other four methods except for EmQL, BSEM improves the accuracy by 13.1%, 12.0%, 5.4% and 10.0%, respectively.
Mingcai Yuan, Qiang Lu 0005, Xianhao Zeng, Jake Luo
IJCNN4
2024 Symbol Graph Genetic Programming for Symbolic Regression
Jinglu Song, Qiang Lu 0005, Bozhou Tian, Jake Luo, Zhiguang Wang
PPSN (1)5
2023 A Survey of Conversational Agents and Their Applications for Self-Management of Chronic Conditions
abstract
Conversational agents have gained their ground in our daily life and various domains including healthcare. Chronic condition self-management is one of the promising healthcare areas in which conversational agents demonstrate significant potential to contribute to alleviating healthcare burdens from chronic conditions. This survey paper introduces and outlines types of conversational agents, their generic architecture and workflow, the implemented technologies, and their application to chronic condition self-management.
Min Sook Park, Paramita Basak Upama, Adib Ahmed Anik, Sheikh Iqbal Ahamed, Jake Luo, Shiyu Tian, Masud Rabbani, Hyungkyoung Oh
COMPSAC5
2022 Towards Developing a Voice-activated Self-monitoring Application (VoiS) for Adults with Diabetes and Hypertension
abstract
The integration of motivational strategies and self-management theory with mHealth tools is a promising approach to changing the behavior of patients with chronic disease. In this manuscript, we describe the development and current architecture of a prototype voice-activated self-monitoring application (VoiS) which is based on these theories. Unlike prior mHealth applications which require textual input, VoiS app relies on the more convenient and adaptable approach of asking users to verbally input markers of diabetes and hypertension control through a smart speaker. The VoiS app can provide real-time feedback based on these markers; thus, it has the potential to serve as a remote, regular, source of feedback to support behavior change. To enhance the usability and acceptability of the VoiS application, we will ask a diverse group of patients to use it in real-world settings and provide feedback on their experience. We will use this feedback to optimize tool performance, so that it can provide patients with an improved understanding of their chronic conditions. The VoiS app can also facilitate remote sharing of chronic disease control with healthcare providers, which can improve clinical efficacy and reduce the urgency and frequency of clinical care encounters. Because the VoiS app will be configured for use with multiple platforms, it will be more robust than existing systems with respect to user accessibility and acceptability.
Masud Rabbani, Shiyu Tian, Adib Ahmed Anik, Jake Luo, Min Sook Park, Jeff Whittle, Sheikh Iqbal Ahamed, Hyunkyoung Oh
COMPSAC4
2022 A Clustering-Aided Approach for Diagnosis Prediction: A Case Study of Elderly Fall
abstract
Data-driven diagnosis prediction has been adopted in clinical decision support systems. However, only a few studies have focused on non-supervised clustering approaches to building a high-quality patient data set. This study focused on a clustering-aided approach to diagnosis prediction. We leveraged clustering-aided machine learning models to predict elderly falls. First, we used patients' risk factors to build a feature set. The feature set showed a clustering-aided approach could aggregate patient factors that shared similar clinical and demographic characteristics. Subsequently, a K-means clustering approach significantly improved the data set quality. Overall, our study demonstrated that clustering approaches improve the prediction performance of elderly falls. A clustering-aided approach can be applied to similar clinical healthcare practices to potentially improve elderly care.
Ling Tong 0002, Jake Luo, Jazzmyne Adams, Kristen Osinski, David R. Friedland
COMPSAC2
2022 Taylor genetic programming for symbolic regression
abstract
Genetic programming (GP) is a commonly used approach to solve symbolic regression (SR) problems. Compared with the machine learning or deep learning methods that depend on the pre-defined model and the training dataset for solving SR problems, GP is more focused on finding the solution in a search space. Although GP has good performance on large-scale benchmarks, it randomly transforms individuals to search results without taking advantage of the characteristics of the dataset. So, the search process of GP is usually slow, and the final results could be unstable. To guide GP by these characteristics, we propose a new method for SR, called Taylor genetic programming (TaylorGP)1. TaylorGP leverages a Taylor polynomial to approximate the symbolic equation that fits the dataset. It also utilizes the Taylor polynomial to extract the features of the symbolic equation: low order polynomial discrimination, variable separability, boundary, monotonic, and parity. GP is enhanced by these Taylor polynomial techniques. Experiments are conducted on three kinds of benchmarks: classical SR, machine learning, and physics. The experimental results show that TaylorGP not only has higher accuracy than the nine baseline methods, but also is faster in finding stable results.
Baihe He, Qiang Lu 0005, Qingyun Yang, Jake Luo, Zhiguang Wang
GECCO4
2022 Exploring hidden semantics in neural networks with symbolic regression
abstract
Many recent studies focus on developing mechanisms to explain the black-box behaviors of neural networks (NNs). However, little work has been done to extract the potential hidden semantics (mathematical representation) of a neural network. A succinct and explicit mathematical representation of a NN model could improve the understanding and interpretation of its behaviors. To address this need, we propose a novel symbolic regression method for neural works (called SRNet) to discover the mathematical expressions of a NN. SRNet creates a Cartesian genetic programming (NNCGP) to represent the hidden semantics of a single layer in a NN. It then leverages a multi-chromosome NNCGP to represent hidden semantics of all layers of the NN. The method uses a (1+λ) evolutionary strategy (called MNNCGP-ES) to extract the final mathematical expressions of all layers in the NN. Experiments on 12 symbolic regression benchmarks and 5 classification benchmarks show that SRNet not only can reveal the complex relationships between each layer of a NN but also can extract the mathematical representation of the whole NN. Compared with LIME and MAPLE, SRNet has higher interpolation accuracy and trends to approximate the real model on the practical dataset1.
Yuanzhen Luo, Qiang Lu 0005, Xilei Hu, Jake Luo, Zhiguang Wang
GECCO4
2021 Enhancing gene expression programming based on space partition and jump for symbolic regression
Qiang Lu 0005, Fan Tao, Jake Luo, Zhiguang Wang
Inf. Sci.4
2019 Challenges to a Data Driven Approach to Population Level Analysis of Hypersensitivity Events in Cancer Clinical Trials
Christina Eldredge, James E. Andrews, Maryam Zolnoori, Timothy B. Patrick, Joel Gallagher, Cesar A. Lam, Jake Luo
AMIA7
2019 Identifying Factors Affecting Drug Discontinuation in Patients with Depression: Text Analysis of Patient Drug Review Posts
Maryam Zolnoori, Che Ngufor, Anthony Faiola, Christina Eldredge, Jake Luo, Sunghwan Sohn, Joyce E. Balls-Berry, Ahmad P. Tafti, Nilay D. Shah, Timothy B. Patrick
AMIA5
2019 Machine Learning-Based Modeling of Big Clinical Trials Data for Adverse Outcome Prediction: A Case Study of Death Events
abstract
It is known that clinical trials have potential risks for participants, which could result in unexpected adverse events. To quantify and predict the risk of adverse outcomes, we leverage a large amount of clinical reports to build machine learning models to predict adverse outcomes. We focused on death events as the predicting target in this study. From Clinicaltrial.gov, we collected 28,340 reports and transformed the data into vectorized machine learning features. These features were harmonized across studies using semantic mapping and feature selection techniques. The resulting selected clinical trial features were used to build five machine learning models for prediction. We evaluated and compared relative model performances for the prediction task. Results show that the logistic regression algorithm achieved the best overall receiver operating characteristic score at 0.7344. This exploratory study showed that it is feasible to use clinical trial factors to predict adverse outcomes. We demonstrated the approach by focusing on building machine learning models to predict death outcomes. Predicting adverse outcomes could help clinical trials estimate harmful risks and design better mechanisms to protect participants. We hope by using our models, a clinical trial expert will be able to assess whether serious adverse events are likely to occur in a clinical trial at the early stage and to estimate what potential trial factors could contribute to the potential serious adverse events.
Ling Tong 0002, Jake Luo, Ron A. Cisler, Michael Cantor
COMPSAC (2)2
2019 A systematic approach for developing a corpus of patient reported adverse drug events: A case study for SSRI and SNRI medications
Maryam Zolnoori, Kin Wah Fung, Timothy B. Patrick, Paul A. Fontelo, Hadi Kharrazi, Anthony Faiola, Yi Shuan Shirley Wu, Christina Eldredge, Jake Luo, Mike Conway, Jiaxi Zhu, Soo Kyung Park, Kelly Xu, Hamideh Moayyed, Somaieh Goudarzvand
J. Biomed. Informatics9
2019 Very large-scale data classification based on K-means clustering and multi-kernel SVM
Tinglong Tang, Shengyong Chen, Meng Zhao 0001, Wei Huang 0015, Jake Luo
Soft Comput.5
2016 Evaluating Acceptability and Efficacy of Antidepressant Medications using Patients Comments in Social Media
Maryam Zolnoori, Timothy B. Patrick, Mike Conway, Anthony Faiola, Jake Luo
AMIA5
2015 Bridging the Representation Gap of Medical Image and Clinical Note through Semantic Association Mining
Jake Luo, Timothy B. Patrick
AMIA1
2013 A human-computer collaborative approach to identifying common data elements in clinical trial eligibility criteria
Jake Luo, Riccardo Miotto, Chunhua Weng
J. Biomed. Informatics1
2011 EliXR: an approach to eligibility criteria extraction and representation
abstract
OBJECTIVE: To develop a semantic representation for clinical research eligibility criteria to automate semistructured information extraction from eligibility criteria text. MATERIALS AND METHODS: An analysis pipeline called eligibility criteria extraction and representation (EliXR) was developed that integrates syntactic parsing and tree pattern mining to discover common semantic patterns in 1000 eligibility criteria randomly selected from http://ClinicalTrials.gov. The semantic patterns were aggregated and enriched with unified medical language systems semantic knowledge to form a semantic representation for clinical research eligibility criteria. RESULTS: The authors arrived at 175 semantic patterns, which form 12 semantic role labels connected by their frequent semantic relations in a semantic network. EVALUATION: Three raters independently annotated all the sentence segments (N=396) for 79 test eligibility criteria using the 12 top-level semantic role labels. Eight-six per cent (339) of the sentence segments were unanimously labelled correctly and 13.8% (55) were correctly labelled by two raters. The Fleiss' κ was 0.88, indicating a nearly perfect interrater agreement. CONCLUSION: This study present a semi-automated data-driven approach to developing a semantic network that aligns well with the top-level information structure in clinical research eligibility criteria text and demonstrates the feasibility of using the resulting semantic role labels to generate semistructured eligibility criteria with nearly perfect interrater reliability.
Chunhua Weng, Xiaoying Wu 0001, Jake Luo, Mary Regina Boland, Dimitri Theodoratos, Stephen B. Johnson
J. Am. Medical Informatics Assoc.3
2011 Dynamic categorization of clinical research eligibility criteria by hierarchical clustering
Jake Luo, Meliha Yetisgen, Chunhua Weng
J. Biomed. Informatics1