Rui Li 0002

dblp:96/4282-2 · DBLP profile ↗
← Back
28ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0001-5096-1553ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorSecurity and privacy · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 HOSL: Hybrid-Order Split Learning for Memory-Constrained Edge Training
Aakriti Lnu, Zhe Li 0083, Dandan Liang, Chao Huang 0028, Rui Li 0002, Haibo Yang 0001
WiOpt5
2025 Minimally Supervised Regression using Topological Projections in Self-Organizing Maps
abstract
Parameter prediction is essential for many applications, facilitating insightful interpretation and decision-making. However, in many real life domains, such as power systems, medicine, and engineering, it can be very expensive to acquire ground truth labels for certain datasets as they may require extensive and expensive laboratory testing. In this work, we introduce a semi-supervised learning approach based on topological projections in self-organizing maps (SOMs), which significantly reduces the required number of labeled data points to perform parameter prediction, effectively exploiting information contained in large unlabeled datasets. While few-shot learning has seen significant advances in recent years, the majority of existing approaches focus on classification tasks, making our regression-based method particularly novel for continuous parameter estimation problems. Our proposed method first trains SOMs on unlabeled data followed by only using a minimal number of available labeled data points to assign targets to key best matching units (BMU). The values estimated for newly-encountered data points are computed utilizing the average of the N closest labeled data points in the SOM’s U-matrix in tandem with a topological shortest path distance calculation scheme. The effectiveness of our approach has been validated through practical application in power engineering, specifically for estimating critical coal property values in coal power plants, where traditional laboratory testing is both time-consuming and costly. Our results indicate that the proposed minimally supervised model significantly outperforms traditional regression techniques, including linear and polynomial regression, Gaussian process regression, K-nearest neighbors, as well as deep neural network models and related clustering schemes.
Zimeng Lyu, Alexander Ororbia, Rui Li 0002, Travis J. Desell
IJCNN3
2025 Towards Straggler-Resilient Split Federated Learning: An Unbalanced Update Approach
abstract
Split Federated Learning (SFL) enables scalable training on edge devices by combining the parallelism of Federated Learning (FL) with the computational offloading of Split Learning (SL). Despite its great success, SFL suffers significantly from the well-known straggler issue in distributed learning systems. This problem is exacerbated by the dependency between Split Server and clients: the Split Server side model update relies on receiving activations from clients. Such synchronization requirement introduces significant time latency, making straggler a critical bottleneck to the scalability and efficiency of the system. To mitigate this problem, we propose *MU-SplitFed*, a straggler-resilient SFL algorithm that decouples training progress from straggler delays via a simple yet effective unbalanced update mechanism. By enabling the server to perform $\tau$ local updates per client round, *MU-SplitFed* achieves convergence rate $\mathcal{O}(\sqrt{d/(\tau T)})$, showing a linear reduction in communication round by a factor of $\tau$. Experiments demonstrate that *MU-SplitFed* consistently outperforms baseline methods with the presence of stragglers and effectively mitigates their impact through adaptive tuning of $\tau$.
Dandan Liang, Zhe Li 0083, Rui Li 0002, Haibo Yang 0001
NeurIPS5
2024 Unsupervised Tumor-Aware Distillation for Multi-Modal Brain Image Translation
abstract
Multi-modal brain images from MRI scans are widely used in clinical diagnosis to provide complementary information from different modalities. However, obtaining fully paired multi-modal images in practice is challenging due to various factors, such as time, cost, and artifacts, resulting in modality-missing brain images. To address this problem, unsupervised multi-modal brain image translation has been extensively studied. Existing methods suffer from the problem of brain tumor deformation during translation, as they fail to focus on the tumor areas when translating the whole images. In this paper, we propose an unsupervised tumor-aware distillation teacher-student network called UTAD-Net, which is capable of perceiving and translating tumor areas precisely. Specifically, our model consists of two parts: a teacher network and a student network. The teacher network learns an end-to-end mapping from source to target modality using unpaired images and corresponding tumor masks first. Then, the translation knowledge is distilled into the student network, enabling it to generate more realistic tumor areas and whole images without masks. Experiments show that our model achieves competitive performance on both quantitative and qualitative evaluations of image quality compared with state-of-the-art methods. Furthermore, we demonstrate the effectiveness of the generated images on downstream segmentation tasks. Our code is available at https://github.com/scut-HC/UTAD-Net.
Jia Wei 0003, Rui Li 0002
IJCNN3
2023 Knowledge Acquisition for Human-In-The-Loop Image Captioning
abstract
Image captioning offers a computational process to understand the semantics of images and convey them using descriptive language. However, automated captioning models may not always generate satisfactory captions due to the complex nature of the images and the quality/size of the training data. We propose an interactive captioning framework to improve machine-generated captions by keeping humans in the loop and performing an online-offline knowledge acquisition (KA) process. In particular, online KA accepts a list of keywords specified by human users and fuses them with the image features to generate a readable sentence that captures the semantics of the image. It leverages a multimodal conditioned caption completion mechanism to ensure the appearance of all user-input keywords in the generated caption. Offline KA further learns from the user inputs to update the model and benefits caption generation for unseen images in the future. It is built upon a Bayesian transformer architecture that dynamically allocates neural resources and supports uncertainty-aware model updates to mitigate overfitting. Our theoretical analysis also proves that Offline KA automatically selects the best model capacity to accommodate the newly acquired knowledge. Experiments on real-world data demonstrate the effectiveness of the proposed framework.
Ervine Zheng, Qi Yu 0001, Rui Li 0002, Anne R. Haake
AISTATS3
2023 SFusion: Self-attention Based N-to-One Multimodal Fusion Block
Zecheng Liu, Jia Wei 0003, Rui Li 0002, Jianlong Zhou
MICCAI (2)3
2023 AdaVAE: Bayesian Structural Adaptation for Variational Autoencoders
abstract
The neural network structures of generative models and their corresponding inference models paired in variational autoencoders (VAEs) play a critical role in the models' generative performance. However, powerful VAE network structures are hand-crafted and fixed prior to training, resulting in a one-size-fits-all approach that requires heavy computation to tune for given data. Moreover, existing VAE regularization methods largely overlook the importance of network structures and fail to prevent overfitting in deep VAE models with cascades of hidden layers. To address these issues, we propose a Bayesian inference framework that automatically adapts VAE network structures to data and prevent overfitting as they grow deeper. We model the number of hidden layers with a beta process to infer the most plausible encoding/decoding network depths warranted by data and perform layer-wise dropout regularization with a conjugate Bernoulli process. We develop a scalable estimator that performs joint inference on both VAE network structures and latent variables. Our experiments show that the inference framework effectively prevents overfitting in both shallow and deep VAE models, yielding state-of-the-art performance. We demonstrate that our framework is compatible with different types of VAE backbone networks and can be applied to various VAE variants, further improving their performance.
Paribesh Regmi, Rui Li 0002
NeurIPS2
2022 Dual-Level Adaptive Information Filtering for Interactive Image Segmentation
abstract
Image segmentation can be performed interactively by accepting user annotations to refine the segmentation. It seeks frequent feedback from humans, and the model is updated with a smaller batch of data in each iteration of the feedback loop. Such a training paradigm requires effective information filtering to guide the model so that it can encode vital information and avoid overfitting due to limited data and inherent heterogeneity and noises thereof. We propose an adaptive interactive segmentation framework to support user interaction while introducing dual-level information filtering to train a robust model. The framework integrates an encoder-decoder architecture with a style-aware augmentation module that applies augmentation to feature maps and customizes the segmentation prediction for different latent styles. It also applies a systematic label softening strategy to generate uncertainty-aware soft labels for model updates. Experiments on both medical and natural image segmentation tasks demonstrate the effectiveness of the proposed framework.
Ervine Zheng, Qi Yu 0001, Rui Li 0002, Anne R. Haake
AISTATS3
2022 Unsupervised Multi-Modal Medical Image Registration via Discriminator-Free Image-to-Image Translation
abstract
In clinical practice, well-aligned multi-modal images, such as Magnetic Resonance (MR) and Computed Tomography (CT), together can provide complementary information for image-guided therapies. Multi-modal image registration is essential for the accurate alignment of these multi-modal images. However, it remains a very challenging task due to complicated and unknown spatial correspondence between different modalities. In this paper, we propose a novel translation-based unsupervised deformable image registration approach to convert the multi-modal registration problem to a mono-modal one. Specifically, our approach incorporates a discriminator-free translation network to facilitate the training of the registration network and a patchwise contrastive loss to encourage the translation network to preserve object shapes. Furthermore, we propose to replace an adversarial loss, that is widely used in previous multi-modal image registration methods, with a pixel loss in order to integrate the output of translation into the target modality. This leads to an unsupervised method requiring no ground-truth deformation or pairs of aligned images for training. We evaluate four variants of our approach on the public Learn2Reg 2021 datasets. The experimental results demonstrate that the proposed architecture achieves state-of-the-art performance. Our code is available at https://github.com/heyblackC/DFMIR.
Zekang Chen, Jia Wei 0003, Rui Li 0002
IJCAI3
2022 Predicting Biomedical Interactions With Higher-Order Graph Convolutional Networks
abstract
Biomedical interaction networks have incredible potential to be useful in the prediction of biologically meaningful interactions, identification of network biomarkers of disease, and the discovery of putative drug targets. Recently, graph neural networks have been proposed to effectively learn representations for biomedical entities and achieved state-of-the-art results in biomedical interaction prediction. These methods only consider information from immediate neighbors but cannot learn a general mixing of features from neighbors at various distances. In this paper, we present a higher-order graph convolutional network (HOGCN)to aggregate information from the higher-order neighborhood for biomedical interaction prediction. Specifically, HOGCN collects feature representations of neighbors at various distances and learns their linear mixing to obtain informative representations of biomedical entities. Experiments on four interaction networks, including protein-protein, drug-drug, drug-target, and gene-disease interactions, show that HOGCN achieves more accurate and calibrated predictions. HOGCN performs well on noisy, sparse interaction networks when feature representations of neighbors at various distances are considered. Moreover, a set of novel interaction predictions are validated by literature-based case studies.
Kishan KC, Rui Li 0002, Anne R. Haake
IEEE ACM Trans. Comput. Biol. Bioinform.2
2021 A Continual Learning Framework for Uncertainty-Aware Interactive Image Segmentation
abstract
Deep learning models have achieved state-of-the-art performance in semantic image segmentation, but the results provided by fully automatic algorithms are not always guaranteed satisfactory to users. Interactive segmentation offers a solution by accepting user annotations on selective areas of the images to refine the segmentation results. However, most existing models only focus on correcting the current image's misclassified pixels, with no knowledge carried over to other images. In this work, we formulate interactive image segmentation as a continual learning problem and propose a framework to effectively learn from user annotations, aiming to improve the segmentation on both the current image and unseen images in future tasks while avoiding deteriorated performance on previously-seen images. It employs a probabilistic mask to control the neural network's kernel activation and extract the most suitable features for segmenting images in each task. We also apply a task-aware embedding to automatically infer the optimal kernel activation for initial segmentation and subsequent refinement. Interactions with users are guided through multi-source uncertainty estimation so that users can focus on the most important areas to minimize the overall manual annotation effort. Experiments are performed on both medical and natural image datasets to illustrate the proposed framework's effectiveness on basic segmentation performance, forward knowledge transfer, and backward knowledge transfer.
Ervine Zheng, Qi Yu 0001, Rui Li 0002, Anne R. Haake
AAAI3
2021 TarGAN: Target-Aware Generative Adversarial Networks for Multi-modality Medical Image Translation
Junxiao Chen, Jia Wei 0003, Rui Li 0002
MICCAI (6)3
2021 Joint Inference for Neural Network Depth and Dropout Regularization
abstract
Dropout regularization methods prune a neural network's pre-determined backbone structure to avoid overfitting. However, a deep model still tends to be poorly calibrated with high confidence on incorrect predictions. We propose a unified Bayesian model selection method to jointly infer the most plausible network depth warranted by data, and perform dropout regularization simultaneously. In particular, to infer network depth we define a beta process over the number of hidden layers which allows it to go to infinity. Layer-wise activation probabilities induced by the beta process modulate neuron activation via binary vectors of a conjugate Bernoulli process. Experiments across domains show that by adapting network depth and dropout regularization to data, our method achieves superior performance comparing to state-of-the-art methods with well-calibrated uncertainty estimates. In continual learning, our method enables neural networks to dynamically evolve their depths to accommodate incrementally available data beyond their initial structures, and alleviate catastrophic forgetting.
Kishan KC, Rui Li 0002, Mahdi Gilany
NeurIPS2
2021 Machine learning predicts nucleosome binding modes of transcription factors
abstract
BACKGROUND: Most transcription factors (TFs) compete with nucleosomes to gain access to their cognate binding sites. Recent studies have identified several TF-nucleosome interaction modes including end binding (EB), oriented binding, periodic binding, dyad binding, groove binding, and gyre spanning. However, there are substantial experimental challenges in measuring nucleosome binding modes for thousands of TFs in different species. RESULTS: We present a computational prediction of the binding modes based on TF protein sequences. With a nested cross-validation procedure, our model outperforms several fine-tuned off-the-shelf machine learning (ML) methods in the multi-label classification task. Our binary classifier for the EB mode performs better than these ML methods with the area under precision-recall curve achieving 75%. The end preference of most TFs is consistent with low nucleosome occupancy around their binding site in GM12878 cells. The nucleosome occupancy data is used as an alternative dataset to confirm the superiority of our EB classifier. CONCLUSIONS: We develop the first ML-based approach for efficient and comprehensive analysis of nucleosome binding modes of TFs.
Kishan KC, Sridevi K. Subramanya, Rui Li 0002
BMC Bioinform.3
2021 Evaluating Technology-Mediated Collaborative Workflows for Telehealth
abstract
GOALS: This paper discusses the need for a predictable method to evaluate gains and gaps of collaborative technology-mediated workflows and introduces an evaluation framework to address this need. METHODS: The Collaborative Space - Analysis Framework (CS-AF), introduced in this research, is a cross-disciplinary evaluation method designed to evaluate technology-mediated collaborative workflows. The 5-step CS-AF meta-process includes: (1) current-state workflow definition, (2) current-state (baseline) workflow assessment, (3) technology-mediated workflow development and deployment, (4) technology-mediated workflow assessment, (5) analysis and conclusions. For this research, a comprehensive, empirical study of hypertension exam workflow for telehealth was conducted using the CS-AF approach. RESULTS: The CS-AF systemized approach reveals critical cross-disciplinary evaluation data concerning gains and gaps of collaborative workflows when technology-mediated enhancements are characterized and compared with a baseline workflow for the goal of continuous workflow improvement. CONCLUSION: The CS-AF is an effective meta-analysis process that can be adapted for use in multiple domains.
Christopher Bondy, Pamela Grover, Vicki L. Hanson, Rui Li 0002
IEEE J. Biomed. Health Informatics5
2020 Interpretable Structured Learning with Sparse Gated Sequence Encoder for Protein-Protein Interaction Prediction
abstract
Predicting protein-protein interactions (PPIs) by learning informative representations from amino acid sequences is a challenging yet important problem in biology. Although various deep learning models in Siamese architecture have been proposed to model PPIs from sequences, these methods are computationally expensive for a large number of PPIs due to the pairwise encoding process. Furthermore, these methods are difficult to interpret because of non-intuitive mappings from protein sequences to their sequence representation. To address these challenges, we present a novel deep framework to model and predict PPIs from sequence alone. Our model incorporates a bidirectional gated recurrent unit to learn sequence representations by leveraging contextualized and sequential information from sequences. We further employ a sparse regularization to model long-range dependencies between amino acids and to select important amino acids (protein motifs), thus enhancing interpretability. Besides, the novel design of the encoding process makes our model computationally efficient and scalable to an increasing number of interactions. Experimental results on up-to-date interaction datasets demonstrate that our model achieves superior performance compared to other state-of-the-art methods. Literature-based case studies illustrate the ability of our model to provide biological insights to interpret the predictions.
Kishan KC, Anne R. Haake, Rui Li 0002
ICPR4
2020 Dynamic Fusion of Eye Movement Data and Verbal Narrations in Knowledge-rich Domains
abstract
We propose to jointly analyze experts' eye movements and verbal narrations to discover important and interpretable knowledge patterns to better understand their decision-making processes. The discovered patterns can further enhance data-driven statistical models by fusing experts' domain knowledge to support complex human-machine collaborative decision-making. Our key contribution is a novel dynamic Bayesian nonparametric model that assigns latent knowledge patterns into key phases involved in complex decision-making. Each phase is characterized by a unique distribution of word topics discovered from verbal narrations and their dynamic interactions with eye movement patterns, indicating experts' special perceptual behavior within a given decision-making stage. A new split-merge-switch sampler is developed to efficiently explore the posterior state space with an improved mixing rate. Case studies on diagnostic error prediction and disease morphology categorization help demonstrate the effectiveness of the proposed model and discovered knowledge patterns.
Ervine Zheng, Qi Yu 0001, Rui Li 0002, Anne R. Haake
NeurIPS3
2019 Multivariate Sparse Coding of Nonstationary Covariances with Gaussian Processes
abstract
This paper studies statistical characteristics of multivariate observations with irregular changes in their covariance structures across input space. We propose a unified nonstationary modeling framework to jointly encode the observation correlations to generate a piece-wise representation with a hyper-level Gaussian process (GP) governing the overall contour of the pieces. In particular, we couple the encoding process with automatic relevance determination (ARD) to promote sparsity to account for the inherent redundancy. The hyper GP enables us to share statistical strength among the observation variables over a collection of GPs defined within the observation pieces to characterize the variables' respective local smoothness. Experiments conducted across domains show superior performances over the state-of-the-art methods.
Rui Li 0002
NeurIPS1
2018 Understanding Social Interpersonal Interaction via Synchronization Templates of Facial Events
abstract
Automatic facial expression analysis in inter-personal communication is challenging. Not only because conversation partners' facial expressions mutually influence each other, but also because no correct interpretation of facial expressions is possible without taking social context into account. In this paper, we propose a probabilistic framework to model interactional synchronization between conversation partners based on their facial expressions. Interactional synchronization manifests temporal dynamics of conversation partners' mutual influence. In particular, the model allows us to discover a set of common and unique facial synchronization templates directly from natural interpersonal interaction without recourse to any predefined labeling schemes. The facial synchronization templates represent periodical facial event coordinations shared by multiple conversation pairs in a specific social context. We test our model on two different dyadic conversations of negotiation and job-interview. Based on the discovered facial event coordination, we are able to predict their conversation outcomes with higher accuracy than HMMs and GMMs.
Rui Li 0002, Jared Curhan, Mohammed E. Hoque 0001
AAAI1
2018 Sparse Covariance Modeling in High Dimensions with Gaussian Processes
Rui Li 0002, Kishan KC, Justin Domke, Anne R. Haake
NeurIPS1
2017 A Behavior-Based Approach for Malware Detection
Rayan Mosli, Rui Li 0002, Bo Yuan 0005
IFIP Int. Conf. Digital Forensics2
2017 Modeling Physicians' Utterances to Explore Diagnostic Decision-making
abstract
Diagnostic error prevention is a long-established but specialized topic in clinical and psychological research. In this paper, we contribute to the field by exploring diagnostic decision-making via modeling physicians' utterances of medical concepts during image-based diagnoses. We conduct experiments to collect verbal narratives from dermatologists while they are examining and describing dermatology images towards diagnoses. We propose a hierarchical probabilistic framework to learn domain-specific patterns from the medical concepts in these narratives. The discovered patterns match the diagnostic units of thought identified by domain experts. These meaningful patterns uncover physicians' diagnostic decision-making processes while parsing the image content. Our evaluation shows that these patterns provide key information to classify narratives by diagnostic correctness levels.
Rui Li 0002, Qi Yu 0001, Anne R. Haake
IJCAI2
2016 An Expert-in-the-loop Paradigm for Learning Medical Image Grouping
Qi Yu 0001, Rui Li 0002, Cecilia O. Alm, Cara Calvelli, Anne R. Haake
PAKDD (1)3
2016 Modeling eye movement patterns to characterize perceptual skill in image-based diagnostic reasoning processes
Rui Li 0002, Jeff B. Pelz, Cecilia O. Alm, Anne R. Haake
Comput. Vis. Image Underst.1
2014 Infusing perceptual expertise and domain knowledge into a human-centered image retrieval system: a prototype application
abstract
Traditional content-based image retrieval techniques, which primarily rely on image content at the pixel level, are not effective in accessing images at the semantic level. Defining approaches to incorporate experts' perceptual and conceptual capabilities of image understanding in their domain of expertise into the retrieval processes promises to help bridge this semantic gap. Towards accomplishing this, we design and implement a novel multimodal interactive system for image retrieval. To incorporate human expertise, the system stores expert-derived information extracted from two human sensor modalities that intuitively relate to image search, eye movements and verbal descriptions, both generated by medical experts. Experimental evaluation of the system shows that by transferring experts' perceptual expertise and domain knowledge into image-based computational procedures, our system can take advantage of the different human-centered modalities' respective strengths and improve the retrieval performance over just using image-based features.
Rui Li 0002, Cecilia O. Alm, Qi Yu 0001, Jeff B. Pelz, Anne R. Haake
ETRA2
2013 Image Understanding from Experts' Eyes by Modeling Perceptual Skill of Diagnostic Reasoning Processes
abstract
Eliciting and representing experts' remarkable perceptual capability of locating, identifying and categorizing objects in images specific to their domains of expertise will benefit image understanding in terms of transferring human domain knowledge and perceptual expertise into image-based computational procedures. In this paper, we present a hierarchical probabilistic framework to summarize the stereotypical and idiosyncratic eye movement patterns shared within 11 board-certified dermatologists while they are examining and diagnosing medical images. Each inferred eye movement pattern characterizes the similar temporal and spatial properties of its corresponding segments of the experts' eye movement sequences. We further discover a subset of distinctive eye movement patterns which are commonly exhibited across multiple images. Based on the combinations of the exhibitions of these eye movement patterns, we are able to categorize the images from the perspective of experts' viewing strategies. In each category, images share similar lesion distributions and configurations. The performance of our approach shows that modeling physicians' diagnostic viewing behaviors informs about medical images' understanding to correct diagnosis.
Rui Li 0002, Anne R. Haake
CVPR1
2012 Learning Image-Derived Eye Movement Patterns to Characterize Perceptual Expertise
Rui Li 0002, Jeff B. Pelz, Anne R. Haake
CogSci1
2012 Learning eye movement patterns for characterization of perceptual expertise
abstract
Human perceptual expertise has significant influence on medical image inspection. However, little is known regarding whether experts differ in their cognitive processing or what effective visual strategies they employ for examining medical images. To remedy this, we conduct an eye tracking experiment and collect both eye movement and verbal description data from three groups of subjects with different medical training levels. Each subject examines and describes 42 photographic dermatological images. We then develop a hierarchical probabilistic framework to extract the common and unique eye movement patterns exhibited among multiple subjects' fixation and saccadic eye movements within each expertise-specific group. Furthermore, experts' annotations of thought units on the transcribed verbal descriptions are time-aligned with these eye movement patterns to identify their semantic meanings. In this work, we are able to uncover the manner in which these subjects alternated their viewing strategies over the course of inspection, and additionally extract their perceptual expertise so that it can be used for advanced medical image understanding.
Rui Li 0002, Jeff B. Pelz, Cecilia O. Alm, Anne R. Haake
ETRA1