Zongyu Li

dblp:289/5513 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Reducing Confounding Bias without Data Splitting for Causal Inference via Optimal Transport
abstract
Causal inference seeks to estimate the effect given a treatment such as a medicine or the dosage of a medication. To reduce the confounding bias caused by the non-randomized treatment assignment, most existing methods reduce the shift between subpopulations receiving different treatments. However, these methods split limited training samples into smaller groups, which cuts down the number of samples in each group, while precise distribution estimation and alignment highly rely on a sufficient number of training samples. In this paper, we propose a distribution alignment paradigm without data splitting, which can be naturally applied in the settings of binary and continuous treatments. To this end, we characterize the confounding bias by considering different probability measures of the same set including all the training samples, and exploit the optimal transport theory to analyze the confounding bias and outcome estimation error. Based on this, we propose to learn balanced representations by reducing the bias between the marginal distribution and the conditional distribution of a treatment. As a result, data reduction caused by splitting is avoided, and the outcome prediction model trained on one treatment group can be generalized to the entire population. The experiments on both binary and continuous treatment settings demonstrate the effectiveness of our method.
Yuguang Yan, Zongyu Li, Zeqin Yang, Ruichu Cai
ICML2
2023 Enhanced RNA Sequence Representation through Sequence Masking and Subsequence Consistency Optimization
abstract
In the burgeoning field of RNA research, accurate and efficient RNA sequence representation remains a pivotal challenge, exacerbated by the complexity and diversity of RNA sequences. Addressing the critical need for enhanced sequence representation and the issues of sequence context and structural alignment, this study introduces a novel, comprehensive approach. The proposed model seamlessly integrates sequence masking and subsequence consistency optimization, offering a robust solution to the intricate problem of RNA sequence representation. Utilizing the filtered RNAStralign dataset, encompassing 20,923 sequences, the model's performance is rigorously evaluated employing a Support Vector Machine (SVM) for subsequent RNA family classification tasks. Despite the inherent imbalance in RNA family sequence distribution, the model demonstrates exemplary performance, achieving high classification accuracy and AUPRC values across diverse RNA sequence groups. This balanced and unbiased assessment, ensured by the use of AUPRC as an evaluation metric, highlights the model's practical utility for comprehensive RNA sequence analysis and classification. In essence, this research presents a method for enhanced RNA sequence representation and laying a robust foundation for future advancements in the nuanced field of RNA sequence analysis.
Yewei Shen, Zongyu Li, Xinmeng Liu, Xuequn Shang 0001, Yongtian Wang
BIBM3
2023 Towards Surgical Context Inference and Translation to Gestures
abstract
Manual labeling of gestures in robot-assisted surgery is labor intensive, prone to errors, and requires expertise or training. We propose a method for automated and explainable generation of gesture transcripts that leverages the abundance of data for image segmentation. Surgical context is detected using segmentation masks by examining the distances and intersections between the tools and objects. Next, context labels are translated into gesture transcripts using knowledge-based Finite State Machine (FSM) and data-driven Long Short Term Memory (LSTM) models. We evaluate the performance of each stage of our method by comparing the results with the ground truth segmentation masks, the consensus context labels, and the gesture labels in the JIGSAWS dataset. Our results show that our segmentation models achieve state-of-the-art performance in recognizing needle and thread in Suturing and we can automatically detect important surgical states with high agreement with crowd-sourced labels (e.g., contact between graspers and objects in Suturing). We also find that the FSM models are more robust to poor segmentation and labeling performance than LSTMs. Our proposed method can significantly shorten the gesture labeling process (~2.8 times).
Kay Hutchinson, Zongyu Li, Ian Reyes, Homa Alemzadeh
ICRA2
2023 Robotic Scene Segmentation with Memory Network for Runtime Surgical Context Inference
abstract
Surgical context inference has recently garnered significant attention in robot-assisted surgery as it can facilitate workflow analysis, skill assessment, and error detection. However, runtime context inference is challenging since it requires timely and accurate detection of the interactions among the tools and objects in the surgical scene based on the segmentation of video data. On the other hand, existing state-of-the-art video segmentation methods are often biased against infrequent classes and fail to provide temporal consistency for segmented masks. This can negatively impact the context inference and accurate detection of critical states. In this study, we propose a solution to these challenges using a Space-Time Correspondence Network (STCN). STCN is a memory network that performs binary segmentation and minimizes the effects of class imbalance. The use of a memory bank in STCN allows for the utilization of past image and segmentation information, thereby ensuring consistency of the masks. Our experiments using the publicly-available JIGSAWS dataset demonstrate that STCN achieves superior segmentation performance for objects that are difficult to segment, such as needle and thread, and improves context inference compared to the state-of-the-art. We also demonstrate that segmentation and context inference can be performed at runtime without compromising performance.
Zongyu Li, Ian Reyes, Homa Alemzadeh
IROS1
2023 Physics-Guided Deep Scatter Estimation by Weak Supervision for Quantitative SPECT
abstract
Accurate scatter estimation is important in quantitative SPECT for improving image contrast and accuracy. With a large number of photon histories, Monte-Carlo (MC) simulation can yield accurate scatter estimation, but is computationally expensive. Recent deep learning-based approaches can yield accurate scatter estimates quickly, yet full MC simulation is still required to generate scatter estimates as ground truth labels for all training data. Here we propose a physics-guided weakly supervised training framework for fast and accurate scatter estimation in quantitative SPECT by using a 100× shorter MC simulation as weak labels and enhancing them with deep neural networks. Our weakly supervised approach also allows quick fine-tuning of the trained network to any new test data for further improved performance with an additional short MC simulation (weak label) for patient-specific scatter modelling. Our method was trained with 18 XCAT phantoms with diverse anatomies / activities and then was evaluated on 6 XCAT phantoms, 4 realistic virtual patient phantoms, 1 torso phantom and 3 clinical scans from 2 patients for 177Lu SPECT with single / dual photopeaks (113, 208 keV). Our proposed weakly supervised method yielded comparable performance to the supervised counterpart in phantom experiments, but with significantly reduced computation in labeling. Our proposed method with patient-specific fine-tuning achieved more accurate scatter estimates than the supervised method in clinical scans. Our method with physics-guided weak supervision enables accurate deep scatter estimation in quantitative SPECT, while requiring much lower computation in labeling, enabling patient-specific fine-tuning capability in testing.
Hanvit Kim, Zongyu Li, Jiye Son, Jeffrey A. Fessler, Yuni K. Dewaraja, Se Young Chun
IEEE Trans. Medical Imaging2
2022 Runtime Detection of Executional Errors in Robot-Assisted Surgery
abstract
Despite significant developments in the design of surgical robots and automated techniques for objective evaluation of surgical skills, there are still challenges in ensuring safety in robot-assisted minimally-invasive surgery (RMIS). This paper presents a runtime monitoring system for the detection of executional errors during surgical tasks through the analysis of kinematic data. The proposed system incorporates dual Siamese neural networks and knowledge of surgical context, including surgical tasks and gestures, their distributional similarities, and common error modes, to learn the differences between normal and erroneous surgical trajectories from small training datasets. We evaluate the performance of the error detection using Siamese networks compared to single CNN and LSTM networks trained with different levels of contextual knowledge and training data, using the dry-lab demonstrations of the Suturing and Needle Passing tasks from the JIGSAWS dataset. Our results show that gesture specific task nonspecific Siamese networks obtain micro F1 scores of 0.94 (Siamese-CNN) and 0.95 (Siamese-LSTM), and perform better than single CNN (0.86) and LSTM (0.87) networks. These Siamese networks also outperform gesture nonspecific task specific Siamese-CNN and Siamese-LSTM models for Suturing and Needle Passing.
Zongyu Li, Kay Hutchinson, Homa Alemzadeh
ICRA1
2022 Volume-awareness and outlier-suppression co-training for weakly-supervised MRI breast mass segmentation with partial annotations
Xianqi Meng, Jingfan Fan, Jinrong Mu, Zongyu Li, Aocai Yang, Kuan Lv, Danni Ai, Yucong Lin, Hong Song 0003, Tianyu Fu 0003, Deqiang Xiao, Guolin Ma, Jian Yang 0009
Knowl. Based Syst.5
2021 Poisson Phase Retrieval With Wirtinger Flow
abstract
This paper discusses algorithms for phase retrieval where the measurements follow independent Poisson distributions. We developed an optimization problem based on maximum likelihood estimation (MLE) for the Poisson model and applied Wirtinger flow algorithm to solve it. Simulation results with a random Gaussian sensing matrix and Poisson measurement noise demonstrated that the Wirtinger flow algorithm based on the Poisson model produced higher quality reconstructions than when algorithms derived from Gaussian noise models (Wirtinger flow, Gerchberg Saxton) are applied to such data, with significantly improved computational efficiency.
Zongyu Li, Kenneth Lange, Jeffrey A. Fessler
ICIP1