Zhengbo Zhang

dblp:26/8352 · DBLP profile ↗
← Back
22ranked-venue papers
4as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Computer networks · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Remaining Useful Life Prediction of CT X-Ray Tubes Based on the Internet of Medical Things
abstract
Accurate Remaining Useful Life (RUL) estimation for CT X-ray tube filaments is critical to reduce unscheduled downtime and optimize maintenance in clinical settings. This work leverages full-lifecycle operational logs collected via an Internet of Medical Things (IoMT) platform and selects the parameter-stable Scout View mode to build physically consistent degradation sequences. A multi-stage preprocessing pipeline—anomaly removal, failure-threshold truncation, normalization, downsampling, and sliding-window segmentation — yields standardized multivariate inputs. To validate the effectiveness of our method, we conducted a comprehensive comparative analysis against five baseline paradigms: Hidden Markov Models (HMM), segmented linear regression, Gated Recurrent Units (GRU), 1D Convolutional Neural Networks (1D-CNN), and Transformers. Based on this, we develop a multivariate prediction framework using filament current, cumulative scan time, and cumulative usage days based on Long Short-Term Memory (LSTM) networks. To further improve temporal consistency, we introduce a novel historical path backtracking correction mechanism. Experiments on 38 complete filament life cycles (7 holdout test samples) demonstrate that the backtracking-corrected LSTM achieves superior performance, yielding a mean absolute percentage error (MAPE) of 9.55% and R² = 0.989, significantly outperforming all statistical and deep learning baselines. This IoMT-driven, deployable approach provides an effective solution for predictive maintenance, enabling proactive tube replacement planning and supporting resource optimization in smart hospital operations.
Jipeng Sun, Chang Liu 0158, Cebing Chu, Yuzhe Lu, Zekun Miao, Chongzhe Zhang, Yonghua Li 0001, Zhengbo Zhang
IEEE Internet Things J.10
2026 Unleashing the Power of Text-to-Image Diffusion Models for Category-Agnostic Pose Estimation
abstract
Category-Agnostic Pose Estimation (CAPE) aims to detect keypoints of unseen object categories in a few-shot setting, where the scarcity of labeled data poses significant challenges to generalization. In this work, we propose Prompt Pose Matching (PPM), a novel framework that unleashes the power of off-the-shelf text-to-image diffusion models for CAPE. PPM learns pseudo prompts from few-shot examples via the text-to-image diffusion model. These learned pseudo prompts capture semantic information of keypoints, which can then be used to locate the same type of keypoints from images. To provide prompts with representative initialization, we introduce a category-agnostic pre-training strategy to capture the foreground prior shared across categories and keypoints. To support the reliable prompt pre-training, we propose a Foreground-Aware Region Aggregation (FARA) module to provide robust and consistent supervision signal. Based on the foreground prior, a Foreground-Guided Attention Refinement (FGAR) module is further proposed to reinforce cross-attention responses for accurate keypoint localization. For efficiency, a Prompt Ensemble Inference (PEI) scheme enables joint keypoint prediction. Unlike previous methods that highly rely on base-category annotated data, our PPM framework can operate in a base-category-free setting while retaining strong performance. Code will be available at: https://github.com/DuoPeng-CVer/Prompt-Pose-Matching.
Duo Peng, Zhengbo Zhang, Ping Hu 0001, Qiuhong Ke, De Wen Soh, Mohammed Bennamoun, Jun Liu 0036
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 Frequency-enhanced diffusion models: curriculum-guided semantic alignment for zero-shot skeleton action recognition
Zhengbo Zhang, Jingyu Pan, Zhigang Tu 0001
Vis. Comput.2
2025 Visual Prompting for One-shot Controllable Video Editing without Inversion
abstract
One-shot controllable video editing (OCVE) is an important yet challenging task, aiming to propagate user edits that are made – using any image editing tool – on the first frame of a video to all subsequent frames, while ensuring content consistency between edited frames and source frames. To achieve this, prior methods employ DDIM inversion to transform source frames into latent noise, which is then fed into a pre-trained diffusion model, conditioned on the user-edited first frame, to generate the edited video. However, the DDIM inversion process accumulates errors, which hinder the latent noise from accurately reconstructing the source frames, ultimately compromising content consistency in the generated edited frames. To overcome it, our method eliminates the need for DDIM inversion by performing OCVE through a novel perspective based on visual prompting. Furthermore, inspired by consistency models that can perform multi-step consistency sampling to generate a sequence of content-consistent images, we propose a content consistency sampling (CCS) to ensure content consistency between the generated edited frames and the source frames. Moreover, we introduce a temporal-content consistency sampling (TCS) based on Stein Variational Gradient Descent to ensure temporal consistency across the edited frames. Extensive experiments validate the effectiveness of our approach.
Zhengbo Zhang, Duo Peng, Joo-Hwee Lim, Zhigang Tu 0001, De Wen Soh, Lin Geng Foo
CVPR1
2025 Dual-Path Model for Pulmonary Artery Segmentation
abstract
The pulmonary artery (PA) is a multi-level vascular system composed of the main pulmonary artery (mPA) and the branch pulmonary arteries (bPA). Accurate segmentation of the PA is of significant importance for the diagnosis of diseases such as pulmonary embolism. However, since the vascular features of the mPA and the bPA are very different, the existing holistic PA segmentation methods will cause the network to pay too much attention to the mPA with significant features and ignore the bPA with weak features, resulting in the imbalance of PA segmentation. To address this issue, we propose a Dual-Path Pulmonary Artery Segmentation Model, which employs two separate paths to learn the features of the mPA and bPA, thereby enhancing the network’s ability to learn features of the bPA. Additionally, we have designed a Skeleton-Optimized Feature Learning Mechanism that optimizes the topological structure of the PA through skeletal guidance, reducing fragmentation and false positives. We conducted experiments on the public dataset provided by PARSE2022, and the results demonstrate that our model indeed improves the segmentation accuracy of the bPA while maintaining the segmentation effectiveness of the mPA. Furthermore, the Skeleton-Optimized Feature Learning Mechanism plays a significant role in reducing vascular fragmentation and false positives.
Yingwen Chen 0001, Zhengbo Zhang
ICASSP4
2025 Performing Defocus Deblurring by Modeling its Formation Process
Zhengbo Zhang, Lin Geng Foo, Hossein Rahmani 0001, Jun Liu 0036, De Wen Soh
ICCV1
2025 InstaInpaint: Instant 3D-Scene Inpainting with Masked Large Reconstruction Model
abstract
Recent advances in 3D scene reconstruction enable real-time viewing in virtual and augmented reality. To support interactive operations for better immersiveness, such as moving or editing objects, 3D scene inpainting methods are proposed to repair or complete the altered geometry. To support users in interacting (such as moving or editing objects) with the scene for the next level of immersiveness, 3D scene inpainting methods are developed to repair the altered geometry. However, current approaches rely on lengthy and computationally intensive optimization, making them impractical for real-time or online applications. We propose InstaInpaint, a reference-based feed-forward framework that produces 3D-scene inpainting from a 2D inpainting proposal within 0.4 seconds. We develop a self-supervised masked-finetuning strategy to enable training of our custom large reconstruction model (LRM) on the large-scale dataset. Through extensive experiments, we analyze and identify several key designs that improve generalization, textural consistency, and geometric correctness. InstaInpaint achieves a 1000$\times$ speed-up from prior methods while maintaining a state-of-the-art performance across two standard benchmarks. Moreover, we show that InstaInpaint generalizes well to flexible downstream applications such as object insertion and multi-region inpainting.
Junqi You, Chieh Hubert Lin, Weijie Lyu, Zhengbo Zhang, Ming-Hsuan Yang 0001
NeurIPS4
2025 FADE: A Dataset for Detecting Falling Objects Around Buildings in Video
abstract
Objects falling from buildings, a frequently occurring event in daily life, can cause severe injuries to pedestrians due to the high impact force they exert. Surveillance cameras are often installed around buildings to detect falling objects, but such detection remains challenging due to the small size and fast motion of the objects. Moreover, the field of falling object detection around buildings (FODB) lacks a large-scale dataset for training learning-based detection methods and for standardized evaluation. To address these challenges, we propose a large and diverse video benchmark dataset named FADE. Specifically, FADE contains 2,611 videos from 25 scenes, featuring 8 falling object categories, 4 weather conditions, and 4 video resolutions. Additionally, we develop a novel detection method for FODB that effectively leverages motion information and generates small-sized yet high-quality detection proposals. The efficacy of our method is evaluated on the proposed FADE dataset by comparing it with state-of-the-art approaches in generic object detection, video object detection, and moving object detection. The dataset and code are publicly available at https://fadedataset.github.io/FADE.github.io/.
Zhigang Tu 0001, Zhengbo Zhang, Zitao Gao, Chunluan Zhou, Junsong Yuan 0001, Bo Du 0001
IEEE Trans. Inf. Forensics Secur.2
2025 Informative Sample Selection Model for Skeleton-Based Action Recognition With Limited Training Samples
abstract
Skeleton-based human action recognition aims to classify human skeletal sequences, which are spatiotemporal representations of actions, into predefined categories. To reduce the reliance on costly annotations of skeletal sequences while maintaining competitive recognition accuracy, the task of 3D Action Recognition with Limited Training Samples, also known as semi-supervised 3D Action Recognition, has been proposed. In addition, active learning, which aims to proactively select the most informative unlabeled samples for annotation, has been explored in semi-supervised 3D Action Recognition for training sample selection. Specifically, researchers adopt an encoder-decoder framework to embed skeleton sequences into a latent space, where clustering information, combined with a margin-based selection strategy using a multi-head mechanism, is utilized to identify the most informative sequences in the unlabeled set for annotation. However, the most representative skeleton sequences may not necessarily be the most informative for the action recognizer, as the model may have already acquired similar knowledge from previously seen skeleton samples. To solve it, we reformulate Semi-supervised 3D action recognition via active learning from a novel perspective by casting it as a Markov Decision Process (MDP). Built upon the MDP framework and its training paradigm, we train an informative sample selection model to intelligently guide the selection of skeleton sequences for annotation. To enhance the representational capacity of the factors in the state-action pairs within our method, we project them from Euclidean space to hyperbolic space. Furthermore, we introduce a meta tuning strategy to accelerate the deployment of our method in real-world scenarios. Extensive experiments on three 3D action recognition benchmarks demonstrate the effectiveness of our method.
Zhigang Tu 0001, Zhengbo Zhang, Jia Gong, Junsong Yuan 0001, Bo Du 0001
IEEE Trans. Image Process.2
2024 Harnessing Text-to-Image Diffusion Models for Category-Agnostic Pose Estimation
Duo Peng, Zhengbo Zhang, Ping Hu 0001, Qiuhong Ke, David K. Y. Yau, Jun Liu 0036
ECCV (13)2
2024 Diff-Tracker: Text-to-Image Diffusion Models are Unsupervised Trackers
Zhengbo Zhang, Duo Peng, Hossein Rahmani 0001, Jun Liu 0036
ECCV (28)1
2023 A Deep Learning Approach Incorporating Data Missing Mechanism in Predicting Acute Kidney Injury in ICU
Zhengbo Zhang, Lei Zha, Fengcong, Xiao-Rui Su 0001, Bo-Wei Zhao, Lun Hu, Pengwei Hu 0001
ICIC (3)2
2023 Tremor detection Transformer: An automatic symptom assessment framework based on refined whole-body pose estimation
Chenbin Ma, Lishuang Guo, Longsheng Pan, Chunyu Yin, Rui Zong, Zhengbo Zhang
Eng. Appl. Artif. Intell.7
2023 Automatic diagnosis of multi-task in essential tremor: Dynamic handwriting analysis using multi-modal fusion neural network
Chenbin Ma, Yulan Ma, Longsheng Pan, Chunyu Yin, Rui Zong, Zhengbo Zhang
Future Gener. Comput. Syst.7
2023 Time Series Anomaly Detection With Adversarial Reconstruction Networks
abstract
Time series data naturally exist in many domains including medical data analysis, infrastructure sensor monitoring, and motion tracking. However, a very small portion of anomalous time series can be observed, comparing to the whole data. Most existing approaches are based on the supervised classification model requiring representative labels for anomaly class(es), which is challenging in real-world problems. So can we learn how to detect anomalous time ticks in an effective yet efficient way, given mostly normal time series data? Therefore, we propose an unsupervised reconstruction model named BeatGAN which learns to detect anomalies based on normal data, or data which majority of samples are normal. BeatGAN provides a framework to adversarially learn to reconstruct, which can cooperate with both 1-d CNN and RNN. Rarely observed anomalies can result in larger reconstruction errors, which are then detected based on extreme value theory. Moreover, data augmentation with dynamic time warping regularizes reconstruction and provides robustness. In the experiments, effectiveness and sensitivity are studied in both synthetic data and various real-world time series. BeatGAN achieves better accuracy and fast inference.
Shenghua Liu, Quan Ding, Bryan Hooi, Zhengbo Zhang, Huawei Shen, Xueqi Cheng 0001
IEEE Trans. Knowl. Data Eng.5
2022 A Quantitative Approach and Preliminary Application in Healthy Subjects and Patients with Valvular Heart Disease for 24-h Breathing Patterns Analysis Using Wearable Devices
abstract
The 24-h breathing patterns may be closely related to health status as well as disease progression. However, there is no consistent and widely accepted approach for mining the potential value in 24-h respiratory signals based on wearable device monitoring. This study presented a reference approach including signal quality assessment, calibration of tidal volume, and breathing patterns parameters based on a wearable continuous physiological parameter monitoring system for 24-h breathing patterns analysis, including time domain, frequency domain and nonlinear domain. 70 healthy subjects and 76 patients undergoing heart valve surgery were enrolled in this study. The normal reference range of breathing patterns was calculated based on healthy subjects. A subgroup study was conducted based on whether patients developed postoperative pulmonary complications (PPCs). Compared with non-PPCs group, the coefficient of variation of breathing rate in the recumbent position was smaller in the PPCs group. During the daytime, the kurtosis of breathing rate and contribution of the abdomen was smaller in PPCs group. During the nighttime, the coefficient of variation of breathing rate and SD2 was smaller in the PPCs group. The quantitative method proposed in this study fills the gap in the field of quantifying 24-h breathing patterns which is effective in discriminating different populations and is expected to be used widely in the context of COVID-19 epidemic.
Yuqiang Wang, Chenbin Ma, Pengming Yu, Yingqiang Guo, Zhengbo Zhang
HealthCom8
2022 Distilling Inter-Class Distance for Semantic Segmentation
abstract
Knowledge distillation is widely adopted in semantic segmentation to reduce the computation cost. The previous knowledge distillation methods for semantic segmentation focus on pixel-wise feature alignment and intra-class feature variation distillation, neglecting to transfer the knowledge of the inter-class distance in the feature space, which is important for semantic segmentation such a pixel-wise classification task. To address this issue, we propose an Inter-class Distance Distillation (IDD) method to transfer the inter-class distance in the feature space from the teacher network to the student network. Furthermore, semantic segmentation is a position-dependent task, thus we exploit a position information distillation module to help the student network encode more position information. Extensive experiments on three popular datasets: Cityscapes, Pascal VOC and ADE20K show that our method is helpful to improve the accuracy of semantic segmentation models and achieves the state-of-the-art performance. E.g. it boosts the benchmark model (``PSPNet+ResNet18") by 7.50% in accuracy on the Cityscapes dataset.
Zhengbo Zhang, Chunluan Zhou, Zhigang Tu 0001
IJCAI1
2022 A feature fusion sequence learning approach for quantitative analysis of tremor symptoms based on digital handwriting
Chenbin Ma, Peng Zhang 0078, Longsheng Pan, Chunyu Yin, Ailing Li, Rui Zong, Zhengbo Zhang
Expert Syst. Appl.8
2020 A Deep Learning Model for Early Prediction of Sepsis from Intensive Care Unit Records
Rui Zhao 0019, Tao Wan 0001, Zhengbo Zhang, Zengchang Qin
ICONIP (4)4
2019 Poster: DeePTOP: Personalized Tachycardia Onset Prediction Using Bi-directional LSTM in Wearable Embedded Systems
Ke Lan, Xiaoli Liu 0003, Peiyao Li, Jiewen Zheng, Desen Cao, Zhengbo Zhang
EWSN10
2018 Automated Sleep Period Estimation in Wearable Multi-sensor Systems
abstract
Sleep period determination is essential to accurate sleep quality analysis. In this paper, we propose an automated algorithm for wearable multi-sensor systems to precisely estimate the sleep period. It leverages the information of accelerometer, and vital signs such as heart rate and breathing rate. Compared to the sleep periods determined by the clinical diagnosing-grade equipment, our algorithm achieves average time differences of 9.0 and 10.4 minutes for healthy subjects and clinical patients, respectively.
Zhengbo Zhang, Xiaoli Liu 0003, Desen Cao, Peiyao Li, Jiewen Zheng, Ke Lan
SenSys3
2018 Breathing Disorder Detection Using Wearable Electrocardiogram And Oxygen Saturation
abstract
Conventional diagnosis using polysomnography (PSG) on breathing disorder is expensive and uncomfortable to patients. In this paper, we present a low-cost portable and wearable multi-sensor system to non-invasively acquire a subject's vital signs, and leverage various machine learning methods on features extracted from Electrocardiogram (ECG) and Blood oxygen saturation (SpO2) signals to detect breathing disorder events. Our preliminary predication accuracies on 110 clinical patients is 90.0%.
Zhengbo Zhang, Peiyao Li, Desen Cao, Xiaoli Liu 0003, Jiewen Zheng, Jianli Pan
SenSys3