Haogang Zhu

dblp:93/9945 · DBLP profile ↗
← Back
39ranked-venue papers
1as first author
31since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 16 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 9 since 2021Artificial intelligence and machine learning · 12 · 10 since 2021Systems, architecture and hardware · 6 · 6 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multimodal visual Test-Time Adaptation in resource constrained environments
Haogang Zhu
Pattern Recognit.3
2026 MMP: Enhancing unsupervised graph anomaly detection with multi-view message passing
Weihu Song, Mengxiao Zhu 0004, Yue Pei, Haogang Zhu
Pattern Recognit.5
2026 ChainOpt: Heterogeneity-Aware Blockchain Performance Optimization for Dynamic Workloads
abstract
As reliable distributed systems, Blockchains have been widely applied in diverse domains, such as the Internet of Things (IoT). Recent studies have explored Deep Reinforcement Learning (DRL) to enhance blockchain performance. However, existing DRL-based blockchain performance optimization methods rely on implicit and idealized assumptions about node behaviors and transactional workloads, limiting their effectiveness and efficiency on the blockchain with dynamic work-loads and heterogeneous nodes. To alleviate this, we propose CHAINOPT, a novel blockchain performance optimization framework devised for optimal parameter configuration to handle dynamic workloads and heterogeneous nodes. Specifically, we first propose an interaction-aware state representation learning module to model both global system-level and local heterogeneous node feature interactions to generate better state representations. Then, a contrastive learning-enhanced workload identification module is designed to extract discriminative workload-specific state representations to improve workload identification accuracy. Finally, we design a workload-similarity guided policy reuse module to produce effective reuse weights to transfer knowledge from history policies based on workload relevance, thereby improving optimization speed and stability. Extensive experiments show the effectiveness of CHAINOPT in improving blockchain performance, achieving 185.87% higher scalability and 1492.16% stronger security with only a marginal 6.26% latency increase. Moreover, it outperforms baselines in static scenarios while maintaining considerable superiority under varying workloads.
Biqi Zhao, Yushan Zeng, Jiejie Zhao, Shan Zhang 0001, Haogang Zhu, Runhe Huang, Weifeng Lv
IEEE Trans. Computers7
2025 Perturbating, Tuning, and Collaborating: Harnessing Vision Foundation Models for Single Domain Generalization on Medical Imaging
abstract
Single Domain Generalization (SDG) is critical in medical imaging applications. Recently, Vision Foundation Models (VFMs) have spearheaded a trend in AI development due to their robust generalizability and versatility. This work aims to fully explore the generalization capabilities of VFMs alongside the domain-specific expertise of specialized models, thoroughly investigating the boundaries of their respective capabilities, thereby collaboratively addressing SDG challenges within medical imaging. We propose a framework for Collaborative reasoning between Specialized and Universal models for Single Domain Generalization (CollaSU-SDG) in medical imaging. Specifically, we first design a model-aware perturbation injection method from the perspective of single-source domain data, enabling differentiated and adaptive perturbation injection for two different scales of models. Then, a domain expansion adapter is designed for the VFM to adapt to the augmented single-source domain medical data. Lastly, we introduce an adaptive hierarchical transfer and dynamic dense prompting method that facilitate collaborative reasoning between the specialized and universal models, eliminating the need for explicit prompts. Through these designs, CollaSU-SDG fully leverages the strengths of both specialized and universal models, achieving robust out-of-distribution generalization capabilities on single-source domain data. Experimental results demonstrate that CollaSU-SDG significantly advances the state-of-the-art performance across a wide range of medical datasets. All the code will be publicly available.
Yichao Cao, YingYing Zhang, Xiu Su, Haogang Zhu
AAAI5
2025 Adversarial Pretrained Language Model for Multivariate Time Series Anomaly Detection
abstract
Multivariate time series anomaly detection plays a vital role in safety-critical domains such as industrial systems, finance, and cybersecurity. However, the scarcity of labeled anomalies poses significant challenges for learning robust normal patterns, often blurring the boundary between normal and abnormal behaviors. To address this challenge, we propose ADLM, an unsupervised adversarial framework that integrates a Language-Model-based Predictor for Time Series (LMPTS) with an autoencoder. To capture normal patterns under limited data, LMPTS repurposes a decoder-only pretrained language model as an autoregressive forecaster, leveraging its strong generative prior to capture temporal dependencies. To model complex cross-sensor dependencies, we incorporate graph structure learning into the framework. Furthermore, we introduce an adversarial training strategy to sharpen the model’s normal-pattern representations and amplify deviations indicative of anomalies. Experiments on six public datasets show that ADLM consistently outperforms state-of-the-art baselines and remains robust under severe data scarcity. By coupling decoder-only language models with an adversarial objective, ADLM offers a label-efficient, structure-aware solution to multivariate time series anomaly detection.
Jianhuan Mao, Mengxiao Zhu 0004, Haogang Zhu
ECAI4
2025 TinyMIG: Transferring Generalization from Vision Foundation Models to Single-Domain Medical Imaging
abstract
Medical imaging faces significant challenges in single-domain generalization (SDG) due to the diversity of imaging devices and the variability among data collection centers. To address these challenges, we propose \textbf{TinyMIG}, a framework designed to transfer generalization capabilities from vision foundation models to medical imaging SDG. TinyMIG aims to enable lightweight specialized models to mimic the strong generalization capabilities of foundation models in terms of both global feature distribution and local fine-grained details during training. Specifically, for global feature distribution, we propose a Global Distribution Consistency Learning strategy that mimics the prior distributions of the foundation model layer by layer. For local fine-grained details, we further design a Localized Representation Alignment method, which promotes semantic alignment and generalization distillation between the specialized model and the foundation model. These mechanisms collectively enable the specialized model to achieve robust performance in diverse medical imaging scenarios. Extensive experiments on large-scale benchmarks demonstrate that TinyMIG, with extremely low computational cost, significantly outperforms state-of-the-art models, showcasing its superior SDG capabilities. All the code and model weights will be publicly available.
Hongyan Xu 0002, Yichao Cao, Xiu Su, Tianfa Li, Shan An, Haogang Zhu
ICML8
2025 Expert Data - Assisted Diagnosis: An INFO - iTransformer - XGBoost Combined Discriminative System for Prenatal Diagnosis of Fetal Congenital Heart Disease
Hao Sheng 0001, Xiaoyan Gu 0005, Jiancheng Han, Da Yang 0001, Xuefei Huang, Yihua He, Haogang Zhu
KSEM (4)10
2025 EVA-S2PLoR: Decentralized Secure 2-Party Logistic Regression with a Subtly Hadamard Product Protocol
abstract
The implementation of accurate nonlinear operators (e.g., sigmoid function) on heterogeneous datasets is a key challenge in privacy-preserving machine learning (PPML). Most existing frameworks approximate it through linear operations, which not only result in significant precision loss but also introduce substantial computational overhead. This paper proposes an efficient, verifiable, and accurate security 2-party logistic regression framework (EVA-S2PLoR), which achieves accurate nonlinear function computation through a subtly secure hadamard product protocol and its derived protocols. All protocols are based on a practical semi-honest security model, which is designed for decentralized privacy-preserving application scenarios that balance efficiency, precision, and security. High efficiency and precision are guaranteed by the asynchronous computation flow on floating point numbers and the few number of fixed communication rounds in the hadamard product protocol, where robust anomaly detection is promised by dimension transformation and Monte Carlo methods. EVA-S2PLoR outperforms many advanced frameworks in terms of precision, improving the performance of the sigmoid function by about 10 orders of magnitude compared to most frameworks. Moreover, EVA-S2PLoR delivers the best overall performance in secure logistic regression experiments with training time reduced by over 47.6 % under WAN settings and a classification accuracy difference of only about 0.5 % compared to the plaintext model.
Tianle Tao, Shizhao Peng, Tianyu Mei, Shoumo Li, Haogang Zhu
SRDS5
2025 Meta reinforcement learning based dynamic tuning for blockchain systems in diverse network environments
abstract
The evolution of blockchain technology across various areas has highlighted the importance of optimizing blockchain systems' performance, especially in fluctuating network bandwidth conditions. We observe that the performance of blockchain systems exhibits variations, and the optimal parameter configuration shifts accordingly when changes in network bandwidth occur. Current methods in blockchain optimization require establishing fixed mappings between various environments and their optimal parameters. However, this process exhibits poor sample efficiency and lacks the ability for fast adaptation to novel bandwidth environments. In this paper, we propose MetaTune, a meta-Reinforcement-Learning (meta-RL)-based dynamic tuning method for blockchain systems. MetaTune can quickly adapt to unknown bandwidth changes and automatically configure optimized parameters. Through empirical evaluations of a real-world blockchain system, ChainMaker, we demonstrate that MetaTune significantly reduces the training samples needed for generalization across different bandwidth environments compared to non-adaptive methods. Our findings suggest that MetaTune offers a promising approach for efficiently optimizing blockchain systems in dynamic network environments.
Yue Pei, Mengxiao Zhu 0004, Weihu Song, Haogang Zhu
Blockchain Res. Appl.7
2025 Transaction Spatio-Temporal Distribution for Permissioned Blockchain Performance Profiling
abstract
ABSTRACT Blockchain systems, characterized by decentralization, irreversibility, and traceability, have attracted widespread adoption across security‐critical domains. However, the discrepancy between claimed and observed performance metrics poses significant challenges for system selection and reliability assurance. Particularly, in multi‐layer blockchain architectures with complex inter‐node interactions, pinpointing abnormal nodes or malfunctioning stages remains a non‐trivial task. While macroscopic metrics, such as transactions per second (TPS) and average latency are essential for measuring system capacity, they lack the granularity to capture fine‐grained operational anomalies. To complement these metrics and provide a new diagnostic dimension, we propose a novel metric system grounded in the spatio‐temporal transition of transaction states, introducing three complementary indicators: overall spatio‐temporal transition cost, spatial transition cost, and temporal transition cost. These metrics characterize system‐wide behavior and enable anomaly detection at both the node and stage levels. We further develop a log‐driven analysis framework that leverages these metrics for anomaly localization and root cause inference. Our system is implemented atop ChainMaker, deployed over 16 containerized nodes using Docker and Kubernetes. Experimental results demonstrate that our proposed metrics exhibit strong stability, sensitivity, and precision in detecting abnormal behaviors under varied execution conditions. These results validate the effectiveness of our methodology in providing fine‐grained insights into the performance reliability of blockchain systems.
Jianhuan Mao, Mengxiao Zhu 0004, Haogang Zhu
Concurr. Comput. Pract. Exp.5
2025 The Devil is in the Frequency: Constrained and Adaptive Fine-Grained Domain Perturbation for Robust Medical Segmentation
abstract
Domain generalization (DG) in medical image analysis is critical for achieving consistent and reliable diagnostics across diverse healthcare systems. However, domain shifts resulting from variations in imaging protocols, devices, and practices hinder accurate anatomical identification. While data augmentation shows promise, it struggles to generate diverse samples that bridge domain gaps and often distorts invariant anatomical features, compromising diagnostic integrity. This paper introduces the Adaptive Dual-Space Spectral Perturbation (AdaDSP) framework to address these issues at both broad and fine-grained levels. At the broad level, AdaDSP injects learnable spectral perturbations into input images and intermediate feature maps, significantly enhancing the diversity of the training data. At the fine-grained level, we propose a Fine-Grained Spectral Perturbation module that utilizes two lightweight attention mechanisms to capture sensitive frequency bands that hinder generalization. By injecting multivariate Gaussian noise within a mini-batch, this module better modulates the distribution of frequencies and accomplishes adaptive perturbation of sensitive frequency bands. Furthermore, we introduce a Universal Triple-stage Semantic Constraint Framework to encourage the networks to learn domain-invariant representations while retaining the discriminabtive capacity. Extensive experiments show that our method outperforms state-of-the-art benchmarks, with improvements of 2.40% and 2.99% in two notable medical imaging tasks, respectively.
Yichao Cao, Haogang Zhu
IEEE J. Biomed. Health Informatics3
2025 AVP-AP: Self-Supervised Automatic View Positioning in 3D Cardiac CT via Atlas Prompting
abstract
Automatic view positioning is crucial for cardiac computed tomography (CT) examinations, including disease diagnosis and surgical planning. However, it is highly challenging due to individual variability and large 3D search space. Existing work needs labor-intensive and time-consuming manual annotations to train view-specific models, which are limited to predicting only a fixed set of planes. However, in real clinical scenarios, the challenge of positioning semantic 2D slices with any orientation into varying coordinate space in arbitrary 3D volume remains unsolved. We thus introduce a novel framework, AVP-AP, the first to use Atlas Prompting for self-supervised Automatic View Positioning in the 3D CT volume. Specifically, this paper first proposes an atlas prompting method, which generates a 3D canonical atlas and trains a network to map slices into their corresponding positions in the atlas space via a self-supervised manner. Then, guided by atlas prompts corresponding to the given query images in a reference CT, we identify the coarse positions of slices in the target CT volume using rigid transformation between the 3D atlas and target CT volume, effectively reducing the search space. Finally, we refine the coarse positions by maximizing the similarity between the predicted slices and the query images in the feature space of a given foundation model. Our framework is flexible and efficient compared to other methods, outperforming other methods by 19.8% average structural similarity (SSIM) in arbitrary view positioning and achieving 9% SSIM in two-chamber view compared to four radiologists. Meanwhile, experiments on a public dataset validate our framework's generalizability.
Yan Wang 0076, Mingkun Bao, Bosen Jia, Jian Cheng 0002, Haogang Zhu
IEEE Trans. Medical Imaging9
2024 DomainVoyager: Embracing The Unknown Domain by Prompting for Automatic Augmentation
abstract
For medical image analysis, domain generalization (DG) faces significant challenges due to variances in data across medical imaging devices. Addressing this, we introduce Prompt Guided Domain Aligning Augmentation (PGDAA), a novel approach that harnesses Large Language Models (LLMs) to iteratively refine data augmentation sequences in DG for medical segmentations. Specifically, by harnessing the LLM’s advanced capabilities in interpreting and responding to specific prompts, PGDAA iteratively voyages the parameter searching space for data augmentation, identifying optimal augmentation sequences. In each searching round, the LLM leverages provided prior data augmentation methods, their parameters, and corresponding evaluation results to acquire a sufficient performance memory bank, thereby proposing more effective augmentation sequences to significantly narrowing inter-domain gaps. Notably, this integration of LLMs in DG represents a pioneering application in the field, which adds minor training parameters and can be easily combined with other DG benchmarks for further improvements. Comprehensive experiments reveal our method outperforms the baseline by 4.47% and 5.08% on Fundus and Prostate datasets, achieving 90.10% and 89.28% accuracy, respectively.
Haogang Zhu, Xiu Su
ICME2
2024 Real-World Visual Navigation for Cardiac Ultrasound View Planning
Mingkun Bao, Xinlong Wei, Bosen Jia, Chuanyu Wang, Haogang Zhu
MICCAI (1)11
2024 Mixed Integer Linear Programming for Discrete Sampling Scheme Design in Diffusion MRI
Si-Miao Zhang, Yi-Xuan Wang, Haogang Zhu
MICCAI (2)5
2024 Universal Frequency Domain Perturbation for Single-Source Domain Generalization
abstract
In this work, we introduce a novel approach to single-source domain generalization (SDG) in medical imaging, focusing on overcoming the challenge of style variation in out-of-distribution (OOD) domains without requiring domain labels or additional generative models. We propose a Universal Frequency Perturbation framework for SDG termed as UniFreqSDG, that performs hierarchical feature-level frequency domain perturbations, facilitating the model's ability to handle diverse OOD styles. Specifically, we design a learnable spectral perturbation module that adaptively learns the frequency distribution range of samples, allowing for precise low-frequency (LF) perturbation. This adaptive approach not only generates stylistically diverse samples but also preserves domain-invariant anatomical features without the need for manual hyperparameter tuning. Then, the frequency features before and after perturbation are decoupled and recombined through the Content Preservation Reconstruction operation, effectively preventing the loss of discriminative content information. Furthermore, we introduce the Active Domain-variance Inducement Loss to encourage effective perturbation in the frequency domain while ensuring the sufficient decoupling of domain-invariant and domain-style features. Extensive experiments demonstrate that UniFreqSDG increases the dice score by an average of 7.47% (from 77.98% to 85.45%) on the fundus dataset and 4.99% (from 71.42% to 76.73%) on the prostate dataset compared to the state-of-the-art approaches.
Yichao Cao, Xiu Su, Haogang Zhu
ACM Multimedia4
2024 MoltDB: Accelerating Blockchain via Ancient State Segregation
abstract
Blockchain store states in Log-Structured Merge (LSM) tree-based database. Due to blockchain traceability, the growing ancient states are inevitably stored in the databases. Unfortunately, by default, this process mixescurrentandancientstates in the data layout, increasing unnecessary disk I/O access and slowing transaction execution. This paper proposes MoltDB, a scalable LSM-based database for efficient transaction execution through a novel idea ofancient state segregation, i.e., to segregate current and ancient states in the data layout. However, the frequently generated and uncertainly accessed characteristics of ancient states make the segregation challenging. Thus, we develop an “extract-compact” mechanism to batch extraction process for frequently generated ancient states and the LSM compaction process to relieve additional disk I/O overhead. Moreover, we design an adaptive LSM-based storage for the uncertainly accessed ancient states extracted for on-demand access. We implement MoltDB as a database engine compatible with many mainstream blockchains and integrate it into Ethereum for evaluation. Experimental results show that MoltDB achieves 1.3 × transaction throughput and 30% disk I/O latency savings over the state-of-the-art works.
Junyuan Liang, Wuhui Chen, Zicong Hong, Haogang Zhu, Wangjie Qiu, Zibin Zheng
IEEE Trans. Parallel Distributed Syst.4
2023 Gererating Twin VFs of Unlabeled and Unpaired VF through Generative Model
abstract
The accurate analysis and denoising of visual field (VF) measurements play a crucial role in the diagnosis and monitoring of glaucoma, a widespread optic disease leading to visual impairment. This study presents a novel approach that harnesses generative models, specifically the Variational Autoencoder (VAE) and conditional Generative Adversarial Network (cGAN), to address the challenges of denoising VFs posed by the scarcity of labeled and paired VF data. By generating twin VFs whose patterns are consistent with the input of generative models, the denoising network can be trained using these generated twin data under the framework of VF2VF. Thus the applicable range of VF2VF framework expands to encompass unlabeled and unpaired VF datasets. Experiments demonstrate that the denoising network, trained on cGAN-generated data, effectively enhances precision while maintaining accuracy. Furthermore, sensitivity analysis using independent validation datasets showcases the improved sensitivity of the transformed VF vectors to changes. Besides, this study delves into the influence of data volume on denoising performance and compares the performance of VAE and cGAN models. While the method requires extensive generative network training, it unveils promising avenues for advancing VF analysis and denoising by capitalizing on the power of generative models in the absence of labeled and paired data.
Zhenyu Zhang 0042, Haogang Zhu
BIBM2
2023 An Open-Source Robotic Chinese Chess Player
abstract
Consumer robots can accompany children growing up, improving their abilities while playing and entertaining. This paper presents an open-source, practical, low-cost robotic Chinese chess player. The proposed system includes an elaborate mechanical structure, a simple kinematic solution, a novel robot operating system, real-time and accurate chess recognition. Regarding its mechanical design, it combines a magnetism structure and mechanical cam drive, while the overall system has just three servo motors. At the same time, its control strategy is simple and effective. Furthermore, a lightweight robot message communication mechanism, entitled TinyROS, is developed for computing resource-limited embedded chips. Concerning the recognition process, our CNNbased object detector determines chess and achieves accurate identification. As a result, our robotic Chinese chess player is exquisite and easy for large-scale promotion while improving users' chess skills. Aiming to facilitate future consumer robot research and popularize customer robots, the model's mechanical and software design and the TinyROS protocol are open-sourced at https://github.com/Star-Robot/chinese-chess-robot.
Shan An, Guangfu Che, Jinghao Guo, Konstantinos A. Tsintotas, Fukai Zhang, Junjie Ye 0004, Changhong Fu 0001, Haogang Zhu, Hong Zhang 0013
IROS10
2023 ARCosmetics: a real-time augmented reality cosmetics try-on system
Shan An, Jianye Chen, Zhaoqi Zhu, Fangru Zhou, Yuxing Yang, Yuqing Ma, Xianglong Liu 0001, Haogang Zhu
Frontiers Comput. Sci.8
2023 An Approach of Combining Convolution Neural Network and Graph Convolution Network to Predict the Progression of Myopia
Haogang Zhu, Longbo Wen, Weizhong Lan, Zhikuan Yang
Neural Process. Lett.2
2023 Learning an insertion region for advertisement embedding on planes
Shan An, Fangru Zhou, Shuang Bai, Chang Tang, Haogang Zhu
Signal Process. Image Commun.6
2023 Improving the Quality of Fetal Heart Ultrasound Imaging With Multihead Enhanced Self-Attention and Contrastive Learning
abstract
Fetal congenital heart disease (FCHD) is a common, serious birth defect affecting ∼1% of newborns annually. Fetal echocardiography is the most effective and important technique for prenatal FCHD diagnosis. The prerequisites for accurate ultrasound FCHD diagnosis are accurate view recognition and high-quality diagnostic view extraction. However, these manual clinical procedures have drawbacks such as, varying technical capabilities and inefficiency. Therefore, the automatic identification of high-quality multiview fetal heart scan images is highly desirable to improve prenatal diagnosis efficiency and accuracy of FCHD. Here, we present a framework for multiview fetal heart ultrasound image recognition and quality assessment that comprises two parts: a multiview classification and localization network (MCLN) and an improved contrastive learning network (ICLN). In the MCLN, a multihead enhanced self-attention mechanism is applied to construct the classification network and identify six accurate and interpretable views of the fetal heart. In the ICLN, anatomical structure standardization and image clarity are considered. With contrastive learning, the absolute loss, feature relative loss and predicted value relative loss are combined to achieve favorable quality assessment results. Experiments show that the MCLN outperforms other state-of-the-art networks by 1.52-13.61% when determining the F1 score in six standard view recognition tasks, and the ICLN is comparable to the performance of expert cardiologists in the quality assessment of fetal heart ultrasound images, reaching 97% on a test set within 2 points for the four-chamber view task. Thus, our architecture offers great potential in helping cardiologists improve quality control for fetal echocardiographic images in clinical practice.
Haogang Zhu, Jian Cheng 0002, Jiancheng Han, Ye Zhang 0025, Ying Zhao 0026, Yihua He
IEEE J. Biomed. Health Informatics2
2023 FVP: Fourier Visual Prompting for Source-Free Unsupervised Domain Adaptation of Medical Image Segmentation
abstract
Medical image segmentation methods normally perform poorly when there is a domain shift between training and testing data. Unsupervised Domain Adaptation (UDA) addresses the domain shift problem by training the model using both labeled data from the source domain and unlabeled data from the target domain. Source-Free UDA (SFUDA) was recently proposed for UDA without requiring the source data during the adaptation, due to data privacy or data transmission issues, which normally adapts the pre-trained deep model in the testing stage. However, in real clinical scenarios of medical image segmentation, the trained model is normally frozen in the testing stage. In this paper, we propose Fourier Visual Prompting (FVP) for SFUDA of medical image segmentation. Inspired by prompting learning in natural language processing, FVP steers the frozen pre-trained model to perform well in the target domain by adding a visual prompt to the input target data. In FVP, the visual prompt is parameterized using only a small amount of low-frequency learnable parameters in the input frequency space, and is learned by minimizing the segmentation loss between the predicted segmentation of the prompted target image and reliable pseudo segmentation label of the target image under the frozen model. To our knowledge, FVP is the first work to apply visual prompts to SFUDA for medical image segmentation. The proposed FVP is validated using three public datasets, and experiments demonstrate that FVP yields better segmentation results, compared with various existing methods.
Yan Wang 0076, Jian Cheng 0002, Shuai Shao 0006, Lanyun Zhu, Zhenzhou Wu, Tao Liu 0067, Haogang Zhu
IEEE Trans. Medical Imaging8
2022 VF2VF: Improving Precision while Maintaining Accuracy
abstract
Visual Field (VF), a measurement of retinal function, is one of the most important references in modern glaucoma diagnosis and treatment. While VF is a typical physical-psychological measurement, the results are often imprecise. How to improve measurement accuracy is always one of the most significant issues in clinical practice. This study proposes a VF2VF training method based on the assumption that the pattern of multiple VF measurements in a short period does not change, and designs a neural network to transform the VF measurements. We analyze the VFs after transformation through experiments. Experiments show that the training method of VF2VF can significantly improve the precision while maintaining the accuracy of original VFs. Besides, the VFs after transformation achieve a significant performance improvement in downstream deterioration detection tasks. When the false positive rate is 5%, the Hit Rate increases by 20%. And the transformed VFs can give a warning 1.8 times ahead of the original VFs.
Zhenyu Zhang 0042, Haogang Zhu
IEEE Big Data3
2022 Segmentation of ten fetal heart components with coarse-to-fine cascading and dynamic feature powering
abstract
Abstract Segmenting heart components in the apical four‐chamber view of fetal echocardiography is of critical significance in clinical practice. However, it is difficult to recognize these components due to small‐scale components and the imbalanced ventricular apex orientation. In this study, a novel segmentation framework is proposed to segment ten general fetal heart components for the first time. This framework consists of a multi‐directional fine‐density (MDFD) data augmentation method and a coarse‐to‐fine cascade network (CFCN). MDFD enhances the apex orientation diversity and balances the orientation distribution. CFCN has two stages including a coarse network and a fine network. These two stages have similar structures that consist of a feature extractor and a feature refined layer named as Element‐Wise Power with Dynamic Exponent layer (EWPDE). EWPDE which is a plug‐and‐play module for segmentation refines the features from the feature extractor to position small components accurately. By adopting EWPDE, the influence of each pixel is adjusted and hard pixels of small components are segmented precisely. Based on the dataset, the method is proved to be effective with the high mean intersection over union (mIoU) value and low missing ratio (MR). With MDFD and EWPDE, CFCN that adopts DeepLabV3+ as the feature extractor outperforms the best segmentation results (mIoU:0.480, MR:0.035). Compared to the original performance (mIoU:0.407, MR:0.085) of DeepLabV3+, the method improves the results significantly.
Tingyang Yang, Mengxiao Zhu 0002, Yan Wang 0076, Shan An, Jiancheng Han, Yihua He, Haogang Zhu
IET Image Process.10
2022 FastHand: Fast monocular hand pose estimation on embedded systems
Shan An, Xiajie Zhang, Haogang Zhu, Jianyu Yang 0002, Konstantinos A. Tsintotas
J. Syst. Archit.4
2021 Real-Time Monocular Human Depth Estimation and Segmentation on Embedded Systems
abstract
Estimating a scene’s depth to achieve collision avoidance against moving pedestrians is a crucial and fundamental problem in the robotic field. This paper proposes a novel, low complexity network architecture for fast and accurate human depth estimation and segmentation in indoor environments, aiming to applications for resource-constrained platforms (including battery-powered aerial, micro-aerial, and ground vehicles) with a monocular camera being the primary perception module. Following the encoder-decoder structure, the proposed framework consists of two branches, one for depth prediction and another for semantic segmentation. Moreover, network structure optimization is employed to improve its forward inference speed. Exhaustive experiments on three self-generated datasets prove our pipeline’s capability to execute in real-time, achieving higher frame rates than contemporary state-of-the-art frameworks (114.6 frames per second on an NVIDIA Jetson Nano GPU with TensorRT) while maintaining comparable accuracy.
Shan An, Fangru Zhou, Haogang Zhu, Changhong Fu 0001, Konstantinos A. Tsintotas
IROS4
2021 ARShoe: Real-Time Augmented Reality Shoe Try-on System on Smartphones
abstract
Virtual try-on technology enables users to try various fashion items using augmented reality and provides a convenient online shopping experience. However, most previous works focus on the virtual try-on for clothes while neglecting that for shoes, which is also a promising task. To this concern, this work proposes a real-time augmented reality virtual shoe try-on system for smartphones, namely ARShoe. Specifically, ARShoe adopts a novel multi-branch network to realize pose estimation and segmentation simultaneously. A solution to generate realistic 3D shoe model occlusion during the try-on process is presented. To achieve a smooth and stable try-on effect, this work further develop a novel stabilization method. Moreover, for training and evaluation, we construct the very first large-scale foot benchmark with multiple virtual shoe try-on task-related labels annotated. Exhaustive experiments on our newly constructed benchmark demonstrate the satisfying performance of ARShoe. Practical tests on common smartphones validate the real-time performance and stabilization of the proposed approach.
Shan An, Guangfu Che, Jinghao Guo, Haogang Zhu, Junjie Ye 0004, Fangru Zhou, Zhaoqi Zhu, Aishan Liu, Wei Zhang 0031
ACM Multimedia4
2021 Trail-Traced Threshold Test (T4) With a Weighted Binomial Distribution for a Psychophysical Test
abstract
Clinical visual field testing is performed with commercial perimetric devices and employs psychophysical techniques to obtain thresholds of the differential light sensitivity (DLS) at multiple retinal locations. Current thresholding algorithms are relatively inefficient and tough to get satisfied test accuracy, stability concurrently. Thus, we propose a novel Bayesian perimetric threshold method called the Trail-Traced Threshold Test (T4), which can better address the dependence of the initial threshold estimation and achieve significant improvement in the test accuracy and variability while also decreasing the number of presentations compared with Zippy Estimation by Sequential Testing (ZEST) and FT. This study compares T4 with ZEST and FT regarding presentation number, mean absolute difference (MAD between the real Visual field result and the simulate result), and measurement variability. T4 uses the complete response sequence with the spatially weighted neighbor responses to achieve better accuracy and precision than ZEST, FT, SWeLZ, and with significantly fewer stimulus presentations. T4 is also more robust to inaccurate initial threshold estimation than other methods, which is an advantage in subjective methods, such as in clinical perimetry. This method also has the potential for using in other psychophysical tests.
Yuxin Gong, Haogang Zhu, Marco Miranda, David P. Crabb, Haolan Yang, Wei Bi, David F. Garway-Heath
IEEE J. Biomed. Health Informatics2
2021 Brain Age Estimation From MRI Using Cascade Networks With Ranking Loss
abstract
Chronological age of healthy people is able to be predicted accurately using deep neural networks from neuroimaging data, and the predicted brain age could serve as a biomarker for detecting aging-related diseases. In this paper, a novel 3D convolutional network, called two-stage-age-network (TSAN), is proposed to estimate brain age from T1-weighted MRI data. Compared with existing methods, TSAN has the following improvements. First, TSAN uses a two-stage cascade network architecture, where the first-stage network estimates a rough brain age, then the second-stage network estimates the brain age more accurately from the discretized brain age by the first-stage network. Second, to our knowledge, TSAN is the first work to apply novel ranking losses in brain age estimation, together with the traditional mean square error (MSE) loss. Third, densely connected paths are used to combine feature maps with different scales. The experiments with 6586 MRIs showed that TSAN could provide accurate brain age estimation, yielding mean absolute error (MAE) of 2.428 and Pearson's correlation coefficient (PCC) of 0.985, between the estimated and chronological ages. Furthermore, using the brain age gap between brain age and chronological age as a biomarker, Alzheimer's disease (AD) and Mild Cognitive Impairment (MCI) can be distinguished from healthy control (HC) subjects by support vector machine (SVM). Classification AUC in AD/HC and MCI/HC was 0.904 and 0.823, respectively. It showed that brain age gap is an effective biomarker associated with risk of dementia, and has potential for early-stage dementia risk screening. The codes and trained models have been released on GitHub: https://github.com/Milan-BUAA/TSAN-brain-age-estimation.
Jian Cheng 0002, Ziyang Liu 0002, Zhenzhou Wu, Haogang Zhu, Jiyang Jiang, Wei Wen 0001, Dacheng Tao, Tao Liu 0067
IEEE Trans. Medical Imaging5
2020 MixedAD: A Scalable Algorithm for Detecting Mixed Anomalies in Attributed Graphs
abstract
Attributed graphs, where nodes are associated with a rich set of attributes, have been widely used in various domains. Among all the nodes, those with patterns that deviate significantly from others are of particular interest. There are mainly two challenges for anomaly detection. For one thing, we often encounter large graphs with lots of nodes and attributes in the real-life scenario, which requires a scalable algorithm. For another, there are anomalies w.r.t. both the structure and attribute in a mixed manner. The algorithm should identify all of them simultaneously. State-of-art algorithms often fail in some respects. In this paper, we propose the scalable algorithm called MixedAD. Theoretical analysis is provided to prove its superiority. Extensive experiments are also conducted on both synthetic and real-life datasets. Specifically, the results show that MixedAD often achieves the F1 scores greater than those of others by at least 25% and runs at least 10 times faster than the others.
Mengxiao Zhu 0002, Haogang Zhu
AAAI2
2020 Learning a Cost-Effective Strategy on Incomplete Medical Data
Mengxiao Zhu 0002, Haogang Zhu
DASFAA (2)2
2020 Brain Age Estimation from MRI Using a Two-Stage Cascade Network with Ranking Loss
Ziyang Liu 0002, Jian Cheng 0002, Haogang Zhu, Jicong Zhang, Tao Liu 0067
MICCAI (7)3
2020 Multi-head enhanced self-attention network for novelty detection
abstract
One-class classification (OCC) is a classical problem in computer vision that can be described as the task of classifying outlier class samples (OC samples) from the OCC model trained on inlier class samples (IC samples) when datasets are highly biased toward one class due to the insufficient sample size of the other class. Currently, the adversarial learning OCC (ALOCC) method has been proven to significantly improve OCC performance. However, its drawbacks include instability issues and non-evident reconstruction between the IC and OC samples. Therefore, we propose multihead enhanced self-attention in the ALOCC network, thereby increasing the difference between the IC and OC samples and significantly increasing OCC accuracy compared with ALOCC accuracy. For training, we propose a new loss, called adversarial-balance loss, that effectively solves the training instability problem, further increasing OCC accuracy. The experiments show the effectiveness of the proposed method compared with state-of-art methods.
Yuxin Gong, Haogang Zhu, Xiao Bai 0001, Wenzhong Tang
Pattern Recognit.3
2020 Fetal Congenital Heart Disease Echocardiogram Screening Based on DGACNN: Adversarial One-Class Classification Combined with Video Transfer Learning
abstract
Fetal congenital heart disease (FHD) is a common and serious congenital malformation in children. In Asia, FHD birth defect rates have reached as high as 9.3%. For the early detection of birth defects and mortality, echocardiography remains the most effective method for screening fetal heart malformations. However, standard echocardiograms of the fetal heart, especially four-chamber view images, are difficult to obtain. In addition, the pathophysiological changes in fetal hearts during different pregnancy periods lead to ever-changing two-dimensional fetal heart structures and hemodynamics, and it requires extensive professional knowledge to recognize and judge disease development. Thus, research on the automatic screening for FHD is necessary. In this paper, we proposed a new model named DGACNN that shows the best performance in recognizing FHD, achieving a rate of 85%. The motivation for this network is to deal with the problem that there are insufficient training datasets to train a robust model. There are many unlabeled video slices, but they are tough and time-consuming to annotate. Thus, how to use these un-annotated video slices to improve the DGACNN capability for recognizing FHD, in terms of both recognition accuracy and robustness, is very meaningful for FHD screening. The architecture of DGACNN comprises two parts, that is, DANomaly and GACNN (Wgan-GP and CNN). DANomaly, similar to the ALOCC network, but incorporates cycle adversarial learning to train an end-to-end one-class classification (OCC) network that is more robust and has a higher accuracy than ALOCC in screening video slices. For the GACNN architecture, we use FCH (four chamber heart) video slices at around the end-systole, as screened by DANomaly, to train a WGAN-GP for the purpose of obtaining ideal low-level features that can robustly improve the FHD recognition accuracy. A few annotated video slices, as screened by DANomaly, can also be used for data augmentation so as to improve the FHD recognition further. The experiments show that the DGACNN outperforms other state-of-the-art networks by 1%-20% in recognizing FHD. A comparison experiment shows that the proposed network already outperforms the performance of expert cardiologists in recognizing FHD, reaching 84% in a test. Thus, the proposed architecture has high potential for helping cardiologists complete early FHD screenings.
Yuxin Gong, Haogang Zhu, Jing Lv, Yihua He, Shuliang Wang 0001
IEEE Trans. Medical Imaging3
2018 On Quantifying Local Geometric Structures of Fiber Tracts
Jian Cheng 0002, Tao Liu 0067, Feng Shi 0001, Ruiliang Bai, Jicong Zhang, Haogang Zhu, Dacheng Tao, Peter J. Basser
MICCAI (3)6
2015 Segmentation of Intra-retinal Layers in 3D Optic Nerve Head Images
Chuang Wang 0005, Yaxing Wang, Djibril Kaba, Haogang Zhu, Zidong Wang 0001, Xiaohui Liu 0001, Yongmin Li 0001
ICIG (3)4
2011 Aligning Scan Acquisition Circles in Optical Coherence Tomography Images of The Retinal Nerve Fibre Layer
abstract
Optical coherence tomography (OCT) is widely used in the assessment of retinal nerve fibre layer thickness (RNFLT) in glaucoma. Images are typically acquired with a circular scan around the optic nerve head. Accurate registration of OCT scans is essential for measurement reproducibility and longitudinal examination. This study developed and evaluated a special image registration algorithm to align the location of the OCT scan circles to the vessel features in the retina using probabilistic modelling that was optimised by an expectation-maximization algorithm. Evaluation of the method on 18 patients undergoing large number of scans indicated improved data acquisition and better reproducibility of measured RNFLT when scanning circles were closely matched. The proposed method enables clinicians to consider the RNFLT measurement and its scan circle location on the retina in tandem, reducing RNFLT measurement variability and assisting detection of real change of RNFLT in the longitudinal assessment of glaucoma.
Haogang Zhu, David P. Crabb, Patricio G. Schlottmann, Gadi Wollstein, David F. Garway-Heath
IEEE Trans. Medical Imaging1