EDBT 2026 Demo / reviewers in the wild / expert
Huaizhong Zhang
dblp:99/7409
· DBLP profile ↗
15ranked-venue papers
7as first author
6since 2021 · last 2026
0000-0001-7867-9453ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Robot manipulation · 80% Segmentation and scene understanding · 20% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-robot interaction · 100% | |
| Computer graphics and multimedia
2 papers |
Virtual and augmented reality · 80% Image and video processing · 20% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Robot manipulation › embodied foundation models
vision-language-action model |
0.9 | 1 | 2025 | SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters · CVPR 2025 |
Virtual and augmented reality
immersive interaction |
0.3 | 1 | 2025 | SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters · CVPR 2025 |
Computer vision › Segmentation and scene understanding › image segmentation › model-based segmentation
deformable model segmentation |
0.2 | 1 | 2015 | Divergence of Gradient Convolution: Deformable Segmentation With Arbitrary Initializations · IEEE Trans. Image Process. 2015 |
Image and video processing › image filtering
nonlinear diffusion |
0.1 | 1 | 2015 | Divergence of Gradient Convolution: Deformable Segmentation With Arbitrary Initializations · IEEE Trans. Image Process. 2015 |
Methods — techniques the papers use, named apart from their topics
synthetic data generation · 2.6multimodal response generation · 2.6VR interface · 2.6gradient descent · 0.4convex relaxation · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Virtual Environments to Real-World Trials: Emerging Trends in Autonomous DrivingabstractAutonomous driving technologies have achieved significant advances in recent years, yet their real-world deployment remains constrained by data scarcity, safety requirements, and the need for generalization across diverse environments. In response, synthetic data and virtual environments have emerged as powerful enablers, offering scalable, controllable, and richly annotated scenarios for training and evaluation. This survey presents a comprehensive review of recent developments at the intersection of autonomous driving, simulation technologies, and synthetic datasets. We organize the landscape across three core dimensions: 1) the use of synthetic data for perception and planning, 2) digital twin-based simulation for system validation, and 3) domain adaptation strategies bridging synthetic and real-world data. We also highlight the role of vision-language models and simulation realism in enhancing scene understanding and generalization. A detailed taxonomy of datasets, tools, and simulation platforms is provided, alongside an analysis of trends in benchmark design. Finally, we discuss critical challenges and open research directions, including Sim2Real transfer, scalable safety validation, cooperative autonomy, and simulation-driven policy learning, that must be addressed to accelerate the path toward safe, generalizable, and globally deployable autonomous driving systems. Aditya Humnabadkar, Arindam Sikdar, Benjamin Cave, Huaizhong Zhang, Nik Bessis, Ardhendu Behera |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous CharactersabstractHuman beings are social animals. How to equip 3D autonomous characters with similar social intelligence that can perceive, understand and interact with humans remains an open yet foundamental problem. In this paper, we introduce SOLAMI, the first end-to-end Social vision-Language-Action (VLA) Modeling framework for Immersive interaction with 3D autonomous characters. Specifically, SOLAMI builds 3D autonomous characters from three aspects: 1) Social VLA Architecture: We propose a unified social VLA framework to generate multi-modal response (speech and motion) based on the user’s multimodal input to drive the character for social interaction. 2) Interactive Multimodal Data: We present Syn-MSI, a synthetic multimodal social interaction dataset generated by an automatic pipeline using only existing motion datasets to address the issue of data scarcity. 3) Immersive VR Interface: We develop a VR interface that enables users to immersively interact with these characters driven by various architectures. Extensive quantitative experiments and user study demonstrate that our framework leads to more precise and natural character responses (in both speech and motion) that align with user expectations with lower latency. Weiye Xiao, Zhengyu Lin, Huaizhong Zhang, Tianxiang Ren, Yang Gao 0042, Zhiqian Lin, Zhongang Cai, Lei Yang 0059, Ziwei Liu 0002 |
CVPR | 4 |
| 2025 | Uncertainty Driven Sampling to Handle Intra-class Imbalance Part Segmentation in WheatabstractWe introduce a novel method to address intra-class imbalance in 3D point cloud segmentation of wheat, focusing on distinguishing between ear and non-ear parts. Variability in plant structure, influenced by factors such as curvature and shape, often leads to data imbalance which complicates segmentation tasks. Our approach utilizes Monte Carlo Dropout to identify and prioritize uncertain samples at the end of each training epoch, employing uncertainty-driven sampling to select samples with the lowest confidence. These samples undergo augmentation through scaling and leaf crossover techniques, enhancing their representation in the training set. Our comparative evaluations demonstrate that this strategy significantly improves the mean Intersection over Union (mIoU) and segmentation accuracy, thereby increasing model robustness for complex 3D plant structures. Reena, John H. Doonan, Huaizhong Zhang, Yonghuai Liu, Fiona M. K. Corke |
IPAS | 3 |
| 2024 | Driving Through Graphs: a Bipartite Graph for Traffic Scene AnalysisabstractWe introduce a novel approach for traffic scene analysis in driving videos by exploring spatio-temporal relationships captured by a temporal frame-to-frame (f2f) bipartite graph, eliminating the need for complex image-level high-dimensional feature extraction. Instead, we rely on object detectors that provide bounding box information. The proposed graph approach efficiently connects objects across frames where nodes represent essential object attributes, and edges signify interactions based on simple spatial metrics such as distance and angles between objects. A key innovation is the integration of dynamic edge attributes, computed using Multilayer Perceptrons (MLP) by exploring this spatial metric. These attributes enhance our Interaction-aware Graph Neural Networks (IA-GNNs) framework by adapting the PageRank-driven approximate personalized propagation of neural predictions (APPNP) scheme and graph attention mechanism in a novel way. This has significantly improved our model’s ability to understand spatio-temporal interactions of multiple objects in traffic scenarios. We have rigorously evaluated our approach on two benchmark datasets, METEOR and INTERACTION, demonstrating its accuracy in analyzing traffic scenarios. This streamlined, graph-based strategy marks a significant shift towards more efficient and insightful traffic scene analysis using video data. Our source code is available at: https://github.com/Addy-1998/Bip_DTG. Aditya Humnabadkar, Arindam Sikdar, Huaizhong Zhang, Ardhendu Behera |
ICIP | 3 |
| 2024 | Clustering-inspired channel selection method for weakly supervised object localization
Xiangru Qiao, Zhiquan Li, Sidong Wu, Yonghuai Liu, Huaizhong Zhang |
Pattern Recognit. Lett. | 10 |
| 2022 | SSP-Regularizer: A Star Shape Prior Based Regularizer for Vessel Lumen Segmentation in Oct ImagesabstractOptical coherence tomography (OCT) is widely used in high-resolution imaging of biological tissues, which can help diagnose coronary heart disease by segmenting the vessel lumen at the pixel-level. However, the lumen shape geometry is not well used in the state-of-the-art techniques for OCT image segmentation, especially the data-driven methods, leaving much room for performance improvement if some geometric features could be exploited to provide prior information. Thanks to the star shape geometry of vessel lumen, in this paper, a new Star Shape Prior based Regularizer (SSP-Regularizer) is proposed to improve segmentation performance. To validate its effectiveness, the proposed SSP-Regularizer is applied to improve the optimization scheme used in Mask-RCNN for vessel lumen segmentation. Experimental results show that superior performance is achieved with SSP-Regularizer, indicating its potentials in OCT imagery and optimization schemes. Huaizhong Zhang, Junyuan Wang 0001, Fuqiang Liu 0001 |
ICIP | 2 |
| 2020 | An Entropy-Based Approach to Real-Time Information Extraction for Industry 4.0abstractIndustry 4.0 has drawn considerable attention from industry and academic research communities. The recent advances in Internet of Things (IoT), Big Data analytics, sensor technology, and artificial intelligence have led to the design and implementation of novel approaches to take full advantage of data-driven solutions applicable to Industry 4.0. With the availability of large datasets, it has become crucially important to identify the appropriate amount of relevant information, which would optimize the overall analysis of the corresponding systems. In this article, specific properties of dynamically evolving data systems are introduced and investigated, which provide framework to assess the appropriate amount of representative information. Marcello Trovati, Huaizhong Zhang, Jeffrey Ray, Xiaolong Xu 0002 |
IEEE Trans. Ind. Informatics | 2 |
| 2020 | Retinal Vascular Network Topology Reconstruction and Artery/Vein Classification via Dominant Set ClusteringabstractThe estimation of vascular network topology in complex networks is important in understanding the relationship between vascular changes and a wide spectrum of diseases. Automatic classification of the retinal vascular trees into arteries and veins is of direct assistance to the ophthalmologist in terms of diagnosis and treatment of eye disease. However, it is challenging due to their projective ambiguity and subtle changes in appearance, contrast, and geometry in the imaging process. In this paper, we propose a novel method that is capable of making the artery/vein (A/V) distinction in retinal color fundus images based on vascular network topological properties. To this end, we adapt the concept of dominant set clustering and formalize the retinal blood vessel topology estimation and the A/V classification as a pairwise clustering problem. The graph is constructed through image segmentation, skeletonization, and identification of significant nodes. The edge weight is defined as the inverse Euclidean distance between its two end points in the feature space of intensity, orientation, curvature, diameter, and entropy. The reconstructed vascular network is classified into arteries and veins based on their intensity and morphology. The proposed approach has been applied to five public databases, namely INSPIRE, IOSTAR, VICAVR, DRIVE, and WIDE, and achieved high accuracies of 95.1%, 94.2%, 93.8%, 91.1%, and 91.0%, respectively. Furthermore, we have made manual annotations of the blood vessel topologies for INSPIRE, IOSTAR, VICAVR, and DRIVE datasets, and these annotations are released for public access so as to facilitate researchers in the community. Yitian Zhao, Yonghuai Liu, Jianyang Xie, Huaizhong Zhang, Yalin Zheng, Yifan Zhao 0001, Yangchun Zhao, Pan Su 0001, Jiang Liu 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2019 | Real-Time Traffic Analysis using Deep Learning Techniques and UAV based VideoabstractIn urban environments there are daily issues of traffic congestion which city authorities need to address. Realtime analysis of traffic flow information is crucial for efficiently managing urban traffic. This paper aims to conduct traffic analysis using UAV-based videos and deep learning techniques. The road traffic video is collected by using a position-fixed UAV. The most recent deep learning methods are applied to identify the moving objects in videos. The relevant mobility metrics are calculated to conduct traffic analysis and measure the consequences of traffic congestion. The proposed approach is validated with the manual analysis results and the visualization results. The traffic analysis process is real-time in terms of the pre-trained model used. Huaizhong Zhang, Mark Liptrott, Nik Bessis, Jianquan Cheng |
AVSS | 1 |
| 2018 | A novel infrared video surveillance system using deep learning based techniquesabstractThis paper presents a new, practical infrared video based surveillance system, consisting of a resolution-enhanced, automatic target detection/recognition (ATD/R) system that is widely applicable in civilian and military applications. To deal with the issue of small numbers of pixel on target in the developed ATD/R system, as are encountered in long range imagery, a super-resolution method is employed to increase target signature resolution and optimise the baseline quality of inputs for object recognition. To tackle the challenge of detecting extremely low-resolution targets, we train a sophisticated and powerful convolutional neural network (CNN) based faster-RCNN using long wave infrared imagery datasets that were prepared and marked in-house. The system was tested under different weather conditions, using two datasets featuring target types comprising pedestrians and 6 different types of ground vehicles. The developed ATD/R system can detect extremely low-resolution targets with superior performance by effectively addressing the low small number of pixels on target, encountered in long range applications. A comparison with traditional methods confirms this superiority both qualitatively and quantitatively. Huaizhong Zhang, Chunbo Luo, Qi Wang 0001, Matthew Kitchin, Andrew Parmley, Jesus Monge-Alvarez, Pablo Casaseca-de-la-Higuera |
Multim. Tools Appl. | 1 |
| 2015 | Divergence of Gradient Convolution: Deformable Segmentation With Arbitrary InitializationsabstractIn this paper, we propose a unified approach to deformable model-based segmentation. The fundamental force field of the proposed method is based on computing the divergence of a gradient convolution field (GCF), which makes the full use of directional information of the image gradient vectors and their interactions across image domain. However, instead of directly using such a vector field for deformable segmentation as in the conventional approaches, we derive a more salient representation for contour evolution, and very importantly, we demonstrate that this representation of image force field not only leads to global minimum through convex relaxation but also can achieve the same result using the conventional gradient descent with an intrinsic regularization. Thus, the proposed method can handle arbitrary initializations. The proposed external force field for deformable segmentation has both edge-based properties in that the GCF is computed from image gradients, and the region-based attributes since its divergence can be treated as a region indication function. Moreover, nonlinear diffusion can be conveniently applied to GCF to improve its performance in dealing with noise interference. We also show the extension of GCF from 2D to 3D. In comparison to the state-of-the-art deformable segmentation techniques, the proposed method shows greater flexibility in model initialization and optimization realization, as well as better performance toward noise interference and appearance variation. Huaizhong Zhang, Xianghua Xie |
IEEE Trans. Image Process. | 1 |
| 2013 | Graph based segmentation with minimal user interactionabstractIn this paper, we present a graph based segmentation method that only requires a single point from user initialization. We incorporate a new image feature into the segmentation scheme. It is derived from a vector field that takes into account gradient vector interactions across the image domain, and has the simplicity of edge based features but also proves to be a useful region indication in two-level segmentation. Effective vector field diffusion is proposed to deal with excessive image noise. Based on a single user point we unravel the image and transfer the object segmentation into a height field segmentation in polar coordinates, which in effect imposes a star shape prior. The search of a minimum closed set on a node weighted, directed graph produces the segmentation result. Comparative analysis on real world images demonstrates promising performances of the proposed method in segmentation accuracy and its simplicity in user interaction. Huaizhong Zhang, Ehab Essa, Xianghua Xie |
ICIP | 1 |
| 2012 | Coupling edge and region-based information for boundary finding in biomedical imagery
Huaizhong Zhang, Philip J. Morrow, Sally I. McClean, Kurt Saetzler |
Pattern Recognit. | 1 |
| 2010 | A stability approach to convergence of curve evolution methods
Huaizhong Zhang, Philip J. Morrow, Sally I. McClean, Kurt Saetzler |
Pattern Recognit. Lett. | 1 |
| 2009 | MCMC-Based Algorithm to Adjust Scale Bias in Large Series of Electron Microscopical Ultrathin Sections
Huaizhong Zhang, E. Patricia Rodriguez, Philip J. Morrow, Sally I. McClean, Kurt Saetzler |
CAIP | 1 |