VLDB 2026 Research / reviewers in the wild / expert
Xiaogang Zhang 0002
dblp:06/3425-2
· DBLP profile ↗
22ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0002-2576-2576ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Temporal-spatial Causal Variational Network for accurate sintering temperature forecasting in rotary kilns
Hua Chen 0008, Xiaogang Zhang 0002, Qianyu Chen 0002, Yuqi Cai |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | EGAS: Enhanced Geometry-aware 3D Asset Generation Using Gaussian SplattingabstractCurrent text-to-3D generation methods that rely solely on 2D diffusion supervision often suffer from incorrect geometry (e.g. geometric collapse) and unrealistic appearance (e.g. Janus issues) due to the inherent ambiguity in 2D lifting methods for scene representation optimization, which also leads to significant time consumption. In this paper, we introduce EGAS, a novel 3DG-based framework with enhanced geometry awareness, comprising geometry generation and texture refinement. Specifically, we extract pseudo-depth constraints and preliminary appearance guidance from coarse 3D priors to modulate the optimization direction and encourage consistent representation. Moreover, the reasonable geomtry formation is further secured through elaborately designed unsupervised constraints leveraging the particle property of 3D Gaussians. Extensive experiments demonstrate that our method achieves an excellent balance between efficiency and effectiveness, resulting in generated 3D assets with reasonable geometry, high fidelity appearance, and intricate details. Shengjie Hu, Xiaogang Zhang 0002, Hua Chen 0008, Wenbin Yan |
ICASSP | 2 |
| 2025 | Consolidating Selective SSM with Spatial-Angular and Bidirectional Structural Fusion Perception for Light Field Semantic SegmentationabstractWith the advancement of neural networks and computational power, significant progress has been made in the research of light field semantic segmentation. However, existing methods are always limited by capturing long-range global dependencies or secondary computational complexities, which restricts performance. In this paper, we investigate the recently proposed selective State Space Model (SSM) and introduce LFSSNet, a novel Light Field full-aperture efficient Semantic Segmentation Network. We design a series of light field visual feature selective scanning modules, enabling parallel and independent decoupled scanning of 4D complex light fields. These modules efficiently extract inherent spatial, angular, and bidirectional structural feature information from independent 2D slices of the 4D light fields. Additionally, we develop an SSM-attention Cross-fusion Enhancement Module, which further integrates the spatial-angular information and structural complementary information, achieving interactive enhancement of the fused features. Extensive experiments conducted on synthetic and real-world datasets validate the state-of-the-art performance of the proposed method. With 20 times the input data, a 6.73% improvement in accuracy, and a 4 times reduction in computational complexity, the proposed LFSSNet demonstrates its capability to achieve high-precision semantic segmentation utilizing full-aperture light field information while maintaining low computational complexity. Our code is available at: https://github.com/HNU-WQW/LFSSNet. Wenbin Yan, Qingwei Wu, Hua Chen 0008, Xiaogang Zhang 0002, Shengjie Hu |
ICME | 4 |
| 2025 | GeoScene: Temporal 3D Semantic Scene Completion with Geometric Correlation between ImagesabstractSemantic Scene Completion (SSC) aims to reconstruct the entire 3D scene in terms of both occupancy and semantics, serving as a fundamental task for autonomous driving and robotic systems. Camera-based methods have seen significant advancements due to their low cost and rich visual cues. However, previous approaches have predominantly focused on semantic recovery. This can lead to inaccurate occupancy predictions and, consequently, the failure of downstream tasks such as trajectory planning. To address this limitation, we propose a novel multi-frame matching framework, GeoScene, which reconstructs spatial structures through inter-frame geometric correlations of temporal images and subsequently infers scene semantic information. Specifically, we extract features from distinct frames in the depth dimension and derive depth features by constructing a cost volume. Following this, dot product and voxelization operations are applied between the extracted features and depth features to correct assignment errors. Furthermore, we introduce a surface normal-based regression loss to preserve fine-grained surface structures. Extensive experiments on the SemanticKITTI dataset demonstrate that GeoScene outperforms existing state-of-the-art methods. Xiaogang Zhang 0002, Hua Chen 0008, Zhiqiang Miao, Yaonan Wang 0001, Kangcheng Liu |
IROS | 2 |
| 2025 | SPLAC: A single-step PTZ camera linear auto-calibration method
Hua Chen 0008, Xiaogang Zhang 0002 |
Neurocomputing | 3 |
| 2025 | LFSSMam: Efficient Aggregation of Multi-Spatial-Angular-Modal Information Using Selective SSM for Light Field Semantic SegmentationabstractEfficiently aggregating 4D light field information to achieve accurate semantic segmentation has always faced challenges in capturing long range dependency information (CNN-based) and the memory limitations of quadratic computational complexity (Transformer-based). Recently, the Mamba architecture, which utilizes the state space model (SSM), has achieved high performance under linear complexity in various vision tasks. However, directly applying Mamba to 4D light field scanning will lead to an inherent loss of multi-spatial-angular information. To address the above challenges, we introduce LFSSMam, a novel Light Field Semantic Segmentation architecture based on the selective state space model (Mamba). Firstly, LFSSMam presents an innovative spatial-angular selective scanning mechanism to decouple and scan 4D multi-dimensional light field data. It separately captures the rich spatial context, complementary angular and structural information of light field 2D slices within the state space. In addition, we design an SSM-attention Cross-Fusion Enhance Module to perform preferential scanning and fusion across multi-spatial-angular-modal light field information, adaptively aggregating and enhancing the central view features. Comprehensive experiments on synthetic and real world datasets demonstrate that LFSSMam achieves leading edge SOTA (State-Of-The-Art) performance (with a 6.97% improvement to LF-based methods) while reducing memory and computational complexity. This work provides valuable guidance for the efficient modeling and application of multi-spatial-angular information in light field semantic segmentation. Our code is available at https://github.com/HNU-WQW/LFSSMam. Wenbin Yan, Hua Chen 0008, Qingwei Wu, Xiaogang Zhang 0002, Qiu Fang, Shengjie Hu, Yaonan Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | A Novel Distributed Orchestration Engine for Time-Sensitive Robotic Service Orchestration Based on Cloud-Edge CollaborationabstractApplying service orchestration to robot application development can significantly accelerate development processes and facilitate online reconstruction. However, the centralized execution of traditional orchestration engines leads to performance and security challenges, making them inadequate for time-sensitive robot applications. To address these challenges, this article proposes a novel distributed orchestration engine based on a cloud-edge collaboration architecture and execution method that migrates orchestration rules to the edge for collaborative execution. Additionally, we propose a dynamic container deployment method based on a minimum communication latency strategy to enhance the performance of the proposed engine in executing complex orchestration tasks involving algorithm services. The proposed engine is implemented based on a two-layer data routing mechanism and compared with traditional orchestration execution through simulation experiments and application cases. The results demonstrate that it has significant advantages in execution efficiency, load pressure reduction, and data privacy. Xiaogang Zhang 0002, Hua Chen 0008, Naizheng Bian, Jinwen Yin |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | Asymptotic stability analysis of time delayed fractional-order replicator dynamics with government's intervention
Zhang Zhe, Toshimitsu Ushio, Yaonan Wang 0001, Jing Zhang 0014, Xiaogang Zhang 0002 |
Neurocomputing | 5 |
| 2024 | Occlusion-Aware Unsupervised Light Field Depth Estimation Based on Multi-Scale GANsabstractThe estimation of depth from 4D light field images is a fundamental problem for perceiving and reconstructing environmental scenes. While learning-based methods have achieved remarkable results in this field, most of them rely on supervised learning, which faces significant challenges in real-world applications due to the lack of sufficient available ground truth depth maps. In this paper, we propose an unsupervised learning architecture based on a generative adversarial learning model for light field image depth estimation(OALFGAN). Specifically, our approach involves a multi-scale deep convolutional generative adversarial network learning system that includes a sparse-to-dense cascaded multi-scale generator and a discriminator, which decomposes the problem of generating high-quality images into more manageable sub-problems. To address the issue of violations of photometric consistency that may be caused by occlusion, we introduce a spatial-angular attention module that adaptively extracts view features with fewer occlusions and richer textures to generate more accurate disparity maps. Furthermore, we design a loss function that incorporates adaptive angular entropy consistency, symmetry loss, and edge-aware loss based on the distribution regularity and self-constraint of light field images to further optimize occlusion and disparity discontinuity issues and improve the reliability of the final depth prediction. Our proposed method demonstrates superior performance over existing methods on synthetic datasets, both quantitatively and qualitatively. Moreover, our proposed method exhibits excellent generalization performance on real-world datasets, demonstrating the effectiveness of our approach. Wenbin Yan, Xiaogang Zhang 0002, Hua Chen 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Spatio-Temporal Graph Attention Network for Sintering Temperature Long-Range Forecasting in Rotary KilnsabstractMonitoring and forecasting of sintering temperature (ST) is vital for safe, stable, and efficient operation of rotary kiln production process. Due to the complex coupling and time-varying characteristics of process data collected by the distributed control system, its long-range prediction remains a challenge. In this article, we propose a multivariate time series forecasting model based on dynamic spatio-temporal graph attention network (GAT) to model time-varying spatio-temporal correlation between the process data and perform long-range forecasting of ST. Aiming at the problem that there is no preset graph structure for multivariate data, we first propose an adaptive adjacency matrix generation algorithm to construct an elementary graph structure for the process data. Then, we design a spatio-temporal graph attention module, which consists of a multihead GAT for extracting time-varying spatial features and a gated dilated convolutional network for temporal features. Finally, considering the different time delay and rhythm of each process variable, we use dynamic system analysis to estimate the delay time and rhythm of each variable to guide the selection of dilation rates in dilated convolutional layers. The application results based on actual data show that the method has high prediction accuracy, and has broad application prospects in industrial processes. Hua Chen 0008, Yu Jiang 0013, Xiaogang Zhang 0002, Yicong Zhou, Lianhong Wang, Jinchao Wei |
IEEE Trans. Ind. Informatics | 3 |
| 2022 | An efficient unsupervised image quality metric with application for condition recognition in kiln
Leyuan Wu, Xiaogang Zhang 0002, Hua Chen 0008, Yicong Zhou, Lianhong Wang, Dingxiang Wang |
Eng. Appl. Artif. Intell. | 2 |
| 2022 | Combustion Condition Recognition of Coal-Fired Kiln Based on Chaotic Characteristics Analysis of Flame VideoabstractKeeping combustion stable and detecting unstable states in time is crucial for coal-fired furnaces such as rotary kilns, boilers, and oxygen furnaces. Because of the interference and complex conditions in the industrial field, recognition of combustion conditions by vision analysis is difficult. In this article, we propose a robust nonlinear dynamic system analysis-based approach for combustion condition recognition by extracting chaotic characteristics from a flame video. We first discover chaotic characteristics in the intensity sequence extracted from a flame video of coal-fired kilns, and then we further find that the underlying chaos rules differ between combustion conditions. Based on this finding, we design a set of trajectory evolution features and morphology distribution features of chaotic attractors for combustion condition recognition. After reconstructing the chaotic attractors from the intensity sequence of a flame video by phase space reconstruction, the quantified features are extracted from the recurrence plot and morphology distribution and put into a decision tree to recognize the combustion condition. The experimental results on real-world data show that the proposed method can recognize the combustion condition in coal-fired kilns effectively and promptly. Compared with other methods, the recognition accuracy is improved more than 5%. Yu Jiang 0013, Hua Chen 0008, Xiaogang Zhang 0002, Yicong Zhou, Lianhong Wang |
IEEE Trans. Ind. Informatics | 3 |
| 2022 | Unsupervised Monocular Depth Estimation Using Attention and Multi-Warp ReconstructionabstractMonocular depth estimation has become one of the most studied topics in computer vision. Most approaches treat depth prediction as a fully supervised regression problem requiring vast amounts of corresponding ground-truth depth and image pairs for training. Unsupervised monocular depth estimation has emerged as a promising alternative that eliminates dataset limitations. This paper proposes an end-to-end unsupervised deep learning framework integrating attention blocks and multi-warp loss for monocular depth estimation. In this framework, to explore more general contextual information among the feature volumes, an attention block that sequentially refines the feature maps along the channel and spatial dimensions is inserted after the first and last stages of the network encoder. Additionally, to further utilize the errors in the original disparity estimation from the network, a novel multi-warp reconstruction strategy is designed for the loss function. The experimental results evaluated on the KITTI, CityScapes and Make3D datasets demonstrate the state-of-the-art performance and satisfactory generalization ability of our proposed method. Chuanwu Ling, Xiaogang Zhang 0002, Hua Chen 0008 |
IEEE Trans. Multim. | 2 |
| 2021 | VP-NIQE: An opinion-unaware visual perception natural image quality evaluator
Leyuan Wu, Xiaogang Zhang 0002, Hua Chen 0008, Dingxiang Wang, Jingfang Deng |
Neurocomputing | 2 |
| 2021 | BGT: A blind image quality evaluator via gradient and texture statistical features
Jingfang Deng, Xiaogang Zhang 0002, Hua Chen 0008, Leyuan Wu |
Signal Process. Image Commun. | 2 |
| 2021 | Multivariate Time-Series Modeling for Forecasting Sintering Temperature in Rotary Kilns Using DCGNetabstractThe sintering temperature (ST) is a critical index for condition monitoring and process control of coal-fired equipment and is widely used in the production of cement, aluminum, electricity, steel, and chemicals. The accurate prediction of the ST is important for control systems to anticipate tragedies. In this article, we propose a deep learning model for forecasting the ST using automatic spatiotemporal feature extraction from multivariate thermal time series. A hybrid deep neural network named deep convolutional neural network and gated recurrent unit network (DCGNet) is designed to extract multivariate coupling and nonlinear dynamic characteristics for forecasting the ST. DCGNet uses convolutional neural networks and gated recurrent unit (GRU) to extract the local spatial-temporal dependence patterns among the multivariates, and another parallel GRU using the historical ST data as input is incorporated to more accurately capture the dynamic characteristics of ST time series. Based on the real-world data, application results show that the proposed approach has high forecasting accuracy and robustness, thus having broad application prospects in industrial processes. Xiaogang Zhang 0002, Yanying Lei, Hua Chen 0008, Yicong Zhou |
IEEE Trans. Ind. Informatics | 1 |
| 2020 | Unsupervised quaternion model for blind colour image quality assessment
Leyuan Wu, Xiaogang Zhang 0002, Hua Chen 0008, Yicong Zhou |
Signal Process. | 2 |
| 2019 | Kernel modified optimal margin distribution machine for imbalanced data classification
Xiaogang Zhang 0002, Dingxiang Wang, Yicong Zhou, Hua Chen 0008, Fanyong Cheng |
Pattern Recognit. Lett. | 1 |
| 2019 | Effective quality metric for contrast-distorted images based on SVD
Leyuan Wu, Xiaogang Zhang 0002, Hua Chen 0008 |
Signal Process. Image Commun. | 2 |
| 2016 | Recognition of the Temperature Condition of a Rotary Kiln Using Dynamic Features of a Series of Blurry Flame ImagesabstractMaintaining a normal burning temperature is essential to ensuring the quality of nonferrous metals and cement clinker in a rotary kiln. Recognition of the temperature condition is an important component of a temperature control system. Because of the interference of smoke and dust in the kiln, the temperature of the burning zone is difficult to be measured accurately using traditional methods. Focusing on blurry images from which only the flame region can be segmented, an image recognition system for the detection of the temperature condition in a rotary kiln is presented. First, the flame region is segmented employing a region-growing method with a dynamic seed point. Seven features, comprising three luminous features and four dynamic features, are then extracted from the flame region. Dynamic features constructed from luminous feature sequences are proposed to overcome the problem of mis-recognition when the temperature of the flame region changes rapidly. Finally, classifiers are trained to recognize the temperature state of the burning zone using its features. Experimental results using real datasets demonstrate that the proposed image-based systems for recognizing the temperature condition are effective and robust. Hua Chen 0008, Xiaogang Zhang 0002, Pengyu Hong, Hongping Hu |
IEEE Trans. Ind. Informatics | 2 |
| 2014 | Recognition of sintering state in rotary kiln using a robust extreme learning machineabstractSintering is a key process for the industrial clinker production. The sintering state estimation in clinker is an essential factor for its process control. In this paper, a feature extraction method from flame image and a robust extreme learning machine (RB-ELM) classifier are provided to recognize sintering process in rotary kiln. After a preprocessing of image denoising and illumination compensation, material region of flame image is segmented by region growing algorithm and a 5-D statistic feature vector is extracted from it for the following classifier. In order to reduce the influence of outliers in training data caused by blurring image and to achieve a real-time application on site, a robust extreme learning machine, which adopted iterative weight least square (IWLS) method based on M-estimator, is used for fast classification of sintering state. Experimental results show that the proposed method can recognize sintering state accurately, quickly and robustly. Hua Chen 0008, Jing Zhang 0014, Hongping Hu, Xiaogang Zhang 0002 |
IJCNN | 4 |
| 2004 | A sintering temperature detection and control method of alumina rotary kiln based on fuzzy data fusionabstractAn algorithm of data fusion based on fuzzy theory is proposed for the temperature's trend judgment of industry rotary kiln in this paper, which is a key for the automation of this complex industrial process. Previous temperature's detection is based on the flame image process that is easy to be disturbed by dust and smog in the sintering strand. The algorithm proposed in this paper can overcome this disadvantage effectively. In the end, a real time expert control system based on the novel temperature detection method is shown .The practice shows that expert control system with this method can achieve high robust and good utility. Xiaogang Zhang 0002, Hua Chen 0008, Jing Zhang 0014 |
ICARCV | 1 |