VLDB 2026 Research / reviewers in the wild / expert
Weidong Zhang 0005
dblp:24/3562-5
· DBLP profile ↗
16ranked-venue papers
8as first author
9since 2021 · last 2026
0000-0003-0081-930XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CVDII: Enhancing One-Shot Skeleton Action Recognition Through Cross-View Dynamic Information InteractionabstractOne-shot 3D skeleton action recognition task struggles with diverse intra-class action execution styles, causing excessive discriminative information to obstruct obtaining separable feature space. We innovatively propose leveraging shared information among intra-class action executions to mitigate the over-influence of discriminative information. To this end, we proposed dynamic information interaction module (DIIM) that enables shared information to effectively weaken excessive discriminative information. Specifically, DIIM facilitates effective information interaction by constructing a guided evolution pool to store execution-related shared information and ensure such information can be retrieved. We devise shared-discriminative projection strategy (SDPS) which adopts different feature extraction strategies for specific skeleton topologies to target mining discriminative and shared information from different views of skeleton data. In summary, our proposed Cross-View Dynamic Information Interaction (CVDII) framework integrates DIIM and SDPS, effectively tackles the problem of discriminative information redundancy caused by diverse intra-class action execution styles. Experiments conducted on NTU 60, NTU 120, PKU-MMD, and Kinetics datasets demonstrate that our proposed CVDII achieves remarkable performance. Youmei Zhang, Weidong Zhang 0005, Zhiheng Li 0005, Bin Li 0042, Wei Zhang 0021 |
IEEE Trans. Image Process. | 4 |
| 2025 | Guided progressive learning for room layout estimation: From pixel-level embeddings to refined depth maps
Weidong Zhang 0005, Ying Liu 0026, Yu Hao 0002 |
Comput. Vis. Image Underst. | 1 |
| 2025 | Deep Learning Based Fine-Grained Image Classification: Recent Advances, Applications and Future OutlookabstractABSTRACT Fine‐grained image classification (FGIC) aims to distinguish visually similar categories by capturing subtle differences, yet the coexistence of large intra‐class variation and small inter‐class differences makes this task highly challenging. This paper provides a systematic review of recent deep learning‐based FGIC methods. According to the type of training data, existing approaches are categorized into four groups: (1) conventional models with large‐scale samples (strongly supervised, weakly supervised, semi‐supervised, and unsupervised); (2) few‐shot learning models for limited‐sample scenarios (e.g. meta‐learning and metric learning); (3) models leveraging external information, including multi‐modal and web‐sourced data; and (4) emerging diffusion‐based models. Representative algorithms in each category are summarized and analysed in terms of their advantages and limitations. The paper also reviews mainstream benchmark datasets and introduces a newly proposed application‐oriented dataset, CIIP‐TPID, to support real‐world tasks. Additionally, practical applications of FGIC in public security, medicine, and commerce are discussed. Finally, future research directions are outlined, including diffusion‐based data augmentation, advanced multi‐modal fusion, transformer architecture optimization, lightweight models for edge deployment, and robustness against noisy labels. This review provides a structured and up‐to‐date reference for researchers and practitioners in the field. Ying Liu 0026, Weidong Zhang 0005, Guojun Lu |
IET Image Process. | 4 |
| 2025 | C2P-Net: Comprehensive Depth Map to Planar Depth Conversion for Room Layout EstimationabstractRoom layout estimation seeks to infer the overall spatial configuration of indoor scenes using perspective or panoramic images. As the layout is determined by the dominant indoor planes, this problem inherently requires the reconstruction of these planes. Some studies reconstruct indoor planes from perspective images by learning pixel-level or instance-level plane parameters. However, directly learning these parameters has the problems of susceptibility to occlusions and position dependency. In this paper, we introduce the Comprehensive depth map to Planar depth (C2P) conversion, which reformulates planar depth reconstruction into the prediction of a comprehensive depth map and planar visibility confidence. Based on the parametric representation of planar depth we propose, the C2P conversion is applicable to both panoramic and perspective images. Accordingly, we present an effective framework for room layout estimation that jointly learns the comprehensive depth map and planar visibility confidence. Due to the differentiability of the C2P conversion, our network autonomously learns planar visibility confidence by constraining the estimated plane parameters and reconstructed planar depth map. We further propose a novel approach for 3D layout generation through sequential planar depth map integration. Experimental results demonstrate the superiority of our method across all evaluated panoramic and perspective datasets. Weidong Zhang 0005, Mengjie Zhou, Jiyu Cheng, Ying Liu 0026, Wei Zhang 0021 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Fast 3D Room Layout Estimation Based on Compact High-Level Representationabstract3D room layout estimation aims to reconstruct the holistic 3D structure from an indoor RGB image. For most of the deep learning-based methods, layout inference is guided by a kind of learned 2D mid-level representation such as pixel-wise surface labels. However, learning such high-resolution 2D representation might suffer from information redundancy and memory consumption, and will increase the runtime of estimation and deployment cost for practical applications. In this paper, we attempt to learn a compact high-level representation with only 29 real numbers for estimating the 3D layout using general regression networks. The learned compact high-level representation contains three components: instance-wise plane parameters, camera intrinsic parameters, and plane location indicators. With the learned representation, the inverse depth map of each plane can be calculated to reconstruct the 3D layout. We further design a set of order-agnostic loss functions to restrict the produced inverse depth maps, with which the model can be trained with either weak 2D layout labels or full 3D layout supervision. Moreover, by jointly learning the plane parameters and locations, the model is benefited from 3D reasoning. Experimental results show that our method is much faster than the existing layout estimation methods and obtains competitive performance on benchmark datasets, showing its potential for real-time applications. Weidong Zhang 0005, Yu Qiao 0001, Ying Liu 0026, Ran Song 0001, Wei Zhang 0021 |
IEEE Trans. Image Process. | 1 |
| 2024 | Image recognition based on lightweight convolutional neural network: Recent advancesabstractImage recognition is an important task in computer vision with broad applications. In recent years, with the advent of deep learning, lightweight convolutional neural network (CNN) has brought new opportunities for image recognition, which allows high-performance recognition algorithms to run on resource-constrained devices with strong representation and generalization capabilities. This paper first presents an overview of several classical lightweight CNN models. Then, a comprehensive review is provided on recent image recognition techniques using lightweight CNN. According to the strategies applied to optimize image recognition performance, existing methods are classified into three categories: (1) model compression, (2) optimization of lightweight network, and (3) combining Transformer with lightweight network. In addition, some representative methods are tested on three commonly used datasets for performance comparison. Finally, technical challenges and future research trends in this field are discussed. Ying Liu 0026, Jiahao Xue, Daxiang Li 0002, Weidong Zhang 0005, Tuan Kiang Chiew, Zhijie Xu |
Image Vis. Comput. | 4 |
| 2023 | Progressive dense feature fusion network for single image deraining
Fuxiang Feng, Youmei Zhang, Weidong Zhang 0005, Bin Li 0042 |
Pattern Recognit. Lett. | 3 |
| 2022 | 3D Layout Estimation via Weakly Supervised Learning of Plane Parameters From 2D SegmentationabstractThe task of 3D layout estimation in an indoor scene is to predict the holistic 3D structural information of the scene from an RGB image. It is costly to obtain the ground truth 3D layout, and this issue severely restricts the learning based 3D layout estimation approaches. In this paper, we present a novel weakly supervised learning framework that is able to learn the 3D layout effectively with 2D layout segmentation mask as supervision. We employ a deep neural network to predict the plane parameters and camera intrinsic parameters in the image. Based on the predicted plane instances, the 3D layout as well as the corresponding depth map and 2D segmentation can be generated. The key objectives for learning meaningful plane parameters are the label consistency of layout segmentation and depth consistency of border pixels from adjacent planes, with which the ground truth 2D layout segmentation is able to supervise the learning of the 3D layout. We further incorporate 3D geometric reasoning and prior knowledge in the learning process to ensure that the learned 3D layout is realistic and reasonable. Experimental results show that our method can produce accurate 3D layout estimates by weakly supervised learning. Weidong Zhang 0005, Youmei Zhang, Ran Song 0001, Ying Liu 0026, Wei Zhang 0021 |
IEEE Trans. Image Process. | 1 |
| 2021 | From Edge to Keypoint: An End-to-End Framework For Indoor Layout EstimationabstractThe task of spatial layout estimation of monocular image is to segment an RGB image of indoor scenes with semantic surface labels (i.e., ceiling, floor, front wall, left wall, and right wall). Most recent methods have to produce layout hypotheses based on the estimated edge map or semantic labels, and then rank the layout hypotheses. In this paper, we present an end-to-end framework that can directly output the layout type and keypoint coordinates (defined in the LSUN challenge). The proposed method takes advantage of transfer learning via learning on the fake samples, i.e., plenty of artificial {type, keypoints, edge map} triplets are generated to learn the mapping from edge maps to keypoint coordinates. Generative adversarial network (GAN) is implemented in this work for domain adaptation of the edge maps. Experimental results show that the proposed method can achieve state-of-the-art layout estimation performance on benchmark datasets. Weidong Zhang 0005, Qian Zhang 0076, Wei Zhang 0021, Jason Gu, Yibin Li 0001 |
IEEE Trans. Multim. | 1 |
| 2020 | GeoLayout: Geometry Driven Room Layout Estimation Based on Depth Maps of Planes
Weidong Zhang 0005, Wei Zhang 0021, Yinda Zhang 0001 |
ECCV (16) | 1 |
| 2020 | Edge-Semantic Learning Strategy for Layout Estimation in Indoor EnvironmentabstractVisual cognition of the indoor environment can benefit from the spatial layout estimation, which is to represent an indoor scene with a 2-D box on a monocular image. In this paper, we propose to fully exploit the edge and semantic information of a room image for layout estimation. More specifically, we present an encoder-decoder network with shared encoder and two separate decoders, which are composed of multiple deconvolution (transposed convolution) layers, to jointly learn the edge maps and semantic labels of a room image. We combine these two network predictions in a scoring function to evaluate the quality of the layouts, which are generated by ray sampling and from a predefined layout pool. Guided by the scoring function, we apply a novel refinement strategy to further optimize the layout hypotheses. Experimental results show that the proposed network can yield accurate estimates of edge maps and semantic labels. By fully utilizing the two different types of labels, the proposed method achieves the state-of-the-art layout estimation performance on the benchmark datasets. Weidong Zhang 0005, Wei Zhang 0021, Jason Gu |
IEEE Trans. Cybern. | 1 |
| 2018 | Long-range terrain perception using convolutional neural networks
Wei Zhang 0021, Qi Chen 0005, Weidong Zhang 0005, Xuanyu He |
Neurocomputing | 3 |
| 2018 | A Feature Descriptor Based on Local Normalized Difference for Real-World Texture ClassificationabstractIn this paper, we propose a normalized difference vector (NDV) for texture representation. Compared to local-binary-pattern-based descriptors, the proposed NDV takes full advantage of the local difference, and the size can be extended flexibly to cover a large local region. We further employ the bag-of-words model to integrate the local descriptors into a global feature representation of an image. In addition, two strategies are introduced for the proposed NDV to achieve rotation invariance. We test the proposed texture descriptor on benchmark datasets, such as AniTex, VehApp, KTH-TIPS2a, OpenSurface, and Kylberg. Classification results demonstrate the superiority of the proposed descriptor over state-of-the-art methods. Wei Zhang 0021, Weidong Zhang 0005, Kan Liu 0001, Jason Gu |
IEEE Trans. Multim. | 2 |
| 2017 | Learning to Predict High-Quality Edge Maps for Room Layout EstimationabstractThe goal of room layout estimation is to predict the three-dimensional box that represents the room spatial structure from a monocular image. In this paper, a deconvolution network is trained first to predict the edge map of a room image. Compared to the previous fully convolutional networks, the proposed deconvolution network has a multilayer deconvolution process that can refine the edge map estimate layer by layer. The deconvolution network also has fully connected layers to aggregate the information of every region throughout the entire image. During the layout generation process, an adaptive sampling strategy is introduced based on the obtained high-quality edge maps. Experimental results prove that the learned edge maps are highly reliable and can produce accurate layouts of room images. Weidong Zhang 0005, Wei Zhang 0021, Kan Liu 0001, Jason Gu |
IEEE Trans. Multim. | 1 |
| 2016 | Pyramid stereo matching for spherical panoramasabstractThis paper presents a novel pyramid stereo matching method to improve the matching accuracy of panoramas. Initial camera parameters and feature correspondences are obtained from Structure From Motion (SFM) with normal images extracted from two panoramas. Then a stereo matching pyramid is constructed to refine the feature correspondences layer by layer, and the correspondence is corrected in the original panoramas. Experimental results show the matching accuracy gains provided by the proposed approach. Jian Weng 0007, Wei Zhang 0021, Weidong Zhang 0005, Jianjie Gao |
VCIP | 3 |
| 2016 | Deep Neural Networks for wireless localization in indoor and outdoor environments
Wei Zhang 0021, Kan Liu 0001, Weidong Zhang 0005, Youmei Zhang, Jason Gu |
Neurocomputing | 3 |