EDBT 2026 Demo / reviewers in the wild / expert
Guanwen Zhang
dblp:127/6135
· DBLP profile ↗
22ranked-venue papers
2as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 15 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Score-Based Model for Low-Rank Tensor RecoveryabstractLow-rank tensor decompositions (TDs) provide an effective framework for multiway data analysis. Traditional TD methods rely on predefined structural assumptions, such as CP or Tucker decompositions. From a probabilistic perspective, these methods effectively model the relationships between latent factors and the low-rank tensor using Dirac delta distributions. However, tensor low-rank decomposition is inherently non-unique, leading to a multimodal distribution over possible solutions. Critically, such prior knowledge is rarely available in practical scenarios, particularly regarding the optimal rank structure and contraction rules. To address this issue, we propose a score-based model that eliminates the need for predefined structural or distributional assumptions, enabling the learning of compatibility between tensors and latent factors. Specifically, a neural network is designed to learn the energy function, which is optimized via score matching to capture the gradient of the joint log-probability of tensor entries and latent factors. Our method allows for modeling structures and distributions beyond the Dirac delta assumption. Moreover, integrating the block coordinate descent (BCD) algorithm with the proposed smooth regularization enables the model to perform both tensor completion and denoising. Experimental results demonstrate significant performance improvements across various tensor types, including sparse and continuous-time tensors, as well as visual data. Zhengyun Cheng, Guanwen Zhang, Yi Xu 0008, Wei Zhou 0020, Xiangyang Ji |
AAAI | 3 |
| 2026 | AMA-ViT: Acoustic-Mechanism-Aware Vision Transformer for Underwater Target Recognition
Zhangjie Cai, Ruiting Sun, Zhenhong Liao, Guanwen Zhang, Wei Zhou 0020 |
ICPR (9) | 4 |
| 2026 | SLIQ-ViT: Sensitivity-aware log-uniform integer quantization for efficient vision transformers
Ruiting Sun, Honglu He, Guanwen Zhang, Wei Zhou 0020 |
Expert Syst. Appl. | 3 |
| 2026 | Low-rank tensor recovery via variational schatten-p quasi-norm and Jacobian regularization
Zhengyun Cheng, Guanwen Zhang, Yi Xu 0008, Xiangyang Ji, Wei Zhou 0020 |
Neurocomputing | 3 |
| 2025 | KPDepth-VO: Self-Supervised Learning of Scale-Consistent Visual Odometry and Depth With Keypoint Features From Monocular VideoabstractMonocular visual odometry (VO) is crucial for the application of various autonomous systems. However, the inherent scale ambiguity issue in monocular methods greatly limits their performance in pose estimation. In this paper, we propose a hybrid monocular VO system named KPDepth-VO, which solves camera pose from monocular video based on sparse keypoints. To estimate the scale-consistent relative pose, we present a novel photometric-sensitive depth uncertainty model that accounts for the depth uncertainty introduced by limitations in the photometric error constraint. We also introduce an uncertainty-aware scale recovery strategy that incorporates depth uncertainty for reliable scale alignment. Additionally, we propose a novel difference attention mechanism to construct a point filter that effectively filters out less distinctive points, ensuring high-quality matches for more accurate and efficient pose estimation in the proposed system. Experimental results on the KITTI dataset and Oxford Robotcar dataset demonstrate that our system can predict scale-consistent trajectories from monocular videos and achieve state-of-the-art performance among similar methods. Meanwhile, the depth network within our system achieves competitive depth estimation performance on KITTI depth benchmark. Guanwen Zhang, Zhengyun Cheng, Wei Zhou 0020 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Parallel Heterogeneous Networks With Adaptive Routing for Online Video Lane DetectionabstractLane detection plays a critical role in the field of autonomous driving. Most previous lane detection methods focus exclusively on analyzing individual images and overlook the inter-frame dynamics, while car-mounted cameras capture continuous streams amenable to leveraging intra-video context. In challenging scenes with blur, occlusion, or illumination variations, considering visible lanes from previous frames can aid current frame interpretation. To tackle above challenges, we introduce a parallel heterogeneous framework, called PHNet, for video lane detection. First, a novel router automatically analyzes multi-level features to determine the inference route of candidates with different visual cues. Then, a cross-frame attention mechanism leverages relationships between current candidates and positive embeddings from past frame to aggregate contextual cues. Unlike existing offline video lane detection methods, which face limitations in handling long video sequences and real-time video streams due to computational constraints, our proposed method operates as an online framework capable of processing clips of any length. Our method demonstrates robust lane detection through temporally association modeling and efficient online inference. The extensive experiments on public benchmark show that PHNet has superior performance versus state-of-the-art video lane detection methods. Zhengyun Cheng, Guanwen Zhang, Wei Zhou 0020 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | A Multi-Scale Feature Fusion Network for Chip Surface Defect DetectionabstractChip surface defect detection plays a crucial role in semiconductor manufacturing and the electronics industry. However, chip surface defects have relatively tiny defect areas and complex defect features make the chip surface defect detection task low in accuracy. In this paper, we propose a multi-scale feature fusion network based on encoder-decoder architecture to solve this challenge. Particularly, we propose a Residual Feature Fusion Attention Module (RFFAM) and an Intersection over Union Loss for Defect with Auxiliary Bounding Box (DA-IoU), and we incorporate a Transformer Prediction Head (TPH) for the network. Additionally, in order to solve the problem of the shortage of chip surface defect datasets, this paper proposes a chip surface defect dataset containing four defect categories. The experimental results show that the method not only improves the detection accuracy but also maintains a small number of parameters, satisfying the engineering requirements of chip surface defect detection. Haoang Ren, Mengke Tian, Guanwen Zhang, Wei Zhou 0020 |
ICIP | 3 |
| 2024 | Video-Based Semi-automatic Drivable Area Segmentation
Zhengyun Cheng, Guanwen Zhang, Wei Zhou 0020 |
ICPR (17) | 2 |
| 2024 | MSDNet: A Multi-scale Dense Network for Chip Surface Defect Segmentation
Guanwen Zhang, Wei Zhou 0020 |
ICPR (25) | 3 |
| 2024 | A Data Augmentation Approach for Well Log Interpretation
Yaobin Wang, Guanwen Zhang, Wei Zhou 0020 |
ICPR (25) | 5 |
| 2024 | Self-Supervised Learning of Monocular Visual Odometry and Depth with Uncertainty-Aware Scale ConsistencyabstractThe inherent scale ambiguity issue greatly limits the performance of monocular visual odometry. In recent years, a variety of methods have been proposed for self-supervised learning of ego-motion and depth estimation, incorporating specifically designed scale-consistency constraints that utilize estimated depth as a reference. However, these existing methods neglect the influence of the depth uncertainty introduced by the dominant photometric loss, which leads to unreliable depth estimation in difficult regions and detrimentally affects scale alignment. To solve these problems, we introduces a feature-based visual odometry learning system with an effective scale recovery strategy in this paper. Additionally, we propose a learning method to estimate the photometric-sensitive depth uncertainty for guiding the scale recovery. The proposed method is evaluated on KITTI odometry, and the experimental results demonstrate that our system can predict scale-consistent trajectories from monocular videos and achieves state-of-the-art performance. Moreover, the proposed method achieves competitive performance on KITTI depth estimation. Guanwen Zhang, Wei Zhou 0020 |
ICRA | 2 |
| 2024 | MonoBooster: Semi-Dense Skip Connection With Cross-Level Attention for Boosting Self-Supervised Monocular Depth Estimation
Guanwen Zhang, Zhengyun Cheng, Wei Zhou 0020 |
IEEE Signal Process. Lett. | 2 |
| 2022 | State Deviation Correction for Offline Reinforcement LearningabstractOffline reinforcement learning aims to maximize the expected cumulative rewards with a fixed collection of data. The basic principle of current offline reinforcement learning methods is to restrict the policy to the offline dataset action space. However, they ignore the case where the dataset's trajectories fail to cover the state space completely. Especially, when the dataset's size is limited, it is likely that the agent would encounter unseen states during test time. Prior policy-constrained methods are incapable of correcting the state deviation, and may lead the agent to its unexpected regions further. In this paper, we propose the state deviation correction (SDC) method to constrain the policy's induced state distribution by penalizing the out-of-distribution states which might appear during the test period. We first perturb the states sampled from the logged dataset, then simulate noisy next states on the basis of a dynamics model and the policy. We then train the policy to minimize the distances between the noisy next states and the offline dataset. In this manner, we allow the trained policy to guide the agent to its familiar regions. Experimental results demonstrate that our proposed method is competitive with the state-of-the-art methods in a GridWorld setup, offline Mujoco control suite, and a modified offline Mujoco dataset with a finite number of valuable samples. Hongchang Zhang, Jianzhun Shao, Yuhang Jiang 0001, Shuncheng He, Guanwen Zhang, Xiangyang Ji |
AAAI | 5 |
| 2022 | DILane: Dynamic Instance-Aware Network for Lane Detection
Zhengyun Cheng, Guanwen Zhang, Wei Zhou 0020 |
ACCV (2) | 2 |
| 2022 | Rethinking Low-Level Features for Interest Point Detection and Description
Guanwen Zhang, Zhengyun Cheng, Wei Zhou 0020 |
ACCV (2) | 2 |
| 2021 | A deep learning method for video-based action recognitionabstractAbstract In this paper, a deep learning method for video‐based action recognition is proposed. On the one hand, boundary compensation on the basis of a deep neural network is performed to achieve action proposal. Boundary compensation considering non‐maximum suppression according to sliding window priority is applied to remove redundant windows. To accurately detect boundaries, a boundary compensation network is established with multiple networks to process different numbers of segments. On the other hand, action recognition based on the resultant action proposals is performed. To further utilise boundary compensation, three methods are introduced for key frame selection. Optical flow and RGB features are combined via a channel fusion to realise feature representation. A two‐stream network with a spatiotemporal structure is adopted for action recognition. The proposed method is evaluated on three public datasets. The experimental results demonstrate that the proposed method achieves a superior performance to that of state‐of‐the‐art methods. Guanwen Zhang, Yukun Rao, Wei Zhou 0020, Xiangyang Ji |
IET Image Process. | 1 |
| 2020 | Recurrent Deep Attention Network for Person Re-IdentificationabstractPerson re-identification (re-id) is an important task in video surveillance. It is challenging due to the appearance of person varying a wide range across non-overlapping camera views. Recent years, attention-based models are introduced to learn discriminative representation. In this paper, we consider the attention selection in a natural way as like human moving attention on different parts of the visual field for person re-id. In concrete, we propose a Recurrent Deep Attention Network (RDAN) with an attention selection mechanism based on reinforcement learning. The proposed RDAN aims to progressively observe the identity-sensitive regions to build up the representation of individuals. Extensive experiments on three person reid benchmarks Market-1501, DukeMTMC-reID, and CUHK03-NP demonstrate the proposed method can achieve competitive performance. Xianfei Duan, Guanwen Zhang, Wei Zhou 0020 |
ICPR | 4 |
| 2020 | TSMSAN: A Three-Stream Multi-Scale Attentive Network for Video Saliency DetectionabstractVideo saliency detection is an important low-level task that has been used in a large range of high-level applications. In this paper, we proposed a three-stream multi-scale attentive network (TSMSAN) for saliency detection in dynamic scenes. TSMSAN integrates motion vector (MV) representation, static saliency map, and RGB information in multi-scales together into one framework on the basis of Fully Convolutional Network (FCN) and spatial attention mechanism. On the one hand, the respective motion features, spatial features, as well as the scene features can provide abundant information for video saliency detection. On the other hand, spatial attention mechanism can combine features with multi-scales to focus on key information in dynamic scenes. In this manner, the proposed TSMSAN can encode the spatiotemporal features of the dynamic scene comprehensively. We evaluate the proposed approach on two public dynamic saliency datasets. The experimental results demonstrate TSMSAN is able to achieve the state-of-the-art performance as well as the excellent generalization ability. Furthermore, the proposed TSMSAN can provide more convincing video saliency information, in line with human perception. Guanwen Zhang, Jiaming Yan, Wei Zhou 0020 |
ICPR | 2 |
| 2020 | Deep Reinforcement Learning for Autonomous Driving by Transferring Visual FeaturesabstractDeep reinforcement learning (DRL) has achieved great success in processing vision-based driving tasks. However, the end-to-end training manner makes DRL agents suffer from overfitting training scenes. The agents easily fail to generalize to unseen environments. In this paper, we propose a deep reinforcement learning for autonomous driving by transferring visual features. We formulate the DRL training as a perception and control module and introduce adversarial training mechanism for autonomous driving. The perception module is able to extract invariant features between different domains through adversarial training. While the DRL agent can then be trained on the basis of low dimensional states. In this manner, the proposed approach enables trained agents to adapt to unseen environments by learning robust features invariant across various scenes. We evaluate the proposed approach by transferring visual features between different simulators. The experimental results demonstrate the driving policy trained in the source domain can be directly applied in the target domain, and achieve great efficient and effective performance for autonomous driving. Guanwen Zhang, Wei Zhou 0020 |
ICPR | 3 |
| 2018 | Multi-Perspective Tracking for Intelligent VehicleabstractThe multi-camera array has drawn attention of researchers in recent years, and has been configured and deployed on intelligent vehicle to capture the panoramic views. Understanding surroundings is crucial for the ego-vehicle. This paper presents a Multi-perspective Tracking (MPT) framework for intelligent vehicle. An iterative search procedure is proposed to associate detections and tracklets in different perspectives. This procedure iteratively assigns determined states and estimates non-determined states for the detections and tracklets. An inherent determined and non-determined graph is utilized to reinforce this procedure. For more reliable associations between perspectives, a Siamese convolutional neural network is employed to learn feature representation. The supervised classification and verification signals are added to train the network. The features in different conventional stages are integrated together as the discriminative appearance model. The experiments are conducted on a MPT data set with five perspectives. The proposed framework is tested in each pair of adjacent perspectives for the ability to associate target objects between perspectives. Xiangyang Ji, Guanwen Zhang, Qi Guo 0009 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2017 | Solving Occlusion Problem in Pedestrian Detection by Constructing Discriminative Part LayersabstractOcclusion handling is one of the most challenging issues for pedestrian detection, and no satisfactory achievement has been found in this issue yet. Using human body parts has been considered as a reasonable way to overcome such an issue. In this paper, we propose a brand new approach based on the fusion of Mid-level body part mining and Convolutional Neural Network (CNN) to solve this problem, named DP-CNN(Discriminative Parts CNN). Two main discussions are included in this paper. First, we take an exhaustive analysis on how to mine useful body parts that contribute to pedestrian detection. Multiple ingredients (e.g. feature representation, pedestrian attributes) are analyzed through a wide range of experiments. Second, we convert the part detectors to the middle layer of CNN and re-train the model to get a better adaption of the dataset. Compare to existing approaches based on fine-tuning CNN models, our method is not only robust to occlusion handling but also has a smaller computational cost. Yu Wang 0018, Jien Kato, Guanwen Zhang, Kenji Mase |
WACV | 4 |
| 2012 | Local Distance Comparison for Multiple-shot People Re-identification
Guanwen Zhang, Yu Wang 0018, Jien Kato, Takafumi Marutani, Kenji Mase |
ACCV (3) | 1 |