Ye Zhang 0037

dblp:147/0497-37 · DBLP profile ↗
← Back
23ranked-venue papers
3as first author
23since 2021 · last 2026
0000-0002-5919-5661ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 2 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 11 since 2021Systems, architecture and hardware · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Triple Spectral Fusion for Sensor-Based Human Activity Recognition
abstract
The field of sensor-based human activity recognition (HAR) mainly uses posture, motion and context data of Inertial Measurement Units (IMUs) to identify daily activities. Despite the advancements in learning-based methods, it is challenging to perform information fusion from the temporal perspective due to the complexities in fusing heterogeneous sensor data and establishing long-term context correlations. This paper proposes a novel triple spectral fusion framework tailored for HAR. First, we develop an adaptive complementary filtering technique for noise suppression and organize each IMU's sensors into posture and motion modality nodes. Given that IMU nodes form a dynamic heterogeneous graph, we then apply adaptive filtering within the graph Fourier domain to merge both homogeneous and heterogeneous node information. Furthermore, an adaptive wavelet frequency selection approach is implemented to suppress context redundancy and shorten the length of features. This approach enhances both timestamp-based graph aggregation and the correlation of long-term contexts. Our framework uses adaptive filtering in the Fourier, graph Fourier, and wavelet domains, enabling effective multi-sensor fusion and context correlation. Extensive experiments on ten benchmark datasets demonstrate the superior performance of our framework.
Ye Zhang 0037, Longguang Wang, Qing Gao 0002, Chaocan Xiang, Mohammed Bennamoun, Yulan Guo
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 GeoStyler: A Generalizable Geometry-Aware Diffusion-Based Approach for Direct 3D Gaussian Style Transfer
abstract
Direct 3D scene stylization from sparse views remains a significant challenge, as existing optimization-based methods are prohibitively slow and require dense inputs to prevent geometric corruption. While recent direct methods accelerate this process, their rigid decoupling of a static geometry from appearance often leads to visual artifacts, where stylistic textures conflict with and distort the underlying scene structure. To address these limitations, we introduce GeoStyler, a direct framework that generates high-fidelity, multi-view consistent stylized 3D scenes in seconds. Our approach reformulates the conventional pipeline by first leveraging a diffusion model to generate a set of geometrically consistent stylized 2D images. The core of this stage is a novel hybrid query formulation for the self-attention mechanism. Specifically, cross-view geometric information is directly embedded into the query to enforce 3D consistency, while style information is independently injected via the key and value to preserve scene structure. This process is further stabilized by a geometrically-aware latent initialization that provides a coherent starting point for the denoising process. Subsequently, a decoupled reconstruction network lifts these 2D stylized images to 3D Gaussians. A geometry branch predicts a robust 3D scaffold from the original content images, while a parallel style branch predicts the final appearance from our generated stylized images, ensuring structural integrity is not compromised. Extensive experiments on large-scale benchmarks, including RealEstate10K and ACID, demonstrate that GeoStyler significantly outperforms prior arts in stylization quality and multi-view consistency, achieving state-of-the-art performance with a dramatic speedup. Our project page: https://huhuhuxiao.github.io/Geo-Styler/.
Qibin Hu, Ye Zhang 0037, Jisheng Dang, Minglin Chen, Longguang Wang, Yulan Guo
IEEE Trans. Image Process.2
2026 EA-GPnP: Efficient and Accurate Generalized-Perspective-n-Point Solution via Optimized Null Space Analysis
Yi Zhang 0130, Baoqiong Wang, Kunhong Li 0001, Xiuqi Wang, Ye Zhang 0037, Yueqiang Zhang, Yulan Guo
IEEE Trans. Robotics6
2025 AIQViT: Architecture-Informed Post-Training Quantization for Vision Transformers
abstract
Post-training quantization (PTQ) has emerged as a promising solution for reducing the storage and computational cost of vision transformers (ViTs). Recent advances primarily target at crafting quantizers to deal with peculiar activations characterized by ViTs. However, most existing methods underestimate the information loss incurred by weight quantization, resulting in significant performance deterioration, particularly in low-bit cases. Furthermore, a common practice in quantizing post-Softmax activations of ViTs is to employ logarithmic transformations, which unfortunately prioritize less informative values around zero. This approach introduces additional redundancies, ultimately leading to suboptimal quantization efficacy. To handle these, this paper proposes an innovative PTQ method tailored for ViTs, termed AIQViT (Architecture-Informed Post-training Quantization for ViTs). First, we design an architecture-informed low-rank compensation mechanism, wherein learnable low-rank weights are introduced to compensate for the degradation caused by weight quantization. Second, we design a dynamic focusing quantizer to accommodate the unbalanced distribution of post-Softmax activations, which dynamically selects the most valuable interval for higher quantization resolution. Extensive experiments on five vision tasks, including image classification, object detection, instance segmentation, point cloud classification, and point cloud part segmentation, demonstrate the superiority of AIQViT over state-of-the-art PTQ methods.
Runqing Jiang, Ye Zhang 0037, Longguang Wang, Pengpeng Yu, Yulan Guo
AAAI2
2025 SaMam: Style-aware State Space Model for Arbitrary Image Style Transfer
abstract
Global effective receptive field plays a crucial role for image style transfer (ST) to obtain high-quality stylized results. However, existing ST backbones (e.g., CNNs and Transformers) suffer huge computational complexity to achieve global receptive fields. Recently, State Space Model (SSM), especially the improved variant Mamba, has shown great potential for long-range dependency modeling with linear complexity, which offers an approach to resolve the above dilemma. In this paper, we develop a Mamba-based style transfer framework, termed SaMam. Specifically, a mamba encoder is designed to efficiently extract content and style information. In addition, a style-aware mamba decoder is developed to flexibly adapt to various styles. Moreover, to address the problems of local pixel forgetting, channel redundancy and spatial discontinuity of existing SSMs, we introduce local enhancement and zigzag scan mechanisms. Qualitative and quantitative results demonstrate that our SaMam outperforms state-of-the-art methods in terms of both accuracy and efficiency.
Hongda Liu 0001, Longguang Wang, Ye Zhang 0037, Ziru Yu, Yulan Guo
CVPR3
2025 Progressive Correspondence Regenerator for Robust 3D Registration
abstract
Obtaining enough high-quality correspondences is crucial for robust registration. Existing correspondence refinement methods mostly follow the paradigm of outlier removal, which either fails to correctly identify the accurate correspondences under extreme outlier ratios, or select too few correct correspondences to support robust registration. To address this challenge, we propose a novel approach named Regor, which is a progressive correspondence regenerator that generates higher-quality matches whist sufficiently robust for numerous outliers. In each iteration, we first apply prior-guided local grouping and generalized mutual matching to generate the local region correspondences. A powerful center-aware three-point consistency is then presented to achieve local correspondence correction, instead of removal. Further, we employ global correspondence refinement to obtain accurate correspondences from a global perspective. Through progressive iterations, this process yields a large number of high-quality correspondences. Extensive experiments on both indoor and outdoor datasets demonstrate that the proposed Regor significantly outperforms existing outlier removal techniques. More critically, our approach obtain 10 times more correct correspondences than outlier removal methods. As a result, our method is able to achieve robust registration even with weak features. The code is available at [Regor].
Guiyu Zhao, Sheng Ao, Ye Zhang 0037, Kai Xu 0004, Yulan Guo
CVPR3
2025 Multi-Modality Test-Time Adaptation for Semantic Segmentation in Robotic Perception
abstract
Test-Time Adaptation (TTA) adjusts pre-trained models in unlabeled unseen environments during the test phase, making it more practical for robotic applications. However, the constant changes of the physical world create significant domain gaps between the received data during robot deployment and the source data used for training. In addition, existing methods mainly focus on a single modality, e.g., RGB images, limiting the application of these methods in multi-modality input scenarios. In this work, we propose a Deep Multi-modality Aggregation Test-time Adaptation (DMATA) method to address the above mentioned issues. To prevent the domain shifts from disrupting the adaptation process, we first propose a Momentum-based Teacher-Student (MTS) framework. Since the teacher model and the student model contain complementary information, we design an Uncertainty-Guide (UG) feature fusion block to fuse the representation of the teacher model and student model of each modality. Finally, we introduce a 3D-Guide-2D (3G2) feature fusion block to leverage spatial information for enhancing 2D feature representation. Extensive experiments across three scenarios, including sensor-to-sensor, day-to-night, and city-to-city, demonstrate the effectiveness of our method in TTA multi-modality semantic segmentation tasks. Notably, under the scenario of sensor-to-sensor adaptation, our proposed DMATA obtains an$m$IoU of 54.2%, which is superior to the state-of-the-art test-time adaptation method.
Yan Liu 0043, Hongyuan Zhu 0002, Ye Zhang 0037, Yinjie Lei, Yulan Guo
ICRA3
2025 $U^2$ Frame: A Unified and Unsupervised Learning Framework for LiDAR-Based Loop Closing
abstract
Loop closing is critically important in Simultaneous Localization and Mapping (SLAM) due to its ability to correct accumulated localization errors. However, existing methods are hindered by the difficulty of acquiring pose labels and the unreliability of ground truth data. In this paper, we propose$U^{2}$Frame, a unified LiDAR-based loop closing framework that handles both loop closure detection and relative pose estimation without any ground truth training data. Specifically, the natural temporal-spatial correlation in point cloud sequences is first leveraged to supervise the network training, where near scans are treated as positives and vice versa as negatives. A new neural architecture is then constructed to jointly learn highly discriminative local and global features for loop closure detection. Additionally, an effective candidate verification module that exploits high-order geometric information is presented to further filter out false loop closures and estimate precise poses. We extensively evaluate$U^{2}$Frame on multiple datasets according to two tasks derived from loop closing: loop closure detection and loop pose estimation. Comparative experiments demonstrate that our method outperforms existing state-of-the-art supervised techniques and has a strong generalization ability across unseen scenarios. Our code is released at https://github.com/yxin-zhang/U2Frame.
Sheng Ao, Ye Zhang 0037, Qingyong Hu, Tao Chang, Yulan Guo
ICRA3
2025 3D Whole-Body Pose Estimation Using Graph High-Resolution Network for Humanoid Robot Teleoperation
abstract
In the realm of robotics, teleoperation plays a pivotal role in performing high-risk or intricate tasks, and obtaining precise 3D whole-body pose is crucial for this purpose. Traditional two-stage methods have limitations in estimating different body parts, leading to complex systems and higher estimation errors. In order to address these issues,the paper introduces a novel framework called Graph High-Resolution Network (GraphHRNet) for accurate 3D whole-body pose estimation, which is essential for the teleoperation of humanoid robots. GraphHRNet effectively captures global structural information and local details by integrating a High-Resolution Module and a Multi-branch Regression Module. The High-Resolution Module utilizes an enhanced graph convolution kernel to fuse multi-scale features, capturing global information, while the Multi-branch Regression Module focuses on refining and predicting accurate 3D coordinates for intricate body parts such as hands and face. Experimental results on the H3WB dataset demonstrate that GraphHRNet surpasses state-of-the-art (SOTA) methods in 3D whole-body pose estimation, significantly improving performance. Furthermore, the paper explores the potential application of this approach in a tele-operation system for humanoid robots, providing an intuitive and high-fidelity solution for remotely executing complex tasks. The code have been publicly available at https://github.com/Z-mingyu/GraphHRNet.git
Qing Gao 0002, Yuanchuan Lai, Ye Zhang 0037, Tao Chang, Yulan Guo
ICRA4
2025 Self-Distilled Stereo Matching: Real-Time Domain Generalization for Robotic Depth Perception
abstract
While human vision inherently achieves robust cross-domain depth estimation through binocular coordination, robotic systems employing stereo matching still confront significant challenges in maintaining robustness across domains when performing real-time environmental depth perception. Furthermore, most stereo matching methods struggle with challenging regions such as object boundaries and non-overlapping areas on the left side of the left image, resulting in disparity maps that are relatively indistinct and lacking fine details. In this paper, we propose Learning More in Challenging Areas (LMC) to alleviate this problem, which enhances the domain generalization of the model through targeted training on challenging regions. LMC is a simple yet effective data-driven training framework primarily based on self-distillation. Specifically, 1) We pre-train models on a high-frequency dataset to improve perception ability on object boundaries; 2) We develop a self-distillation training strategy to benefit learning in non-overlapping areas on the left side of the left image; 3) We design an adaptive difficult area mask to balance the loss weight on other undefined challenging regions. Under our proposed training framework, GwcNet achieves 33% and 23% performance improvements in autonomous driving benchmarks KITTI 2012 and KITTI 2015 respectively, while preserving real-time inference efficiency without computational overhead.
Xuxin Zhang, Kunhong Li 0001, Runqing Jiang, Ye Zhang 0037, Yulan Guo
IROS6
2025 3D skeleton aware driver behavior recognition framework for autonomous driving system
Rongtian Huo, Junkang Chen, Ye Zhang 0037, Qing Gao 0002
Neurocomputing3
2025 Differentiable Prior-Driven Data Augmentation for Sensor-Based Human Activity Recognition
abstract
Sensor-based human activity recognition (HAR) usually suffers from the problem of insufficient annotated data, due to the difficulty in labeling the intuitive signals of wearable sensors. To this end, recent advances have adopted handcrafted operations or generative models for data augmentation. The handcrafted operations are driven by some physical priors of human activities, e.g., action distortion and strength fluctuations. However, these approaches may face challenges in maintaining semantic data properties. Although the generative models have better data adaptability, it is difficult for them to incorporate important action priors into data generation. This article proposes a differentiable prior-driven data augmentation framework for HAR. First, we embed the handcrafted augmentation operations into a differentiable module, which adaptively selects and optimizes the operations to be combined together. Then, we construct a generative module to add controllable perturbations to the data derived by the handcrafted operations and further improve the diversity of data augmentation. By integrating the handcrafted operation module and the generative module into one learnable framework, the generalization performance of the recognition models is enhanced effectively. Extensive experimental results with three different classifiers on five public datasets demonstrate the effectiveness of the proposed framework. Project page:https://github.com/crocodilegogogo/DriveData-Under-Review.
Ye Zhang 0037, Qing Gao 0002, Qingtang Ding, Boyang Li 0007, Yulan Guo
IEEE Trans. Comput. Soc. Syst.1
2025 Enhancing Event-Based Video Reconstruction With Bidirectional Temporal Information
abstract
Event-based video reconstruction has emerged as an appealing research direction to break through the limitations of traditional cameras to better record dynamic scenes. Most existing methods reconstruct each frame from its corresponding event subset in chronological order. Since the temporal information contained in the whole event sequence is not fully exploited, these methods suffer inferior reconstruction quality. In this paper, we propose to enhance event-based video reconstruction by leveraging the bidirectional temporal information in event sequences. The proposed model processes event sequences in a bidirectional fashion, allowing for exploiting bidirectional information in the whole sequence. Furthermore, a transformer-based temporal information fusion module is introduced to aggregate long-range information in both temporal and spatial dimensions. Additionally, we propose a new dataset for the event-based video reconstruction task which contains a variety of objects and movement patterns. Extensive experiments demonstrate that the proposed model outperforms existing state-of-the-art event-based video reconstruction methods both quantitatively and qualitatively.
Pinghai Gao, Longguang Wang, Sheng Ao, Ye Zhang 0037, Yulan Guo
IEEE Trans. Multim.4
2025 Hierarchical Distortion Learning for Fast Lossy Compression of Point Clouds
abstract
The growth of 3D point cloud applications requires efficient compression techniques for high-quality and low-latency services. Recently, learning-based point cloud compression models have made significant progress. However, geometric distortion resulting from downsampling limits the feature depth within large-scale point clouds, thereby constraining the receptive field and suppressing the redundant removal. Moreover, the issues of computational efficiency and reconstruction quality still persist in the compression of large-scale point clouds. To address these challenges, we propose a hierarchical distortion learning framework for end-to-end lossy compression of point clouds. First, we design a feature residual compression module to efficiently transmit shallow semantics between the encoder and the decoder, which enables a lightweight design of our framework. Second, we introduce a geometry residual compression module to progressively complement the reconstruction distortion, avoiding the accumulation of geometric distortion. By integrating these two modules and employing sufficient downsampling processes, we develop a high-performance framework with a significantly enlarged receptive field and low computational cost. Extensive experiments demonstrate that our method achieves state-ofthe- art performance in geometry lossy compression, while delivering competitive performance in joint geometry and color lossy compression with fast running speed. Code is available athttps://github.com/pengpeng-yu/FastPCC.
Pengpeng Yu, Ye Zhang 0037, Fan Liang 0001, Haoran Li 0009, Yulan Guo
IEEE Trans. Multim.2
2025 Energy-guided test-time adaptation for data shifts in multi-modal perception
Yun Pei 0001, Lingbo Liu, Runqing Jiang, Ye Zhang 0037, Pengpeng Yu, Liang Lin 0004, Yulan Guo
Vis. Comput.4
2024 Pluggable Style Representation Learning for Multi-style Transfer
Hongda Liu 0001, Longguang Wang, Weijun Guan, Ye Zhang 0037, Yulan Guo
ACCV (6)4
2024 LoS: Local Structure-Guided Stereo Matching
abstract
Estimating disparities in challenging areas is difficult and limits the performance of stereo matching models. In this paper, we exploit local structure information (LSI) to better handle these areas. Specifically, our LSI comprises a series of key elements, including the slant plane (parameterised by disparity gradients), disparity offset details and neighbouring relations. This LSI empowers our method to effectively handle intricate structures, including object boundaries and curved surfaces. We bootstrap the LSI from monocular depth and subsequently refine it to bet-ter capture the underlying scene geometry constraints in an iterative manner. Building upon the LSI, we introduce the Local Structure-Guided Propagation (LSGP), which enhances the disparity initialization, optimization, and refinement processes. By combining LSGP with a Gated Re-current Unit (GRU), we present our novel stereo matching method, referred to as Local Structure-guided stereo matching (LoS). Remarkably, LoS achieves top-ranking results on four widely recognized public benchmark datasets (ETH3D, Middlebury, KITTI 15 & 12) and robust vision challenge, demonstrating the superior capabilities of our model.
Kunhong Li 0001, Longguang Wang, Ye Zhang 0037, Shunbo Zhou, Yulan Guo
CVPR3
2024 ICPR 2024 Competition on Moving Object Detection and Tracking in Satellite Videos: Methods and Results
Yulan Guo, Qingyong Hu, Feng Zhang 0046, Ye Zhang 0037, Hanyun Wang, Han Wang 0049, Furui Chen, Silei Liu, Xiaomin Huang, Shining Wang, Ying Li 0017, Peng Wang 0015, Shiyong Peng, Xiaokai Bi, Renbin Zou, Wenjing Deng, Zhen Cui 0001
ICPR (34)6
2024 AIP-Net: An anchor-free instance-level human part detection network
Ye Zhang 0037, Yuquan Leng, Qing Gao 0002
Neurocomputing2
2024 GRLoR: A Unified Global Retrieval and Local Reranking Framework for 3-D Place Recognition
abstract
Three-dimensional place recognition aims to search point cloud in a large database that matches the query. It is an essential task in remote sensing applications, such as smart city management and disaster monitoring. The existing methods commonly leverage global descriptors to perform point cloud retrieval for place recognition. However, these methods rely on spatial aggregation to obtain global descriptors, which are neither discriminative nor general. In this letter, we propose a unified global retrieval and local reranking (namely, GRLoR) framework for 3-D place recognition. Specifically, we first utilize a self-attention mechanism to capture the channel dependencies of local features and design a spatial-fusion pooling (SFP) approach to obtain a discriminative global descriptor for retrieval. We then construct a feature correlation module for local reranking, which uses a cross-attention mechanism to determine whether the point cloud pair matches correctly by predicting the similarity of local regions. Experiments conducted on several public benchmarks validate the superiority performance of our method. For instance, it outperforms the strongest model by an average of about 1% on the public datasets in terms of AR@1.
Wenshuo Liu, Sheng Ao, Ye Zhang 0037, Hanyun Wang, Yulan Guo
IEEE Geosci. Remote. Sens. Lett.3
2023 Knowledge transfer via distillation from time and frequency domain for time series classification
Kewei Ouyang, Ye Zhang 0037, Chao Ma 0021, Shilin Zhou 0001
Appl. Intell.3
2022 The First Challenge on Moving Object Detection and Tracking in Satellite Videos: Methods and Results
abstract
In this paper, we briefly summarize the first challenge on moving object detection and tracking in satellite videos (SatVideoDT). This challenge has three tracks related to satellite video analysis, including moving object detection (Track 1), single object tracking (Track 2), and multiple-object tracking (Track 3). 123, 89, and 70 participants successfully registered, while 37, 42, and 29 teams submitted their final results on the test datasets for Tracks 1-3, respectively. The top-performing methods and their results in each track are described with details. This challenge establishes a new benchmark for satellite video analysis.
Yulan Guo, Qingyong Hu, Feng Zhang 0046, Ye Zhang 0037, Hanyun Wang, Chenguang Dai, Weilong Guo, Xiyu Qi, Kelong Tu, Shudan Zhu, Lai Chen, Bin Lin 0013, Chaocan Xue, Jinlei Zheng, Limei Qin, Ying Li 0017, Manqi Zhao, Lu Ruan 0003, Mingpeng Cui, Guanchen Ding, Guangwei Jiang, Zhenzhong Chen 0001, Kaiyang Cao, Lingyu Kong, Shaodong Chen, Zhicheng Zhao 0001, Qin Shen, Lei Liu 0049, Chenglong Li 0002, Yun Xiao 0003
ICPR6
2022 Multi-scale signed recurrence plot based time series classification using inception architectural networks
abstract
Inspired by the great success of deep neural networks in image classification, recent works use Recurrence Plots (RP) to encode time series as images for classification. RP provide rich texture information and construct long-term time correlations, which are effective supplements to the networks. However, RP cannot handle the scale and length variability of sequences. Moreover, RP have serious tendency confusion problem. They cannot represent the upward and downward trends of sequences effectively. In addition to the defects of RP, existing time series classification (TSC) networks cannot adapt to the various scales of discriminative regions of time series effectively. To tackle these problems, this paper proposes a method, named MSRP-IFCN. It is composed of two submodules, the Multi-scale Signed RP (MSRP) and the Inception Fully Convolutional Network (IFCN). MSRP are proposed to handle the defects of RP. They comprise three components, namely the multi-scale RP, the asymmetric RP and the signed RP. We first use the multi-scale RP to enrich the scales of images. Then, the asymmetric RP are constructed to represent long sequences. Finally, the signed RP images are obtained by multiplying the designed sign masks to remove the tendency confusion. Besides, IFCN is proposed to enhance the existing TSC networks in multi-scale feature extraction. By introducing the modified Inception modules, IFCN obtains extensive receptive fields and better extracts multi-scale features from the MSRP images. Experimental results on 85 UCR datasets indicate the superior performance of MSRP-IFCN. The visualization results further demonstrate the effectiveness of our method.
Ye Zhang 0037, Kewei Ouyang, Shilin Zhou 0001
Pattern Recognit.1