Hanqi Wang

dblp:136/1035 · DBLP profile ↗
← Back
22ranked-venue papers
9as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 AD-LIO: Adaptive Voxel LiDAR-Inertial Odometry Based on Scene Degradation Factor
abstract
Accurate and stable self-localization across the full domain scenarios is critical for autonomous driving and the Internet of Vehicle (IoV). When autonomous vehicles enter degraded scenarios such as tunnels, self-localization may fail due to the sparse geometric features and high similarity structures. This paper introduces AD-LIO, an adaptive voxel LiDAR inertial odometry framework designed for IoVs, which can quantitatively evaluate the degradation of scenarios to guide the framework in dealing with challenging scenarios. Points-based registration is utilized to guarantee the full details of the feature-lack scene can be described. To ensure this robust registration, a novel point-to-plane strategy, which employs the LiDAR geometry and intensity constraints, is incorporated into the point-to-plane matching. The framework can adaptively change voxel size to adjust the scenario condition in realtime by proposing a lightweight scene degradation factor. This factor can be simply calculated and attached to the scan-to-map optimization process. Extensive experiments under degradation and non-degradation scenarios show that the proposed AD-LIO framework outperforms state-of-the-art methods in robustness and estimation precision. The framework also achieves a favorable balance between precision and computational efficiency across LiDAR configurations from 16 to 128 beams. Furthermore, AD-LIO enables reliable large-scale pose estimation over trajectories exceeding 60 km. Our demonstration video is available now at https://youtu.be/Qy2mD4347mQ.
Hanqi Wang, Huiran Wang, Yucheng Tao
IEEE Internet Things J.1
2026 A knowledge-driven self-supervised learning method for enhancing EEG-based emotion recognition
Hanqi Wang, Peng Ye 0006, Kun Yang 0010, Jichuan Xiong, Tao Chen 0003
Neural Networks1
2026 APDiff: An Adaptive Physics-Guided Diffusion Framework for efficient unpaired image dehazing
Li Zhao 0005, Hanqi Wang, Chenxiang Fan, Haigen Hu, Wenqi Ren, Zhonglong Zheng
Pattern Recognit.2
2026 FourierMask: Explain EEG-Based End-to-End Deep Learning Models in the Frequency Domain
abstract
The rise of EEG-based end-to-end deep learning models has underscored the need to elucidate how these models process time-series raw EEG signals to generate predictions. The frequency domain provides a more suitable perspective for this task due to two key advantages: the strong correlation with cognitive states and the inherent capacity to model long-range temporal dependencies. However, this perspective remains underexplored in existing research. To bridge this gap, we propose FourierMask, the first mask perturbation framework specifically designed for frequency-domain explanation of EEG-based end-to-end models. Our method introduces three key innovations. First, the Fourier-based domain transformation enables direct manipulation of spectral components. Second, A learnable mask mechanism jointly models the spectral-spatial couplings relationship for EEG explanation. Third, a perturbation generator constrained by a target alignment loss ensures natural perturbations by minimizing distribution shift via cluster-aware regularization. We validate our method through experiments on an EEG benchmark dataset across EEGNet, TSCeption, and DeepConvNet models. Our method reaches a 36.0% average accuracy drop gap (vs. 8.6% for LIME and 6.6% for easyPEASI) at the group-level. And, it reaches a 17.8% average accuracy drop gap (vs. 8.9% for LIME and 9.9% for easyPEASI) at the instance-level. Our model-agnostic framework provides a plug-and-play solution for enhancing transparency of EEG-based end-to-end deep learning models. It links model decisions to frequency biomarkers, with potential applications in neuromedicine and brain-computer interfaces.
Hanqi Wang, Kun Yang 0010, Jichuan Xiong, Tao Chen 0003
IEEE J. Biomed. Health Informatics1
2026 CD-HCP: A Heterogeneous Cooperative Perception Framework Utilizing Cross-Modality Dual-Attention
abstract
Cooperative perception is essential for enhancing the perceptual capabilities of intelligent transportation systems. However, existing methods predominantly focus on homogeneous agents, whereas real-world scenarios often involve agents equipped with diverse sensor types. This diversity leads to significant discrepancies in the representation and informational content of shared heterogeneous features, making direct fusion susceptible to information loss or redundancy. To address this challenge, we propose a novel heterogeneous cooperative perception method based on a cross-modality dual-attention mechanism. This mechanism combines global cross-modality attention with local dense spatial attention, capturing inter-modal interactions and fine-grained spatial correlations for robust fusion of heterogeneous features. Additionally, we introduce a multi-scale feature distillation strategy that employs LiDAR bird’s-eye-view features from multiple agents to guide cross-modal alignment, transferring valuable information to image features and reducing modality discrepancies. Quantitative and qualitative experiments conducted on the OPV2V and DAIR-V2X datasets demonstrate that the proposed method achieves state-of-the-art performance in heterogeneous cooperative perception tasks. The framework shows strong potential for deployment in real-world intelligent transportation systems, exhibiting superior detection accuracy, robustness, and computational efficiency, thereby validating its effectiveness and advancement.
Linglong Lin, Han Li 0007, Hanqi Wang
IEEE Trans. Intell. Transp. Syst.5
2025 Physics-Guided Diffusion Model for Unpaired Real-World Dehazing
Hanqi Wang, Chenxiang Fan, Haigen Hu, Li Zhao 0005, Xiaoqin Zhang 0002
PRCV (9)1
2025 V2V-APG: Adversarial Progressive Generalization for Vehicle-to-Vehicle Cooperative Perception
abstract
Vehicle-to-Vehicle (V2V) cooperative perception markedly broadens scene comprehension by aggregating observations from multiple connected vehicles. However, existing methods often suffer from performance degradation when confronted with domain gaps between training and deployment environments. To bridge this domain gap, we introduce V2V-APG, an Adversarial Progressive Generalization framework for cross-domain V2V cooperative perception. V2V-APG first leverages a Spatial-Interaction Perception Attention network to simultaneously model intra-vehicle spatial dependencies and inter-vehicle relationships, yielding rich, discriminative features. Additionally, we propose a Domain-Aware Dynamic Gradient Reversal Layer to dynamically align feature distributions based on real-time domain discrepancies. Finally, a Confidence-Guided Progressive Pseudo-label Refinement strategy iteratively refines target-domain pseudo-labels through adaptive thresholding and curriculum learning, boosting supervision quality without ground-truth annotations. Extensive experiments on the OPV2V and V2V4Real datasets demonstrate that V2V-APG outperforms state-of-the-art methods in cross-domain tasks, achieving superior generalization and robustness.
Chuan Hu 0003, Hanqi Wang
IEEE Internet Things J.5
2025 Efficient and Robust Collaborative Perception via Cross-Vehicle Spatio-Temporal Feature Selecting
abstract
Collaborative perception systems enhance the perception capabilities of individual vehicles by facilitating information exchange between neighbouring vehicles. This approach effectively addresses challenges like occlusions and long-range perceptions that single vehicle cannot manage alone. However, practical applications often face difficulties due to constraints in wireless communication resources and reliability, which limit the effectiveness of latency-sensitive collaborative perception. To overcome these barriers, we introduce CERCP, a Communication Efficient and Robust Collaborative Perception framework. CERCP comprises two core modules: a cross-vehicle spatio-temporal feature selection module, which minimizes communication by transmitting only essential sensor regions with spatio-temporal complementarity, and a global-aware feature synchronization module, which mitigates data delays due to communication latency. To our knowledge, CERCP is the first general collaborative perception framework designed for efficient communication and is applicable across various tasks and modalities. We comprehensively evaluate CERCP on three datasets from real-world and simulated scenarios, using two sensor modalities (LiDAR and camera) and two perception tasks (3D object detection and BEV semantic segmentation). Extensive experiments demonstrate the superior performance of our method.
Kun Yang 0010, Hanqi Wang, Peng Sun 0007
IEEE Trans. Intell. Transp. Syst.4
2024 ERMVP: Communication-Efficient and Collaboration-Robust Multi-Vehicle Perception in Challenging Environments
abstract
Collaborative perception enhances perception performance by enabling autonomous vehicles to exchange complementary information. Despite its potential to revolutionize the mobile industry, challenges in various environments, such as communication bandwidth limitations, localization errors and information aggregation inefficiencies, hinder its implementation in practical applications. In this work, we propose ERMVP, a communication-Efficient and collaboration-Robust Multi-Vehicle Perception method in challenging environments. Specifically, ERMVP has three distinct strengths: i) It utilizes the hierarchical feature sampling strategy to abstract a representative set of feature vectors, using less communication overhead for efficient communication; ii) It employs the sparse consensus features to execute precise spatial location calibrations, effectively mitigating the implications of vehicle localization errors; iii) A pioneering feature fusion and interaction paradigm is introduced to integrate holistic spatial semantics among different vehicles and data sources. To thoroughly validate our method, we conduct extensive experiments on real-world and simulated datasets. The results demonstrate that the proposed ERMVP is significantly superior to the state-of-the-art collaborative perception methods.
Kun Yang 0010, Hanqi Wang, Peng Sun 0007
CVPR4
2024 InLIOM: Tightly-Coupled Intensity LiDAR Inertial Odometry and Mapping
abstract
State estimation and mapping are vital prerequisites for autonomous vehicle intelligent navigation. However, maintaining high accuracy in urban environments remains challenging, especially when the satellite signal is unavailable. This paper proposes a novel framework, InLIOM, which tightly couples LiDAR intensity measurements into the system to improve mapping performance in various challenging environments. The proposed framework introduces a stable intensity LiDAR odometry based on scan-to-scan optimization. By extracting features pairwise from intensity information of consecutive frames, this method tackles the instability issue of LiDAR intensity. To ensure the odometry’s robustness, a training-free residual-based dynamic objects filter module is further integrated into the scan-to-scan registration process. The obtained intensity LiDAR odometry solution is incorporated into the factor graph with other multi-sensors relative and absolute measurements, obtaining global optimization estimation. Experiments in indoor and outdoor urban environments show that the proposed framework achieves superior accuracy to state-of-the-art methods. Our approach can robustly adapt to high-dynamic roads, tunnels, underground parking, and large-scale urban scenarios.
Hanqi Wang, Huawei Liang
IEEE Trans. Intell. Transp. Syst.1
2023 A Novel Efficient Multi-View Traffic-Related Object Detection Framework
abstract
With the rapid development of intelligent transportation system applications, a tremendous amount of multi-view video data has emerged to enhance vehicle perception. However, performing video analytics efficiently by exploiting the spatial-temporal redundancy from video data remains challenging. Accordingly, we propose a novel traffic-related framework named CEVAS to achieve efficient object detection using multi-view video data. Briefly, a fine-grained input filtering policy is introduced to produce a reasonable region of interest from the captured images. Also, we design a sharing object manager to manage the information of objects with spatial redundancy and share their results with other vehicles. We further derive a content-aware model selection policy to select detection methods adaptively. Experimental results show that our framework significantly reduces response latency while achieving the same detection accuracy as the state-of-the-art methods.
Kun Yang 0010, Jing Liu 0050, Dingkang Yang, Hanqi Wang, Peng Sun 0007
ICASSP4
2023 Spatio-Temporal Domain Awareness for Multi-Agent Collaborative Perception
abstract
Multi-agent collaborative perception as a potential application for vehicle-to-everything communication could significantly improve the perception performance of autonomous vehicles over single-agent perception. However, several challenges remain in achieving pragmatic information sharing in this emerging research. In this paper, we propose SCOPE, a novel collaborative perception frame-work that aggregates the spatio-temporal awareness characteristics across on-road agents in an end-to-end manner. Specifically, SCOPE has three distinct strengths: i) it considers effective semantic cues of the temporal context to enhance current representations of the target agent; ii) it aggregates perceptually critical spatial information from heterogeneous agents and overcomes localization errors via multi-scale feature interactions; iii) it integrates multi-source representations of the target agent based on their complementary contributions by an adaptive fusion paradigm. To thoroughly evaluate SCOPE, we consider both real-world and simulated scenarios of collaborative 3D object detection tasks on three datasets. Extensive experiments show the superiority of our approach and the necessity of the proposed components. The project link is https://ydk122024.github.io/SCOPE/.
Kun Yang 0010, Dingkang Yang, Mingcheng Li, Yang Liu 0246, Jing Liu 0050, Hanqi Wang, Peng Sun 0007
ICCV7
2023 What2comm: Towards Communication-efficient Collaborative Perception via Feature Decoupling
abstract
Multi-agent collaborative perception has received increasing attention recently as an emerging application in driving scenarios. Despite advancements in previous approaches, challenges remain due to redundant communication patterns and vulnerable collaboration processes. To address these issues, we propose What2comm, an end-to-end collaborative perception framework to achieve a trade-off between perception performance and communication bandwidth. Our novelties lie in three aspects. First, we design an efficient communication mechanism based on feature decoupling to transmit exclusive and common feature maps among heterogeneous agents to provide perceptually holistic messages. Secondly, a spatio-temporal collaboration module is introduced to integrate complementary information from collaborators and temporal ego cues, leading to a robust collaboration procedure against transmission delay and localization errors. Ultimately, we propose a common-aware fusion strategy to refine final representations with informative common features. Comprehensive experiments in real-world and simulated scenarios demonstrate the effectiveness of What2comm.
Kun Yang 0010, Dingkang Yang, Hanqi Wang, Peng Sun 0007
ACM Multimedia4
2023 DSDCLA: driving style detection via hybrid CNN-LSTM with multi-level attention fusion
Jing Liu 0050, Yang Liu 0246, Hanqi Wang
Appl. Intell.4
2023 A fast coarse-to-fine point cloud registration based on optical flow for autonomous vehicles
Hanqi Wang, Huawei Liang, Liangji Chen
Appl. Intell.1
2023 Dynamic vehicle pose estimation and tracking based on motion feedback for LiDARs
Fengyu Xu 0002, Hanqi Wang, Linglong Lin, Huawei Liang
Appl. Intell.3
2023 A multi-modal vehicle trajectory prediction framework via conditional diffusion model: A coarse-to-fine approach
Huawei Liang, Hanqi Wang
Knowl. Based Syst.3
2021 Stack Multiple Shallow Autoencoders into a Strong One: A New Reconstruction-Based Method to Detect Anomaly
Hanqi Wang, Xing Hu 0006, Yang Liu 0246, Jing Liu 0050, Linhua Jiang
ICONIP (1)1
2018 Identifying Objective and Subjective Words via Topic Modeling
abstract
It is observed that distinct words in a given document have either strong or weak ability in delivering facts (i.e., the objective sense) or expressing opinions (i.e., the subjective sense) depending on the topics they associate with. Motivated by the intuitive assumption that different words have varying degree of discriminative power in delivering the objective sense or the subjective sense with respect to their assigned topics, a model named as dentified bjective- ubjective latent Dirichlet allocation (LDA) ( osLDA) is proposed in this paper. In the osLDA model, the simple Pólya urn model adopted in traditional topic models is modified by incorporating it with a probabilistic generative process, in which the novel "Bag-of-Discriminative-Words" (BoDW) representation for the documents is obtained; each document has two different BoDW representations with regard to objective and subjective senses, respectively, which are employed in the joint objective and subjective classification instead of the traditional Bag-of-Topics representation. The experiments reported on documents and images demonstrate that: 1) the BoDW representation is more predictive than the traditional ones; 2) osLDA boosts the performance of topic modeling via the joint discovery of latent topics and the different objective and subjective power hidden in every word; and 3) osLDA has lower computational complexity than supervised LDA, especially under an increasing number of topics.
Hanqi Wang, Fei Wu 0001, Weiming Lu 0001, Yi Yang 0001, Xi Li 0001, Xuelong Li 0001, Yueting Zhuang
IEEE Trans. Neural Networks Learn. Syst.1
2017 Bag-of-Discriminative-Words (BoDW) Representation via Topic Modeling
abstract
Many of the words in a given document either deliver facts (objective) or express opinions (subjective), respectively, depending on the topics they are involved in. For example, given a bunch of documents, the word “bug” assigned to the topic “order Hemiptera” apparently remarks one object (i.e., one kind of insects), while the same word assigned to the topic “software” probably conveys a negative opinion. Motivated by the intuitive assumption that different words have varying degrees of discriminative power in delivering the objective sense or the subjective sense with respect to their assigned topics, a model named as discriminatively objective-subjective LDA (dosLDA) is proposed in this paper. The essential idea underlying the proposed dosLDA is that a pair of objective and subjective selection variables are explicitly employed to encode the interplay between topics and discriminative power for the words in documents in a supervised manner. As a result, each document is appropriately represented as “bag-of-discriminativewords” (BoDW). The experiments reported on documents and images demonstrate that dosLDA not only performs competitively over traditional approaches in terms of topic modeling and document classification, but also has the ability to discern the discriminative power of each word in terms of its objective or subjective sense with respect to its assigned topic.
Yueting Zhuang, Hanqi Wang, Jun Xiao 0001, Fei Wu 0001, Yi Yang 0001, Weiming Lu 0001, Zhongfei Zhang
IEEE Trans. Knowl. Data Eng.2
2014 Jointly Discovering Fine-grained and Coarse-grained Sentiments via Topic Modeling
abstract
The ever-increasing user-generated contents in social media and other web services make it highly desirable to discover opinions of users on all kinds of topics. Motivated by the assumption that individual word and paragraph in documents will deliver fine-grained (e.g., "laudatory", "annoyed" or "boring") and coarse-grained (e.g., positive, negative or neutral) sentiments about certain topics respectively, this paper focuses on a deeper thematic level to jointly disentangle fine-grained and coarse-grained opinions towards topics in terms of sentiment analysis, named as LDA with multi-grained sentiments (MgS-LDA). As a result, the proposed MgS-LDA not only discovers the topics in social media, but also identifies opinions about a given topic in terms of fine-grained and coarse-grained sentiment. Results of several experiments show that our proposed MgS-LDA achieves better performance on both sentimental classification and topic modeling than related methods.
Hanqi Wang, Fei Wu 0001, Xi Li 0001, Siliang Tang, Jian Shao 0001, Yueting Zhuang
ACM Multimedia1
2013 πLDA: document clustering with selective structural constraints
abstract
Segments, such as sentence boundaries in texts or annotated regions in images, can be considered as useful structural constraints (i.e., priors) for unsupervised topic modeling. However, some segment units (e.g., words in texts or visual words in images) inside a given segment may be irrelevant to the topic of this segment due to their characteristics. This paper proposes a model called πLDA, which introduces a latent variable π into LDA, a traditional topic model, to capture the characteristic of each segment unit. That is to say, the πLDA model is conducted to determine whether a segment unit is assigned (or selected) to the topic embedded in its corresponding segment. Compared with other approaches that assume all the segment units in one segment to share a common topic, our proposed πLDA has the selective ability to discover the discriminative segment units (e.g., informative words or visual words). Experimental results and interpretations of them are presented for demonstrating the promising performance of our method.
Siliang Tang, Hanqi Wang, Jian Shao 0001, Fei Wu 0001, Yueting Zhuang
ACM Multimedia2