Weihang Kong

dblp:187/2991 · DBLP profile ↗
← Back
33ranked-venue papers
11as first author
25since 2021 · last 2026
0000-0003-1453-0603ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 8 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Learnable multi-scale temporal convolutional network for multichannel time-series classification
He Li 0053, Weihang Kong
Eng. Appl. Artif. Intell.5
2026 BoostCount: Diffusion-based position-sensitive adversarial purification for crowd counting
Yiru Du, Linya Huang, He Li 0053, Weihang Kong
Expert Syst. Appl.4
2026 Multi-view pedestrian detection via residual mask fusion and cosine similarity-based passive sampler for video surveillance systems
abstract
Multi-view pedestrian detection aims to generate a bird’s-eye view occupancy map of pedestrians from multiple calibrated camera views. Multi-view methods offer advantages over single-view approaches: they can mitigate occlusions, expand scene coverage, and improve robustness. However, existing multi-view detection methods still face two critical challenges: mixing heterogeneous cross-view information in the fused representation and feature misalignment in the world coordinate system caused by various scales across views. To solve these issues, we develop a novel multi-view pedestrian detection framework that includes a residual mask fusion module and a cosine similarity-based passive sampler. Specifically, the residual mask fusion module enables adaptive feature selection and compensation across views, yielding an optimal fusion under geometric redundancy. Moreover, the cosine similarity-based passive sampler computes dynamic coordinate offsets by evaluating feature consistency. This reduces the impact of unavoidable biases introduced during projection. Experimental results on Wildtrack, MultiviewX and CityStreet demonstrate the effectiveness and reliability of the developed framework for multi-view pedestrian detection. Our code is available at https://github.com/guixiaojia/improve-shot .
He Li 0053, Jiajia Gui, Weihang Kong
Future Gener. Comput. Syst.3
2026 Haze-prior guided frequency-embedded attention learning for single-stage hazy-weather crowd counting
abstract
Conventional two-stage hazy crowd counting suffers from error propagation between separate dehazing and counting pipelines, leading to degraded performance. To address this, we propose an end-to-end single-stage framework that jointly optimizes haze-invariant feature learning and crowd density estimation, achieving state-of-the-art accuracy through two key innovations. Frequency-embedded Hybrid Attention Aggregation (FHAA): This module uses frequency-domain attention to explore frequency features in hazy images, thereby enhancing key feature capture and improving feature learning. Experiments show it reduces Mean Absolute Error (MAE) by 34.48% compared to the model without it, proving its effectiveness in boosting performance. Haze-prior Guided Learning Mechanism: It explicitly models haze distortion, understands haze’s impact on images, and adaptively mitigates interference without manual dehazing annotations, reducing annotation cost and difficulty. Comparative experiments reveal a further 13.64% MAE reduction compared to the model without this mechanism, validating its anti-interference capability. The FHAA module focuses on key frequency features for crowd counting, suppressing haze noise and improving robustness in hazy weather. The haze-prior mechanism uses predicted haze distribution maps to adjust feature learning based on haze intensity, adapting to complex hazy scenes. To support research, we release two synthetic hazy crowd counting datasets at https://github.com/312524/Hazy-CC-extended . These datasets, with the same scale as Hazy-ShanghaiTechRGBD but higher haze densities, address the lack of haze-intensity diversity in existing benchmarks. Extensive ablation studies and performance comparisons on four datasets demonstrate the feasibility and superiority of our method for hazy-weather crowd counting.
Weihang Kong, Jienan Shen, Liangang Tong, He Li 0053
Inf. Process. Manag.1
2026 DualTTrack: Improving information flow consistency in multi-object tracking via dual-track matching and recovery
abstract
• Introduce a Potential Trajectory pool to manage ambiguous low-confidence detections. • Design a Dual-Track Matching mechanism considering main and potential trajectories. • Propose a Dual-Track Recovery mechanism for reactivating lost main trajectories. • Develop a complete and robust tracking framework for multi-object tracking.
He Li 0053, Weihang Kong
Inf. Process. Manag.3
2026 Three-stage progressive modality reweighting and interaction framework for robust RGB-T tracking
abstract
RGB and thermal infrared (TIR) modalities provide complementary visual cues, enabling more robust tracking in adverse lighting and weather conditions. However, most existing RGB-T trackers focus on static feature aggregation, lacking awareness of how modality reliability evolves over time and across scenes. This often leads to unstable performance when the visibility of one modality fluctuates or degrades. To address this issue, we propose a Three-stage Progressive Modality Reweighting and Interaction (TPMRI) framework that progressively models temporal dynamics, semantic interaction, and global fusion consistency. Specifically, a temporal modality reweighting strategy captures short-term temporal variations and dynamically adjusts modality reliability; a cross-modality interaction mechanism establishes fine-grained semantic alignment between RGB and TIR features, facilitating TIR-guided information compensation; and a modality-aware fusion stage performs adaptive global aggregation for stable and consistent representations. These three modules are integrated into the core component, namely the Three-stage Progressive Interaction (TPI) structure, which unifies temporal, semantic, and global modeling within a single framework. By incorporating TPI into the tracker, TPMRI effectively mitigates temporal bias accumulation and enhances robustness under dynamic conditions. Extensive experiments on GTOT, RGBT234, and LasHeR benchmarks demonstrate that our method achieves superior tracking accuracy and success rate compared with state-of-the-art RGB-T trackers.
He Li 0053, Weihang Kong
Knowl. Based Syst.3
2026 Towards a robust adversarial patch attack against RGB-T crowd counting
abstract
Deep learning-based RGB-T crowd counting has attracted increasing attention due to its robustness under diverse environmental conditions, supporting applications in public safety monitoring and smart city systems. However, the vulnerability of such multimodal models to adversarial attacks remains largely unexplored, posing significant threats to their reliability and security in real-world deployments. To address this gap, we propose AttackCounting, a novel adversarial patch attack specifically designed for RGB-T crowd counting models. Unlike existing attack methods that target single-modal inputs, our approach jointly leverages multi-modal saliency from both RGB and thermal domains to identify cross-modal critical regions. By injecting adversarial perturbations into these regions, the attack effectively disrupts the cooperative fusion process between modalities. Furthermore, a local distortion incentive module is introduced to concentrate perturbations in crowd-dense areas while suppressing irrelevant noise. Extensive experiments on several state-of-the-art RGB-T crowd counting models demonstrate that AttackCounting significantly degrades counting accuracy and system reliability. This study exposes fundamental security weaknesses in current RGB-T crowd counting systems and highlights the urgent need for developing attack-resilient multimodal perception models.
Weihang Kong, Yiru Du, He Li 0053
Pattern Recognit.1
2026 AttackCC-CDM: Stealthy backdoor attacks on crowd counting via crowd distribution manipulation
He Li 0053, Mimi Gu, Weihang Kong
Pattern Recognit.4
2026 GFAF-Net: Global-Local Feature Extraction and Adaptive Fusion Network for RGB-T Crowd Counting in Video Surveillance Systems
Liangang Tong, He Li 0053, Weihang Kong
IEEE Trans Autom. Sci. Eng.3
2025 MSASW: Multi-view scale-aware pedestrian detection with shared-weight supervision
He Li 0053, Taiyu Liao, Weihang Kong
Adv. Eng. Informatics3
2025 Multi-modal multi-level feature representation learning for flow pattern identification of oil-water two-phase flow
Weihang Kong, Yaohan Chi, Hongbao Tang, He Li 0053
Eng. Appl. Artif. Intell.1
2025 PromptHC: Multi-attention prompt guided haze-weather crowd counting
Jienan Shen, Liangang Tong, Weihang Kong
Expert Syst. Appl.4
2025 An Adversarial Perturbation Attack Method for Crowd Counting With Scope-Restricted Perturbation and Label-Setting Adaption
abstract
With the rapid development of Internet of Things (IoT) techniques, crowd counting serves as one of the important components in the intelligent surveillance systems for smart city, not only effectively ensuring public safety but also significantly enhancing the efficiency of resource allocation. Existing crowd counting works based on deep neural network are sensitive to small perturbations in their input, leading to concerns about task security. Taking into full account the vulnerability of gradient information to attack, this article develops an adversarial example generation method to realize an effective attack on crowd counting. Specifically, the developed attack method restricts the global perturbation range to generate adversarial examples with high stealthiness, during the gradient descent process; subsequently, the effective attacks are conducted on crowd-counting models with transparent testing samples. Further, since the detailed annotations in testing samples might not be accessible or visible to the attacker, an annotation information alternative strategy is developed under this more practical condition, to expand the scope of practical application scenarios for the proposed method. Experimental results indicate that the generated adversarial examples from the proposed method exhibit significant attack performance against various crowd counting models across multiple public datasets, especially at most +1072.4 mean absolute error (MAE) and +1265.3 root-mean-squared error (RMSE) when attacking CHSNet model on UCF-QNRF. When applied to real-world scenarios, the proposed method also demonstrates significant attack effectiveness. This work presents a covert yet effective adversarial attack method with wide-ranging application scenarios, inspiring current intelligent monitoring fields, such as crowd counting to pay attention to the vulnerabilities of artificial intelligence algorithms and prompts researchers and developers to place greater emphasis on the security of intelligent surveillance systems. This work also provides a solution reference to enhance the robustness of subsequent crowd counting models.
Weihang Kong, Linya Huang, Yiru Du, He Li 0053
IEEE Internet Things J.1
2025 RestoCrowd: Image restoration coadjutant learning for hazy-weather crowd counting without weather labels
Weihang Kong, Jienan Shen, He Li 0053
Pattern Recognit.1
2024 Cross-modal misalignment-robust feature fusion for crowd counting
Weihang Kong, Zepeng Yu, He Li 0053, Junge Zhang
Eng. Appl. Artif. Intell.1
2024 Cross-modal collaborative feature representation via Transformer-based multimodal mixers for RGB-T crowd counting
Weihang Kong, Yao Hong, He Li 0053, Jienan Shen
Expert Syst. Appl.1
2024 CrowdAlign: Shared-weight dual-level alignment fusion for RGB-T crowd counting
Weihang Kong, Zepeng Yu, He Li 0053, Liangang Tong, Fengda Zhao
Image Vis. Comput.1
2023 Multi-level learning counting via pyramid vision transformer and CNN
He Li 0053, Weihang Kong
Eng. Appl. Artif. Intell.3
2023 Direction-aware attention aggregation for single-stage hazy-weather crowd counting
Weihang Kong, Jienan Shen, He Li 0053, Junge Zhang
Expert Syst. Appl.1
2023 CSA-Net: Cross-modal scale-aware attention-aggregated network for RGB-T crowd counting
He Li 0053, Junge Zhang, Weihang Kong, Jienan Shen, Yuguang Shao
Expert Syst. Appl.3
2023 Multi-Scale Geometric Consistency Guided and Planar Prior Assisted Multi-View Stereo
abstract
In this paper, we propose some efficient multi-view stereo methods for accurate and complete depth map estimation. We first present our basic methods with Adaptive Checkerboard sampling and Multi-Hypothesis joint view selection (ACMH & ACMH+). Based on our basic models, we develop two frameworks to deal with the depth estimation of ambiguous regions (especially low-textured areas) from two different perspectives: multi-scale information fusion and planar geometric clue assistance. For the former one, we propose a multi-scale geometric consistency guidance framework (ACMM) to obtain the reliable depth estimates for low-textured areas at coarser scales and guarantee that they can be propagated to finer scales. For the latter one, we propose a planar prior assisted framework (ACMP). We utilize a probabilistic graphical model to contribute a novel multi-view aggregated matching cost. At last, by taking advantage of the above frameworks, we further design a multi-scale geometric consistency guided and planar prior assisted multi-view stereo (ACMMP). This greatly enhances the discrimination of ambiguous regions and helps their depth sensing. Experiments on extensive datasets show our methods achieve state-of-the-art performance, recovering the depth estimation not only in low-textured areas but also in details. Related codes are available at https://github.com/GhiXu.
Qingshan Xu 0001, Weihang Kong, Wenbing Tao, Marc Pollefeys
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 RGB-D Crowd Counting With Cross-Modal Cycle-Attention Fusion and Fine-Coarse Supervision
abstract
To tackle the negative effect of the arbitrary crowd distribution on the counting task, in this article, we propose a novel RGB-D crowd counting approach, including a cross-modal cycle-attention fusion (CmCaF) model and a novel fine-coarse (FC) supervision. In the feature level, the CmCaF model combines the RGB feature and depth feature in a cycle-attention way so as to model the crowd distribution effectively. In the supervision level, the novel design of FC supervision could optimize the counting model from both the fine pixel-aware level and coarse region-aware level to enhance its sensitivity to the whole crowd distribution and the instance location. Extensive evaluations on benchmarks well illustrate the feasibility of the proposed approach for the RGB-D crowd counting, as well as RGB and RGB-T counting. And the ablation study demonstrates the effectiveness of its main components on both the feature representation of cross-modal data and the accurate estimation of the crowd distribution.
He Li 0053, Weihang Kong
IEEE Trans. Ind. Informatics3
2023 LGP-MVS: combined local and global planar priors guidance for indoor multi-view stereo
Weihang Kong, Qingshan Xu 0001, Wanjuan Su, Wenbing Tao
Vis. Comput.1
2022 Learning the cross-modal discriminative feature representation for RGB-T crowd counting
He Li 0053, Weihang Kong
Knowl. Based Syst.3
2021 A cross-modal fusion based approach with scale-aware deep representation for RGB-D crowd counting and density estimation
He Li 0053, Weihang Kong
Expert Syst. Appl.3
2020 Effective crowd counting using multi-resolution context and image quality assessment-guided training
He Li 0053, Weihang Kong
Comput. Vis. Image Underst.2
2020 Object counting method based on dual attention network
abstract
The challenging problem that the authors solved in this study is to precisely estimate the number of objects in an image. Combining the spatial attention mechanism and pyramid structure, a novel atrous pyramid attention module is introduced to extract precise dense multi‐scale features for object counting. Also, a global attention feature module is designed to enhance the ability of the network to learn feature representation based on channel attention mechanism. Combining the proposed atrous pyramid attention module and global attention feature module, a novel object counting method based on a dual attention network is established in this study. The experiments on public vehicle counting dataset including TRANCOS and crowd counting dataset including Mall and Shanghitech_A datasets demonstrate the proposed method achieves competitive performance, and the ablation study verifies the structure rationality of the designed modules.
He Li 0053, Weihang Kong
IET Image Process.3
2020 A multi-context representation approach with multi-task learning for object counting
Weihang Kong, He Li 0053, Xi Zhang 0008, Gongda Zhao
Knowl. Based Syst.1
2020 Deeply scale aggregation network for object counting
He Li 0053, Weihang Kong
Knowl. Based Syst.2
2020 An attention-guided and prior-embedded approach with multi-task learning for shadow detection
He Li 0053, Weihang Kong, Weidong Ren
Knowl. Based Syst.3
2020 Bilateral counting network for single-image object counting
He Li 0053, Weihang Kong
Vis. Comput.3
2019 Crowd counting using a self-attention multi-scale cascaded network
abstract
Recent developments of crowd analysis and behaviour prediction have attracted much attention. Crowd counting, as the essential and challenging task in crowd analysis, is riddled with many issues, such as large scale variations, serious occlusion, and so on. In this study, a self‐attention‐based multi‐scale cascaded network called SAMC‐Net to estimate density map for crowd counting, especially for high congested scene, is proposed. The proposed SAMC‐Net consists of two components: a classification sub‐network for density estimation and an end‐to‐end multi‐scale convolution neural network for crowd counting. In order to reduce the negative effect of multi‐scale issue on crowd counting task, the main network is designed as a multi‐scale structure similar to U‐Net. In order to enhance the crowd feature representation, this study proposes a self‐attention‐based crowd feature extraction way and uses it in the proposed SAMC‐Net . Extensive experiments demonstrate the feasibility, effectiveness and robustness of the proposed SAMC‐Net .
He Li 0053, Weihang Kong
IET Comput. Vis.3
2019 An object counting network based on hierarchical context and feature fusion
He Li 0053, Weihang Kong, Xiaofang Niu
J. Vis. Commun. Image Represent.3