Shengnan Hu

dblp:193/7121 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Face, body and person analysis · 67% Graph learning · 33%
Network and information security
1 paper
Privacy and data protection · 100%
Databases, data mining, and information retrieval
1 paper
Spatial and temporal data management · 77% Query processing and optimization · 23%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Face, body and person analysis
human pose estimation
1.012026
Learning Topology-Aware Dynamic Associations for Robust Multi-Person Pose Estimation · AAAI 2026
Computer vision › Face, body and person analysis › human pose estimation
multi-person pose estimation
1.012026
Learning Topology-Aware Dynamic Associations for Robust Multi-Person Pose Estimation · AAAI 2026
Spatial and temporal data management › spatial query processing
range count query
0.912025
U-DPAP: Utility-aware Efficient Range Counting on Privacy-preserving Spatial Data Federation · Proc. ACM Manag. Data 2025
Privacy and data protection
differential privacy
0.912025
U-DPAP: Utility-aware Efficient Range Counting on Privacy-preserving Spatial Data Federation · Proc. ACM Manag. Data 2025
Privacy and data protection
privacy-preserving query processing
0.912025
U-DPAP: Utility-aware Efficient Range Counting on Privacy-preserving Spatial Data Federation · Proc. ACM Manag. Data 2025
Query processing and optimization
approximate query processing
0.312025
U-DPAP: Utility-aware Efficient Range Counting on Privacy-preserving Spatial Data Federation · Proc. ACM Manag. Data 2025

Methods — techniques the papers use, named apart from their topics

secure multiparty computation · 1.7grouping-based perturbation · 1.7differential privacy · 1.7topology-aware dynamic associations · 1.0
YearPublicationVenuePosition
2026 Learning Topology-Aware Dynamic Associations for Robust Multi-Person Pose Estimation
Shengnan Hu, Yandong Liu 0004, Jiangnan Liu, Yahong Chen
AAAI1
2026 CGAGF-Net: Cross-Guided and Adaptive Gated Fusion Network for Medical Image Segmentation
Xinyue Liao, Jinzhe Li, Shengnan Hu
ICIC (10)4
2026 FA-Det: Adaptive Multiscale Feature Fusion with Efficient Linear Attention for Real-Time Small Object Detection
Xinyue Liao, Jinzhe Li, Shengnan Hu
ICIC (10)4
2025 SlimPose: Lightweight Multi-Person Pose Estimation via Multi-Scale Structural Feature Fusion and Selective Attention
abstract
As a foundational technology for human-centered visual understanding, pose estimation has broad applications in human-computer interaction, behavior recognition, and video surveillance. However, existing methods often incur high computational costs, limiting their deployment on resource-constrained edge devices. In this paper, we propose SlimPose, a lightweight multi-person pose estimation framework that enhances the model’s ability to perceive human poses across scales by integrating multi-scale features with directional structural cues of the human body. We introduce a compact Selective Attentional Feature Gating module that adaptively emphasizes human regions while suppressing background features, thereby reducing unnecessary computational overhead. Additionally, we design a Deconvolutional Shift-Channel Mixer that enhances the model’s ability to infer occluded keypoints with minimal increase in parameters or computational cost. Extensive experiments on the COCO and CrowdPose datasets demonstrate that our approach achieves state-of-the-art accuracy among lightweight bottom-up methods, particularly excelling in crowded scenes and under resource-limited conditions.
Yandong Liu 0004, Jiangnan Liu, Shengnan Hu
SMC4
2025 BTDS: Blockchain-Enabled Trusted Vehicle Violation Detection by Self-Supervision
abstract
In recent years, the accelerated advancement of Internet of Vehicles (IoV) technology has significantly enhanced user experiences by providing intelligent services, such as multimedia entertainment and autonomous driving in vehicles. However, the enforcement of regulations concerning vehicle violations in IoV environments predominantly relies on manual methods, which are both expensive and challenging. Moreover, the inherent constraints in existing surveillance systems result in regulatory blind spots. Consequently, it is imperative to develop intelligent IoV-based surveillance mechanisms to improve the efficiency of detecting and rectifying violations. In this article, we propose a blockchain-based self-supervision model for vehicle violations that utilizes intervehicle reporting and voting mechanisms to enhance the detection rate of violations and reduce regulatory pressure. A forensic blockchain is introduced in the model to enable a review of the reporting results, which improves the security and reliability of the system. Additionally, more vehicles are incentivized to participate in the system through reputation-based rewards, punishments, and incentives. The system was deployed on the Hyperledger Fabric platform. Simulation experiments were conducted using Veins, SUMO, and OMNeT++. The experimental results verify the effectiveness of the model. The reporting and voting mechanism significantly inhibit violations, and the reward and reputation mechanism effectively promote the participation of vehicles.
Rui Zhu 0009, Shengnan Hu, Abdelsalam Helal, Junqiao Song, Jishu Wang, Yeting Chen
IEEE Internet Things J.2
2025 U-DPAP: Utility-aware Efficient Range Counting on Privacy-preserving Spatial Data Federation
abstract
Range counting is a fundamental operation in spatial data applications. There is a growing demand to facilitate this operation over a data federation, where spatial data are separately held by multiple data providers (a.k.a., data silos). Most existing data federation schemes employ Secure Multiparty Computation (SMC) to protect privacy, but this approach is computationally expensive and leads to high latency. Consequently, private data federations are often impractical for typical database workloads.This challenge highlights the need for a private data federation scheme capable of providing fast and accurate query responses while maintaining strong privacy. To address this issue, we propose U-DPAP, a utility-aware efficient privacy-preserving method. It is the first scheme to exclusively use differential privacy for privacy protection in spatial data federation, without employing SMC. Moreover, it combines approximate query processing to further enhance efficiency. Our experimental results indicate that a straightforward combination of the two techniques results in unacceptable impacts on data utility. Thus, we design two novel algorithms: one to make differential privacy practical by optimizing the privacy-utility trade-off, and another to address the efficiency-utility trade-off in approximate query processing. The grouping-based perturbation algorithm reduces noise by grouping similar data and applying noise to the groups. The representative data silos selection algorithm minimizes approximate error by selecting representative silos using the similarity between data silos. We rigorously prove the privacy guarantees of U-DPAP. Moreover, experimental results demonstrate that U-DPAP enhances data utility by an order of magnitude while maintaining high communication efficiency.
Yahong Chen, Xiaoyi Pang, Ben Niu 0001, Shengnan Hu
Proc. ACM Manag. Data6
2024 Multimodal Fusion Networks for Workload Modeling
abstract
The advent of low cost sensors for measuring gaze, heart rate, EEG, and galvanic skin response have made it feasible to cheaply collect physiological data from human operators. However, leveraging this data for machine learning problems requires a good multimodal fusion architecture. When dealing with multimodal features, uncovering the correlations between different modalities is as crucial as identifying effective unimodal features. This paper proposes a hybrid multimodal tensor fusion network that is effective at learning both unimodal and bimodal dynamics for cognitive workload modeling. Our architecture comprises two parts: (1) intra-modality for learning high-level representations of each signal modality (2) inter-modality for modeling bimodal interactions using a tensor fusion layer created from the Cartesian product of modality embeddings. We compare this architecture to the usage of a cross-modal transformer fusion module that learns an inter-modality embedding. Experimental results conducted on the HP Omnicept Cognitive Load Database (HPO-CLD) show that both techniques outperform the most commonly used techniques used for multimodal fusion of physio-logical data and that the cross-modal transformer fusion module is especially effective.
Shengnan Hu, Gita Reese Sukthankar
ICMLA1
2024 Strategic Analysis of the Parameter Servers and Participants in Federated Learning: An Evolutionary Game Perspective
abstract
Federated learning (FL) is a new decentralized deep learning paradigm developed for collaborative model training and solving the problem of data privacy and has received extensive attention from both the academic and business worlds. However, FL still faces challenges in encouraging participants to contribute private data and computational resources. Although many studies have applied game theory models to improve the incentive mechanism design of FL, they assume that the players are absolutely rational and that the game models are static. In this study, a mathematical model based on evolutionary game theory (EGT) is established to analyze the interaction between parameter servers and participants, considering that the participants are not completely rational in the long-term dynamic decision-making process. The evolutionarily stable status of the FL system and the strategies of the parameter servers and participants were analyzed under eight different scenarios. Based on the model analysis and results of the numerical experiments, managerial insights for maintaining a sustainable FL system are summarized.
Zhongliang Zhang 0001, Shengnan Hu
IEEE Trans. Comput. Soc. Syst.4
2023 The Potential of Vision-Language Models for Content Moderation of Children's Videos
abstract
Natural language supervision has been shown to be effective for zero-shot learning in many computer vision tasks, such as object detection and activity recognition. However, generating informative prompts can be challenging for more subtle tasks, such as video content moderation. This can be difficult, as there are many reasons why a video might be inappropriate, beyond violence and obscenity. For example, scammers may attempt to create junk content that is similar to popular educational videos but with no meaningful information. This paper evaluates the performance of several CLIP variations for content moderation of children's cartoons in both the supervised and zero-shot setting. We show that our proposed model (Vanilla CLIP with Projection Layer) outperforms previous work conducted on the Malicious or Benign (MOB) benchmark for video content moderation. This paper presents an in depth analysis of how context-specific language prompts affect content moderation performance. Our results indicate that it is important to include more context in content moderation prompts, particularly for cartoon videos as they are not well represented in the CLIP training data.
Syed Hammad Ahmed, Shengnan Hu, Gita Reese Sukthankar
ICMLA2
2023 LAMP: Leveraging Language Prompts for Multi-Person Pose Estimation
abstract
Human-centric visual understanding is an important desideratum for effective human-robot interaction. In order to navigate crowded public places, social robots must be able to interpret the activity of the surrounding humans. This paper addresses one key aspect of human-centric visual understanding, multi-person pose estimation. Achieving good performance on multi-person pose estimation in crowded scenes is difficult due to the challenges of occluded joints and instance separation. In order to tackle these challenges and overcome the limitations of image features in representing invisible body parts, we propose a novel prompt-based pose inference strategy called LAMP (Language Assisted Multi-person Pose estimation). By utilizing the text representations generated by a well-trained language model (CLIP), LAMP can facilitate the understanding of poses on the instance and joint levels, and learn more robust visual representations that are less susceptible to occlusion. This paper demonstrates that language-supervised training boosts the performance of single-stage multi-person pose estimation, and both instance-level and joint-level prompts are valuable for training. The code is available at https://github.com/shengnanh20/LAMP.
Shengnan Hu, Chen Chen 0001, Gita Reese Sukthankar
IROS1
2022 Predicting Team Performance with Spatial Temporal Graph Convolutional Networks
abstract
This paper presents a new approach for predicting team performance from the behavioral traces of a set of agents. This spatiotemporal forecasting problem is very relevant to sports analytics challenges such as coaching and opponent modeling. We demonstrate that our proposed model, Spatial Temporal Graph Convolutional Networks (ST-GCN), outperforms other classification techniques at predicting game score from a short segment of player movement and game features. Our proposed architecture uses a graph convolutional network to capture the spatial relationships between team members and Gated Recurrent Units to analyze dynamic motion information. An ablative evaluation was performed to demonstrate the contributions of different aspects of our architecture.
Shengnan Hu, Gita Reese Sukthankar
ICPR1
2020 CCA: Exploring the Possibility of Contextual Camouflage Attack on Object Detection
abstract
Deep neural network based object detection has become the cornerstone of many real-world applications. Along with this success comes concerns about its vulnerability to malicious attacks. To gain more insight into this issue, we propose a contextual camouflage attack (CCA for short) algorithm to influence the performance of object detectors. In this paper, we use an evolutionary search strategy and adversarial machine learning in interactions with a photo-realistic simulated environment to find camouflage patterns that are effective over a huge variety of object locations, camera poses, and lighting conditions. The proposed camouflages are validated effective to most of the state-of-the-art object detectors.
Shengnan Hu, Yang Zhang 0035, Sumit Laha, Ankit Kumar Sharma, Hassan Foroosh
ICPR1
2020 Near-Infrared Depth-Independent Image Dehazing using Haar Wavelets
abstract
We propose a fusion algorithm for haze removal that combines color information from an RGB image and edge information extracted from its corresponding NIR image using Haar wavelets. The proposed algorithm is based on the key observation that NIR edge features are more prominent in the hazy regions of the image than the RGB edge features in those same regions. To combine the color and edge information, we introduce a haze-weight map which proportionately distributes the color and edge information during the fusion process. Because NIR images are, intrinsically, nearly haze-free, our work makes no assumptions like existing works that rely on a scattering model and essentially designing a depth-independent method. This helps in minimizing artifacts and gives a more realistic sense to the restored haze-free image. Extensive experiments show that the proposed algorithm is both qualitatively and quantitatively better on several key metrics when compared to existing state-of-the-art methods.
Sumit Laha, Ankit Kumar Sharma, Shengnan Hu, Hassan Foroosh
ICPR3
2019 Learning Compact Appearance Representation for Video-Based Person Re-Identification
abstract
This paper presents a novel approach for video-based person re-identification using multiple convolutional neural networks (CNNs). Unlike the previous work, we intend to extract a compact yet discriminative appearance representation from several frames rather than the whole sequence. Specifically, given a video, the representative frames are selected based on the walking profile of consecutive frames. A multiple CNN architecture incorporated with feature pooling is proposed to learn and compile the features of the selected representative frames into a compact description about the pedestrian for identification. Experiments are conducted on benchmark data sets to demonstrate the superiority of the proposed method over existing person re-identification approaches.
Wei Zhang 0021, Shengnan Hu, Kan Liu 0001, Zhengjun Zha
IEEE Trans. Circuits Syst. Video Technol.2
2017 Exploiting patch-based correlation for ghost removal in exposure fusion
abstract
In this paper, we present a robust exposure fusion algorithm to tackle the problems of motion removal and detail preserving in dynamic scenes. With one exposure as reference, the motion appeared in the exposure stack can be detected by comparing the structural consistency, which is extracted by measuring the degree of linear correlation between the patches of the reference image and the other source images. Then, a stack of latent images with consistent contents can be synthesized after motion removal. For detail preserving, a contrast criterion is introduced to measure the exposedness and generate visibility maps of each latent image. Guided by the visibility maps, a tonemapped-like HDR image which is ghost-free and with all details preserved could be produced by seamlessly merging the latent images. Exposure fusion tests on various dynamic scenes demonstrate the superiority of the proposed method over existing state-of-the-art approaches.
Shengnan Hu, Wei Zhang 0021
ICME1
2017 Patch-Based correlation for deghosting in exposure fusion
Wei Zhang 0021, Shengnan Hu, Kan Liu 0001
Inf. Sci.2
2017 Motion-free exposure fusion based on inter-consistency and intra-consistency
Wei Zhang 0021, Shengnan Hu, Kan Liu 0001
Inf. Sci.2