Yonghao Dang

dblp:234/0603 · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
19since 2021 · last 2026
0000-0002-1118-5587ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 2 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Systems, architecture and hardware · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Exploring Position Encoding Mechanism in Diffusion U-Net for Training-free High-resolution Image Generation
abstract
Denoising higher-resolution latents using a pre-trained U-Net often results in repetitive and disordered image patterns. In this work, we are motivated to reveal the intrinsic cause of such pattern disruption in high-resolution image generation. Through theoretical analysis and empirical studies, we reveal that the pre-trained U-Net fails to provide sufficient positional information for tokens at high-resolution. Specifically, 1) zero-padding serves as a critical mechanism for position encoding but lacks robustness across varying resolutions; and 2) tokens located farther from the feature map boundaries have increasing difficulty acquiring positional awareness, leading to pattern disruptions. Inspired by these findings, we propose a novel training-free approach for high-resolution generation, introducing a Progressive Boundary Complement (PBC) method. It creates dynamic virtual image boundaries inside the feature map to supplement position information at high resolution, enabling high-quality and rich-content high-resolution image synthesis. Extensive experiments show that our method significantly improves high-resolution image synthesis in terms of visual quality and content richness, achieving state-of-the-art performance.
Pu Cao, Yiyang Ma, Lu Yang 0006, Yonghao Dang, Jianqin Yin
AAAI5
2026 ActivityCLIP: Enhancing group activity recognition by mining complementary information from text to supplement image modality
Jianqin Yin, Yonghao Dang
Pattern Recognit.4
2026 Facial 3D Regional Structural Motion Representation Using Lightweight Point Cloud Networks for Micro-Expression Recognition
abstract
Human-computer interaction (HCI) relies on understanding and adapting to users' emotional states. Micro-expressions (MEs), a critical component of emotional perception, are characterized by their spontaneity, rapidity, subtlety, and difficulty to control. They often reveal an individual's true emotions. A comprehensive and detailed representation of motion is necessary to capture the nuances of facial dynamics effectively. Presently, motion representation methods are predominantly confined to 2D analysis within RGB images, overlooking the critical role of facial structure and its movements in conveying emotions. To overcome this limitation, we introduce an innovative facial motion representation that encompasses 3D facial structure, regionalized RGB and structural motion features. Furthermore, we segment the face into eight distinct regions, selecting only the most significant motion points to delineate the primary motion characteristics of each area. To model the interactions among crucial facial motion regions, we employ an advanced, lightweight point cloud and graph convolution network (Lite-Point-GCN). Comprehensive testing on the$\mathrm{CAS(ME)^{3}}$dataset, using leave-one-subject-out (LOSO), demonstrates that our method outperforms existing state-of-the-art methods.
Jianqin Yin, Yonghao Dang, Huaping Liu 0001
IEEE Trans. Affect. Comput.4
2026 ESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event Localization
abstract
Dense audio-visual event localization (DAVE) aims to identify event categories and locate the temporal boundaries in untrimmed videos. For such challenging task settings, most studies only employ audio-visual event semantic constraints on the final outputs, lacking progressive cross-modal semantic bridging in intermediate layers. This causes semantic gaps that hinder alignment between representations in audio and visual features, making it difficult to distinguish between event-related and irrelevant background content. Moreover, they rarely consider the correlations between events, which limits the model to infer co-occurring events among complex scenarios. In this paper, we incorporate multi-stage semantic guidance and multi-event relationship modeling, which respectively enable progressive semantic understanding of audio-visual events and adaptive extraction of event dependencies, thereby better focusing on event-related information. Specifically, our event-aware semantic guided network (ESG-Net) includes a early semantic interaction (ESI) module and a mixture of dependency experts (MoDE) module. ESI applys multi-stage semantic guidance to explicitly constrain the model in learning semantic information through multi-stage feature fusion and several classification loss functions, ensuring multi-stage understanding of event-related content. MoDE promotes the extraction of multi-event dependencies through multiple serial mixture of experts with adaptive weight allocation. Extensive experiments demonstrate that our method significantly surpasses the state-of-the-art methods, while greatly reducing parameters and computational load. Our code will be released on https://github.com/uchiha99999/ESG-Net.
Huilai Li, Yonghao Dang, Jianqin Yin
IEEE Trans. Multim.2
2026 Multi-Granularity Query Network With Adaptive Category Feature Embedding for Behavior Recognition
abstract
Behavior recognition is a highly challenging task, particularly in scenarios requiring unified recognition across both human and animal subjects. Most existing approaches primarily focus on single-species datasets or rely heavily on prior information such as species labels, positional annotations, or skeletal keypoints, which limits their applicability in real-world scenarios where species labels may be ambiguous or annotations are insufficient. To address these limitations, we propose a query-based Multi-Granularity Behavior Recognition Network that directly mines cross-species shared spatiotemporal behavior patterns from raw video inputs. Specifically, we design a Multi-Granularity Query module to effectively fuse fine-grained and coarse-grained features, thereby enhancing the model's capability in capturing spatiotemporal dynamics at different granularities. Additionally, we introduce a Category Query Decoder that leverages learnable category query vectors to achieve explicit behavior category modeling and mapping. Without relying on any extra annotations, the proposed method achieves unified recognition of multi-species and multi-category behaviors, setting a new state-of-the-art on the Animal Kingdom dataset and demonstrating strong generalization ability on the Charades dataset.
Nuoer Long, Yonghao Dang, Chengpeng Xiong, Shaobin Chen, Tao Tan 0002, Wei Ke 0001, Chan-Tong Lam, Jianqin Yin, Peter H. N. de With, Yue Sun 0001
IEEE Trans. Multim.2
2025 Quart-Online: Latency-Free Multimodal Large Language Model for Quadruped Robot Learning
abstract
This paper addresses the inherent inference latency challenges associated with deploying multimodal large language models (MLLM) in quadruped vision-language-action (QUAR-VLA) tasks. Our investigation reveals that conventional parameter reduction techniques ultimately impair the performance of the language foundation model during the action instruction tuning phase, making them unsuitable for this purpose. We introduce a novel latency-free quadruped MLLM model, dubbed QUARTOnline, designed to enhance inference efficiency without degrading the performance of the language foundation model. By incorporating Action Chunk Discretization (ACD), we compress the original action representation space, mapping continuous action values onto a smaller set of discrete representative vectors while preserving critical information. Subsequently, we fine-tune the MLLM to integrate vision, language, and compressed actions into a unified semantic space. Experimental results demonstrate that QUART-Online operates in tandem with the existing MLLM system, achieving real-time inference at 50 Hz in sync with the underlying controller frequency, significantly boosting the success rate across various tasks by 65 %. Our project page is https://quart-online.github.io.
Xinyang Tong, Pengxiang Ding, Yiguo Fan, Can Cui 0008, Han Zhao 0008, Hongyin Zhang 0001, Yonghao Dang, Siteng Huang, Shangke Lyu
ICRA10
2025 Towards Physically Realizable Adversarial Attacks in Embodied Vision Navigation
abstract
The significant advancements in embodied vision navigation have raised concerns about its susceptibility to adversarial attacks exploiting deep neural networks. Investigating the adversarial robustness of embodied vision navigation is crucial, especially given the threat of 3D physical attacks that could pose risks to human safety. However, existing attack methods for embodied vision navigation often lack physical feasibility due to challenges in transferring digital perturbations into the physical world. Moreover, current physical attacks for object detection struggle to achieve both multi-view effectiveness and visual naturalness in navigation scenarios. To address this, we propose a practical attack method for embodied navigation by attaching adversarial patches to objects, where both opacity and textures are learnable. Specifically, to ensure effectiveness across varying viewpoints, we employ a multi-view optimization strategy based on object-aware sampling, which optimizes the patch’s texture based on feedback from the vision-based perception model used in navigation. To make the patch inconspicuous to human observers, we introduce a two-stage opacity optimization mechanism, in which opacity is fine-tuned after texture optimization. Experimental results demonstrate that our adversarial patches decrease the navigation success rate by an average of 22.39%, outperforming previous methods in practicality, effectiveness, and naturalness. Code is available at: github.com/chen37058/Physical-Attacks-in-Embodied-Nav.
Jiawei Tu, Yonghao Dang, Jianqin Yin
IROS4
2025 3DWSNet: A Novel 3D Wavelet Spiking Neural Network for Event-based Action Recognition
abstract
In robotics applications, event cameras provide low-latency and high-dynamic-range sensing by asynchronously detecting brightness changes, making them well-suited for capturing fast motions and subtle cues in dynamic environments. However, most existing Spiking Neural Network (SNN)-based methods enhance spatial information by stacking multiple frames of events, while neglecting the explicit modeling of high-and low-frequency components in the event stream. To address this limitation, we proposes a 3D Wavelet Spiking Neural Network (3DWSNet), which integrates a 3D wavelet transform with a cascaded Wavelet Spiking Convolution (WSC) module as its core. Specifically, the 3D wavelet transform decomposes input data into eight frequency sub-bands across spatial and temporal dimensions, enabling the model to preserve fine-grained high-frequency details while enriching low-frequency motion representations. The cascaded WSC architecture further improves the extraction of multi-scale spatio-temporal features by integrating information from feature maps at different resolutions. Extensive experiments show that our 3DWSNet significantly outperforms SOTA SNN performances on the CIFAR-10, CIFAR-100, DVS128 Gesture, and CIFAR10-DVS datasets. The source code will be publicly released soon.
Junkang Fang, Yonghao Dang, Wending Zhao, Jianqin Yin
IROS2
2025 MaskSem: Semantic-Guided Masking for Learning 3D Hybrid High-Order Motion Representation
abstract
Human action recognition is a crucial task for intelligent robotics, particularly within the context of human-robot collaboration research. In self-supervised skeleton-based action recognition, the mask-based reconstruction paradigm learns the spatial structure and motion patterns of the skeleton by masking joints and reconstructing the target from unlabeled data. However, existing methods focus on a limited set of joints and low-order motion patterns, limiting the model’s ability to understand complex motion patterns. To address this issue, we introduce MaskSem, a novel semantic-guided masking method for learning 3D hybrid high-order motion representations. This novel framework leverages Grad-CAM based on relative motion to guide the masking of joints, which can be represented as the most semantically rich temporal orgions. The semantic-guided masking process can encourage the model to explore more discriminative features. Furthermore, we propose using hybrid high-order motion as the reconstruction target, enabling the model to learn multi-order motion patterns. Specifically, low-order motion velocity and high-order motion acceleration are used together as the reconstruction target. This approach offers a more comprehensive description of the dynamic motion process, enhancing the model’s understanding of motion patterns. Experiments on the NTU60, NTU120, and PKU-MMD datasets show that MaskSem, combined with a vanilla transformer, improves skeleton-based action recognition, making it more suitable for applications in human-robot interaction. The source code of our MaskSem is available at https://github.com/JayEason66/MaskSem.
Shaojie Zhang 0004, Yonghao Dang, Jianqin Yin
IROS3
2025 A generically Contrastive Spatiotemporal Representation Enhancement for 3D skeleton action recognition
Shaojie Zhang 0004, Jianqin Yin, Yonghao Dang
Pattern Recognit.3
2024 Full-Frequency Dynamic Convolution: A Physical Frequency-Dependent Convolution for Sound Event Detection
Haobo Yue, Da Mu, Yonghao Dang, Jianqin Yin, Jin Tang 0007
ICPR (20)4
2024 Physics-constrained attack against convolution-based human motion prediction
Chengxu Duan, Yonghao Dang, Jianqin Yin
Neurocomputing4
2024 DHRNet: A Dual-path Hierarchical Relation Network for multi-person pose estimation
Yonghao Dang, Jianqin Yin, Pengxiang Ding
Knowl. Based Syst.1
2024 Kinematics modeling network for video-based human pose estimation
Yonghao Dang, Jianqin Yin, Shaojie Zhang 0004, Jiping Liu
Pattern Recognit.1
2024 SiT-MLP: A Simple MLP With Point-Wise Topology Feature Learning for Skeleton-Based Action Recognition
abstract
Graph convolution networks (GCNs) have achieved remarkable performance in skeleton-based action recognition. However, previous GCN-based methods rely on elaborate human priors excessively and construct complex feature aggregation mechanisms, which limits the generalizability and effectiveness of networks. To solve these problems, we propose a novel Spatial Topology Gating Unit (STGU), an MLP-based variant without extra priors, to capture the co-occurrence topology features that encode the spatial dependency across all joints. In STGU, to learn the point-wise topology features, a new gate-based feature interaction mechanism is introduced to activate the features point-to-point by the attention map generated from the input sample. Based on the STGU, we propose the first MLP-based model, SiT-MLP, for skeleton-based action recognition in this work. Compared with previous methods on three large-scale datasets, SiT-MLP achieves competitive performance. In addition, SiT-MLP reduces the parameters significantly with favorable results. The code will be available at https://github.com/BUPTSJZhang/SiT-MLP.
Shaojie Zhang 0004, Jianqin Yin, Yonghao Dang, Jiajun Fu
IEEE Trans. Circuits Syst. Video Technol.3
2024 Leveraging the Video-Level Semantic Consistency of Event for Audio-Visual Event Localization
abstract
Audio-visual event (AVE) localization has attracted much attention in recent years. Most existing methods are often limited to independently encoding and classifying each video segment separated from the full video (which can be regarded as the segment-level representations of events). However, they ignore the semantic consistency of the event within the same full video (which can be considered as the video-level representations of events). In contrast to existing methods, we propose a novel video-level semantic consistency guidance network for the AVE localization task. Specifically, we propose an event semantic consistency modeling (ESCM) module to explore video-level semantic information for semantic consistency modeling. It consists of two components: a cross-modal event representation extractor (CERE) and an intra-modal semantic consistency enhancer (ISCE). CERE is proposed to obtain the event semantic information at the video level. Furthermore, ISCE takes video-level event semantics as prior knowledge to guide the model to focus on the semantic continuity of an event within each modality. Moreover, we propose a new negative pair filter loss to encourage the network to filter out the irrelevant segment pairs and a new smooth loss to further increase the gap between different categories of events in the weakly-supervised setting. We perform extensive experiments on the public AVE dataset and outperform the state-of-the-art methods in both fully- and weakly-supervised settings, thus verifying the effectiveness of our method.
Jianqin Yin, Yonghao Dang
IEEE Trans. Multim.3
2024 Learning Constrained Dynamic Correlations in Spatiotemporal Graphs for Motion Prediction
abstract
Human motion prediction is challenging due to the complex spatiotemporal feature modeling. Among all methods, graph convolution networks (GCNs) are extensively utilized because of their superiority in explicit connection modeling. Within a GCN, the graph correlation adjacency matrix drives feature aggregation, and thus, is the key to extracting predictive motion features. State-of-the-art methods decompose the spatiotemporal correlation into spatial correlations for each frame and temporal correlations for each joint. Directly parameterizing these correlations introduces redundant parameters to represent common relations shared by all frames and all joints. Besides, the spatiotemporal graph adjacency matrix is the same for different motion samples, and thus, cannot reflect samplewise correspondence variances. To overcome these two bottlenecks, we propose dynamic spatiotemporal decompose GC (DSTD-GC), which only takes 28.6% parameters of the state-of-the-art GC. The key of DSTD-GC is constrained dynamic correlation modeling, which explicitly parameterizes the common static constraints as a spatial/temporal vanilla adjacency matrix shared by all frames/joints and dynamically extracts correspondence variances for each frame/joint with an adjustment modeling function. For each sample, the common constrained adjacency matrices are fixed to represent generic motion patterns, while the extracted variances complete the matrices with specific pattern adjustments. Meanwhile, we mathematically reformulate GCs on spatiotemporal graphs into a unified form and find that DSTD-GC relaxes certain constraints of other GC, which contributes to a better representation capability. Moreover, by combining DSTD-GC with prior knowledge like body connection and temporal context, we propose a powerful spatiotemporal GCN called DSTD-GCN. On the Human3.6M, Carnegie Mellon University (CMU) Mocap, and 3D Poses in the Wild (3DPW) datasets, DSTD-GCN outperforms state-of-the-art methods by 3.9%-8.7% in prediction accuracy with 55.0%-96.9% fewer parameters. Codes are available at https://github.com/Jaakk0F/DSTD-GCN.
Jiajun Fu, Fuxing Yang, Yonghao Dang, Jianqin Yin
IEEE Trans. Neural Networks Learn. Syst.3
2022 Relation-Based Associative Joint Location for Human Pose Estimation in Videos
abstract
Video-based human pose estimation (VHPE) is a vital yet challenging task. While deep learning algorithms have made tremendous progress for the VHPE, lots of these approaches to this task implicitly model the long-range interaction between joints by expanding the receptive field of the convolution or designing a graph manually. Unlike prior methods, we design a lightweight and plug-and-play joint relation extractor (JRE) to explicitly and automatically model the associative relationship between joints. The JRE takes the pseudo heatmaps of joints as input and calculates their similarity. In this way, the JRE can flexibly learn the correlation between any two joints, allowing it to learn the rich spatial configuration of human poses. Furthermore, the JRE can infer invisible joints according to the correlation between joints, which is beneficial for locating occluded joints. Then, combined with temporal semantic continuity modeling, we propose a Relation-based Pose Semantics Transfer Network (RPSTN) for video-based human pose estimation. Specifically, to capture the temporal dynamics of poses, the pose semantic information of the current frame is transferred to the next with a joint relation guided pose semantics propagator (JRPSP). The JRPSP can transfer the pose semantic features from the non-occluded frame to the occluded frame. The proposed RPSTN achieves state-of-the-art or competitive results on the video-based Penn Action, Sub-JHMDB, PoseTrack2018, and HiEve datasets. Moreover, the proposed JRE improves the performance of backbones on the image-based COCO2017 dataset. Code is available at https://github.com/YHDang/pose-estimation.
Yonghao Dang, Jianqin Yin, Shaojie Zhang 0004
IEEE Trans. Image Process.1
2021 Energy-Based Periodicity Mining With Deep Features for Action Repetition Counting in Unconstrained Videos
abstract
Action repetition counting is to estimate the occurrence times of the repetitive motion in one action, which is a relatively new, significant, but challenging problem. To solve this problem, we propose a new method superior to the traditional ways in two aspects, without preprocessing and applicable for arbitrary periodicity actions. Without preprocessing, the proposed model makes our scheme convenient for real applications; processing the arbitrary periodicity action makes our model more suitable for the actual circumstance. In terms of methodology, firstly, we extract action features using ConvNets and then use Principal Component Analysis algorithm to generate the intuitive periodic information from the chaotic high-dimensional features; secondly, we propose an energy-based adaptive feature mode selection scheme to adaptively select proper deep feature mode according to the background of the video; thirdly,we construct the periodic waveform of the action based on the high-energy rules by filtering the irrelevant information. Finally, we detect the peaks to obtain the times of the action repetition. Our work features two-fold: 1) We give a significant insight that features extracted by ConvNets for action recognition can well model the self-similarity periodicity of the repetitive action. 2) A high-energy based periodicity mining rule using features from ConvNets is presented, which can process arbitrary actions without preprocessing. Experimental results show that our method achieves superior or comparable performance on the three benchmark datasets, i.e. YT_Segments, QUVA, and RARV.
Jianqin Yin, Yanchun Wu, Chaoran Zhu, Zijin Yin, Huaping Liu 0001, Yonghao Dang, Zhiyi Liu, Jun Liu 0007
IEEE Trans. Circuits Syst. Video Technol.6
2018 Estimating Cement Compressive Strength from Microstructure Images Using Broad Learning System
abstract
The microstructure images of cement are often used as the main data source for estimating compressive strength. They contain ample physical properties during the hydration process. Different gray values represent different substances in the grayscale image of cement. Deep learning algorithm based on microstructure images have been proposed to estimate cement compressive strength (CCS). However, there are a large number of parameters that need to be adjusted in deep structure. The high-efficiency system named broad learning system (BLS) is tried to use to estimate the cement compressive strength. The original cement microstructure images and the extracted features are used as input respectively, the connection weights can be obtained directly by calculating pseudo inverse matrix of feature matrix of microstructure image. If the structure is not sufficient to gain suitable result, BLS only calculates the pseudo inverse matrix of additional nodes to improve accuracy. The experiment shows that the broad learning structure (BLS) is an effective and efficient method on estimating cement compressive strength by contrasting with deep learning structure.
Yonghao Dang, Lin Wang 0004, Jianqin Yin, Xuehui Zhu, Zhiquan Feng, Jifeng Guo 0002
SMC1