Yongji Zhang

dblp:58/7715 · DBLP profile ↗
← Back
15ranked-venue papers
1as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Computer networks · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 HyperSign: Hierarchical Hypergraph-based Co-occurrence Modeling for Sign Language Recognition and Translation
abstract
Effectively capturing co-occurrence signals, such as hand shapes, facial expressions, and body postures, is critical for semantic understanding in sign language recognition (SLR) and translation (SLT). Although skeleton data offer greater efficiency and robustness than RGB inputs, existing methods typically rely on pairwise graph structures, limiting their ability to model complex high-order interactions across body regions. To address this limitation, we propose HyperSign, a hierarchical hypergraph neural network that systematically captures high-order co-occurrence patterns among diverse body parts. The Co-occurrence Graph Perception Module jointly learns relational structures via three complementary pathways: (1) traditional graph convolutions for modeling physical joint connections, (2) dynamic geometric hypergraphs constructed via k-nearest neighbors to encode local spatial patterns, and (3) soft hypergraphs generated by learnable prototypes to reveal latent semantic associations. To further enhance structural modeling and semantic consistency, a Meta-Part Hypergraph Fusion Module abstracts feature streams from the hands, face, and body into unified hypergraph nodes, while leveraging empirically derived co-occurrence priors to model high-order cross-part dependencies. Moreover, an uncertainty-aware collaborative distillation mechanism guides the model to focus on critical body regions. Extensive experiments on standard SLR and SLT benchmarks (e.g., PHOENIX-2014, PHOENIX-2014T, and CSL-Daily) demonstrate that HyperSign not only outperforms existing skeleton-based approaches in both speed and accuracy but also achieves competitive or superior results compared to several state-of-the-art RGB-based methods across multiple evaluation metrics.
Qianren Guo, Yuehang Wang, Yongji Zhang, Qi Chu 0010, Yu Jiang 0006
AAAI3
2026 Mixture-of-Experts for Hybrid Channel Prediction
Ningyan Guo, Yuanhao Cui, Haozhe Gu, Yongji Zhang, Zhiyong Feng 0001
IEEE Internet Things J.5
2026 Hyper-BTS: Brain tumor segmentation based on hypergraph guidance
Qianren Guo, Yuehang Wang, Yongji Zhang, Yuhua Hu, Yu Jiang 0006
Pattern Recognit.3
2026 Event-based facial expression recognition via large vision-language models
Siqi Li 0001, Yongji Zhang, Yue Gao 0002
Pattern Recognit.2
2026 TextSLR: Learning Text-Aware Representations for Sign Language Recognition
Qi Chu 0010, Yuehang Wang, Qianren Guo, Yongji Zhang, Yu Jiang 0006
IEEE Trans. Ind. Informatics5
2026 Event-Based Image Deblurring via Cross-Modal Interaction Fusion
abstract
Image deblurring aims to restore sharp images from degraded single-frame images, enabling industrial vision systems to operate reliably in complex industrial environments, such as high-speed and low-light conditions. Event cameras offer motion and edge information at high temporal resolution, allowing us to overcome the challenges red, green, and blue (RGB) cameras face with fast movements and lighting changes—common in industrial environments—by incorporating these cameras. However, considering the significant differences between different modalities, how to construct a multimodal vision system based on events and RGB images to jointly enhance the quality of multimodal features remains a topic worth exploring. In this article, we propose event-based deblurring network (EDNet), an event-based deblurring network that achieves efficient deblurring performance through cross-modal and cross-stage information interaction. We introduce the image-event interaction fusion module, which enhances the color information of images and the motion information of events through cross-modal attention, achieving higher quality feature fusion. To address the lack of low-level information during the image reconstruction stage, we designed the cross-stage information integration module, which effectively enhances the fidelity of image details and overall visual quality by integrating texture and fine-grained information from multiple feature layers. In addition, we provide an event-based deblurring test benchmark with authentic events and real blur in both indoor and outdoor scenes to evaluate the algorithm's performance and generalization capabilities. Experimental results demonstrate that EDNet outperforms other algorithms, restoring more detailed textures and exhibiting superior generalization capabilities, making it highly suitable for industrial applications.
Qi Chu 0010, Yuehang Wang, Yongji Zhang, Yu Jiang 0006
IEEE Trans. Ind. Informatics3
2026 H3Former: Hypergraph-Based Semantic-Aware Aggregation via Hyperbolic Hierarchical Contrastive Loss for Fine-Grained Visual Classification
abstract
Fine-Grained Visual Classification (FGVC) remains a challenging task due to subtle inter-class differences and large intra-class variations. Existing approaches typically rely on feature-selection mechanisms or region-proposal strategies to localize discriminative regions for semantic analysis. However, these methods often fail to capture discriminative cues comprehensively while introducing substantial category-agnostic redundancy. To address these limitations, we propose $\text {H}^{3}$ Former, a novel token-to-region framework that leverages high-order semantic relations to aggregate local fine-grained representations with structured region-level modeling. Specifically, we propose the Semantic-Aware Aggregation Module (SAAM), which exploits multi-scale contextual cues to dynamically construct a weighted hypergraph among tokens. By applying hypergraph convolution, SAAM captures high-order semantic dependencies and progressively aggregates token features into compact region-level representations. Furthermore, we introduce the Hyperbolic Hierarchical Contrastive Loss (HHCL), which enforces hierarchical semantic constraints in a non-Euclidean embedding space. The HHCL enhances inter-class separability and intra-class consistency while preserving the intrinsic hierarchical relationships among fine-grained categories. Comprehensive experiments conducted on four standard FGVC benchmarks validate the superiority of our $\text {H}^{3}$ Former framework. Code is available at https://github.com/xiaozhangfangyang/H3Former.
Yongji Zhang, Siqi Li 0001, Kuiyang Huang, Yue Gao 0002, Yu Jiang 0006
IEEE Trans. Image Process.1
2026 SkiTrack: An Aerial Skiing Benchmark for Human-Centric Object Tracking
abstract
Aerial skiing is a challenging human-centric sport characterized by rapid motion, large-scale variations, and frequent occlusions. Its extensive spatial range is typically captured by cameras or drones from multiple perspectives, resulting in frequent and complex viewpoint shifts. These challenges encompass nearly all difficulties inherent in human-centric tracking tasks. In this article, we introduce SkiTrack , the first dataset explicitly designed for tracking in aerial skiing. SkiTrack enhances the performance of existing tracking algorithms across a range of human-centric scenarios by providing precise annotations. We observe distinct characteristics in the tracked components, with the skis being rigid and low in visibility and the athlete’s body highly deformable but more visible. To leverage these differences, we propose a components decoupled loss that applies separate constraints to the tracking of the athlete and skis, thereby improving tracking accuracy in skiing scenes. Our experimental results validate the effectiveness of both the SkiTrack dataset and the proposed decoupled loss function, demonstrating consistent improvements in the performance of established models on human-centric tracking tasks. Data are available at https://github.com/xiaozhangfangyang/FineSkiing .
Yu Jiang 0006, Yongji Zhang, Siqi Li 0001, Yuehang Wang, Yue Gao 0002
ACM Trans. Multim. Comput. Commun. Appl.2
2025 A Real-World Animation Super-Resolution Benchmark With Color Degradation and Multi-Scale Multi-Frequency Alignment
abstract
Animation super-resolution (SR) aims to generate high-resolution (HR) animation frames from degraded low-resolution (LR) inputs, constituting an important task in real-world SR. Existing animation SR methods typically follow a photorealistic real-world SR computational paradigm. However, digital animation frames commonly suffer from compression and transmission-related degradation, distinct from degradations in camera-captured real-world images. In this paper, we introduce a novel real-world animation super-resolution benchmark designed explicitly for animation frames, named ADASR, featuring both 2D and modern 3D animation content to facilitate industry applications. Additionally, we propose a Color-Aware Animation Super-Resolution (CAASR) method. CAASR, for the first time, incorporates a color degradation simulation mechanism tailored for animations, addressing color banding, blocking, and color shift. Furthermore, we develop a multi-scale multi-frequency alignment mechanism to robustly extract degradation-invariant features. Extensive experiments conducted on both the existing AVC dataset and our newly constructed ADASR dataset demonstrate that our proposed CAASR achieves state-of-the-art performance in restoring HR frames for both 2D and 3D animations. Code and data are available at https://github.com/huangyang-666/CAASR.
Yu Jiang 0006, Yongji Zhang, Siqi Li 0001, Yuehang Wang, Yutong Yao, Yue Gao 0002
IEEE Trans. Image Process.2
2025 EvCSLR: Event-Guided Continuous Sign Language Recognition and Benchmark
abstract
Classical continuous sign language recognition (CSLR) suffers from some main challenges in real-world scenarios: accurate inter-frame movement trajectories may fail to be captured by traditional RGB cameras due to the motion blur, and valid information may be insufficient under low-illumination scenarios. In this paper, we for the first time leverage an event camera to overcome the above-mentioned challenges. Event cameras are bio-inspired vision sensors that could efficiently record high-speed sign language movements under low-illumination scenarios and capture human information while eliminating redundant background interference. To fully exploit the benefits of the event camera for CSLR, we propose a novel event-guided multi-modal CSLR framework, which could achieve significant performance under complex scenarios. Specifically, a time redundancy correction (TRCorr) module is proposed to rectify redundant information in the temporal sequences, directing the model to focus on distinctive features. A multi-modal cross-attention interaction (MCAI) module is proposed to facilitate information fusion between events and frame domains. Furthermore, we construct the first event-based CSLR dataset, namedEvCSLR, which will be released as the first event-based CSLR benchmark. Experimental results demonstrate that our proposed method achieves state-of-the-art performance on EvCSLR and PHOENIX-2014 T datasets.
Yu Jiang 0006, Yuehang Wang, Siqi Li 0001, Yongji Zhang, Qianren Guo, Qi Chu 0010, Yue Gao 0002
IEEE Trans. Multim.4
2024 A Feature-Enhanced and Adaptive Routing Framework for Fish School Detection on AUVs for Degraded Underwater Imaging Environments
abstract
Underwater biological detection technology holds considerable promise for delving into the abundance of marine species and resources. It can seamlessly integrate into Autonomous Underwater Vehicles (AUVs), providing robust support for the application and enhancement of the Internet of Underwater Things (IoUT) systems. Recently, data-driven artificial intelligence technologies have exhibited considerable potential in underwater object detection. Nonetheless, the degraded underwater imaging environments present challenges for imaging devices on AUVs or other underwater devices. In this paper, we propose an advanced fish school detection framework designed to significantly contribute to the IoUT systems. This framework enhances performance in challenging conditions like inadequate illumination, blurriness, and high fish population density. The proposed framework demonstrates the capability to promptly identify changes in underwater species, such as migration or aggregation of large fish schools. The timely recognition of such changes offers vital information for IoUT applications, particularly in the management of marine biological resources. Furthermore, our framework can be flexibly deployed on AUVs and in-situ observation stations. Additionally, degraded underwater conditions and high acquisition costs hinder the scalability of data-driven methods. Therefore, we create a novel dense fish school detection dataset named DUFish, expertly annotated with high-quality bounding boxes. The proposed detection framework showcases exceptional performance on DUFish, outperforming state-of-the-art target detection algorithms. This has the potential to augment the capabilities of biological recognition and applications within the IoUT systems.
Yu Jiang 0006, Yuehang Wang, Yongji Zhang, Qianren Guo, Minghao Zhao 0003, Hongde Qin
IEEE Internet Things J.3
2024 Efficient Vision Transformer With Token-Selective and Merging Strategies for Autonomous Underwater Vehicles
abstract
Underwater fine-grained classification technology is crucial for discerning subtle differences among marine life classes, playing a pivotal role in marine resource exploration and the discovery of new species. Autonomous underwater vehicles equipped with this technology can enhance their environmental interaction and perception, providing critical data for Internet of Underwater Things (IoUT) systems. However, popular vision transformer (ViT)-based methods encounter challenges in complex marine environments, particularly due to limited computational resources. In this article, we introduce an efficient ViT with token-selective and merging strategies (TSMVTs), which significantly improves underwater fine-grained classification performance while reducing the number of processed tokens. TSMVT can be flexibly integrated into various IoUT systems, promoting the discovery of new species and the sustainable development of marine ecology. First, we propose a dynamic token filtering mechanism that effectively retains important tokens, merges low-information tokens, and discards irrelevant background tokens, significantly reducing computational demands. Second, we propose the multihead attention weighting token-selective (MAWTS) module, which dynamically adjusts attention weights. MAWTS enables the network to focus on key features, such as fin shape, head structure, and body proportions, thereby improving classification accuracy. With a 30% reduction in tokens, TSMVT achieves superior precision in classifying marine species, enhancing its applicability on various underwater mobile platforms. Extensive experiments conducted on four marine and three terrestrial data sets demonstrate the outstanding accuracy and efficiency of the proposed TSMVT.
Yu Jiang 0006, Yongji Zhang, Yuehang Wang, Qianren Guo, Minghao Zhao 0003, Hongde Qin
IEEE Internet Things J.2
2024 RNVE: A Real Nighttime Vision Enhancement Benchmark and Dual-Stream Fusion Network
abstract
Images captured under challenging low-light conditions often suffer from myriad issues, including diminished contrast and obscured details, stemming from factors such as constrained lighting conditions or pervasive noise interference. Existing learning-based methods struggle in extreme low-light scenarios due to a lack of diverse paired datasets. In this letter, we meticulously curate a challenging real nighttime vision enhancement dataset called RNVE. RNVE comprises diverse data from various devices, including cameras and smartphones, available in both RGB and RAW formats. To enhance data diversity and enable comprehensive algorithm validation, we integrate synthetically generated low-light data, showcasing a spectrum of low-light effects. Additionally, we propose a low-light vision enhancement pipeline based on a dual-stream fusion network, proficiently improving the reconstruction quality of real nighttime scenes and restoring their authentic colors and contrast. Numerous experiments consistently demonstrate that the proposed pipeline excels in low-light enhancement and exhibits robust generalization capabilities across different datasets.
Yuehang Wang, Yongji Zhang, Qianren Guo, Minghao Zhao 0003, Yu Jiang 0006
IEEE Signal Process. Lett.2
2024 Event-Based Low-Illumination Image Enhancement
abstract
Event cameras are bio-inspired vision sensors with a high dynamic range (140 dB for event camerasvs.60 dB for traditional cameras) and can be used to tackle the image degradation problem under extremely low-illumination scenarios, which is still not well-explored yet. In this article, we propose a joint framework to compose the underexposed frames and event streams captured by the event camera to reconstruct clear images with detailed textures under almost dark conditions. A residual fusion module is proposed to reduce the domain gap between event streams and frames by using the residuals of both modalities. A multi-level reconstruction loss based on the variability of the contrast distribution is proposed to reduce the perceptual errors of the output image. In addition, we construct the first real-world low-illumination image enhancement dataset (mainly under 2 lux illumination scenes), named LIE, containing event streams and frames collected under indoor and outdoor low-light scenarios together with the ground truth clear images. Experimental results on our LIE dataset demonstrate that our proposed method could achieve significant improvements compared with existing methods.
Yu Jiang 0006, Yuehang Wang, Siqi Li 0001, Yongji Zhang, Minghao Zhao 0003, Yue Gao 0002
IEEE Trans. Multim.4
2023 Dynamic Spatial-temporal Hypergraph Convolutional Network for Skeleton-based Action Recognition
abstract
Skeleton-based action recognition relies on the extraction of spatial-temporal topological information. Hypergraphs can establish prior unnatural dependencies for the skeleton. However, the existing methods only focus on the construction of spatial topology and ignore the time-point dependence. This paper proposes a dynamic spatial-temporal hypergraph convolutional network (DST-HCN) to capture spatial-temporal information for skeleton-based action recognition. DST-HCN introduces a time-point hypergraph (TPH) to learn relationships at time points. With multiple spatial static hypergraphs and dynamic TPH, our network can learn more complete spatial-temporal features. In addition, we use the high-order information fusion module (HIF) to fuse spatial-temporal information synchronously. Extensive experiments on NTU RGB+D, NTU RGB+D 120, and NW-UCLA datasets show that our model achieves state-of-the-art, especially compared with hypergraph methods.
Shengqin Wang, Yongji Zhang, Minghao Zhao 0003, Yu Jiang 0006
ICME2