Yu Jiang 0006

dblp:21/4633-6 · DBLP profile ↗
← Back
25ranked-venue papers
9as first author
21since 2021 · last 2026
0000-0001-9025-3375ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Computer networks · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HyperSign: Hierarchical Hypergraph-based Co-occurrence Modeling for Sign Language Recognition and Translation
abstract
Effectively capturing co-occurrence signals, such as hand shapes, facial expressions, and body postures, is critical for semantic understanding in sign language recognition (SLR) and translation (SLT). Although skeleton data offer greater efficiency and robustness than RGB inputs, existing methods typically rely on pairwise graph structures, limiting their ability to model complex high-order interactions across body regions. To address this limitation, we propose HyperSign, a hierarchical hypergraph neural network that systematically captures high-order co-occurrence patterns among diverse body parts. The Co-occurrence Graph Perception Module jointly learns relational structures via three complementary pathways: (1) traditional graph convolutions for modeling physical joint connections, (2) dynamic geometric hypergraphs constructed via k-nearest neighbors to encode local spatial patterns, and (3) soft hypergraphs generated by learnable prototypes to reveal latent semantic associations. To further enhance structural modeling and semantic consistency, a Meta-Part Hypergraph Fusion Module abstracts feature streams from the hands, face, and body into unified hypergraph nodes, while leveraging empirically derived co-occurrence priors to model high-order cross-part dependencies. Moreover, an uncertainty-aware collaborative distillation mechanism guides the model to focus on critical body regions. Extensive experiments on standard SLR and SLT benchmarks (e.g., PHOENIX-2014, PHOENIX-2014T, and CSL-Daily) demonstrate that HyperSign not only outperforms existing skeleton-based approaches in both speed and accuracy but also achieves competitive or superior results compared to several state-of-the-art RGB-based methods across multiple evaluation metrics.
Qianren Guo, Yuehang Wang, Yongji Zhang, Qi Chu 0010, Yu Jiang 0006
AAAI6
2026 Hyper-BTS: Brain tumor segmentation based on hypergraph guidance
Qianren Guo, Yuehang Wang, Yongji Zhang, Yuhua Hu, Yu Jiang 0006
Pattern Recognit.6
2026 TextSLR: Learning Text-Aware Representations for Sign Language Recognition
Qi Chu 0010, Yuehang Wang, Qianren Guo, Yongji Zhang, Yu Jiang 0006
IEEE Trans. Ind. Informatics7
2026 Event-Based Image Deblurring via Cross-Modal Interaction Fusion
abstract
Image deblurring aims to restore sharp images from degraded single-frame images, enabling industrial vision systems to operate reliably in complex industrial environments, such as high-speed and low-light conditions. Event cameras offer motion and edge information at high temporal resolution, allowing us to overcome the challenges red, green, and blue (RGB) cameras face with fast movements and lighting changes—common in industrial environments—by incorporating these cameras. However, considering the significant differences between different modalities, how to construct a multimodal vision system based on events and RGB images to jointly enhance the quality of multimodal features remains a topic worth exploring. In this article, we propose event-based deblurring network (EDNet), an event-based deblurring network that achieves efficient deblurring performance through cross-modal and cross-stage information interaction. We introduce the image-event interaction fusion module, which enhances the color information of images and the motion information of events through cross-modal attention, achieving higher quality feature fusion. To address the lack of low-level information during the image reconstruction stage, we designed the cross-stage information integration module, which effectively enhances the fidelity of image details and overall visual quality by integrating texture and fine-grained information from multiple feature layers. In addition, we provide an event-based deblurring test benchmark with authentic events and real blur in both indoor and outdoor scenes to evaluate the algorithm's performance and generalization capabilities. Experimental results demonstrate that EDNet outperforms other algorithms, restoring more detailed textures and exhibiting superior generalization capabilities, making it highly suitable for industrial applications.
Qi Chu 0010, Yuehang Wang, Yongji Zhang, Yu Jiang 0006
IEEE Trans. Ind. Informatics4
2026 H3Former: Hypergraph-Based Semantic-Aware Aggregation via Hyperbolic Hierarchical Contrastive Loss for Fine-Grained Visual Classification
abstract
Fine-Grained Visual Classification (FGVC) remains a challenging task due to subtle inter-class differences and large intra-class variations. Existing approaches typically rely on feature-selection mechanisms or region-proposal strategies to localize discriminative regions for semantic analysis. However, these methods often fail to capture discriminative cues comprehensively while introducing substantial category-agnostic redundancy. To address these limitations, we propose $\text {H}^{3}$ Former, a novel token-to-region framework that leverages high-order semantic relations to aggregate local fine-grained representations with structured region-level modeling. Specifically, we propose the Semantic-Aware Aggregation Module (SAAM), which exploits multi-scale contextual cues to dynamically construct a weighted hypergraph among tokens. By applying hypergraph convolution, SAAM captures high-order semantic dependencies and progressively aggregates token features into compact region-level representations. Furthermore, we introduce the Hyperbolic Hierarchical Contrastive Loss (HHCL), which enforces hierarchical semantic constraints in a non-Euclidean embedding space. The HHCL enhances inter-class separability and intra-class consistency while preserving the intrinsic hierarchical relationships among fine-grained categories. Comprehensive experiments conducted on four standard FGVC benchmarks validate the superiority of our $\text {H}^{3}$ Former framework. Code is available at https://github.com/xiaozhangfangyang/H3Former.
Yongji Zhang, Siqi Li 0001, Kuiyang Huang, Yue Gao 0002, Yu Jiang 0006
IEEE Trans. Image Process.5
2026 SkiTrack: An Aerial Skiing Benchmark for Human-Centric Object Tracking
abstract
Aerial skiing is a challenging human-centric sport characterized by rapid motion, large-scale variations, and frequent occlusions. Its extensive spatial range is typically captured by cameras or drones from multiple perspectives, resulting in frequent and complex viewpoint shifts. These challenges encompass nearly all difficulties inherent in human-centric tracking tasks. In this article, we introduce SkiTrack , the first dataset explicitly designed for tracking in aerial skiing. SkiTrack enhances the performance of existing tracking algorithms across a range of human-centric scenarios by providing precise annotations. We observe distinct characteristics in the tracked components, with the skis being rigid and low in visibility and the athlete’s body highly deformable but more visible. To leverage these differences, we propose a components decoupled loss that applies separate constraints to the tracking of the athlete and skis, thereby improving tracking accuracy in skiing scenes. Our experimental results validate the effectiveness of both the SkiTrack dataset and the proposed decoupled loss function, demonstrating consistent improvements in the performance of established models on human-centric tracking tasks. Data are available at https://github.com/xiaozhangfangyang/FineSkiing .
Yu Jiang 0006, Yongji Zhang, Siqi Li 0001, Yuehang Wang, Yue Gao 0002
ACM Trans. Multim. Comput. Commun. Appl.1
2025 ERetinex: Event Camera Meets Retinex Theory for Low-Light Image Enhancement
abstract
Low-light image enhancement aims to restore the under-exposure image captured in dark scenarios. Under such scenarios, traditional frame-based cameras may fail to capture the structure and color information due to the exposure time limitation. Event cameras are bio-inspired vision sensors that respond to pixel-wise brightness changes asynchronously. Event cameras' high dynamic range is pivotal for visual perception in extreme low-light scenarios, surpassing traditional cameras and enabling applications in challenging dark environments. In this paper, inspired by the success of the retinex theory for traditional frame-based low-light image restoration, we introduce the first methods that combine the retinex theory with event cameras and propose a novel retinex-based lowlight image restoration framework named ERetinex. Among our contributions, the first is developing a new approach that leverages the high temporal resolution data from event cameras with traditional image information to estimate scene illumination accurately. This method outperforms traditional image-only techniques, especially in low-light environments, by providing more precise lighting information. Additionally, we propose an effective fusion strategy that combines the high dynamic range data from event cameras with the color information of traditional images to enhance image quality. Through this fusion, we can generate clearer and more detailrich images, maintaining the integrity of visual information even under extreme lighting conditions. The experimental results indicate that our proposed method outperforms state-of-theart (SOTA) methods, achieving a gain of 1.0613 dB in PSNR while reducing FLOPS by 84.28 %. The code is available at https://github.com/lodew920/ERetinex.
Xuejian Guo, Yuehang Wang, Siqi Li 0001, Yu Jiang 0006, Shaoyi Du, Yue Gao 0002
ICRA5
2025 A Real-World Animation Super-Resolution Benchmark With Color Degradation and Multi-Scale Multi-Frequency Alignment
abstract
Animation super-resolution (SR) aims to generate high-resolution (HR) animation frames from degraded low-resolution (LR) inputs, constituting an important task in real-world SR. Existing animation SR methods typically follow a photorealistic real-world SR computational paradigm. However, digital animation frames commonly suffer from compression and transmission-related degradation, distinct from degradations in camera-captured real-world images. In this paper, we introduce a novel real-world animation super-resolution benchmark designed explicitly for animation frames, named ADASR, featuring both 2D and modern 3D animation content to facilitate industry applications. Additionally, we propose a Color-Aware Animation Super-Resolution (CAASR) method. CAASR, for the first time, incorporates a color degradation simulation mechanism tailored for animations, addressing color banding, blocking, and color shift. Furthermore, we develop a multi-scale multi-frequency alignment mechanism to robustly extract degradation-invariant features. Extensive experiments conducted on both the existing AVC dataset and our newly constructed ADASR dataset demonstrate that our proposed CAASR achieves state-of-the-art performance in restoring HR frames for both 2D and 3D animations. Code and data are available at https://github.com/huangyang-666/CAASR.
Yu Jiang 0006, Yongji Zhang, Siqi Li 0001, Yuehang Wang, Yutong Yao, Yue Gao 0002
IEEE Trans. Image Process.1
2025 EvCSLR: Event-Guided Continuous Sign Language Recognition and Benchmark
abstract
Classical continuous sign language recognition (CSLR) suffers from some main challenges in real-world scenarios: accurate inter-frame movement trajectories may fail to be captured by traditional RGB cameras due to the motion blur, and valid information may be insufficient under low-illumination scenarios. In this paper, we for the first time leverage an event camera to overcome the above-mentioned challenges. Event cameras are bio-inspired vision sensors that could efficiently record high-speed sign language movements under low-illumination scenarios and capture human information while eliminating redundant background interference. To fully exploit the benefits of the event camera for CSLR, we propose a novel event-guided multi-modal CSLR framework, which could achieve significant performance under complex scenarios. Specifically, a time redundancy correction (TRCorr) module is proposed to rectify redundant information in the temporal sequences, directing the model to focus on distinctive features. A multi-modal cross-attention interaction (MCAI) module is proposed to facilitate information fusion between events and frame domains. Furthermore, we construct the first event-based CSLR dataset, namedEvCSLR, which will be released as the first event-based CSLR benchmark. Experimental results demonstrate that our proposed method achieves state-of-the-art performance on EvCSLR and PHOENIX-2014 T datasets.
Yu Jiang 0006, Yuehang Wang, Siqi Li 0001, Yongji Zhang, Qianren Guo, Qi Chu 0010, Yue Gao 0002
IEEE Trans. Multim.1
2025 Hierarchical Set-to-Set Representation for 3-D Cross-Modal Retrieval
abstract
Three-dimensional in-domain retrieval has recently achieved significant success, but 3-D cross-modal retrieval still faces problems and challenges. Existing methods only rely on a simple global feature (GF), which overlooks the local information of complex 3-D objects and the connections between similar local features across complex multimodal instances. To tackle this issue, we propose a hierarchical set-to-set representation (HSR) and a corresponding hierarchical similarity that incorporates global-to-global and local-to-local similarity metrics. Specifically, we employ feature extractors for each modality to learn both GFs and local feature sets. We then project these features into their respective common space and use bilinear pooling to generate compact-set features that maintain the invariant for set-to-set similarity measurement. To facilitate effective hierarchical similarity measurement, we design an operation to combine the GF and the compact-set feature to generate the hierarchical representation for 3-D cross-modal retrieval, which preserves hierarchical similarity measurement. To optimize the framework, we adopt the joint loss functions, including cross-modal center loss (CMCL), mean square loss, and cross-entropy loss, to reduce the cross-modal discrepancy for each instance and minimize the distances between the instances in the same category. Experimental results demonstrate that our method outperforms the state-of-the-art methods on the 3-D cross-modal retrieval task on both ModelNet10 and ModelNet40 datasets.
Yu Jiang 0006, Cong Hua, Yifan Feng 0001, Yue Gao 0002
IEEE Trans. Neural Networks Learn. Syst.1
2024 Structure-aware Residual-center Representation for Self-Supervised Open-set 3D Cross-modal Retrieval
abstract
Existing methods of 3D cross-modal retrieval heavily lean on category distribution priors within the training set, which diminishes their efficacy when tasked with unseen categories under open-set environments. To tackle this problem, we propose the Structure-Aware Residual-Center Representation (SRCR) framework for self-supervised open-set 3D cross-modal retrieval. To address the center deviation due to category distribution differences, we utilize the Residual-Center Embedding (RCE) for each object by nested auto-encoders, rather than directly mapping them to the modality or category centers. Besides, we perform the Hierarchical Structure Learning (HSL) approach to leverage the high-order correlations among objects for generalization, by constructing a heterogeneous hypergraph structure based on hierarchical inter-modality, intra-object, and implicit-category correlations. Extensive experiments and ablation studies on four benchmarks demonstrate the superiority of our proposed framework compared to state-of-the-art methods.
Yang Xu 0064, Yifan Feng 0001, Yu Jiang 0006
ICME3
2024 A Feature-Enhanced and Adaptive Routing Framework for Fish School Detection on AUVs for Degraded Underwater Imaging Environments
abstract
Underwater biological detection technology holds considerable promise for delving into the abundance of marine species and resources. It can seamlessly integrate into Autonomous Underwater Vehicles (AUVs), providing robust support for the application and enhancement of the Internet of Underwater Things (IoUT) systems. Recently, data-driven artificial intelligence technologies have exhibited considerable potential in underwater object detection. Nonetheless, the degraded underwater imaging environments present challenges for imaging devices on AUVs or other underwater devices. In this paper, we propose an advanced fish school detection framework designed to significantly contribute to the IoUT systems. This framework enhances performance in challenging conditions like inadequate illumination, blurriness, and high fish population density. The proposed framework demonstrates the capability to promptly identify changes in underwater species, such as migration or aggregation of large fish schools. The timely recognition of such changes offers vital information for IoUT applications, particularly in the management of marine biological resources. Furthermore, our framework can be flexibly deployed on AUVs and in-situ observation stations. Additionally, degraded underwater conditions and high acquisition costs hinder the scalability of data-driven methods. Therefore, we create a novel dense fish school detection dataset named DUFish, expertly annotated with high-quality bounding boxes. The proposed detection framework showcases exceptional performance on DUFish, outperforming state-of-the-art target detection algorithms. This has the potential to augment the capabilities of biological recognition and applications within the IoUT systems.
Yu Jiang 0006, Yuehang Wang, Yongji Zhang, Qianren Guo, Minghao Zhao 0003, Hongde Qin
IEEE Internet Things J.1
2024 Efficient Vision Transformer With Token-Selective and Merging Strategies for Autonomous Underwater Vehicles
abstract
Underwater fine-grained classification technology is crucial for discerning subtle differences among marine life classes, playing a pivotal role in marine resource exploration and the discovery of new species. Autonomous underwater vehicles equipped with this technology can enhance their environmental interaction and perception, providing critical data for Internet of Underwater Things (IoUT) systems. However, popular vision transformer (ViT)-based methods encounter challenges in complex marine environments, particularly due to limited computational resources. In this article, we introduce an efficient ViT with token-selective and merging strategies (TSMVTs), which significantly improves underwater fine-grained classification performance while reducing the number of processed tokens. TSMVT can be flexibly integrated into various IoUT systems, promoting the discovery of new species and the sustainable development of marine ecology. First, we propose a dynamic token filtering mechanism that effectively retains important tokens, merges low-information tokens, and discards irrelevant background tokens, significantly reducing computational demands. Second, we propose the multihead attention weighting token-selective (MAWTS) module, which dynamically adjusts attention weights. MAWTS enables the network to focus on key features, such as fin shape, head structure, and body proportions, thereby improving classification accuracy. With a 30% reduction in tokens, TSMVT achieves superior precision in classifying marine species, enhancing its applicability on various underwater mobile platforms. Extensive experiments conducted on four marine and three terrestrial data sets demonstrate the outstanding accuracy and efficiency of the proposed TSMVT.
Yu Jiang 0006, Yongji Zhang, Yuehang Wang, Qianren Guo, Minghao Zhao 0003, Hongde Qin
IEEE Internet Things J.1
2024 RNVE: A Real Nighttime Vision Enhancement Benchmark and Dual-Stream Fusion Network
abstract
Images captured under challenging low-light conditions often suffer from myriad issues, including diminished contrast and obscured details, stemming from factors such as constrained lighting conditions or pervasive noise interference. Existing learning-based methods struggle in extreme low-light scenarios due to a lack of diverse paired datasets. In this letter, we meticulously curate a challenging real nighttime vision enhancement dataset called RNVE. RNVE comprises diverse data from various devices, including cameras and smartphones, available in both RGB and RAW formats. To enhance data diversity and enable comprehensive algorithm validation, we integrate synthetically generated low-light data, showcasing a spectrum of low-light effects. Additionally, we propose a low-light vision enhancement pipeline based on a dual-stream fusion network, proficiently improving the reconstruction quality of real nighttime scenes and restoring their authentic colors and contrast. Numerous experiments consistently demonstrate that the proposed pipeline excels in low-light enhancement and exhibits robust generalization capabilities across different datasets.
Yuehang Wang, Yongji Zhang, Qianren Guo, Minghao Zhao 0003, Yu Jiang 0006
IEEE Signal Process. Lett.5
2024 Event-Based Low-Illumination Image Enhancement
abstract
Event cameras are bio-inspired vision sensors with a high dynamic range (140 dB for event camerasvs.60 dB for traditional cameras) and can be used to tackle the image degradation problem under extremely low-illumination scenarios, which is still not well-explored yet. In this article, we propose a joint framework to compose the underexposed frames and event streams captured by the event camera to reconstruct clear images with detailed textures under almost dark conditions. A residual fusion module is proposed to reduce the domain gap between event streams and frames by using the residuals of both modalities. A multi-level reconstruction loss based on the variability of the contrast distribution is proposed to reduce the perceptual errors of the output image. In addition, we construct the first real-world low-illumination image enhancement dataset (mainly under 2 lux illumination scenes), named LIE, containing event streams and frames collected under indoor and outdoor low-light scenarios together with the ground truth clear images. Experimental results on our LIE dataset demonstrate that our proposed method could achieve significant improvements compared with existing methods.
Yu Jiang 0006, Yuehang Wang, Siqi Li 0001, Yongji Zhang, Minghao Zhao 0003, Yue Gao 0002
IEEE Trans. Multim.1
2023 Dynamic Spatial-temporal Hypergraph Convolutional Network for Skeleton-based Action Recognition
abstract
Skeleton-based action recognition relies on the extraction of spatial-temporal topological information. Hypergraphs can establish prior unnatural dependencies for the skeleton. However, the existing methods only focus on the construction of spatial topology and ignore the time-point dependence. This paper proposes a dynamic spatial-temporal hypergraph convolutional network (DST-HCN) to capture spatial-temporal information for skeleton-based action recognition. DST-HCN introduces a time-point hypergraph (TPH) to learn relationships at time points. With multiple spatial static hypergraphs and dynamic TPH, our network can learn more complete spatial-temporal features. In addition, we use the high-order information fusion module (HIF) to fuse spatial-temporal information synchronously. Extensive experiments on NTU RGB+D, NTU RGB+D 120, and NW-UCLA datasets show that our model achieves state-of-the-art, especially compared with hypergraph methods.
Shengqin Wang, Yongji Zhang, Minghao Zhao 0003, Yu Jiang 0006
ICME5
2023 Dolphin movement direction recognition using contour-skeleton information
Mingzhu Xue, Xianglong Peng, Chong Wang 0015, Yu Jiang 0006
Multim. Tools Appl.5
2023 End-to-End Entity Detection with Proposer and Regressor
Xueru Wen, Changjiang Zhou, Haotian Tang, Luguang Liang, Yu Jiang 0006
Neural Process. Lett.6
2022 SHREC'22 track: Open-Set 3D Object Retrieval
Yifan Feng 0001, Yue Gao 0002, Xibin Zhao, Yandong Guo, Nihar Bagewadi, Nhat-Tan Bui, Hieu Dao, Shankar Gangisetty, Ripeng Guan, Xie Han 0001, Cong Hua, Chidambar Hunakunti, Yu Jiang 0006, Shichao Jiao, Yuqi Ke, Liqun Kuang, Anan Liu, Dinh-Huan Nguyen, Hai-Dang Nguyen, Weizhi Nie, Bang-Dang Pham, Karthik Raikar, Qingmei Tang, Minh-Triet Tran, Jialong Wan, Chenggang Yan 0001, Haoxuan You, Difei Zhu
Comput. Graph.13
2021 Event Stream Super-Resolution via Spatiotemporal Constraint Learning
abstract
Event cameras are bio-inspired sensors that respond to brightness changes asynchronously and output in the form of event streams instead of frame-based images. They own outstanding advantages compared with traditional cameras: higher temporal resolution, higher dynamic range, and lower power consumption. However, the spatial resolution of existing event cameras is insufficient and challenging to be enhanced at the hardware level while maintaining the asynchronous philosophy of circuit design. Therefore, it is imperative to explore the algorithm of event stream super-resolution, which is a non-trivial task due to the sparsity and strong spatio-temporal correlation of the events from an event camera. In this paper, we propose an end-to-end framework based on spiking neural network for event stream super-resolution, which can generate high-resolution (HR) event stream from the input low-resolution (LR) event stream. A spatiotemporal constraint learning mechanism is proposed to learn the spatial and temporal distributions of the event stream simultaneously. We validate our method on four large-scale datasets and the results show that our method achieves state-of-the-art performance. The satisfying results on two downstream applications, i.e. object classification and image reconstruction, further demonstrate the usability of our method. To prove the application potential of our method, we deploy it on a mobile platform. The high-quality HR event stream generated by our real-time system demonstrates the effectiveness and efficiency of our method.
Siqi Li 0001, Yutong Feng, Yu Jiang 0006, Changqing Zou, Yue Gao 0002
ICCV4
2021 Spirit Distillation: A Model Compression Method with Multi-domain Knowledge Transfer
Yu Jiang 0006, Minghao Zhao 0003, Chupeng Cui, Zongmin Yang, Xinhui Xue
KSEM2
2019 Deep Learning for Intelligent Train Driving with Augmented BLSTM
Jin Huang 0002, Siguang Huang, Yukun Hu, Yu Jiang 0006
PRICAI (2)5
2019 A parallel FP-growth algorithm on World Ocean Atlas data with multi-core CPU
Yu Jiang 0006, Minghao Zhao 0003, Chengquan Hu, Lili He 0002, Hongtao Bai, Jin Wang 0001
J. Supercomput.1
2018 Revised simplex algorithm for linear programming on GPUs with CUDA
Lili He 0002, Hongtao Bai, Yu Jiang 0006, Dantong Ouyang
Multim. Tools Appl.3
2016 A Low Power Balanced Security Control Protocol of WSN
Yu Jiang 0006, Jin Wang 0001, Lili He 0002, Yuanbo Xu, Hongtao Bai
QSHINE1