Yunan Li 0001

dblp:130/6591-1 · DBLP profile ↗
← Back
43ranked-venue papers
13as first author
34since 2021 · last 2026
0000-0001-7316-4354ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 10 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 7 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PriorRG: Prior-Guided Contrastive Pre-training and Coarse-to-Fine Decoding for Chest X-ray Report Generation
abstract
Chest X-ray report generation aims to reduce radiologists' workload by automatically producing high-quality preliminary reports. A critical yet underexplored aspect of this task is the effective use of patient-specific prior knowledge---including clinical context (e.g., symptoms, medical history) and the most recent prior image---which radiologists routinely rely on for diagnostic reasoning. Most existing methods generate reports from single images, neglecting this essential prior information and thus failing to capture diagnostic intent or disease progression. To bridge this gap, we propose PriorRG, a novel chest X-ray report generation framework that emulates real-world clinical workflows via a two-stage training pipeline. In Stage 1, we introduce a prior-guided contrastive pre-training scheme that leverages clinical context to guide spatiotemporal feature extraction, allowing the model to align more closely with the intrinsic spatiotemporal semantics in radiology reports. In Stage 2, we present a prior-aware coarse-to-fine decoding for report generation that progressively integrates patient-specific prior knowledge with the vision encoder's hidden states. This decoding allows the model to align with diagnostic focus and track disease progression, thereby enhancing the clinical accuracy and fluency of the generated reports. Extensive experiments on MIMIC-CXR and MIMIC-ABN datasets demonstrate that PriorRG outperforms state-of-the-art methods, achieving a 3.6% BLEU-4 and 3.8% F1 score improvement on MIMIC-CXR, and a 5.9% BLEU-1 gain on MIMIC-ABN.
Kang Liu 0025, Zhuoqi Ma, Zikang Fang, Yunan Li 0001, Kun Xie 0011, Qiguang Miao
AAAI4
2026 Patient-specific multimodal learning with multi-view contrastive alignment for chest X-ray report generation
abstract
MOTIVATION: Radiology reports play a pivotal role in guiding treatment planning and enabling effective doctor-patient communication. However, their manual composition imposes a substantial workload on radiologists. Although automatic radiology report generation has emerged as a promising alternative, existing approaches predominantly rely on single-view chest X-rays and fail to adequately leverage patient-specific context, thereby limiting diagnostic accuracy. RESULTS: To address this challenge, we propose EVOKE, a novel chest X-ray report generation framework that incorporates multi-view contrastive learning and patient-specific knowledge. Specifically, we introduce a multi-view contrastive learning method that captures semantic correspondences both among multi-view radiographs within a study and between these radiographs and their associated report, thereby improving visual representation learning. We further present a knowledge-guided report generation module that integrates available patient-specific knowledge (i.e., indication, which includes symptom descriptions) to facilitate the generation of accurate and coherent radiology reports. To support research in multi-view report generation, we construct Multiview CXR and Two-view CXR datasets using publicly available sources. Our proposed EVOKE surpasses recent state-of-the-art methods across multiple datasets, achieving a 2.9% F1 RadGraph improvement on MIMIC-CXR, a 5.0% BLEU-1 improvement on MIMIC-ABN, a 1.5% BLEU-4 improvement on Multi-view CXR, and an 8.2% F1,mic-14 CheXbert improvement on Two-view CXR. AVAILABILITY: Code is publicly available at https://github.com/mk-runner/EVOKE, with an archived release available on Zenodo (doi:10.5281/zenodo.21000219).
Qiguang Miao, Kang Liu 0025, Zhuoqi Ma, Yunan Li 0001, Xiaolu Kang, Ruixuan Liu, Kun Xie 0011
Bioinform.4
2026 Holistic co-speech motion generation via cross-gated attention and cross-limb interaction
Zixiang Lu, Zhixiang Sheng, Zhitong He, Yunan Li 0001, Qiguang Miao
Expert Syst. Appl.5
2026 DESformer: Disentangling emotion and style for co-speech body-motion synthesis
Zhitong He, Zixiang Lu, Qiguang Miao, Yunan Li 0001, Kun Xie 0011, Bujia Tian
Neurocomputing5
2026 A two-stage sign language generation framework with self-supervised latent representation learning
Qiguang Miao, Guanwen Feng, Junwei Jing, Yilin Zhang 0007, Yunan Li 0001, Chi-Man Pun
Knowl. Based Syst.6
2026 No blind alignment but generation: A different view of continuous sign language recognition based on diffusion
Xi Geng, Yunan Li 0001, Zhuoqi Ma, Zixiang Lu, Qiguang Miao
Pattern Recognit.2
2026 LES-Talker: Fine-Grained Emotion Editing for Talking Head Generation in Linear Emotion Space
abstract
While existing one-shot talking head generation models have achieved progress in coarse-grained emotion editing, there is still a lack of fine-grained emotion editing models with high interpretability. We argue that for an approach to be considered fine-grained, it needs to provide clear definitions and sufficiently detailed differentiation. We present LES-Talker, a novel one-shot talking head generation model with high interpretability, to achieve fine-grained emotion editing across emotion types, emotion levels, and facial units. We propose a Linear Emotion Space (LES) definition based on Facial Action Units to characterize emotion transformations as vector transformations. We design the Cross-Dimension Attention Net (CDAN) to deeply mine the correlation between LES representation and 3D model representation. Through mining multiple relationships across different feature and structure dimensions, we enable LES representation to guide the controllable deformation of 3D model. In order to adapt the multimodal data with deviations to the LES and enhance visual quality, we utilize specialized network design and training strategies. Experiments show that our method provides high visual quality along with multilevel and inter pretable fine-grained emotion editing, outperforming mainstream methods. Project page: https://peterfanfan.github.io/LES-Talker/
Guanwen Feng, Zhihao Qian 0001, Yunan Li 0001, Qiguang Miao, Chi-Man Pun
IEEE Trans. Affect. Comput.3
2026 EmoSpeaker: One-Shot Fine-Grained Emotion-Controlled Talking Face Generation
abstract
Implementing fine-grained emotion control is crucial for emotion generation tasks because it enhances the expressive capability of the generative model, allowing it to accurately and comprehensively capture and express various nuanced emotional states, thereby improving the emotional quality and personalization of generated content. Generating fine-grained facial animations that accurately portray emotional expressions using only a portrait and an audio recording presents a challenge. In order to address this challenge, we propose a visual attribute-guided audio decoupler. This enables the obtention of content vectors solely related to the audio content, enhancing the stability of subsequent lip movement coefficient predictions. To achieve more precise emotional expression, we introduce a fine-grained emotion coefficient prediction module. Additionally, we propose an emotion intensity control method using a fine-grained emotion matrix. Through these, effective control over emotional expression in the generated videos and finer classification of emotion intensity are accomplished. Subsequently, a series of 3DMM coefficient generation networks are designed to predict 3D coefficients, followed by the utilization of a rendering network to generate the final video. Our experimental results demonstrate that our proposed method, EmoSpeaker, outperforms existing emotional talking face generation methods in terms of expression variation and lip synchronization. Project page:https://peterfanfan.github.io/EmoSpeaker/
Guanwen Feng, Yunan Li 0001, Chaoneng Li, Zhihao Qian 0001, Qiguang Miao, Chi-Man Pun
IEEE Trans. Multim.3
2025 Enhanced Contrastive Learning with Multi-view Longitudinal Data for Chest X-ray Report Generation
abstract
Automated radiology report generation offers an effective solution to alleviate radiologists’ workload. However, most existing methods focus primarily on single or fixed-view images to model current disease conditions, which limits diagnostic accuracy and overlooks disease progression. Although some approaches utilize longitudinal data to track disease progression, they still rely on single images to analyze current visits. To address these issues, we propose enhanced contrastive learning with Multi-view Longitudinal data to facilitate chest X-ray Report Generation, named MLRG. Specifically, we introduce a multi-view longitudinal contrastive learning method that integrates spatial information from current multi-view images and temporal information from longitudinal data. This method also utilizes the inherent spatiotemporal information of radiology reports to supervise the pre-training of visual and textual representations. Subsequently, we present a tokenized absence encoding technique to flexibly handle missing patient-specific prior knowledge, allowing the model to produce more accurate radiology reports based on available prior knowledge. Extensive experiments on MIMIC-CXR, MIMIC-ABN, and Two-view CXR datasets demonstrate that our MLRG outperforms recent state-of-the-art methods, achieving a 2.3% BLEU-4 improvement on MIMIC-CXR, a 5.5% F1 score improvement on MIMIC-ABN, and a 2.7% F1 RadGraph improvement on Two-view CXR.
Kang Liu 0025, Zhuoqi Ma, Xiaolu Kang, Yunan Li 0001, Kun Xie 0011, Zhicheng Jiao, Qiguang Miao
CVPR4
2025 Learning to Suppress Backgrounds and Bidirectionally Fuse Modalities for RGB-D Gesture Recognition
abstract
RGB-D video-based gesture recognition is a fundamental task in computer vision, yet it remains challenging due to the small size of hand regions and the presence of redundant background noise. Existing methods often fail to effectively suppress irrelevant background features and inadequately exploit the complementary nature of RGB and depth modalities, leading to suboptimal semantic alignment in feature fusion. To address these issues, we propose a novel end-to-end RGB-D gesture recognition framework that incorporates Spatiotemporal Background Suppression (STBS) and a Bidirectional Modality Fusion Adapter (BMFA). STBS leverages Vision Transformers to construct region-wise tokens and adaptively merges them based on responses scores, suppressing irrelevant background content while preserving fine-grained gesture-related spatiotemporal features. Meanwhile, BMFA enables deep, bidirectional interaction between RGB and depth features across encoder layers, enhancing cross-modal semantic consistency. Extensive experiments on three public RGB-D gesture datasets validate the effectiveness of our method. Experiments conducted on three public RGB-D gesture datasets validate the effectiveness of our approach and demonstrate significant improvements in recognition performance. Our code is available at https://github.com/caicai211/SB-BFM
Yunan Li 0001, Yulang Xu, Yilin Zhang 0007, Zixiang Lu, Qiguang Miao
ECAI2
2025 KAN-Face: Efficient Resource Usage and Precision Lip-Sync in Talking Head Generation
abstract
Despite significant progress in NeRF-based talking head generation, problems like poor lip synchronization and inefficient resource usage remain. To solve these, we propose KANFace, a lightweight framework. In preprocessing, we introduce a Lip-Sync Enhancement Module that uses Wav2Lip to extract high-resolution audio features and map them to an explicit intermediate representation, ensuring precise lip movement alignment with the speaker’s identity. These predicted lip-sync features are combined with fundamental audio-extracted lip features and injected into the rendering module to improve synchronization. For rendering, we introduce FastKAN to map spatial points to color values. As a variant of KAN, FastKAN’s sensitivity to 3D scenes and efficient structure enable precise, fast color prediction. Our framework reduces resource consumption while enhancing lip-sync accuracy and facial reconstruction, making it ideal for talking head generation tasks in resource-limited settings. Project: https://peterfanfan.github.io/KAN-Face/
Guanwen Feng, Zhihao Qian 0001, Yunan Li 0001, Qiguang Miao
ICASSP4
2025 Gaussian-Face: Talking Head Generation with Hybrid Density via 3D Gaussian Splatting
abstract
In recent years, audio-driven neural radiance field (NeRF)-based talking head generation techniques have achieved impressive results. However, these methods still have some limitations, such as unsynchronized lip movements and visual jitter. Recently, 3D Gaussian splatting has gradually replaced NeRF. Compared to NeRF, 3D Gaussian offers notable advantages, including higher efficiency and better reconstruction quality. Based on this, we propose Gaussian-Face, an audio-driven Gaussian-based facial avatar. With just a few minutes of monocular video and audio, a high-fidelity, driveable facial avatar can be reconstructed within hours. To achieve this, we first use FLAME to obtain the 3D representation of the face and then design a Lip Motion Translator to map audio to 3D lip representations. To model higher-quality facial details, we propose a hybrid density modeling method that balances rendering speed and quality, enabling our approach to render high-fidelity facial avatars at more than 160 FPS. Project page: https://peterfanfan.github.io/Gaussian-Face/
Guanwen Feng, Yilin Zhang 0007, Yunan Li 0001, Qiguang Miao
ICASSP3
2025 Sign-Mamba: Advanced Mamba-Based Sign Language Generation
abstract
In the field of sign language generation, Transformer-based models have been widely studied, but their quadratic computational complexity poses challenges. State Space Models (SSMs), like Mamba, offer a promising alternative with efficient long-range interaction modeling and linear complexity. In this study, we propose a two-stage generative framework for sign language generation, called Sign-Mamba, based on the state selection mechanism of SSM. In the first stage, we designed a Mamba-based encoder-decoder architecture, where the encoder captures the latent space representation and a symmetric decoder reconstructs it back into sign language skeletal points. In the second stage, the sign language latent space predicted by the Mamba-based latent predictor is used as a condition to guide the reconstruction network in generating more precise skeletal point sequences. We conducted comprehensive experiments on the PHOENIX and How2Sign datasets, and the results indicate that Sign-Mamba demonstrates competitive performance in sign language generation tasks. Project page: https://peterfanfan.github.io/Sign-Mamba/
Guanwen Feng, Yilin Zhang 0007, Yunan Li 0001, Qiguang Miao
ICASSP3
2025 RetouchDiffusion: Unsupervised Personalized Image Retouching via Diffusion Models
abstract
Image retouching aims to enhance the visual quality of images, but existing methods based on style-specific learning often lack the flexibility to cater to individual preferences. To address this limitation, we propose RetouchDiffusion, a personalized image retouching framework that utilizes a user-selected reference image as a guide and leverages the powerful representational capabilities of a pre-trained diffusion model to refine the retouching process. To enable more precise and adaptable adjustments, we introduce the Retouch Network, a dedicated retouching network that preprocesses brightness and tonality, providing controllable auxiliary guidance for the diffusion procedure. Experimental results demonstrate that our method produces high-quality, personalized retouched images more closely aligned with users’ aesthetic preferences across various scenarios. Our code is available at https://github.com/SuperOptimalZ/RetouchDiffusion.
Zhuoqi Ma, Zejun You, Yunan Li 0001, Qiguang Miao
ICME4
2025 Multi-scale information sharing and selection network with boundary attention for polyp segmentation
Xiaolu Kang, Zhuoqi Ma, Kang Liu 0025, Yunan Li 0001, Qiguang Miao
Eng. Appl. Artif. Intell.4
2025 One-shot handwriting imitation via self-supervised cross spatial transformer networks
Bocheng Zhao, Guanwen Feng, Wenxing Zhang, Yunan Li 0001, Qiguang Miao, Xiangzeng Liu, Ruyi Liu 0001
Neurocomputing4
2025 Modeling multi-scale uncertainty with evidence integration for reliable polyp segmentation
Xiaolu Kang, Zhuoqi Ma, Kang Liu 0025, Yunan Li 0001, Qiguang Miao
Neural Networks4
2025 Integrated multi-local and global dynamic perception structure for sign language recognition
Siyu Liang 0002, Yunan Li 0001, Huizhou Chen, Qiguang Miao
Pattern Anal. Appl.2
2025 CompNET: Boosting image recognition and writer identification via complementary neural network post-processing
abstract
In current classification tasks, an important method to improve accuracy is to pre-train the model using a large-scale domain-specific dataset. However, many tasks such as writer identification (writerID) lack suitable large-scale datasets in practical scenarios. To address this issue, this paper proposes a method that can improve prediction accuracy without relying on significant pre-training but leveraging the diversity of probability distributions predicted by multiple networks, and enhancing the top-1 accuracy through complementary post-processing. Specifically, top-k distributions are sampled from the multiple probability mass functions separately. When the distribution differences of top-k are maximized, the intersection other than the correct category can be narrowed down. Finally, the correct target with suboptimal probability can be rectified by the only intersection. Furthermore, our method has exhibited an intriguing trait during experimentation. Its prediction accuracy enhances concurrently with the incorporation of novel SOTA methods, ultimately surpassing the performance of these new methods.
Bocheng Zhao, Xuan Cao, Wenxing Zhang, Xujie Liu, Qiguang Miao, Yunan Li 0001
Pattern Recognit.6
2025 Boundary-Aware Sentence-Gloss Alignment With Semantic Similarity Measurement for Continuous Sign Language Recognition
abstract
Continuous sign language recognition (CSLR) plays a crucial role in facilitating communication between deaf and hearing individuals. A key aspect of achieving precise CSLR is the alignment of the video segment of each sign with its gloss, namely its corresponding text representation in natural language. However, the coarticulation phenomenon, where contextual dependencies between adjacent signs blur the boundaries of individual signs, poses a significant challenge to this task. In this paper, we propose a novel boundary-aware sentence-gloss alignment network for CSLR to address this challenge. Our network first designs a task-relevant boundary-aware similarity measurement, evaluating sign frames by both appearance and their recognition contribution, mitigating coarticulation-induced transition noise to restore precise boundaries. For enhanced alignment, we propose a hierarchical sentence-gloss alignment: coarse sentence-level alignment reduces cross-modal disparity, while fine-grained gloss-level alignment refines video-to-token mapping. Finally, an adaptive class-divergence loss sharpens gloss decoding by maximizing inter-class discrimination. Our proposed framework provides a simple and effective solution to mitigate the boundary ambiguity caused by coarticulation, optimizing continuous sign language recognition algorithms from a new perspective. Extensive experiments conducted on four public sign language recognition (SLR) datasets demonstrate that our proposed boundary-aware sentence-gloss alignment network learns precise alignments and achieves state-of-the-art performance.
Yunan Li 0001, Xi Geng, Zhuoqi Ma, Qiguang Miao, Chi-Man Pun
IEEE Trans. Circuits Syst. Video Technol.1
2025 Seeking a Hierarchical Prototype for Multimodal Gesture Recognition
abstract
Gesture recognition has drawn considerable attention from many researchers owing to its wide range of applications. Although significant progress has been made in this field, previous works always focus on how to distinguish between different gesture classes, ignoring the influence of inner-class divergence caused by gesture-irrelevant factors. Meanwhile, for multimodal gesture recognition, feature or score fusion in the final stage is a general choice to combine the information of different modalities. Consequently, the gesture-relevant features in different modalities may be redundant, whereas the complementarity of modalities is not exploited sufficiently. To handle these problems, we propose a hierarchical gesture prototype framework to highlight gesture-relevant features such as poses and motions in this article. This framework consists of a sample-level prototype and a modal-level prototype. The sample-level gesture prototype is established with the structure of a memory bank, which avoids the distraction of gesture-irrelevant factors in each sample, such as the illumination, background, and the performers' appearances. Then the modal-level prototype is obtained via a generative adversarial network (GAN)-based subnetwork, in which the modal-invariant features are extracted and pulled together. Meanwhile, the modal-specific attribute features are used to synthesize the feature of other modalities, and the circulation of modality information helps to leverage their complementarity. Extensive experiments on three widely used gesture datasets demonstrate that our method is effective to highlight gesture-relevant features and can outperform the state-of-the-art methods.
Yunan Li 0001, Tianyu Qi, Zhuoqi Ma, Dou Quan, Qiguang Miao
IEEE Trans. Neural Networks Learn. Syst.1
2024 Watching it in Dark: A Target-Aware Representation Learning Framework for High-Level Vision Tasks in Low Illumination
Yunan Li 0001, Shoude Li, Dou Quan, Chaoneng Li, Qiguang Miao
ECCV (75)1
2024 MGR-Dark: A Large Multimodal Video Dataset and RGB-IR Benchmark for Gesture Recognition in Darkness
Yunan Li 0001, Siyu Liang 0002, Huizhou Chen, Qiguang Miao
ACM Multimedia2
2024 CISampler: Correlated Information Guided Frame Sampling for Gesture Recognition in Video
Yunan Li 0001, Huizhou Chen, Siyu Liang 0002, Qiguang Miao
MMAsia2
2024 DiffTAD: Denoising diffusion probabilistic models for vehicle trajectory anomaly detection
Chaoneng Li, Guanwen Feng, Yunan Li 0001, Ruyi Liu 0001, Qiguang Miao, Liang Chang 0003
Knowl. Based Syst.3
2024 F3Net: Adaptive Frequency Feature Filtering Network for Multimodal Remote Sensing Image Registration
abstract
Multimodal remote sensing image registration is crucial for multimodal information fusion and applications. The significant nonlinear appearance difference between multimodal images caused by the various imaging mechanisms dramatically increases the challenge of image registration. This article proposes an adaptive frequency feature filtering network (F3Net) for cross-modal remote sensing image registration. On the one hand, F3Net explicitly explores the useful frequency components across modal images based on multilevel deep features. On the other hand, F3Net can take advantage of the nonlocal receptive fields by frequency modulation for feature learning and boosting image registration performances. F3Net inserts frequency feature filtering (F3) modules in multilevel deep features. Specifically, F3Net first performs the fast Fourier transform (FFT) for deep features. Then, F3Net designs a frequency attention (FA) module to adaptive enhance the shared and discriminative frequency features between multimodal images while suppressing the frequency components that hinder the cross-modal image registration. In addition, F3Net adopts multiscale frequency filtering fusion to facilitate discriminative feature learning, including global frequency feature filtering (GF3) based on the global image spectrum and local frequency feature filtering (LF3) based on the spectrum of stacked image regions. Experimental results on many remote sensing images have demonstrated the efficiency of the F3Net on multimodal image registration.
Dou Quan, Shuang Wang 0001, Yunan Li 0001, Bo Ren 0001, Mengte Kang, Jocelyn Chanussot, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2024 A Semi-Supervised Underexposed Image Enhancement Network With Supervised Context Attention and Multi-Exposure Fusion
abstract
Recently, image enhancement approaches yield impressive progress. However, most methods are still based supervised-learning, which requires plenty of paired data. Meanwhile, owing to the complex illumination condition in a real-world scenario, those methods trained on synthetic images cannot restore details in extremely dark or bright areas and lead to exposure errors. The traditional losses that deem all pixels the same in training also produce blurry edges in the result. To handle these problems, in this article, we present an effective semi-supervised framework for severely underexposed image enhancement. Our network consists of a supervised and an unsupervised branch, which shares weights and can make full use of paired data and plenty of unpaired data. Meanwhile, a multi-exposure fusion module is designed to adaptively fuse the corrected images to address the low contrast and color bias issues occurring in some extreme situations. Moreover, we propose a supervised context attention module to better use the edge information as supervision to recover fine image details. Extensive experiments have proved that the proposed method outperforms state-of-the-art approaches in enhancing exposure images.
Xiaolong Fu, Yunan Li 0001, Kaibin Miao, Xiangzeng Liu, Bocheng Zhao, Qiguang Miao
IEEE Trans. Multim.3
2023 Learning Robust Representations with Information Bottleneck and Memory Network for RGB-D-based Gesture Recognition
abstract
Although previous RGB-D-based gesture recognition methods have shown promising performance, researchers often overlook the interference of task-irrelevant cues like illumination and background. These unnecessary factors are learned together with the predictive ones by the network and hinder accurate recognition. In this paper, we propose a convenient and analytical framework to learn a robust feature representation that is impervious to gesture-irrelevant factors. Based on the Information Bottleneck theory, two rules of Sufficiency and Compactness are derived to develop a new information-theoretic loss function, which cultivates a more sufficient and compact representation from the feature encoding and mitigates the impact of gesture-irrelevant information. To highlight the predictive information, we further integrate a memory network. Using our proposed content-based and contextual memory addressing scheme, we weaken the nuisances while preserving the task-relevant information, providing guidance for refining the feature representation. Experiments conducted on three public datasets demonstrate that our approach leads to a better feature representation and achieves better performance than state-of-the-art methods. The code of our method is available at: https://github.com/Carpumpkin/InBoMem.
Yunan Li 0001, Huizhou Chen, Guanwen Feng, Qiguang Miao
ICCV1
2023 Semi-Supervised Domain Alignment Learning for Single Image Dehazing
abstract
Convolutional neural networks (CNNs) have attracted much research attention and achieved great improvements in single-image dehazing. However, previous learning-based dehazing methods are mainly trained on synthetic data, which greatly degrades their generalization capability on natural hazy images. To address this issue, this article proposes a semi-supervised learning approach for single-image dehazing, where both synthetic and realistic images are leveraged during training. Considering the situation that it is hard to obtain the realistic pairs of hazy and haze-free images, how to utilize the realistic data is not a trivial work. In this article, a domain alignment module is introduced to narrow the distribution distance between synthetic data and realistic hazy images in a latent feature space. Meanwhile, a haze-aware attention module is designed to describe haze densities of different regions in the image, thus adaptively responds for different hazy areas. Furthermore, the dark channel prior is introduced to the framework to improve the quality of the unsupervised learning results by considering the statistical characters of haze-free images. Such a semi-supervised design can significantly address the domain shift issue between the synthetic and realistic data, and improve generalization performance in the real world. Experiments indicate that the proposed method obtains state-of-the-art performance on both public synthetic and realistic hazy images with better visual results.
Yunan Li 0001, He Zhang 0004, Shifeng Chen
IEEE Trans. Cybern.2
2023 Image Hazing and Dehazing: From the Viewpoint of Two-Way Image Translation With a Weakly Supervised Framework
abstract
Image dehazing is an important task since it is the prerequisite for many downstream high-level computer vision tasks. Previous dehazing methods depend on either the hand-designed priors/assumptions or supervised learning with plenty of data, which are not easy to implement in practice. Meanwhile, synthesizing hazy images is also significant in many scenes like multi-weather image generation. In this paper, we change the viewpoint of this task to image translation and develop a weakly supervised framework to achieve it. Instead of simply considering the hazy image as the source domain and the haze-free image as the target domain for translation, we design a feature representation scheme that generates a domain indicator, and embed it into the decoder to achieve both hazing and dehazing within one network. This design significantly reduces the complexity of network and can be more easily extended to multi-domain translation tasks than the previous methods, which need one pair of generator-discriminator for each direction of the translation. Meanwhile, aiming at solving the haze-relevant task, we design a haze attention module, which takes the local entropy map as the input. Unlike the previous weakly supervised dehazing methods, our approach only requires unpaired hazy and haze-free images rather than any intermediate supervising data like the transmission map or atmospheric light defined in the atmospheric scattering model. Experimental results on synthetic datasets show our method can achieve competitive results when compared with the state-of-the-art methods and yield more appealing dehazing and hazing results on real-world images.
Yunan Li 0001, Huizhou Chen, Qiguang Miao, Siyu Liang 0002, Zhuoqi Ma, Bocheng Zhao
IEEE Trans. Multim.1
2022 Fidelity Evaluation of Virtual Traffic Based on Anomalous Trajectory Detection
abstract
Measuring the fidelity of synthesized virtual traffic has become an important and fundamental concern for evaluating the performance of different traffic simulation techniques and applications of autonomous vehicle testing. In this work, we propose a novel method to evaluate the fidelity of any trajectory data from the perspective of anomalous trajectory detection. First, given the trajectory data to be evaluated as input, the method learns spatio-temporal traffic features and reconstructs the input trajectory through a Long Short-Term Memory (LSTM)-based autoencoder architecture. Then, the anomalous trajectories are detected by comparing the reconstructed trajectories and the input ones using the reconstruction error as the benchmark. Our method can detect eight different kinds of anomalous trajectory in terms of changes in velocity and moving direction. In order to evaluate the fidelity of the input trajectory, we design a perceptual evaluation on virtual traffic fidelity and derive a mapping from the reconstruction error to the evaluation score. We demonstrated the effectiveness and robustness of our metric through many experiments on real-world and synthetic trajectory data containing different types of motion anomalies.
Chaoneng Li, Qianwen Chao, Guanwen Feng, Qiongyan Wang, Yunan Li 0001, Qiguang Miao
IROS6
2022 ChaLearn Looking at People: IsoGD and ConGD Large-Scale RGB-D Gesture Recognition
abstract
The ChaLearn large-scale gesture recognition challenge has run twice in two workshops in conjunction with the International Conference on Pattern Recognition (ICPR) 2016 and International Conference on Computer Vision (ICCV) 2017, attracting more than 200 teams around the world. This challenge has two tracks, focusing on isolated and continuous gesture recognition, respectively. It describes the creation of both benchmark datasets and analyzes the advances in large-scale gesture recognition based on these two datasets. In this article, we discuss the challenges of collecting large-scale ground-truth annotations of gesture recognition and provide a detailed analysis of the current methods for large-scale isolated and continuous gesture recognition. In addition to the recognition rate and mean Jaccard index (MJI) as evaluation metrics used in previous challenges, we introduce the corrected segmentation rate (CSR) metric to evaluate the performance of temporal segmentation for continuous gesture recognition. Furthermore, we propose a bidirectional long short-term memory (Bi-LSTM) method, determining video division points based on skeleton points. Experiments show that the proposed Bi-LSTM outperforms state-of-the-art methods with an absolute improvement of 8.1% (from 0.8917 to 0.9639) of CSR.
Jun Wan 0001, Chi Lin 0002, Longyin Wen, Yunan Li 0001, Qiguang Miao, Sergio Escalera, Gholamreza Anbarjafari, Isabelle Guyon, Guodong Guo, Stan Z. Li
IEEE Trans. Cybern.4
2021 Regional Attention with Architecture-Rebuilt 3D Network for RGB-D Gesture Recognition
abstract
Human gesture recognition has drawn much attention in the area of computer vision. However, the performance of gesture recognition is always influenced by some gesture-irrelevant factors like the background and the clothes of performers. Therefore, focusing on the regions of hand/arm is important to the gesture recognition. Meanwhile, a more adaptive architecture-searched network structure can also perform better than the block-fixed ones like ResNet since it increases the diversity of features in different stages of the network better. In this paper, we propose a regional attention with architecture-rebuilt 3D network (RAAR3DNet) for gesture recognition. We replace the fixed Inception modules with the automatically rebuilt structure through the network via Neural Architecture Search (NAS), owing to the different shape and representation ability of features in the early, middle, and late stage of the network. It enables the network to capture different levels of feature representations at different layers more adaptively. Meanwhile, we also design a stackable regional attention module called Dynamic-Static Attention (DSA), which derives a Gaussian guidance heatmap and dynamic motion map to highlight the hand/arm regions and the motion information in the spatial and temporal domains, respectively. Extensive experiments on two recent large-scale RGB-D gesture datasets validate the effectiveness of the proposed method and show it outperforms state-of-the-art methods. The codes of our method are available at: https://github.com/zhoubenjia/RAAR3DNet.
Benjia Zhou, Yunan Li 0001, Jun Wan 0001
AAAI2
2021 Review of dynamic gesture recognition
abstract
In recent years, gesture recognition has been widely used in the fields of intelligent driving, virtual reality, and human-computer interaction. With the development of artificial intelligence, deep learning has achieved remarkable success in computer vision. To help researchers better understanding the development status of gesture recognition in video, this article provides a detailed survey of the latest developments in gesture recognition technology for videos based on deep learning. The reviewed methods are broadly categorized into three groups based on the type of neural networks used for recognition: twostream convolutional neural networks, 3D convolutional neural networks, and Long-short Term Memory (LSTM) networks. In this review, we discuss the advantages and limitations of existing technologies, focusing on the feature extraction method of the spatiotemporal structure information in a video sequence, and consider future research directions.
Yunan Li 0001, Xiaolong Fu, Kaibin Miao, Qiguang Miao
Virtual Real. Intell. Hardw.2
2020 CR-Net: A Deep Classification-Regression Network for Multimodal Apparent Personality Analysis
Yunan Li 0001, Jun Wan 0001, Qiguang Miao, Sergio Escalera, Huijuan Fang, Huizhou Chen, Xiangda Qi, Guodong Guo
Int. J. Comput. Vis.1
2019 LAP-Net: Level-Aware Progressive Network for Image Dehazing
abstract
In this paper, we propose a level-aware progressive network (LAP-Net) for single image dehazing. Unlike previous multi-stage algorithms that generally learn in a coarse-to-fine fashion, each stage of LAP-Net learns different levels of haze with different supervision. Then the network can progressively learn the gradually aggravating haze. With this design, each stage can focus on a region with specific haze level and restore clear details. To effectively fuse the results of varying haze levels at different stages, we develop an adaptive integration strategy to yield the final dehazed image. This strategy is achieved by a hierarchical integration scheme, which is in cooperation with the memory network and the domain knowledge of dehazing to highlight the best-restored regions of each stage. Extensive experiments on both real-world images and two dehazing benchmarks validate the effectiveness of our proposed method.
Yunan Li 0001, Qiguang Miao, Wanli Ouyang, Zhenxin Ma, Huijuan Fang, Chao Dong 0005, Yi-Ning Quan
ICCV1
2019 Multiscale road centerlines extraction from high-resolution aerial imagery
Ruyi Liu 0001, Qiguang Miao, Jianfeng Song, Yi-Ning Quan, Yunan Li 0001, Pengfei Xu 0003
Neurocomputing5
2019 A spatiotemporal attention-based ResC3D model for large-scale gesture recognition
Yunan Li 0001, Qiguang Miao, Xiangda Qi, Zhenxin Ma, Wanli Ouyang
Mach. Vis. Appl.1
2019 Large-scale gesture recognition with a fusion of RGB-D data based on optical flow and the C3D model
Yunan Li 0001, Qiguang Miao, Kuan Tian, Xin Xu 0001, Zhenxin Ma, Jianfeng Song
Pattern Recognit. Lett.1
2018 A multi-scale fusion scheme based on haze-relevant features for single image dehazing
Yunan Li 0001, Qiguang Miao, Ruyi Liu 0001, Jianfeng Song, Yi-Ning Quan, Yuhui Huang
Neurocomputing1
2018 Large-Scale Gesture Recognition With a Fusion of RGB-D Data Based on Saliency Theory and C3D Model
abstract
Gesture recognition has raised wide attention in computer vision owing to its many applications. However, the task of video-based large-scale gesture recognition yet faces many challenges, since many gesture-irrelevant factors like the background may disturb the recognition accuracy. To better recognize gestures with large-scale videos, we propose a method based on RGB-D data in this paper, where the “RGB-D” means RGB and depth data captured simultaneously by specific devices like Kinect. To learn gesture details better, we first use an adaptive frame unification strategy to unify the frame number of inputs, and then the RGB and depth data are sent to the C3D model to extract spatiotemporal features, respectively. In order to alleviate the interference of gesture-irrelevant factors, the saliency theory is also employed to generate auxiliary data. Next the features of these data are combined to boost the performance, which can also avoid unreasonable synthetic data, since the dimension of C3D features is uniform. Finally the performances of several classifiers are tested and the best one of SVM classifier is selected to output the ultimate accuracy. Our approach achieves 52.04% and 59.43% accuracy on the validation and testing subset of the Chalearn LAP IsoGD, respectively, both of which outperform our results in the chalearn LAP Large-scale Gesture Recognition Challenge as reported in ICPR 2016.
Yunan Li 0001, Qiguang Miao, Kuan Tian, Xin Xu 0001, Jianfeng Song
IEEE Trans. Circuits Syst. Video Technol.1
2016 Large-scale gesture recognition with a fusion of RGB-D data based on the C3D model
abstract
The gesture recognition has raised attention in computer vision owing to its many applications. However, video-based large-scale gesture recognition still faces many challenges, since many factors like background may disturb the accuracy. To achieve gesture recognition with large-scale videos, we propose a method based on RGB-D data. To learn gesture details better, the inputs are expanded into 32-frame videos first, and then the RGB and depth videos are sent to the C3D model to extract spatiotemporal features respectively. Next these features are combined to boost the performance, which can also avoid unreasonable synthetic data due to the uniform dimension of C3D features. Our approach achieves 49.2% accuracy on the validation subset of the Chalearn LAP IsoGD Database just with a linear SVM classifier. It also outperforms the baseline and other methods in the challenge and wins the first place at 56.9% on testing set.
Yunan Li 0001, Qiguang Miao, Kuan Tian, Xin Xu 0001, Jianfeng Song
ICPR1
2016 Single image haze removal based on haze physical characteristics and adaptive sky region detection
Yunan Li 0001, Qiguang Miao, Jianfeng Song, Yi-Ning Quan, Weisheng Li 0001
Neurocomputing1