VLDB 2026 Research / reviewers in the wild / expert
Xin Liu 0012
dblp:76/1820-12
· DBLP profile ↗
70ranked-venue papers
12as first author
45since 2021 · last 2026
0000-0002-2242-6139ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 48 · 10 first-author · 30 since 2021Artificial intelligence and machine learning · 32 · 4 first-author · 22 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EMO-LLaMA: Enhancing Facial Emotion Understanding with Instruction TuningabstractAbstract Facial expression recognition (FER) has emerged as an important research topic in recent years. However, current FER paradigms face challenges in generalization, lack semantic information aligned with natural language, and struggle to process both images and videos within a unified framework. Multimodal Large Language Models (MLLMs) have recently achieved success, offering advantages in addressing these issues and potentially overcoming the limitations of current FER paradigms. Nonetheless, directly applying pre-trained MLLMs to FER remains challenging due to insufficient instruction datasets and the inability of vision encoders to extract fine-grained facial information. Our zero-shot evaluations of existing open-source MLLMs on FER reveal a significant performance gap compared to GPT-4V and state-of-the-art supervised methods. In this paper, we aim to enhance MLLMs’ capabilities in understanding facial expressions. We first introduce a facial expression recognition instruction dataset ( FERID ), which has 376k category instructions and 339k conversational instructions. We then propose a novel MLLM, named EMO-LLaMA , which incorporates facial priors from a pretrained facial analysis network to enhance its understanding of human facial information. Specifically, we design a Face Info Mining module to extract both global and local facial information. Furthermore, we utilize a handcrafted prompt to introduce age-gender-race attributes, considering the emotional differences across diverse human groups. Extensive experiments show that EMO-LLaMA achieves results comparable to or competitive with SOTA on both static and dynamic FER datasets. The instruction dataset and code will be available at https://github.com/xxtars/EMO-LLaMA . Bohao Xing, Zitong Yu, Xin Liu 0012, Kaishen Yuan, Qilang Ye, Weicheng Xie 0001, Huanjing Yue, Heikki Kälviäinen |
Int. J. Comput. Vis. | 3 |
| 2026 | Identity-Free Artificial Emotional Intelligence via Micro-Gesture UnderstandingabstractIn this work, we focus on a special group of human body language — themicro-gesture (MG), which differs from the range of ordinary illustrative gestures in that they are not intentional behaviors performed to convey information to others, but rather unintentional behaviors driven by inner feelings. This characteristic introduces two novel challenges regarding micro-gestures that are worth rethinking. The first is whether strategies designed for other action recognition are entirely applicable to micro-gestures. The second is whether micro-gestures, as supplementary data, can provide additional insights for emotional understanding. In recognizing micro-gestures, we explore various augmentation strategies that take into account the subtle spatial and brief temporal characteristics of micro-gestures, often accompanied by repetitiveness, to determine more suitable augmentation methods. Considering the significance of temporal domain information for micro-gestures, we introduce a simple and efficient spatiotemporal balancing fusion method. We not only study our method on the considered micro-gesture dataset but also conduct experiments on mainstream gesture/action datasets. The results show that our approach performs well in micro-gesture recognition and on other datasets, achieving state-of-the-art performance compared to previous micro-gesture recognition methods. For emotional understanding based on micro-gestures, we construct complex emotional reasoning scenarios. Our evaluation, conducted with large language models, shows that micro-gestures play a significant and positive role in enhancing comprehensive emotional understanding. We confirm that our new insights contribute to advancing research in micro-gesture and emotional artificial intelligence. Rong Gao 0005, Xin Liu 0012, Bohao Xing, Zitong Yu, Björn W. Schuller, Heikki Kälviäinen |
IEEE Trans. Affect. Comput. | 2 |
| 2026 | HemNet: Hemoglobin-Assistant Network for Video-Based Remote Photoplethysmography MeasurementabstractTraditional skin-contact physical sensors typically detect changes of blood volume to predict the periodicity of heartbeat by analyzing the absorption spectra of hemoglobin. However, the contact on human skin may cause uncomfortable feeling and induce difficulty for long-term monitoring. Recently, video-based remote photoplethysmography (rPPG) estimation approaches analyze the periodic facial color changes for matching cardiac cycle in a contactless manner. Nevertheless, the inherent relationship between the changes of facial color and blood volume is not fully exploited. Besides the influence of blood volume (i.e., hemoglobin), there are also other factors such as lighting and reflection that cause the change on facial color. We exploit the physical principles that cause skin color variations to separate the hemoglobin factor driven by blood volume. Based on the physical prior of the reflection of human skin, we introduce an rPPG estimation network assisted by decoupled hemoglobin sequence, named HemNet, which first explicitly leverages hemoglobin to assist rPPG signal estimation. To obtain meaningful hemoglobin from facial video, we design a human skin color disentangler that decouples the facial color variations into four significant features, i.e., hemoglobin, melanin, shading, and specular. We then present a multi-modality rPPG estimator that utilizes cross-covariance attention to extract fused feature from hemoglobin and RGB video inputs. Finally, an adaptive negative Pearson loss is proposed to effectively address phase misalignment between the blood volume in the finger and facial region during the training phase. We evaluate our HemNet on four widely used public benchmark datasets. The superiority of our method is demonstrated in both intra-dataset and cross-dataset test settings. The code is available at https://github.com/jingang-cv/hemnet. Ruize Wu, Jingang Shi, Xin Liu 0012, LinLin Shen, Yihong Gong, Guoying Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Multi-Granularity Facial Emotional Representation With Unlabeled Data and Textual SupervisionabstractFacial expressions (FEs) and action units (AUs) are facial emotional representations at different levels of granularity. In the past, recognizing them has often been treated as two separate tasks. There are also some methods that use the knowledge of one to aid in recognizing the other, but currently, unified models capable of recognizing both FEs and AUs simultaneously remain rare. In this paper, we construct a unified model with strong generalization capability to jointly perform facial expression recognition (FER) and action unit detection (AUD). Considering the extremely limited training samples annotated with both FEs and AUs, we introduce a large amount of unlabeled facial data from the wild. We carefully design category-specific confidence margins and leverage the correspondences between FEs and AUs to assign credible pseudo-labels to the unlabeled facial data. Furthermore, we incorporate semantically richer textual descriptions as supervision and refine them through visual perception, leveraging the inherent correlations between AUs and between FEs and AUs to enhance their precision. Extensive experiments demonstrate the superiority of the proposed method from various perspectives, including a unified zero-shot benchmark for exploring the model's comprehensive generalization capability to recognize facial emotional representations across multiple datasets, as well as within-domain and cross-domain evaluations after fine-tuning. The code for the proposed method is available at https://github.com/yuankaishen2001/MGFER. Kaishen Yuan, Zitong Yu, Xin Liu 0012, Bohao Xing, Yuting Zhang 0008, Weicheng Xie 0001, LinLin Shen, Björn W. Schuller |
IEEE Trans. Image Process. | 3 |
| 2026 | DIC-DDA: Learned Asymmetric Distributed Image Compression via Dual Domain AlignmentabstractMulti-view or stereo image compression is an essential technology in 3D related applications. Due to the overlap between different views, exploring their correlations can help improve the compression rate. However, the computing complexity of joint encoding at the encoding side is a heavy burden for terminal encoders. To solve this problem, the learned Distributed Image Coding (DIC), which only uses the correlated view (namely the side image, SI) in the decoder side, has gained much attention in recent years. In this work, we explore asymmetric DIC where one view is selected as the SI and is losslessly compressed. The key problem in learned asymmetric DIC is alignment between the transmitted low-quality target image and high-quality SI. Previous methods usually adopt patch-level alignment with the offset index obtained from degraded (via re-encoded and decoded) SI and the decoded target image, which hinders the alignment accuracy. In this work, we propose a dual domain alignment strategy, which includes degraded domain and fused domain pixel-wise offset estimation. For the degraded domain alignment, we estimate the offset between the degraded SI feature and the degraded target image feature, which eliminates the difficulties in cross-domain matching. For the fused-domain alignment, we observe that the fusion result of degraded target feature and aligned side image feature implicitly contains fine-scale disparity information. Therefore, we estimate the fine-scale offset from the fusion result, which helps refine the degraded domain offsets. We further propose a selective enhancement module to repair the mismatched region in the aligned feature. Extensive experiments on three datasets demonstrate the superiority of our proposed method, outperforming the second-best method by 16% in terms of average BD-rate reduction on the KITTI Stereo dataset. Our code is available at https://github.com/lixianghuitju/DIC-DDA. Huanjing Yue, Gaosheng Liu, Xin Liu 0012, Jing-Yu Yang 0002 |
IEEE Trans. Image Process. | 5 |
| 2026 | MSF-Mamba: Motion-Aware State Fusion Mamba for Efficient Micro-Gesture RecognitionabstractMicro-gesture recognition (MGR) targets the identification of subtle and fine-grained human motions and requires accurate modeling of both long-range and local spatiotemporal dependencies. While convolutional neural networks (CNNs) are effective at capturing local patterns, they struggle with long-range dependencies due to their limited receptive fields. Transformer-based models address this limitation through self-attention mechanisms but suffer from high computational costs. Recently, Mamba has shown promise as an efficient model, leveraging state space models (SSMs) to enable linear-time processing. However, directly applying the vanilla Mamba to MGR may not be optimal. This is because Mamba processes inputs as 1D sequences, with state updates relying solely on the previous state, and thus lacks the ability to model local spatiotemporal dependencies. In addition, previous methods lack a design of motion-awareness, which is crucial in MGR. To overcome these limitations, we propose motion-aware state fusion mamba (MSF-Mamba), which enhances Mamba with local spatiotemporal modeling by fusing local contextual neighboring states. Our design introduces a motion-aware state fusion module based on central frame difference (CFD). Furthermore, a multiscale version named MSF-Mamba$^{+}$has been proposed. Specifically, MSF-Mamba$^{+}$supports multiscale motion-aware state fusion, as well as an adaptive scale weighting module that dynamically weighs the fused states across different scales. These enhancements explicitly address the limitations of vanilla Mamba by enabling motion-aware local spatiotemporal modeling, allowing MSF-Mamba and MSF-Mamba$^{+}$to effectively capture subtle motion cues for MGR. Experiments on two public MGR datasets (i.e., SMG and iMiGUE) demonstrate that even the lightweight version, namely, MSF-Mamba, achieves state-of-the-art performance, outperforming existing CNN-, Transformer-, and SSM-based models while maintaining high efficiency. For example, MSF-Mamba improves Top-1 accuracy by +2.2% and +1.5% over VideoMamba on SMG and iMiGUE, respectively. MSF-Mamba$^{+}$outperforms VideoMamba on SMG and iMiGUE, achieving Top-1 accuracy improvements of 2.9% and 3.0%, respectively. The code will be accessible onhttps://github.com/Leedeng/MSF-Mamba Deng Li 0002, Bohao Xing, Rong Gao 0005, Bihan Wen, Heikki Kälviäinen, Xin Liu 0012 |
IEEE Trans. Multim. | 7 |
| 2025 | Zero-shot Video Restoration and Enhancement Using Pre-Trained Image Diffusion ModelabstractDiffusion-based zero-shot image restoration and enhancement models have achieved great success in various tasks of image restoration and enhancement. However, directly applying them to video restoration and enhancement results in severe temporal flickering artifacts. In this paper, we propose the first framework for zero-shot video restoration and enhancement based on the pre-trained image diffusion model. By replacing the spatial self-attention layer with the proposed short-long-range (SLR) temporal attention layer, the pre-trained image diffusion model can take advantage of the temporal correlation between frames. We further propose temporal consistency guidance, spatial-temporal noise sharing, and an early stopping sampling strategy to improve temporally consistent sampling. Our method is a plug-and-play module that can be inserted into any diffusion-based image restoration or enhancement methods to further improve their performance. Experimental results demonstrate the superiority of our proposed method. Cong Cao 0005, Huanjing Yue, Xin Liu 0012, Jing-Yu Yang 0002 |
AAAI | 3 |
| 2025 | FSBench: A Figure Skating Benchmark for Advancing Artistic Sports UnderstandingabstractFigure skating, known as the "Art on Ice," is among the most artistic sports, challenging to understand due to its blend of technical elements (like jumps and spins) and overall artistic expression. Existing figure skating datasets mainly focus on single tasks, such as action recognition or scoring, lacking comprehensive annotations for both technical and artistic evaluation. Current sports research is largely centered on ball games, with limited relevance to artistic sports like figure skating. To address this, we introduce FSAnno, a large-scale dataset advancing artistic sports understanding through figure skating. FSAnno includes an open-access training and test dataset, alongside a benchmark dataset, FSBench, for fair model evaluation. FSBench consists of FSBench-Text, with multiple-choice questions and explanations, and FSBench-Motion, containing multimodal data and Question and Answer (QA) pairs, supporting tasks from technical analysis to performance commentary. Initial tests on FSBench reveal significant limitations in existing models’ understanding of artistic sports. We hope FSBench will become a key tool for evaluating and enhancing model comprehension of figure skating. All data, models, and more details are available at: https://github.com/Moomin-Fin/Ano. Rong Gao 0005, Xin Liu 0012, Zhuozhao Hu, Bohao Xing, Baiqiang Xia, Zitong Yu, Heikki Kälviäinen |
CVPR | 2 |
| 2025 | CATANet: Efficient Content-Aware Token Aggregation for Lightweight Image Super-ResolutionabstractTransformer-based methods have demonstrated impressive performance in low-level visual tasks such as Image Super-Resolution (SR). However, its computational complexity grows quadratically with the spatial resolution. A series of works attempt to alleviate this problem by dividing Low-Resolution images into local windows, axial stripes, or dilated windows. SR typically leverages the redundancy of images for reconstruction, and this redundancy appears not only in local regions but also in long-range regions. However, these methods limit attention computation to content-agnostic local regions, limiting directly the ability of attention to capture long-range dependency. To address these issues, we propose a lightweight Content-Aware Token Aggregation Network (CATANet). Specifically, we propose an efficient Content-Aware Token Aggregation module for aggregating long-range content-similar tokens, which shares token centers across all image tokens and updates them only during the training phase. Then we utilize intra-group self-attention to enable long-range information interaction. Moreover, we design an inter-group cross-attention to further enhance global information interaction. The experimental results show that, compared with the state-of-the-art cluster-based method SPIN, our method achieves superior performance, with a maximum PSNR improvement of 0.33dB and nearly double the inference speed. Xin Liu 0012, Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
CVPR | 1 |
| 2025 | AutoLUT: LUT-Based Image Super-Resolution with Automatic Sampling and Adaptive Residual LearningabstractIn recent years, the increasing popularity of Hi-DPI screens has driven a rising demand for high-resolution images. However, the limited computational power of edge devices poses a challenge in deploying complex super-resolution neural networks, highlighting the need for efficient methods. While prior works have made significant progress, they have not fully exploited pixel-level information. Moreover, their reliance on fixed sampling patterns limits both accuracy and the ability to capture fine details in low-resolution images. To address these challenges, we introduce two plug-and-play modules designed to capture and leverage pixel information effectively in Look-Up Table (LUT) based super-resolution networks. Our method introduces Automatic Sampling (AutoSample), a flexible LUT sampling approach where sampling weights are automatically learned during training to adapt to pixel variations and expand the receptive field without added inference cost. We also incorporate Adaptive Residual Learning (AdaRL) to enhance inter-layer connections, enabling detailed information flow and improving the network’s ability to reconstruct fine details. Our method achieves significant performance improvements on both MuLUT and SPF-LUT while maintaining similar storage sizes. Specifically, for MuLUT, we achieve a PSNR improvement of approximately +0.20 dB improvement on average across five datasets. For SPF-LUT, with more than a 50% reduction in storage space and about a 2/3 reduction in inference time, our method still maintains performance comparable to the original. The code is available at https://github.com/SuperKenVery/AutoLUT. Yuheng Xu, Xin Liu 0012, Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
CVPR | 3 |
| 2025 | Period-LLM: Extending the Periodic Capability of Multimodal Large Language ModelabstractPeriodic or quasi-periodic phenomena reveal intrinsic characteristics in various natural processes, such as weather patterns, movement behaviors, traffic flows, and biological signals. Given that these phenomena span multiple modalities, the capabilities of Multimodal Large Language Models (MLLMs) offer promising potential to effectively capture and understand their complex nature. However, current MLLMs struggle with periodic tasks due to limitations in: 1) lack of temporal modelling and 2) conflict between short and long periods. This paper introduces Period-LLM, a multimodal large language model designed to enhance the performance of periodic tasks across various modalities, and constructs a benchmark of various difficulty for evaluating the cross-modal periodic capabilities of large models. Specially, We adopt an "Easy to Hard Generalization" paradigm, starting with relatively simple text-based tasks and progressing to more complex visual and multimodal tasks, ensuring that the model gradually builds robust periodic reasoning capabilities. Additionally, we propose a "Resisting Logical Oblivion" optimization strategy to maintain periodic reasoning abilities during semantic alignment. Extensive experiments demonstrate the superiority of the proposed Period-LLM over existing MLLMs in periodic tasks. The code is available at https: //github.com/keke-nice/Period-LLM. Yuting Zhang 0008, Hao Lu 0009, Qingyong Hu, Yin Wang 0004, Kaishen Yuan, Xin Liu 0012, Kaishun Wu |
CVPR | 6 |
| 2025 | Dynamic Collaboration of Multi-Language Models based on Minimal Complete Semantic UnitsabstractThis paper investigates the enhancement of reasoning capabilities in language models through token-level multi-model collaboration.Our approach selects the optimal tokens from the next token distributions provided by multiple models to perform autoregressive reasoning.Contrary to the assumption that more models yield better results, we introduce a distribution distance-based dynamic selection strategy (DDS) to optimize the multi-model collaboration process.To address the critical challenge of vocabulary misalignment in multi-model collaboration, we propose the concept of minimal complete semantic units (MCSU), which is simple yet enables multiple language models to achieve natural alignment within the linguistic space.Experimental results across various benchmarks demonstrate the superiority of our method.The code will be available at https://github.com/Fanye12/DDS. Chao Hao, Yanhua Huang, Ruiwen Xu, Wenzhe Niu, Xin Liu 0012, Zitong Yu |
EMNLP | 6 |
| 2025 | DADM: Dual Alignment of Domain and Modality for Face Anti-SpoofingabstractWith the availability of diverse sensor modalities (i.e., RGB, Depth, Infrared) and the success of multi-modal learning, multi-modal face anti-spoofing (FAS) has emerged as a prominent research focus. The intuition behind it is that leveraging multiple modalities can uncover more intrinsic spoofing traces. However, this approach presents more risk of misalignment. We identify two main types of misalignment: (1) \textbf{Intra-domain modality misalignment}, where the importance of each modality varies across different attacks. For instance, certain modalities (e.g., Depth) may be non-defensive against specific attacks (e.g., 3D mask), indicating that each modality has unique strengths and weaknesses in countering particular attacks. Consequently, simple fusion strategies may fall short. (2) \textbf{Inter-domain modality misalignment}, where the introduction of additional modalities exacerbates domain shifts, potentially overshadowing the benefits of complementary fusion. To tackle (1), we propose a alignment module between modalities based on mutual information, which adaptively enhances favorable modalities while suppressing unfavorable ones. To address (2), we employ a dual alignment optimization method that aligns both sub-domain hyperplanes and modality angle margins, thereby mitigating domain gaps. Our method, dubbed \textbf{D}ual \textbf{A}lignment of \textbf{D}omain and \textbf{M}odality (DADM), achieves state-of-the-art performance in extensive experiments across four challenging protocols demonstrating its robustness in multi-modal domain generalization scenarios. The codes will be released soon. Xun Lin, Zitong Yu, Liepiao Zhang, Xin Liu 0012, Hui Li 0089, Xiaochen Yuan, Xiaochun Cao |
ICCV | 5 |
| 2025 | AU-TTT: Vision Test-Time Training model for Facial Action Unit DetectionabstractFacial Action Units (AUs) detection is a cornerstone of objective facial expression analysis and a critical focus in affective computing. Despite its importance, AU detection faces significant challenges, such as the high cost of AU annotation and the limited availability of datasets. These constraints often lead to overfitting in existing methods, resulting in substantial performance degradation when applied across diverse datasets. Addressing these issues is essential for improving the reliability and generalizability of AU detection methods. Moreover, many current approaches leverage Transformers for their effectiveness in long-context modeling, but they are hindered by the quadratic complexity of self-attention. Recently, Test-Time Training (TTT) layers have emerged as a promising solution for long-sequence modeling. Additionally, TTT applies self-supervised learning for iterative updates during both training and inference, offering a potential pathway to mitigate the generalization challenges inherent in AU detection tasks. In this paper, we propose a novel vision backbone tailored for AU detection, incorporating bidirectional TTT blocks, named AU-TTT. Our approach introduces TTT Linear to the AU detection task and optimizes image scanning mechanisms for enhanced performance. Additionally, we design an AU-specific Region of Interest (RoI) scanning mechanism to capture fine-grained facial features critical for AU detection. Experimental results demonstrate that our method achieves competitive performance in both within-domain and cross-domain scenarios. Bohao Xing, Kaishen Yuan, Zitong Yu, Xin Liu 0012, Heikki Kälviäinen |
ICME | 4 |
| 2025 | Human Body Restoration with One-Step Diffusion Model and A New BenchmarkabstractHuman body restoration, as a specific application of image restoration, is widely applied in practice and plays a vital role across diverse fields. However, thorough research remains difficult, particularly due to the lack of benchmark datasets. In this study, we propose a high-quality dataset automated cropping and filtering (HQ-ACF) pipeline. This pipeline leverages existing object detection datasets and other unlabeled images to automatically crop and filter high-quality human images. Using this pipeline, we constructed a person-based restoration with sophisticated objects and natural activities (PERSONA) dataset, which includes training, validation, and test sets. The dataset significantly surpasses other human-related datasets in both quality and content richness. Finally, we propose OSDHuman, a novel one-step diffusion model for human body restoration. Specifically, we propose a high-fidelity image embedder (HFIE) as the prompt generator to better guide the model with low-quality human image information, effectively avoiding misleading prompts. Experimental results show that OSDHuman outperforms existing methods in both visual quality and quantitative metrics. The dataset and code are available at: https://github.com/gobunu/OSDHuman. Jue Gong, Jingkai Wang 0003, Zheng Chen 0014, Xin Liu 0012, Yulun Zhang 0001, Xiaokang Yang 0001 |
ICML | 4 |
| 2025 | DEEMO: De-identity Multimodal Emotion Recognition and ReasoningabstractEmotion understanding is a critical yet challenging task. Most existing approaches rely heavily on identity-sensitive information, such as facial expressions and speech, which raises concerns about personal privacy. To address this, we introduce the De-identity Multimodal Emotion Recognition and Reasoning ( DEEMO ), a novel task designed to enable emotion understanding using de-identified video and audio inputs. The DEEMO dataset consists of two subsets: DEEMO-NFBL , which includes rich annotations of Non-Facial Body Language (NFBL), and DEEMO-MER , an instruction dataset for Multimodal Emotion Recognition and Reasoning using identity-free cues. This design supports emotion understanding without compromising identity privacy. In addition, we propose DEEMO-LLaMA, a Multimodal Large Language Model (MLLM) that integrates de-identified audio, video, and textual information to enhance both emotion recognition and reasoning. Extensive experiments show that DEEMO-LLaMA achieves state-of-the-art performance on both tasks, outperforming existing MLLMs by a significant margin, achieving 74.49% accuracy and 74.45% F1-score in de-identity emotion recognition, and 6.20 clue overlap and 7.66 label overlap in de-identity emotion reasoning. Our work contributes to ethical AI by advancing privacy-preserving emotion understanding and promoting responsible affective computing. The dataset and codes will be available at https://github.com/Leedeng/DEEMO. Deng Li 0002, Bohao Xing, Xin Liu 0012, Baiqiang Xia, Bihan Wen, Heikki Kälviäinen |
ACM Multimedia | 3 |
| 2025 | MER 2025: When Affective Computing Meets Large Language ModelsabstractMER2025 is the third year of our MER series of challenges. Previously, MER2023 (http://merchallenge.cn/mer2023) focused on multi-label learning, noise robustness, and semi-supervised learning, while MER2024 (https://zeroqiaoba.github.io/MER2024-website) introduced a new track dedicated to open-vocabulary emotion recognition. This year, MER2025 centers on the theme ''When Affective Computing Meets Large Language Models (LLMs)''. We aim to shift the paradigm from traditional categorical frameworks reliant on predefined emotion taxonomies to LLM-driven generative methods, offering innovative solutions for more accurate and reliable emotion understanding. The challenge contains four tracks: MER-SEMI focuses on fixed categorical emotion recognition enhanced by semi-supervised learning; MER-FG explores fine-grained emotions, expanding recognition from basic to nuanced emotional states; MER-DES incorporates multimodal cues (beyond emotion words) into predictions to enhance model interpretability; MER-PR reveals whether emotion prediction results can improve personality recognition performance. For the first three tracks, the baseline code is available at MERTools (https://github.com/zeroQiaoba/MERTools) and datasets can be accessed via Hugging Face (https://huggingface.co/datasets/MERChallenge/MER2025). For the last track, the dataset and baseline code are available on GitHub (https://github.com/cai-cong/MER25_personality). Zheng Lian 0004, Rui Liu 0008, Kele Xu, Bin Liu 0041, Xuefei Liu, Yazhou Zhang 0001, Xin Liu 0012, Yong Li 0032, Zebang Cheng, Haolin Zuo, Ziyang Ma 0001, Xiaojiang Peng, Xie Chen 0001, Ya Li 0001, Erik Cambria, Guoying Zhao 0001, Björn W. Schuller, Jianhua Tao 0001 |
ACM Multimedia | 7 |
| 2025 | To Remember, To Adapt, To Preempt: A Stable Continual Test-Time Adaptation Framework for Remote Physiological Measurement in Dynamic Domain ShiftsabstractRemote photoplethysmography (rPPG) aims to extract non-contact physiological signals from facial videos and has shown great potential. However, existing rPPG approaches struggle to bridge the gap between source and target domains. Recent test-time adaptation (TTA) solutions typically optimize rPPG model for the incoming test videos using self-training loss under an unrealistic assumption that the target domain remains stationary. However, time-varying factors like weather and lighting in dynamic environments often cause continual domain shifts. The erroneous gradients accumulation from these shifts may corrupt the model's key parameters for physiological information, leading to catastrophic forgetting. Therefore, We propose a physiology-related parameters freezing strategy to retain such knowledge. It isolates physiology-related and domain-related parameters by assessing the model's uncertainty to current domain and freezes the physiology-related parameters during adaptation to prevent catastrophic forgetting. Moreover, the dynamic domain shifts with various non-physiological characteristics may lead to conflicting optimization objectives during TTA, which is manifested as the over-adapted model losing its adaptability to future domains. To fix over-adaptation, we propose a preemptive gradient modification strategy. It preemptively adapts to future domains and uses the acquired gradients to modify current adaptation, thereby preserving the model's adaptability. In summary, we propose a stable continual test-time adaptation (CTTA) framework for rPPG measurement, called PhysRAP, which Remembers the past, Adapts to the present, and Preempts the future. Extensive experiments show its state-of-the-art performance, especially in domain shifts. The code is available at https://github.com/xjtucsy/PhysRAP. Shuyang Chu, Jingang Shi, Xu Cheng 0003, Haoyu Chen 0001, Xin Liu 0012, Guoying Zhao 0001 |
ACM Multimedia | 5 |
| 2025 | FEALLM: Advancing Facial Emotion Analysis in Multimodal Large Language Models with Emotional Synergy and Reasoning
Zhuozhao Hu, Kaishen Yuan, Xin Liu 0012, Zitong Yu, Yuan Zong, Jingang Shi, Huanjing Yue, Jing-Yu Yang 0002 |
ACM Multimedia | 3 |
| 2025 | HAODiff: Human-Aware One-Step Diffusion via Dual-Prompt GuidanceabstractHuman-centered images often suffer from severe generic degradation during transmission and are prone to human motion blur (HMB), making restoration challenging. Existing research lacks sufficient focus on these issues, as both problems often coexist in practice. To address this, we design a degradation pipeline that simulates the coexistence of HMB and generic noise, generating synthetic degraded data to train our proposed HAODiff, a human-aware one-step diffusion. Specifically, we propose a triple-branch dual-prompt guidance (DPG), which leverages high-quality images, residual noise (LQ minus HQ), and HMB segmentation masks as training targets. It produces a positive–negative prompt pair for classifier‑free guidance (CFG) in a single diffusion step. The resulting adaptive dual prompts let HAODiff exploit CFG more effectively, boosting robustness against diverse degradations. For fair evaluation, we introduce MPII‑Test, a benchmark rich in combined noise and HMB cases. Extensive experiments show that our HAODiff surpasses existing state-of-the-art (SOTA) methods in terms of both quantitative metrics and visual quality on synthetic and real-world datasets, including our introduced MPII-Test. Code is available at: https://github.com/gobunu/HAODiff. Jue Gong, Tingyu Yang, Jingkai Wang 0003, Zheng Chen 0014, Xin Liu 0012, Yulun Zhang 0001, Xiaokang Yang 0001 |
NeurIPS | 5 |
| 2025 | CAT+: Investigating and Enhancing Audio-Visual Understanding in Large Language ModelsabstractMultimodal Large Language Models (MLLMs) have gained significant attention due to their rich internal implicit knowledge for cross-modal learning. Although advances in bringing audio-visuals into LLMs have resulted in boosts for a variety of Audio-Visual Question Answering (AVQA) tasks, they still face two crucial challenges: 1) audio-visual ambiguity, and 2) audio-visual hallucination. Existing MLLMs can respond to audio-visual content, yet sometimes fail to describe specific objects due to the ambiguity or hallucination of responses. To overcome the two aforementioned issues, we introduce the CAT+, which enhances MLLM to ensure more robust multimodal understanding. We first propose the Sequential Question-guided Module (SQM), which combines tiny transformer layers and cascades Q-Formers to realize a solid audio-visual grounding. After feature alignment and high-quality instruction tuning, we introduce Ambiguity Scoring Direct Preference Optimization (AS-DPO) to correct the problem of CAT+ bias toward ambiguous descriptions. To explore the hallucinatory deficits of MLLMs in dynamic audio-visual scenes, we build a new Audio-visual Hallucination Benchmark, named AVHbench. This benchmark detects the extent of MLLM's hallucinations across three different protocols in the perceptual object, counting, and holistic description tasks. Extensive experiments across video-based understanding, open-ended, and close-ended AVQA demonstrate the superior performance of our method. The AVHbench is released at https://github.com/rikeilong/Bay-CAT. Qilang Ye, Zitong Yu, Rui Shao 0001, Yawen Cui, Xiangui Kang, Xin Liu 0012, Philip Torr 0001, Xiaochun Cao |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Multi-Scale Promoted Self-Adjusting Correlation Learning for Facial Action Unit DetectionabstractFacial Action Unit (AU) detection is a crucial task in affective computing and social robotics as it helps to identify emotions expressed through facial expressions. Anatomically, there are innumerable correlations between AUs, which contain rich information and are vital for AU detection. Previous methods used fixed AU correlations based on expert experience or statistical rules on specific benchmarks, but it is challenging to comprehensively reflect complex correlations between AUs via hand-crafted settings. There are alternative methods that employ a fully connected graph to learn these dependencies exhaustively. However, these approaches can result in a computational explosion and high dependency with a large dataset. To address these challenges, this paper proposes a novel self-adjusting AU-correlation learning (SACL) method with less computation for AU detection. This method adaptively learns and updates AU correlation graphs by efficiently leveraging the characteristics of different levels of AU motion and emotion representation information extracted in different stages of the network. Moreover, this paper explores the role of multi-scale learning in correlation information extraction, and design a simple yet effective multi-scale feature learning (MSFL) method to promote better performance in AU detection. By integrating AU correlation information with multi-scale features, the proposed method obtains a more robust feature representation for the final AU detection. Extensive experiments show that the proposed method outperforms the state-of-the-art methods on widely used AU detection benchmark datasets, with only 28.7% and 12.0% of the parameters and FLOPs of the best method, respectively. Xin Liu 0012, Kaishen Yuan, Xuesong Niu, Jingang Shi, Zitong Yu, Huanjing Yue, Jing-Yu Yang 0002 |
IEEE Trans. Affect. Comput. | 1 |
| 2025 | A Simple Yet Effective Network Based on Vision Transformer for Camouflaged Object and Salient Object DetectionabstractCamouflaged object detection (COD) and salient object detection (SOD) are two distinct yet closely-related computer vision tasks widely studied during the past decades. Though sharing the same purpose of segmenting an image into binary foreground and background regions, their distinction lies in the fact that COD focuses on concealed objects hidden in the image, while SOD concentrates on the most prominent objects in the image. Building universal segmentation models is currently a hot topic in the community. Previous works achieved good performance on certain task by stacking various hand-designed modules and multi-scale features. However, these careful task-specific designs also make them lose their potential as general-purpose architectures. Therefore, we hope to build general architectures that can be applied to both tasks. In this work, we propose a simple yet effective network (SENet) based on vision Transformer (ViT), by employing a simple design of an asymmetric ViT-based encoder-decoder structure, we yield competitive results on both tasks, exhibiting greater versatility than meticulously crafted ones. To enhance the performance of universal architectures on both tasks, we propose some general methods targeting some common difficulties of the two tasks. First, we use image reconstruction as an auxiliary task during training to increase the difficulty of training, forcing the network to have a better perception of the image as a whole to help with segmentation tasks. In addition, we propose a local information capture module (LICM) to make up for the limitations of the patch-level attention mechanism in pixel-level COD and SOD tasks and a dynamic weighted loss (DW loss) to solve the problem that small target samples are more difficult to locate and segment in both tasks. Finally, we also conduct a preliminary exploration of joint training, trying to use one model to complete two tasks simultaneously. Extensive experiments on multiple benchmark datasets demonstrate the effectiveness of our method. The code is available at https://github.com/linuxsino/SENet. Chao Hao, Zitong Yu, Xin Liu 0012, Jun Xu 0019, Huanjing Yue, Jing-Yu Yang 0002 |
IEEE Trans. Image Process. | 3 |
| 2025 | Advancing Generalizable Remote Physiological Measurement Through the Integration of Explicit and Implicit Prior KnowledgeabstractRemote photoplethysmography (rPPG) is a promising technology for capturing physiological signals from facial videos, with potential applications in medical health, affective computing, and biometric recognition. The demand for rPPG tasks has evolved from achieving high performance in intra-dataset testing to excelling in cross-dataset testing (i.e., domain generalization). However, most existing methods have overlooked the incorporation of prior knowledge specific to rPPG, leading to limited generalization capabilities. In this paper, we propose a novel framework that effectively integrates both explicit and implicit prior knowledge into the rPPG task. Specifically, we conduct a systematic analysis of noise sources (e.g., variations in cameras, lighting conditions, skin types, and motion) across different domains and embed this prior knowledge into the network design. Furthermore, we employ a two-branch network to disentangle physiological feature distributions from noise through implicit label correlation. Extensive experiments demonstrate that the proposed method not only surpasses state-of-the-art approaches in RGB cross-dataset evaluation but also exhibits strong generalization from RGB datasets to NIR datasets. The code is publicly available at https://github.com/keke-nice/Greip. Yuting Zhang 0008, Hao Lu 0009, Xin Liu 0012, Ying-Cong Chen, Kaishun Wu |
IEEE Trans. Image Process. | 3 |
| 2025 | CodePhys: Robust Video-Based Remote Physiological Measurement Through Latent Codebook QueryingabstractRemote photoplethysmography (rPPG) aims to measure non-contact physiological signals from facial videos, which has shown great potential in many applications. Most existing methods directly extract video-based rPPG features by designing neural networks for heart rate estimation. Although they can achieve acceptable results, the recovery of rPPG signal faces intractable challenges when interference from real-world scenarios takes place on facial video. Specifically, facial videos are inevitably affected by non-physiological factors (e.g., camera device noise, defocus, and motion blur), leading to the distortion of extracted rPPG signals. Recent rPPG extraction methods are easily affected by interference and degradation, resulting in noisy rPPG signals. In this paper, we propose a novel method named CodePhys, which innovatively treats rPPG measurement as a code query task in a noise-free proxy space (i.e., codebook) constructed by ground-truth PPG signals. We consider noisy rPPG features as queries and generate high-fidelity rPPG features by matching them with noise-free PPG features from the codebook. Our approach also incorporates a spatial-aware encoder network with a spatial attention mechanism to highlight physiologically active areas and uses a distillation loss to reduce the influence of non-periodic visual interference. Experimental results on four benchmark datasets demonstrate that CodePhys outperforms state-of-the-art methods in both intra-dataset and cross-dataset settings. Shuyang Chu, Menghan Xia, Mengyao Yuan, Xin Liu 0012, Tapio Seppänen, Guoying Zhao 0001, Jingang Shi |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | DiffFAS: Face Anti-spoofing via Generative Diffusion Models
Xinxu Ge, Xin Liu 0012, Zitong Yu, Jingang Shi, Chun Qi, Heikki Kälviäinen |
ECCV (54) | 2 |
| 2024 | AUFormer: Vision Transformers Are Parameter-Efficient Facial Action Unit Detectors
Kaishen Yuan, Zitong Yu, Xin Liu 0012, Weicheng Xie 0001, Huanjing Yue, Jing-Yu Yang 0002 |
ECCV (50) | 3 |
| 2024 | Adversarial Robustness in RGB-Skeleton Action Recognition: Leveraging Attention Modality ReweighterabstractDeep neural networks (DNNs) have been applied in many computer vision tasks and achieved state-of-the-art (SOTA) performance. However, misclassification will occur when DNNs predict adversarial examples which are created by adding human-imperceptible adversarial noise to natural examples. This limits the application of DNN in security-critical fields. In order to enhance the robustness of models, previous research has primarily focused on the unimodal domain, such as image recognition and video understanding. Although multi-modal learning has achieved advanced performance in various tasks, such as action recognition, research on the robustness of RGB-skeleton action recognition models is scarce. In this paper, we systematically investigate how to improve the robustness of RGB-skeleton action recognition models. We initially conducted empirical analysis on the robustness of different modalities and observed that the skeleton modality is more robust than the RGB modality. Motivated by this observation, we propose the Attention-based Modality Reweighter (AMR), which utilizes an attention layer to re-weight the two modalities, enabling the model to learn more robust features. Our AMR is plug-and-play, allowing easy integration with multimodal models. To demonstrate the effectiveness of AMR, we conducted extensive experiments on various datasets. For example, compared to the SOTA methods, AMR exhibits a 43.77% improvement against PGD20 attacks on the NTURGB+D 60 dataset. Furthermore, it effectively balances the differences in robustness between different modalities. Xin Liu 0012, Zitong Yu, Yonghong Hou, Huanjing Yue, Jing-Yu Yang 0002 |
IJCB | 2 |
| 2024 | 4K-Resolution Photo Exposure Correction at 125 FPS with ~8K ParametersabstractThe illumination of improperly exposed photographs has been widely corrected using deep convolutional neural networks or Transformers. Despite with promising performance, these methods usually suffer from large parameter amounts and heavy computational FLOPs on high-resolution photographs. In this paper, we propose extremely light-weight (with only ~8K parameters) Multi-Scale Linear Transformation (MSLT) networks under the multi-layer perception architecture, which can process 4K-resolution sRGB images at 125 Frame-Per-Second (FPS) by a Titan RTX GPU. Specifically, the proposed MSLT networks first decompose an input image into high and low frequency layers by Laplacian pyramid techniques, and then sequentially correct different layers by pixel-adaptive linear transformation, which is implemented by efficient bilateral grid learning or 1 × 1 convolutions. Experiments on two benchmark datasets demonstrate the efficiency of our MSLTs against the state-of-the-arts on photo exposure correction. Extensive ablation studies validate the effectiveness of our contributions. The code is available at https://github.com/Zhou-Yijie/MSLTNet. Xin Liu 0012, Jun Xu 0019 |
WACV | 5 |
| 2024 | Rethinking Vision Transformer and Masked Autoencoder in Multimodal Face Anti-SpoofingabstractAbstract Recently, vision transformer (ViT) based multimodal learning methods have been proposed to improve the robustness of face anti-spoofing (FAS) systems. However, there are still no works to explore the fundamental natures (e.g., modality-aware inputs, suitable multimodal pre-training, and efficient finetuning) in vanilla ViT for multimodal FAS. In this paper, we investigate three key factors (i.e., inputs, pre-training, and finetuning) in ViT for multimodal FAS with RGB, Infrared (IR), and Depth. First, in terms of the ViT inputs, we find that leveraging local feature descriptors (such as histograms of oriented gradients) benefits the ViT on IR modality but not RGB or Depth modalities. Second, in consideration of the task (FAS vs. generic object classification) and modality (multimodal vs. unimodal) gaps, ImageNet pre-trained models might be sub-optimal for the multimodal FAS task. Finally, in observation of the inefficiency on direct finetuning the whole or partial ViT, we design an adaptive multimodal adapter (AMA), which can efficiently aggregate local multimodal features while freezing majority of ViT parameters. To bridge these gaps, we propose the modality-asymmetric masked autoencoder (M $$^{2}$$ 2 A $$^{2}$$ 2 E) for multimodal FAS self-supervised pre-training without costly annotated labels. Compared with the previous modality-symmetric autoencoder, the proposed M $$^{2}$$ 2 A $$^{2}$$ 2 E is able to learn more intrinsic task-aware representation and compatible with modality-agnostic (e.g., unimodal, bimodal, and trimodal) downstream settings. Extensive experiments with both unimodal (RGB, Depth, IR) and multimodal (RGB+Depth, RGB+IR, Depth+IR, RGB+Depth+IR) settings conducted on multimodal FAS benchmarks demonstrate the superior performance of the proposed methods. One highlight is that the proposed method is robust under various missing-modality cases where previous multimodal FAS models suffer serious performance drops. We hope these findings and solutions can facilitate the future research for ViT-based multimodal FAS. Zitong Yu, Rizhao Cai, Yawen Cui, Xin Liu 0012, Yongjian Hu, Alex Chichung Kot |
Int. J. Comput. Vis. | 4 |
| 2024 | Enhancing Micro Gesture Recognition for Emotion Understanding via Context-Aware Visual-Text Contrastive LearningabstractPsychological studies have shown that Micro Gestures (MG) are closely linked to human emotions. MG-based emotion understanding has attracted much attention because it allows for emotion understanding through nonverbal body gestures without relying on identity information (e.g., facial and electrocardiogram data). Therefore, it is essential to recognize MG effectively for advanced emotion understanding. However, existing Micro Gesture Recognition (MGR) methods utilize only a single modality (e.g., RGB or skeleton) while overlooking crucial textual information. In this letter, we propose a simple but effective visual-text contrastive learning solution that utilizes text information for MGR. In addition, instead of using handcrafted prompts for visual-text contrastive learning, we propose a novel module called Adaptive prompting to generate contextaware prompts. The experimental results show that the proposed method achieves state-of-the-art performance on two public datasets. Furthermore, based on an empirical study utilizing the results of MGR for emotion understanding, we demonstrate that using the textual results of MGR significantly improves performance by 6%+ compared to directly using video as input. Deng Li 0002, Bohao Xing, Xin Liu 0012 |
IEEE Signal Process. Lett. | 3 |
| 2024 | Unsupervised HDR Image and Video Tone Mapping via Contrastive LearningabstractCapturing high dynamic range (HDR) images (videos) is attractive because it can reveal the details in both dark and bright regions. Since the mainstream screens only support low dynamic range (LDR) content, tone mapping algorithm is required to compress the dynamic range of HDR images (videos). Although image tone mapping has been widely explored, video tone mapping is lagging behind, especially for the deep-learning-based methods, due to the lack of HDR-LDR video pairs. In this work, we propose a unified framework (IVTMNet) for unsupervised image and video tone mapping. To improve unsupervised training, we propose domain and instance based contrastive learning loss. Instead of using a universal feature extractor, such as VGG to extract the features for similarity measurement, we propose a novel latent code, which is an aggregation of the brightness and contrast of extracted features, to measure the similarity of different pairs. We totally construct two negative pairs and three positive pairs to constrain the latent codes of tone mapped results. For the network structure, we propose a spatial-feature-enhanced (SFE) module to enable information exchange and transformation of nonlocal regions. For video tone mapping, we propose a temporal-feature-replaced (TFR) module to efficiently utilize the temporal correlation and improve the temporal consistency of video tone-mapped results. We construct a large-scale unpaired HDR-LDR video dataset to facilitate the unsupervised training process for video tone mapping. Experimental results demonstrate that our method outperforms state-of-the-art image and video tone mapping methods. Our code and dataset are available athttps://github.com/cao-cong/UnCLTMO. Cong Cao 0005, Huanjing Yue, Xin Liu 0012, Jing-Yu Yang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | rPPG-MAE: Self-Supervised Pretraining With Masked Autoencoders for Remote Physiological MeasurementsabstractRemote photoplethysmography (rPPG) is an important technique for detecting human vital signs and has received extensive attention. For a long time, researchers have focused attention on supervised methods that rely on large amounts of labeled data. These methods are limited by their need for large amounts of data and the difficulty of acquiring ground truth physiological signals. To address these issues, several self-supervised methods based on contrastive learning have been proposed. However, they focus on contrastive learning between samples, which neglects inherent self-similar priors in physiological signals and seems to have a limited ability to cope with noise. In this paper, a linear self-supervised reconstruction task was designed for extracting the inherent self-similar priors in physiological signals. In addition, a specific noise-insensitive strategy was explored for reducing the interference of motion and illumination. The framework proposed in this paper, rPPG-MAE, demonstrates excellent performance even on the challenging VIPL-HR dataset. We also evaluate the proposed method on two public datasets, namely, PURE and UBFC-rPPG. The results show that our method not only outperforms existing self-supervised methods but also outperforms state-of-the-art (SOTA) supervised methods. One important observation is that the quality of the dataset appears to be more important than the size of the dataset used in self-supervised pretraining of the rPPG. The source code is available athttps://github.com/linuxsino/rPPG-MAE. Xin Liu 0012, Yuting Zhang 0008, Zitong Yu, Hao Lu 0009, Huanjing Yue, Jing-Yu Yang 0002 |
IEEE Trans. Multim. | 1 |
| 2024 | From Recognition to Prediction: Leveraging Sequence Reasoning for Action AnticipationabstractThe action anticipation task refers to predicting what action will happen based on observed videos, which requires the model to have a strong ability to summarize the present and then reason about the future. Experience and common sense suggest that there is a significant correlation between different actions, which provides valuable prior knowledge for the action anticipation task. However, previous methods have not effectively modeled this underlying statistical relationship. To address this issue, we propose a novel end-to-end video modeling architecture that utilizes attention mechanisms, named Anticipation via Recognition and Reasoning (ARR). ARR decomposes the action anticipation task into action recognition and sequence reasoning tasks and effectively learns the statistical relationship between actions by next action prediction (NAP). In comparison to existing temporal aggregation strategies, ARR is able to extract more effective features from observable videos to make more reasonable predictions. In addition, to address the challenge of relationship modeling that requires extensive training data, we propose an innovative approach for the unsupervised pre-training of the decoder, which leverages the inherent temporal dynamics of video to enhance the reasoning capabilities of the network. Extensive experiments on the Epic-kitchen-100, EGTEA Gaze+, and 50salads datasets demonstrate the efficacy of the proposed methods. The code is available at https://github.com/linuxsino/ARR . Xin Liu 0012, Chao Hao, Zitong Yu, Huanjing Yue, Jing-Yu Yang 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | Recaptured Raw Screen Image and Video Demoiréing via Channel and Spatial ModulationsabstractCapturing screen contents by smartphone cameras has become a common way for information sharing. However, these images and videos are often degraded by moiré patterns, which are caused by frequency aliasing between the camera filter array and digital display grids. We observe that the moiré patterns in raw domain is simpler than those in sRGB domain, and the moiré patterns in raw color channels have different properties. Therefore, we propose an image and video demoiréing network tailored for raw inputs. We introduce a color-separated feature branch, and it is fused with the traditional feature-mixed branch via channel and spatial modulations. Specifically, the channel modulation utilizes modulated color-separated features to enhance the color-mixed features. The spatial modulation utilizes the feature with large receptive field to modulate the feature with small receptive field. In addition, we build the first well-aligned raw video demoiréing (RawVDemoiré) dataset and propose an efficient temporal alignment method by inserting alternating patterns. Experiments demonstrate that our method achieves state-of-the-art performance for both image and video demoiréing. Our dataset and code will be released after the acceptance of this work. Yijia Cheng, Xin Liu 0012, Jing-Yu Yang 0002 |
NeurIPS | 2 |
| 2023 | SMG: A Micro-gesture Dataset Towards Spontaneous Body Gestures for Emotional Stress State AnalysisabstractAbstract We explore using body gestures for hidden emotional state analysis. As an important non-verbal communicative fashion, human body gestures are capable of conveying emotional information during social communication. In previous works, efforts have been made mainly on facial expressions, speech, or expressive body gestures to interpret classical expressive emotions. Differently, we focus on a specific group of body gestures, called micro-gestures (MGs), used in the psychology research field to interpret inner human feelings. MGs are subtle and spontaneous body movements that are proven, together with micro-expressions, to be more reliable than normal facial expressions for conveying hidden emotional information. In this work, a comprehensive study of MGs is presented from the computer vision aspect, including a novel spontaneous micro-gesture (SMG) dataset with two emotional stress states and a comprehensive statistical analysis indicating the correlations between MGs and emotional states. Novel frameworks are further presented together with various state-of-the-art methods as benchmarks for automatic classification, online recognition of MGs, and emotional stress state recognition. The dataset and methods presented could inspire a new way of utilizing body gestures for human emotion understanding and bring a new direction to the emotion AI community. The source code and dataset are made available: https://github.com/mikecheninoulu/SMG . Haoyu Chen 0001, Henglin Shi, Xin Liu 0012, Guoying Zhao 0001 |
Int. J. Comput. Vis. | 3 |
| 2023 | Forgetting memristor based STDP learning circuit for neural networks
Shiping Wen 0001, Xin Liu 0012, Ling Chen 0010 |
Neural Networks | 5 |
| 2023 | Long-term and short-term memory networks based on forgetting memristors
Ling Chen 0010, Chuandong Li 0001, Xin Liu 0012 |
Soft Comput. | 4 |
| 2023 | Consistency Regularization for Deep Face Anti-SpoofingabstractFace anti-spoofing (FAS) plays a crucial role in securing face recognition systems. Empirically, given an image, a model with more consistent output on different views (i.e., augmentations) of this image usually performs better. Motivated by this exciting observation, we conjecture that encouraging feature consistency of different views may be a promising way to boost FAS models. In this paper, we explore this way thoroughly by enhancing both Embedding-level and Prediction-level Consistency Regularization (EPCR) in FAS. Specifically, at the embedding level, we design a dense similarity loss to maximize the similarities between all positions of two intermediate feature maps in a self-supervised fashion; while at the prediction level, we optimize the mean square error between the predictions of two views. Notably, our EPCR is free of annotations and can directly integrate into semi-supervised learning schemes. Considering different application scenarios, we further design five diverse semi-supervised protocols to measure semi-supervised FAS techniques. We conduct extensive experiments to show that EPCR can significantly improve the performance of several supervised and semi-supervised tasks on benchmark datasets. The codes and protocols are available athttps://github.com/clks-wzz/EPCR. Zezheng Wang 0002, Zitong Yu, Yunxiao Qin, Xin Liu 0012, Zhen Lei 0001 |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2022 | CdCLR: Clip-Driven Contrastive Learning for Skeleton-Based Action RecognitionabstractIn this study, we propose a Clip-Driven Contrastive Learning for Skeleton-Based Action Recognition (CdCLR). In-stead of considering sequences as instances, CdCLR extracts clips from the sequences as new instances. Aim to implement inherent supervision-guided contrastive learning through joint optimal training of sequences discrimination, clips discrimination, and order verification. Mining abundant positive/negative pairs inside sequence while learning inter-and intra-sequence semantic repre-sentations. Extensive experiments on the NTU RGB+D 60, UCLA and iMiGUE datasets present that CdCLR exhibits superior performance under various evaluation protocols and reaches state-of-the-art. Our code is available at https://github.com/Erich-G/CdCLRI. Rong Gao 0005, Xin Liu 0012, Jing-Yu Yang 0002, Huanjing Yue |
VCIP | 2 |
| 2022 | A Robust GAN-Generated Face Detection Method Based on Dual-Color Spaces and an Improved XceptionabstractIn recent years, generative adversarial networks (GANs) have been widely used to generate realistic fake face images, which can easily deceive human beings. To detect these images, some methods have been proposed. However, their detection performance will be degraded greatly when the testing samples are post-processed. In this paper, some experimental studies on detecting post-processed GAN-generated face images find that (a) both the luminance component and chrominance components play an important role, and (b) the RGB and YCbCr color spaces achieve better performance than the HSV and Lab color spaces. Therefore, to enhance the robustness, both the luminance component and chrominance components of dual-color spaces (RGB and YCbCr) are considered to utilize color information effectively. In addition, the convolutional block attention module and multilayer feature aggregation module are introduced into the Xception model to enhance its feature representation power and aggregate multilayer features, respectively. Finally, a robust dual-stream network is designed by integrating dual-color spaces RGB and YCbCr and using an improved Xception model. Experimental results demonstrate that our method outperforms some existing methods, especially in its robustness against different types of post-processing operations, such as JPEG compression, Gaussian blurring, gamma correction, and median filtering. Beijing Chen, Xin Liu 0012, Yuhui Zheng, Guoying Zhao 0001, Yun Q. Shi 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Unsupervised Domain-Invariant Feature Learning for Cloud Detection of Remote Sensing ImagesabstractThe detection of clouds in remote sensing (RS) images is an important task, and convolutional neural networks (CNNs) have been used to perform it. However, supervised cloud detection CNNs rely heavily on a large number of samples annotated at pixel level to tune their parameter. Annotating RS images is a labor-intensive procedure and requires expert-level human knowledge. To reduce the labeling cost, we propose an unsupervised domain adaptation (UDA) approach to enable the model trained on labeled source satellite images to generalize to unlabeled target satellite images. Specifically, we propose a fine-grained feature alignment (FGFA) domain adaptation strategy that encourages a cloud detection network to extract domain-invariant representations, which improves the accuracy of cloud detection in unlabeled target satellite images. The proposed FGFA strategy consists of two steps: 1) fine-grained class-relevant feature selection based on an attention-guided mechanism and 2) class-relevant feature alignment (FA) based on a proposed grouped FA approach. Experimental results on the “Landsat-$8~\rightarrow $ZY-3” and “GF-$1\rightarrow $ZY-3” domain adaptation tasks demonstrate the effectiveness of our method and its superiority to existing state-of-the-art UDA approaches. Jianhua Guo 0002, Jing-Yu Yang 0002, Huanjing Yue, Xin Liu 0012, Kun Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | iMiGUE: An Identity-Free Video Dataset for Micro-Gesture Understanding and Emotion AnalysisabstractWe introduce a new dataset for the emotional artificial intelligence research: identity-free video dataset for Micro-Gesture Understanding and Emotion analysis (iMiGUE). Different from existing public datasets, iMiGUE focuses on nonverbal body gestures without using any identity information, while the predominant researches of emotion analysis concern sensitive biometric data, like face and speech. Most importantly, iMiGUE focuses on micro-gestures, i.e., unintentional behaviors driven by inner feelings, which are different from ordinary scope of gestures from other gesture datasets which are mostly intentionally performed for illustrative purposes. Furthermore, iMiGUE is designed to evaluate the ability of models to analyze the emotional states by integrating information of recognized micro-gesture, rather than just recognizing prototypes in the sequences separately (or isolatedly). This is because the real need for emotion AI is to understand the emotional states behind gestures in a holistic way. Moreover, to counter for the challenge of imbalanced sample distribution of this dataset, an unsupervised learning method is proposed to capture latent representations from the micro-gesture sequences themselves. We systematically investigate representative methods on this dataset, and comprehensive experimental results reveal several interesting insights from the iMiGUE, e.g., micro-gesture-based analysis can promote emotion understanding. We confirm that the new iMiGUE dataset could advance studies of micro-gesture and emotion AI. Xin Liu 0012, Henglin Shi, Haoyu Chen 0001, Zitong Yu, Guoying Zhao 0001 |
CVPR | 1 |
| 2021 | Searching Multi-Rate and Multi-Modal Temporal Enhanced Networks for Gesture RecognitionabstractGesture recognition has attracted considerable attention owing to its great potential in applications. Although the great progress has been made recently in multi-modal learning methods, existing methods still lack effective integration to fully explore synergies among spatio-temporal modalities effectively for gesture recognition. The problems are partially due to the fact that the existing manually designed network architectures have low efficiency in the joint learning of multi-modalities. In this paper, we propose the first neural architecture search (NAS)-based method for RGB-D gesture recognition. The proposed method includes two key components: 1) enhanced temporal representation via the proposed 3D Central Difference Convolution (3D-CDC) family, which is able to capture rich temporal context via aggregating temporal difference information; and 2) optimized backbones for multi-sampling-rate branches and lateral connections among varied modalities. The resultant multi-modal multi-rate network provides a new perspective to understand the relationship between RGB and depth modalities and their temporal dynamics. Comprehensive experiments are performed on three benchmark datasets (IsoGD, NvGesture, and EgoGesture), demonstrating the state-of-the-art performance in both single- and multi-modality settings. The code is available at https://github.com/ZitongYu/3DCDC-NAS. Zitong Yu, Benjia Zhou, Jun Wan 0001, Pichao Wang, Haoyu Chen 0001, Xin Liu 0012, Stan Z. Li, Guoying Zhao 0001 |
IEEE Trans. Image Process. | 6 |
| 2021 | 3D Skeletal Gesture Recognition via Discriminative Coding on Time-Warping Invariant Riemannian TrajectoriesabstractLearning 3D skeleton-based representation for gesture recognition has progressively stood out because of its invariance to the viewpoint and background dynamics of video. Typically, existing techniques use absolute coordinates to determine human motion features. The recognition of gestures, however, is irrespective of the position of the performer, and the extracted features should be invariant to body size. In addition, when comparing and classifying gestures, the problem of temporal dynamics can greatly distort the distance metric. In this paper, we represent a 3D skeleton as a point in the special orthogonal group$SO(3)$product space that expressly models the 3D geometric relationships between body parts. As such, a gesture skeletal sequence can be described by a trajectory on a Riemannian manifold. Following that, we propose to generalize the transported square-root vector field to obtain a time-warping invariant metric for comparing these trajectories (identifying these gestures). Moreover, by specifically considering the labeling information with encoding, a sparse coding scheme of skeletal trajectories is presented to enforce the discriminant validity of atoms in the dictionary. Experimental results indicate that the proposed approach has achieved state-of-the-art performance on many challenging gesture recognition benchmarks. Xin Liu 0012, Guoying Zhao 0001 |
IEEE Trans. Multim. | 1 |
| 2020 | An efficient volume repairing method by using a modified Allen-Cahn equation
Yibao Li, Shouren Lan, Xin Liu 0012, Bingheng Lu, Lisheng Wang |
Pattern Recognit. | 3 |
| 2020 | Temporal Hierarchical Dictionary Guided Decoding for Online Gesture Segmentation and RecognitionabstractOnline segmentation and recognition of skeleton- based gestures are challenging. Compared with offline cases, the inference of online settings can only rely on the current few frames and always completes before whole temporal movements are performed. However, incompletely performed gestures are ambiguous and their early recognition is easy to fall into local optimum. In this work, we address the problem with a temporal hierarchical dictionary to guide the hidden Markov model (HMM) decoding procedure. The intuition is that, gestures are ambiguous with high uncertainty at early performing phases, and only become discriminate after certain phases. This uncertainty naturally can be measured by entropy. Thus, we propose a measurement called "relative entropy map" (REM) to encode this temporal context to guide HMM decoding. Furthermore, we introduce a progressive learning strategy with which neural networks could learn a robust recognition of HMM states in an iterative manner. The performance of our method is intensively evaluated on three challenging databases and achieves state-of-the-art results. Our method shows the abilities of both extracting the discriminate connotations and reducing large redundancy in the HMM transition process. It is verified that our framework can achieve online recognition of continuous gesture streams even when they are halfway performed. Haoyu Chen 0001, Xin Liu 0012, Jingang Shi, Guoying Zhao 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | 3D Skeletal Gesture Recognition via Hidden States ExplorationabstractTemporal dynamics is an open issue for modeling human body gestures. A solution is resorting to the generative models, such as the hidden Markov model (HMM). Nevertheless, most of the work assumes fixed anchors for each hidden state, which make it hard to describe the explicit temporal structure of gestures. Based on the observation that a gesture is a time series with distinctly defined phases, we propose a new formulation to build temporal compositions of gestures by the low-rank matrix decomposition. The only assumption is that the gesture's "hold" phases with static poses are linearly correlated among each other. As such, a gesture sequence could be segmented into temporal states with semantically meaningful and discriminative concepts. Furthermore, different to traditional HMMs which tend to use specific distance metric for clustering and ignore the temporal contextual information when estimating the emission probability, we utilize the long short-term memory to learn probability distributions over states of HMM. The proposed method is validated on multiple challenging datasets. Experiments demonstrate that our approach can effectively work on a wide range of gestures, and achieve state-of-the-art performance. Xin Liu 0012, Henglin Shi, Xiaopeng Hong, Haoyu Chen 0001, Dacheng Tao, Guoying Zhao 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Advancing Image Understanding in Poor Visibility Environments: A Collective Benchmark StudyabstractExisting enhancement methods are empirically expected to help the high-level end computer vision task: however, that is observed to not always be the case in practice. We focus on object or face detection in poor visibility enhancements caused by bad weathers (haze, rain) and low light conditions. To provide a more thorough examination and fair comparison, we introduce three benchmark sets collected in real-world hazy, rainy, and low-light conditions, respectively, with annotated objects/faces. We launched the UG2+ challenge Track 2 competition in IEEE CVPR 2019, aiming to evoke a comprehensive discussion and exploration about whether and how low-level vision techniques can benefit the high-level automatic visual recognition in various scenarios. To our best knowledge, this is the first and currently largest effort of its kind. Baseline results by cascading existing enhancement and detection models are reported, indicating the highly challenging nature of our new data as well as the large room for further technical innovations. Thanks to a large participation from the research community, we are able to analyze representative team solutions, striving to better identify the strengths and limitations of existing mindsets as well as the future directions. Wenhan Yang, Ye Yuan 0012, Wenqi Ren, Jiaying Liu 0001, Walter J. Scheirer, Zhangyang Wang, Taiheng Zhang, Qiaoyong Zhong, Di Xie, Shiliang Pu, Yuqiang Zheng, Yanyun Qu, Yuhong Xie, Hao Jiang 0014, Siyuan Yang 0001, Yan Liu 0041, Xiaochao Qu, Pengfei Wan 0001, Shuai Zheng 0005, Minhui Zhong, Taiyi Su, Lingzhi He, Yandong Guo, Yao Zhao 0001, Zhenfeng Zhu, Jinxiu Liang, Jingwen Wang 0003, Yuhui Quan, Yong Xu 0007, Bo Liu 0112, Xin Liu 0012, Tingyu Lin 0003, Xiaochuan Li 0001, Feng Lu 0005, Lin Gu 0003, Shengdi Zhou, Cong Cao 0005, Cheng Chi 0003, Chubin Zhuang, Zhen Lei 0001, Stan Z. Li, Shizheng Wang, Ruizhe Liu, Dong Yi, Zheming Zuo, Jianning Chi, Huan Wang 0014, Kai Wang 0036, Yixiu Liu, Xingyu Gao 0001, Zhenyu Chen 0003, Yongzhou Li, Huicai Zhong, Jing Huang 0017, Heng Guo 0003, Jianfei Yang 0001, Wenjuan Liao, Jiangang Yang, Liguo Zhou, Mingyue Feng, Likun Qin |
IEEE Trans. Image Process. | 35 |
| 2019 | Analyze Spontaneous Gestures for Emotional Stress State Recognition: A Micro-gesture Dataset and Analysis with Deep LearningabstractEmotions are central for human intelligence and should have a similar role in AI. When it comes to emotion recognition, however, analysis cues for robots were mostly limited to human facial expressions and speech. As an alternative important non-verbal communicative fashion, the body gesture is proved to be capable of conveying emotional information which should gain more attention. Inspired by recent researches on micro-expressions, in this paper, we try to explore a specific group of gestures which are spontaneously and unconsciously elicited by inner feelings. These gestures are different from common gestures for facilitating communications or to express feelings on ones own initiative and always ignored in our daily life. This kind of subtle body movements is known as `micro-gestures' (MGs). Work of interpreting the human hidden emotions via these specific gestural behaviors in unconstrained situations, however, is limited. It is because of an unclear correspondence between body movements and emotional states which need multidisciplinary efforts from computer science, psychology, and statistic researchers. To fill the gap, we built a novel Spontaneous Micro-Gesture (SMG) dataset containing 3,692 manually labeled gesture clips. The data collection from 40 participants was conducted through a story-telling game with two emotional state settings. In this paper, we explored the emotional gestures with a sign-based measurement. To verify the latent relationship between emotional states and MGs, we proposed a framework that encodes the objective gestures to a Bayesian network to infer the subjective emotional states. Our experimental results revealed that, most of the participants would do `micro-gestures' spontaneously to relieve their mental strains. We also carried out a human test on ordinary and trained people for comparison. The performance of both our framework and human beings was evaluated on 142 testing instances (71 for each emotional state) by subject-independent testing. To authors' best knowledge, this is the first presented MG dataset. Results showed that the proposed MG recognition method achieved promising performance. We also showed that MGs could be helpful cues for the recognition of hidden emotional states. Haoyu Chen 0001, Xin Liu 0012, Henglin Shi, Guoying Zhao 0001 |
FG | 2 |
| 2019 | 3D Skeletal Gesture Recognition via Sparse Coding of Time-Warping Invariant Riemannian Trajectories
Xin Liu 0012, Guoying Zhao 0001 |
MMM (1) | 1 |
| 2019 | Hidden States Exploration for 3D Skeleton-Based Gesture Recognitionabstract3D skeletal data has recently attracted wide attention in human behavior analysis for its robustness to variant scenes, while accurate gesture recognition is still challenging. The main reason lies in the high intra-class variance caused by temporal dynamics. A solution is resorting to the generative models, such as the hidden Markov model (HMM). However, existing methods commonly assume fixed anchors for each hidden state, which is hard to depict the explicit temporal structure of gestures. Based on the observation that a gesture is a time series with distinctly defined phases, we propose a new formulation to build temporal compositions of gestures by the low-rank matrix decomposition. The only assumption is that the gesture's "hold" phases with static poses are linearly correlated among each other. As such, a gesture sequence could be segmented into temporal states with semantically meaningful and discriminative concepts. Furthermore, different to traditional HMMs which tend to use specific distance metric for clustering and ignore the temporal contextual information when estimating the emission probability, the Long Short-Term Memory (LSTM) is utilized to learn probability distributions over states of HMM. The proposed method is validated on two challenging datasets. Experiments demonstrate that our approach can effectively work on a wide range of gestures and actions, and achieve state-of-the-art performance. Xin Liu 0012, Henglin Shi, Xiaopeng Hong, Haoyu Chen 0001, Dacheng Tao, Guoying Zhao 0001 |
WACV | 1 |
| 2019 | Discriminative Spatiotemporal Local Binary Pattern with Revisited Integral Projection for Spontaneous Facial Micro-Expression RecognitionabstractRecently, there have been increasing interests in inferring mirco-expression from facial image sequences. Due to subtle facial movement of micro-expressions, feature extraction has become an important and critical issue for spontaneous facial micro-expression recognition. Recent works used spatiotemporal local binary pattern (STLBP) for micro-expression recognition and considered dynamic texture information to represent face images. However, they miss the shape attribute of face images. On the other hand, they extract the spatiotemporal features from the global face regions while ignore the discriminative information between two micro-expression classes. The above-mentioned problems seriously limit the application of STLBP to micro-expression recognition. In this paper, we propose a discriminative spatiotemporal local binary pattern based on an integral projection to resolve the problems of STLBP for micro-expression recognition. First, we revisit an integral projection for preserving the shape attribute of micro-expressions by using robust principal component analysis. Furthermore, a revisited integral projection is incorporated with local binary pattern across spatial and temporal domains. Specifically, we extract the novel spatiotemporal features incorporating shape attributes into spatiotemporal texture features. For increasing the discrimination of micro-expressions, we propose a new feature selection based on Laplacian method to extract the discriminative information for facial micro-expression recognition. Intensive experiments are conducted on three availably published micro-expression databases including CASME, CASME2 and SMIC databases. We compare our method with the state-of-the-art algorithms. Experimental results demonstrate that our proposed method achieves promising performance for micro-expression recognition. Xiaohua Huang 0003, Xin Liu 0012, Guoying Zhao 0001, Xiaoyi Feng, Matti Pietikäinen |
IEEE Trans. Affect. Comput. | 3 |
| 2019 | Saliency Integration: An Arbitrator ModelabstractSaliency integration has attracted much attention on unifying saliency maps from multiple saliency models. Previous offline integration methods usually face two challenges: 1) if most of the candidate saliency models misjudge the saliency on an image, the integration result will lean heavily on those inferior candidate models; and 2) an unawareness of the ground truth saliency labels brings difficulty in estimating the expertise of each candidate model. To address these problems, in this paper, we propose an arbitrator model (AM) for saliency integration. First, we incorporate the consensus of multiple saliency models and the external knowledge into a reference map to effectively rectify the misleading by candidate models. Second, our quest for ways of estimating the expertise of the saliency models without ground truth labels gives rise to two distinct online model-expertise estimation methods. Finally, we derive a Bayesian integration framework to reconcile the saliency models of varying expertise and the reference map. To extensively evaluate the proposed AM model, we test 27 state-of-the-art saliency models, covering both traditional and deep learning ones, on various combinations over four datasets. The evaluation results show that the AM model improves the performance substantially compared to the existing state-of-the-art integration methods, regardless of the chosen candidate saliency models. Yingyue Xu, Xiaopeng Hong, Fatih Porikli, Xin Liu 0012, Jie Chen 0001, Guoying Zhao 0001 |
IEEE Trans. Multim. | 4 |
| 2018 | Bidirectional Long Short-Term Memory Variational Autoencoder
Henglin Shi, Xin Liu 0012, Xiaopeng Hong, Guoying Zhao 0001 |
BMVC | 2 |
| 2018 | Temporal Hierarchical Dictionary with HMM for Fast Gesture RecognitionabstractIn this paper, we propose a novel temporal hierarchical dictionary with hidden Markov model (HMM) for gesture recognition task. Dictionaries with spatio-temporal elements have been commonly used for gesture recognition. However, the existing spatio-temporal dictionary based methods need the whole pre-segmented gestures for inference, thus are hard to deal with nonstationary sequences. The proposed method combines HMM with Deep Belief Networks (DBN) to tackle both gesture segmentation and recognition by the inference at the frame level. Besides, we investigate the redundancy in dictionaries and introduce the relative entropy to measure the information richness of a dictionary. Furthermore, when inferring an element, a temporal hierarchy-flat dictionary will be searched entirely every time in which the temporal structure of gestures isn't utilized sufficiently. The proposed temporal hierarchical dictionary is organized in HMM states and can limit the search range to distinct states. Our framework includes three key novel properties: (1) a temporal hierarchical structure with HMM, which makes both the HMM transition and Viterbi decoding more efficient; (2) a relative entropy model to compress the dictionary with less redundancy; (3) an unsupervised hierarchical clustering algorithm to build a hierarchical dictionary automatically. Our method is evaluated on two gesture datasets and consistently achieves state-of-the-art performance. The results indicate that the dictionary redundancy has a significant impact on the performance which can be tackled by a temporal hierarchy and an entropy model. Haoyu Chen 0001, Xin Liu 0012, Guoying Zhao 0001 |
ICPR | 2 |
| 2018 | Group sparsity residual constraint for image denoising with external nonlocal self-similarity prior
Zhiyuan Zha, Xinggan Zhang, Qiong Wang 0002, Yechao Bai, Lan Tang, Xin Liu 0012 |
Neurocomputing | 7 |
| 2018 | Group-based sparse representation for image compressive sensing reconstruction with non-convex regularization
Zhiyuan Zha, Xinggan Zhang, Qiong Wang 0002, Lan Tang, Xin Liu 0012 |
Neurocomputing | 5 |
| 2018 | Non-convex weighted ℓp nuclear norm based ADMM framework for image restoration
Zhiyuan Zha, Xinggan Zhang, Yu Wu 0022, Qiong Wang 0002, Xin Liu 0012, Lan Tang, Xin Yuan 0002 |
Neurocomputing | 5 |
| 2018 | Saliency detection via bi-directional propagation
Yingyue Xu, Xiaopeng Hong, Xin Liu 0012, Guoying Zhao 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2018 | Background Subtraction Using Spatio-Temporal Group Sparsity RecoveryabstractBackground subtraction is a key step in a wide spectrum of video applications, such as object tracking and human behavior analysis. Compressive sensing-based methods, which make little specific assumptions about the background, have recently attracted wide attention in background subtraction. Within the framework of compressive sensing, background subtraction is solved as a decomposition and optimization problem, where the foreground is typically modeled as pixel-wised sparse outliers. However, in real videos, foreground pixels are often not randomly distributed, but instead, group clustered. Moreover, due to costly computational expenses, most compressive sensing-based methods are unable to process frames online. In this paper, we take into account the group properties of foreground signals in both spatial and temporal domains, and propose a greedy pursuit-based method called spatio-temporal group sparsity recovery, which prunes data residues in an iterative process, according to both sparsity and group clustering priors, rather than merely sparsity. Furthermore, a random strategy for background dictionary learning is used to handle complex background variations, while foreground-free training is not required. Finally, we propose a two-pass framework to achieve online processing. The proposed method is validated on multiple challenging video sequences. Experiments demonstrate that our approach effectively works on a wide range of complex scenarios and achieves a state-of-the-art performance with far fewer computations. Xin Liu 0012, Jiawen Yao, Xiaopeng Hong, Xiaohua Huang 0003, Ziheng Zhou 0003, Chun Qi, Guoying Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Hallucinating Face Image by Regularization Models in High-Resolution Feature SpaceabstractIn this paper, we propose two novel regularization models in patch-wise and pixel-wise respectively, which are efficient to reconstruct high-resolution (HR) face image from low-resolution (LR) input. Unlike the conventional patch-based models which depend on the assumption of local geometry consistency in LR and HR spaces, the proposed method directly regularizes the relationship between the target patch and corresponding training set in the HR space. It avoids to deal with the tough problem of preserving local geometry in various resolutions. Taking advantage of kernel function in efficiently describing intrinsic features, we further conduct the patch-based reconstruction model in the high-dimensional kernel space for capturing nonlinear characteristics. Meanwhile, a pixel-based model is proposed to regularize the relationship of pixels in the local neighborhood, which can be employed to enhance the fuzzy details in the target HR face image. It privileges the reconstruction of pixels along the dominant orientation of structure, which is useful for preserving high-frequency information on complex edges. Finally, we combine the two reconstruction models into a unified framework. The output HR face image can be finally optimized by performing an iterative procedure. Experimental results demonstrate that the proposed face hallucination method produces superior performance than the state-of-the-art methods. Jingang Shi, Xin Liu 0012, Yuan Zong, Chun Qi, Guoying Zhao 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Compressed sensing image reconstruction via adaptive sparse nonlocal regularization
Zhiyuan Zha, Xin Liu 0012, Xinggan Zhang, Lan Tang, Yechao Bai, Qiong Wang 0002, Zhenhong Shang |
Vis. Comput. | 2 |
| 2017 | Image denoising via group sparsity residual constraintabstractGroup sparsity has shown great potential in various low-level vision tasks (e.g, image denoising, deblurring and inpainting). In this paper, we propose a new prior model for image denoising via group sparsity residual constraint (GSRC). To enhance the performance of group sparse-based image denoising, the concept of group sparsity residual is proposed, and thus, the problem of image denoising is translated into one that reduces the group sparsity residual. To reduce the residual, we first obtain some good estimation of the group sparse coefficients of the original image by the first-pass estimation of noisy image, and then centralize the group sparse coefficients of noisy image to the estimation. Experimental results have demonstrated that the proposed method not only outperforms many state-of-the-art denoising methods such as BM3D and WNNM, but results in a faster speed. Zhiyuan Zha, Xin Liu 0012, Ziheng Zhou 0003, Xiaohua Huang 0003, Jingang Shi, Zhenhong Shang, Lan Tang, Yechao Bai, Qiong Wang 0002, Xinggan Zhang |
ICASSP | 2 |
| 2017 | Estimation of signal-dependent noise level function using multi-column convolutional neural networkabstractTo estimate the levels of signal-dependent noise (SDN) from a single image is challenging. This paper proposes a novel method to estimate the noise level function (NLF) from a single image using a Multi-column Convolutional Neural Network (MC-Net) with an end-to-end architecture. The MC-Net is trained on a synthesized dataset containing noisy images with known NLFs, and it allows to learn rich hierarchical features using three sub-networks. Moreover, this method performs end-to-end training to retain more details for pixel-wise noise level estimation. Experimental results indicate that our method is accurate and robust to estimate NLFs of SDN for various types of images. Jing-Yu Yang 0002, Xin Liu 0012, Kun Li 0001 |
ICIP | 2 |
| 2017 | Analyzing the group sparsity based on the rank minimization methodsabstractSparse coding has achieved a great success in various image processing studies. However, there is not any benchmark to measure the sparsity of image patch/group because sparse discriminant conditions cannot keep unchanged. This paper analyzes the sparsity of group based on the strategy of the rank minimization. Firstly, an adaptive dictionary for each group is designed. Then, we prove that group-based sparse coding is equivalent to the rank minimization problem, and thus the sparse coefficients of each group are measured by estimating the singular values of each group. Based on that measurement, the weighted Schatten p-norm minimization (WSNM) has been found to be the closest solution to the real singular values of each group. Thus, WSNM can be equivalently transformed into a non-convex ℓp-norm minimization problem in group-based sparse coding. Experimental results on two applications: image in painting and image compressive sensing (CS) recovery show that the proposed scheme outperforms many state-of-the-art methods. Zhiyuan Zha, Xin Liu 0012, Xiaohua Huang 0003, Henglin Shi, Yingyue Xu, Qiong Wang 0002, Lan Tang, Xinggan Zhang |
ICME | 2 |
| 2015 | Background Subtraction Based on Low-Rank and Structured Sparse DecompositionabstractLow rank and sparse representation based methods, which make few specific assumptions about the background, have recently attracted wide attention in background modeling. With these methods, moving objects in the scene are modeled as pixel-wised sparse outliers. However, in many practical scenarios, the distributions of these moving parts are not truly pixel-wised sparse but structurally sparse. Meanwhile a robust analysis mechanism is required to handle background regions or foreground movements with varying scales. Based on these two observations, we first introduce a class of structured sparsity-inducing norms to model moving objects in videos. In our approach, we regard the observed sequence as being constituted of two terms, a low-rank matrix (background) and a structured sparse outlier matrix (foreground). Next, in virtue of adaptive parameters for dynamic videos, we propose a saliency measurement to dynamically estimate the support of the foreground. Experiments on challenging well known data sets demonstrate that the proposed approach outperforms the state-of-the-art methods and works effectively on a wide range of complex videos. Xin Liu 0012, Guoying Zhao 0001, Jiawen Yao, Chun Qi |
IEEE Trans. Image Process. | 1 |
| 2014 | Foreground detection using low rank and structured sparsityabstractIn this paper, a novel foreground detection method based on two-stage framework is presented. In the first stage, a class of structured sparsity-inducing norms is introduced to model moving objects in videos and thus regard the observed sequence as being made up of the sum of a low-rank matrix and a structured sparse outlier matrix. In virtue of adaptive parameters, the proposed method includes a motion saliency measurement to dynamically estimate the support of the foreground in the second stage. Experiments on challenging datasets demonstrate that the proposed approach outperforms the state-of-the-art methods and works effectively on a wide range of complex videos. Jiawen Yao, Xin Liu 0012, Chun Qi |
ICME | 2 |
| 2014 | Global consistency, local sparsity and pixel correlation: A unified framework for face hallucination
Jingang Shi, Xin Liu 0012, Chun Qi |
Pattern Recognit. | 2 |
| 2013 | Future-data driven modeling of complex backgrounds using mixture of Gaussians
Xin Liu 0012, Chun Qi |
Neurocomputing | 1 |