VLDB 2026 Research / reviewers in the wild / expert
Haoliang Li
dblp:121/0867
· DBLP profile ↗
118ranked-venue papers
15as first author
93since 2021 · last 2026
0000-0002-8723-8112ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 56 · 7 first-author · 45 since 2021Graphics, computer vision, multimedia, augmented reality and games · 42 · 3 first-author · 29 since 2021Security and privacy · 17 · 2 first-author · 14 since 2021Databases, data management, data science and information retrieval · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 7 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Variation-Bounded Loss for Noise-Tolerant LearningabstractMitigating the negative impact of noisy labels has been a perennial issue in supervised learning. Robust loss functions have emerged as a prevalent solution to this problem. In this work, we introduce the Variation Ratio as a novel property related to the robustness of loss functions, and propose a new family of robust loss functions, termed Variation-Bounded Loss (VBL), which is characterized by a bounded variation ratio. We provide theoretical analyses of the variation radio, proving that a smaller variation ratio would lead to better robustness. Furthermore, we reveal that the variation ratio provides a feasible method to relax the symmetric condition and offers a more concise path to achieve the asymmetric condition. Based on the variation ratio, we reformulate several commonly used loss functions into a variation-bounded form for pract ical applications. Positive experiments on various datasets exhibit the effectiveness and flexibility of our approach. Jialiang Wang 0003, Xianming Liu 0005, Gangfeng Hu, Deming Zhai, Junjun Jiang, Haoliang Li |
AAAI | 7 |
| 2026 | Task-Aware LLM Routing with Multi-Level Task-Profile-Guided Data Synthesis for Cold-Start ScenariosabstractLarge language models (LLMs) exhibit substantial variability in performance and computational cost across tasks and queries, motivating routing systems that select models to meet user-specific cost--performance trade-offs. However, existing routers generalize poorly in cold-start scenarios where in-domain training data is unavailable. We address this limitation with a multi-level task-profile--guided data synthesis framework that constructs a hierarchical task taxonomy and produces diverse question--answer pairs to approximate the test-time query distribution. Building on this, we introduce TRouter, a task-type--aware router approach that models query-conditioned cost and performance via latent task-type variables, with prior regularization derived from the synthesized task taxonomy. This design enhances TRouter's routing utility under both cold-start and in-domain settings. Across multiple benchmarks, we show that our synthesis framework alleviates cold-start issues and that TRouter delivers effective LLM routing. ©2026 Association for Computational Linguistics Hui Liu 0036, Kecheng Chen, Jie Liu 0044, Wenya Wang 0001, Haoliang Li |
ACL (1) | 6 |
| 2026 | A Collaborative Adversarial Purification Framework with Perturbation-Insensitive Semantics-Guidance
Kaifeng Chen, Peisong He, Haoliang Li, Ke Xu 0003 |
ISCAS | 4 |
| 2026 | Learning Dynamic Graph Embeddings With Neural Controlled Differential EquationsabstractThis paper focuses on representation learning for dynamic graphs with temporal interactions. A fundamental issue is that both the graph structure and the nodes own their own dynamics, and their blending induces intractable complexity in the temporal evolution over graphs. Drawing inspiration from the recent progress of physical dynamic models in deep neural networks, we propose Graph Neural Controlled Differential Equations (GN-CDEs), a continuous-time framework that jointly models node embeddings and structural dynamics by incorporating a graph enhanced neural network vector field with a time-varying graph path as the control signal. Our framework exhibits several desirable characteristics, including the ability to express dynamics on evolving graphs without piecewise integration, the capability to calibrate trajectories with subsequent data, and robustness to missing observations. Empirical evaluation on a range of dynamic graph representation learning tasks demonstrates the effectiveness of our proposed approach in capturing the complex dynamics of dynamic graphs. Tiexin Qin, Benjamin Walker 0001, Terry J. Lyons, Hong Yan 0001, Haoliang Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Adversarially robust multimedia watermarking via data-centric optimization
Ziyuan Luo, Qi Song 0003, Haoliang Li, Anderson Rocha 0001, Renjie Wan |
Pattern Recognit. | 3 |
| 2026 | Exploring transferable inconsistencies with regional guidance for reference-based deepfake detection
Liyue Ming, Peisong He, Haoliang Li, Xinghao Jiang |
Pattern Recognit. | 4 |
| 2026 | When Video Compression Meets Multimodal Large Language Models: A Unified Paradigm for Cross-Modality Video CompressionabstractTraditional video compression methods perform well at high bitrates but struggle to preserve fine-grained semantic information at low bitrates. Recently, with the blossoming of Multimodal Large Language Models (MLLMs), Cross-modal compression techniques offer prospective solutions for improving video compression under low-bitrate conditions. In this paper, we propose a unified Cross-Modality Video Compression (CMVC) framework that integrates multimodal representations and video generative models. The encoder disentangles video into spatial and temporal components, which are mapped to compact cross modal representations using MLLMs. During decoding, different encoding-decoding modes are employed to acquire various video reconstruction qualities, including Text-Text-to-Video (TT2V) for semantic preservation and Image-Text-to-Video (IT2V) for perceptual consistency. Additionally, we elaborate on an efficient frame interpolation model using Low-Rank Adaptation (LoRA) to improve the perceptual quality. Experimental results demon strate that TT2V achieves effective semantic reconstruction, while IT2V ensures competitive perceptual consistency. These findings suggest the potential of leveraging multimodal priors to improve video compression, offering promising future research directions. Jinlong Li 0003, Kecheng Chen, Meng Wang 0017, Long Xu 0001, Haoliang Li, Nicu Sebe, Sam Kwong, Shiqi Wang 0001 |
IEEE Signal Process. Lett. | 6 |
| 2026 | Open-Set Deepfake Detection: A Parameter-Efficient Adaptation Method With Forgery Style MixtureabstractOpen-set face forgery detection poses significant security threats and presents substantial challenges for existing detection models. These detectors primarily have two limitations: they cannot generalize across unknown forgery domains or inefficiently adapt to new data. To address these issues, we introduce an approach that is both general and parameter-efficient for face forgery detection. Our method builds on the assumption that different forgery source domains exhibit distinct style statistics. Specifically, we design a forgery-style-mixture formulation that augments the diversity of forgery source domains, enhancing the model’s generalizability across unseen domains. In addition, previous methods typically require fully fine-tuning pretrained networks, consuming substantial time and computational resources. Drawing on recent advancements in vision transformers (ViT) for face forgery detection, we develop a parameter-efficient ViT-based detection model that includes lightweight forgery feature extraction modules and enables the model to extract global and local forgery clues simultaneously. We only optimize the inserted lightweight modules during training, maintaining the original ViT structure with its pre-trained weights. This training strategy effectively preserves the informative pre-trained knowledge while flexibly adapting the model to the task of Deepfake detection. Extensive experimental results demonstrate that the designed model achieves state-of-the-art generalizability with significantly reduced trainable parameters, representing an important step toward open-set Deepfake detection in the wild. Chenqi Kong, Anwei Luo, Peijun Bao, Haoliang Li, Renjie Wan, Zengwei Zheng, Anderson Rocha 0001, Alex Chichung Kot |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Dynamically Perceived Forgery Conditional Diffusion Model for Scientific Image Tampering LocalizationabstractRecently, image tampering localization techniques for scientific publications have attracted increasing attention due to the prevalence of data manipulation and the integrity issue of image content. However, existing methods are still inefficient to expose tampering traces in scientific images due to their unique properties, such as acquisition noise and ambiguous edges. To address these limitations, we propose a Dynamically Perceived Forgery Conditional Diffusion Model, which formulates the prediction of the localization mask as a noise-state aware denoising process. This process progressively localizes the tampered regions by involving time-step guidance to dynamically perceive tampering traces under the variation of diffusion noise, which is jointly controlled by two conditions, including a forgery condition with hierarchically aggregated forensic clues and an enhanced edge condition with multilevel spatial attention. To conduct dynamic controls efficiently, two conditions are fused and then applied to the denoising process via a channel-cross attention module. Furthermore, in the inference stage, a salient element ensemble-based sampling strategy is developed to further improve the reliability against undesired factors of scientific images. Extensive experiments have been conducted on several scientific image tampering datasets, compared with state-of-the-art methods, which demonstrates our superiority in aspects of intra-/cross-dataset evaluations and robustness against post-processing operations. Jialing Xu, Peisong He, Haoliang Li, Shiqi Wang 0001, Yi Zhang 0018, Xinghao Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | MoE-FFD: Mixture of Experts for Generalized and Parameter-Efficient Face Forgery DetectionabstractDeepfakes have recently raised significant trust issues and security concerns among the public. Compared to CNN-based face forgery detectors, ViT-based methods take advantage of the expressivity of transformers, achieving superior detection performance. However, these approaches still exhibit the following limitations: (1) Fully fine-tuning ViT-based models from ImageNet weights demands substantial computational and storage resources; (2) ViT-based methods struggle to capture local forgery clues, leading to model bias; (3) These methods limit their scope on only one or few face forgery features, resulting in limited generalizability. To tackle these challenges, this work introduces Mixture-of-Experts modules for Face Forgery Detection (MoE-FFD), a generalized yet parameter-efficient ViT-based approach. MoE-FFD only updates lightweight Low-Rank Adaptation (LoRA) and Adapter layers while keeping the ViT backbone frozen, thereby achieving parameter-efficient training. Moreover, MoE-FFD leverages the expressivity of transformers and local priors of CNNs to simultaneously extract global and local forgery clues. Additionally, novel MoE modules are designed to scale the model's capacity and smartly select optimal forgery experts, further enhancing forgery detection performance. Our proposed learning scheme can be seamlessly adapted to various transformer backbones in a plug-and-play manner. Extensive experimental results demonstrate that the proposed method achieves state-of-the-art face forgery detection performance with significantly reduced parameter overhead in cross-dataset, cross-manipulation, and robustness evaluations. Our ablation studies further validate the effectiveness of the designed components and the proposed learning scheme. The code is available at: https://github.com/LoveSiameseCat/MoE-FFD. Chenqi Kong, Anwei Luo, Peijun Bao, Yi Yu 0011, Haoliang Li, Zengwei Zheng, Shiqi Wang 0001, Alex Chichung Kot |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2026 | MantleMark: Migrating Watermarks From Multi-View Images to Radiance Fields via Frequency ModulationabstractMulti-view images are essential for modern radiance field reconstruction methods like Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS). While image watermarking is a crucial data protection and ownership verification technique, it faces unprecedented challenges in multi-view scenarios. Traditional 2D watermarking techniques often fail to maintain detectability in rendered views, while existing 3D watermarking methods are typically limited to specific reconstruction methods and require access to the reconstruction process. To address these limitations, we propose MantleMark, a watermarking framework that migrates watermarks from multi-view images to radiance fields via frequency modulation. Our key insight is constructing a mantle-like Frequency-domain Watermarking Representation in 3D frequency space, which can be projected to create view-dependent watermarking patterns. Relying upon the Fourier Projection-Slice Theorem, we embed these patterns through magnitude spectrum modulation in the image frequency domain, enabling watermarks to migrate into 3D representations. This approach ensures watermark detectability in rendered views regardless of the reconstruction methods used by adversaries. Extensive experiments demonstrate that our method achieves robust watermark detection while maintaining high visual quality across various radiance field-based reconstruction methods. Ziyuan Luo, Jun Liu 0036, Haoliang Li, Anderson Rocha 0001, Renjie Wan |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2026 | Propose and Rectify: A Forensics-Driven MLLM Framework for Image Manipulation LocalizationabstractThe increasing sophistication of image manipulation techniques demands robust forensic solutions that can both reliably detect alterations and precisely localize tampered regions. Recent Multimodal Large Language Models (MLLMs) show promise by leveraging world knowledge and semantic understanding for context-aware detection, yet they struggle with perceiving subtle, low-level forensic artifacts crucial for accurate manipulation localization. This paper presents a novel Propose-Rectify framework that effectively bridges semantic reasoning with forensic-specific analysis. In the proposal stage, our approach utilizes a forensic-adapted LLaVA model to generate initial manipulation analysis and preliminary localization of suspicious regions based on semantic understanding and contextual reasoning. In the rectification stage, we introduce a Forensics Rectification Module that systematically validates and refines these initial proposals through multi-scale forensic feature analysis, integrating technical evidence from several specialized filters. Additionally, we present an Enhanced Segmentation Module that incorporates critical forensic cues into SAM’s encoded image embeddings, thereby overcoming inherent semantic biases to achieve precise delineation of manipulated regions. By synergistically combining advanced multimodal reasoning with established forensic methodologies, our framework ensures that initial semantic proposals are systematically validated and enhanced through concrete technical evidence, resulting in comprehensive detection accuracy and localization precision. Extensive experimental validation demonstrates state-of-the-art performance across diverse datasets with exceptional robustness and generalization capabilities. Keyang Zhang, Chenqi Kong, Hui Liu 0036, Bo Ding 0006, Xinghao Jiang, Haoliang Li |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2026 | Generalizable Dynamic Representation Learning for Source Identification in Sequential DataabstractSource identification is a foundational task in multimedia forensics, enabling the attribution and verification of digital content. While existing methods have achieved significant progress for static data, they often fail to generalize effectively on sequential data, which exhibit unique challenges such as temporal dependencies and dynamic variations caused by environmental and transmission factors. These challenges are further exacerbated in real-world scenarios, where crossdomain variations-spanning devices, software, and transmission protocols-significantly degrade the performance of traditional approaches. To address these limitations, we propose VoVAE, a probabilistic variational framework tailored for generalizable source identification in sequential data. VoVAE explicitly models temporal dependencies while disentangling dynamic variations (e.g., transmission distortions) from static source-specific features (e.g., device patterns) within a decoupled but complementary feature space. By separating these factors, VoVAE enables the extraction of robust and transferable representations, ensuring accurate source attribution across diverse and unseen conditions. We evaluate VoVAE on two challenging forensic applications: cross-domain VoIP phone call identification and cross-domain video source camera identification, using the VPCID and QUFVD datasets. Experimental results demonstrate that VoVAE outperforms state-of-the-art methods, achieving significant improvements in generalization across cross-device, cross-software, and cross-brand scenarios. Comprehensive ablation studies further highlight the importance of dynamic representation learning and feature disentanglement in capturing temporal patterns and enhancing robustness to domain shifts. These findings establish VoVAE as a scalable and robust solution for source identification in sequential data across diverse forensic scenarios. Bo Ding 0006, Tiexin Qin, Renjie Wan, Haoliang Li |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2025 | Unraveling the Mechanics of Learning-Based Demonstration Selection for In-Context LearningabstractHui Liu, Wenya Wang, Hao Sun, Chris Xing Tian, Chenqi Kong, Xin Dong, Haoliang Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Hui Liu 0036, Wenya Wang 0001, Chris Xing Tian, Chenqi Kong, Haoliang Li |
ACL (1) | 7 |
| 2025 | Q-PART: Quasi-Periodic Adaptive Regression with Test-time Training for Pediatric Left Ventricular Ejection Fraction RegressionabstractIn this work, we address the challenge of adaptive pediatric Left Ventricular Ejection Fraction (LVEF) assessment. While Test-time Training (TTT) approaches show promise for this task, they suffer from two significant limitations. Existing TTT works are primarily designed for classification tasks rather than continuous value regression, and they lack mechanisms to handle the quasi-periodic nature of cardiac signals. To tackle these issues, we propose a novel Quasi-Periodic Adaptive Regression with Test-time Training (Q-PART) framework. In the training stage, the proposed Quasi-Period Network decomposes the echocardiogram into periodic and aperiodic components within latent space by combining parameterized helix trajectories with Neural Controlled Differential Equations. During inference, our framework further employs a variance minimization strategy across image augmentations that simulate common quality issues in echocardiogram acquisition, along with differential adaptation rates for periodic and aperiodic components. Theoretical analysis is provided to demonstrate that our variance minimization objective effectively bounds the regression error under mild conditions. Furthermore, extensive experiments across three pediatric age groups demonstrate that Q-PART not only significantly outperforms existing approaches in pediatric LVEF prediction, but also exhibits strong clinical screening capability with high mAUROC scores (up to 0.9747) and maintains gender-fair performance across all metrics, validating its robustness and practical utility in pediatric echocardiography analysis. The project can be found in Q-PART. Jie Liu 0044, Tiexin Qin, Hui Liu 0036, Yilei Shi, Lichao Mou, Xiao Xiang Zhu 0001, Shiqi Wang 0001, Haoliang Li |
CVPR | 8 |
| 2025 | Test-time Adaptation for Foundation Medical Segmentation Model without Parametric UpdatesabstractFoundation medical segmentation models, with MedSAM being the most popular, have achieved promising performance across organs and lesions. However, MedSAM still suffers from compromised performance on specific lesions with intricate structures and appearance, as well as bounding box prompt-induced perturbations. Although current test-time adaptation (TTA) methods for medical image segmentation may tackle this issue, partial (e.g., batch normalization) or whole parametric updates restrict their effectiveness due to limited update signals or catastrophic forgetting in large models. Meanwhile, these approaches ignore the computational complexity during adaptation, which is particularly significant for modern foundation models. To this end, our theoretical analyses reveal that directly refining image embeddings is feasible to approach the same goal as parametric updates under the MedSAM architecture, which enables us to realize high computational efficiency and segmentation performance without the risk of catastrophic forgetting. Under this framework, we propose to encourage maximizing factorized conditional probabilities of the posterior prediction probability using a proposed distribution-approximated latent conditional random field loss combined with an entropy minimization loss. Experiments show that we achieve about 3\% Dice score improvements across three datasets while reducing computational complexity by over 7 times. Kecheng Chen, Xinyu Luo, Tiexin Qin, Jie Liu 0044, Hui Liu 0036, Victor Ho-fun Lee, Hong Yan 0001, Haoliang Li |
ICCV | 8 |
| 2025 | Temporal Unlearnable Examples: Preventing Personal Video Data from Unauthorized Exploitation by Object TrackingabstractWith the rise of social media, vast amounts of user-uploaded videos (e.g., YouTube) are utilized as training data for Visual Object Tracking (VOT). However, the VOT community has largely overlooked video data-privacy issues, as many private videos have been collected and used for training commercial models without authorization. To alleviate these issues, this paper presents the first investigation on preventing personal video data from unauthorized exploitation by deep trackers. Existing methods for preventing unauthorized data use primarily focus on image-based tasks (e.g., image classification), directly applying them to videos reveals several limitations, including inefficiency, limited effectiveness, and poor generalizability. To address these issues, we propose a novel generative framework for generating Temporal Unlearnable Examples (TUEs), and whose efficient computation makes it scalable for usage on large-scale video datasets. The trackers trained w/ TUEs heavily rely on unlearnable noises for temporal matching, ignoring the original data structure and thus ensuring training video data-privacy. To enhance the effectiveness of TUEs, we introduce a temporal contrastive loss, which further corrupts the learning of existing trackers when using our TUEs for training. Extensive experiments demonstrate that our approach achieves state-of-the-art performance in video data-privacy protection, with strong transferability across VOT models, datasets, and temporal matching tasks. Qiangqiang Wu, Yi Yu 0011, Chenqi Kong, Ziquan Liu, Jia Wan 0001, Haoliang Li, Alex Chichung Kot, Antoni B. Chan |
ICCV | 6 |
| 2025 | Test-time Adaptation for Image Compression with Distribution RegularizationabstractCurrent test- or compression-time adaptation image compression (TTA-IC) approaches, which leverage both latent and decoder refinements as a two-step adaptation scheme, have potentially enhanced the rate-distortion (R-D) performance of learned image compression models on cross-domain compression tasks, \textit{e.g.,} from natural to screen content images. However, compared with the emergence of various decoder refinement variants, the latent refinement, as an inseparable ingredient, is barely
tailored to cross-domain scenarios. To this end, we are interested in developing an advanced latent refinement method by extending the effective hybrid latent refinement (HLR) method, which is designed for \textit{in-domain} inference improvement but shows noticeable degradation of the rate cost in \textit{cross-domain} tasks. Specifically, we first provide theoretical analyses, in a cue of marginalization approximation from in- to cross-domain scenarios, to uncover that the vanilla HLR suffers from an underlying mismatch between refined Gaussian conditional and hyperprior distributions, leading to deteriorated joint probability approximation of marginal distribution with increased rate consumption. To remedy this issue, we introduce a simple Bayesian approximation-endowed \textit{distribution regularization} to encourage learning a better joint probability approximation in a plug-and-play manner. Extensive experiments on six in- and cross-domain datasets demonstrate that our proposed method not only improves the R-D performance compared with other latent refinement counterparts, but also can be flexibly integrated into existing TTA-IC methods with incremental benefits. Kecheng Chen, Tiexin Qin, Shiqi Wang 0001, Hong Yan 0001, Haoliang Li |
ICLR | 6 |
| 2025 | Deep Signature: Characterization of Large-Scale Molecular DynamicsabstractUnderstanding protein dynamics are essential for deciphering protein functional mechanisms and developing molecular therapies. However, the complex high-dimensional dynamics and interatomic interactions of biological processes pose significant challenge for existing computational techniques. In this paper, we approach this problem for the first time by introducing Deep Signature, a novel computationally tractable framework that characterizes complex dynamics and interatomic interactions based on their evolving trajectories. Specifically, our approach incorporates soft spectral clustering that locally aggregates cooperative dynamics to reduce the size of the system, as well as signature transform that collects iterated integrals to provide a global characterization of the non-smooth interactive dynamics. Theoretical analysis demonstrates that Deep Signature exhibits several desirable properties, including invariance to translation, near invariance to rotation, equivariance to permutation of atomic coordinates, and invariance under time reparameterization. Furthermore, experimental results on three benchmarks of biological processes verify that our approach can achieve superior performance compared to baseline methods. Tiexin Qin, Mengxu Zhu, Terry Lyons, Hong Yan 0001, Haoliang Li |
ICLR | 6 |
| 2025 | Remote Sensing Change Detection via Graph Compression and Adaptive Feature FusionabstractRemote sensing change detection identifies and locates surface changes from multi-temporal remote sensing images, supporting applications such as environmental monitoring, disaster assessment, and land use management. Current deep learning methods for high-resolution remote sensing change detection face challenges, including high computational costs for bi-temporal interactions and inadequate utilization of feature semantics. To address these challenges, we propose a novel method that deeply considers the physical significance of bi-temporal remote sensing images for accurate change detection. Specifically, we introduce a bi-temporal cross-perception approach based on graph compression. By leveraging spatial compaction in graph structures, this method employs cross-attention-based bi-temporal interactions to enhance cross-perception and reduce pseudo-changes caused by isolated semantic understanding. Additionally, we efficiently integrate features from hierarchical CNNs by separately considering low- and high-level features. For low-level features, we apply a contrastive learning-based semantic enhancement strategy to clearly differentiate change regions from background. For high-level features, we propose an adaptive bi-temporal feature fusion method to avoid the generalization issues of fixed learnable parameters. Experimental results on the CLCD and Google datasets demonstrate that our method outperforms baseline methods in both IoU and F1 scores. Jingxing Zhong, Haoliang Li, Renda Han, Wenxin Zhang 0005, Yutian You, Junhao Xiao 0002 |
IJCNN | 2 |
| 2025 | Large Language Models for Lossless Image Compression: Next-Pixel Prediction in Language Space is All You NeedabstractWe have recently witnessed that ''Intelligence" and `''Compression" are the two sides of the same coin, where the language large model (LLM) with unprecedented intelligence is a general-purpose lossless compressor for various data modalities. This attribute is particularly appealing to the lossless image compression community, given the increasing need to compress high-resolution images in the current streaming media era. Consequently, a spontaneous envision emerges: Can the compression performance of the LLM elevate lossless image compression to new heights? However, our findings indicate that the naive application of LLM-based lossless image compressors suffers from a considerable performance gap compared with existing state-of-the-art (SOTA) codecs on common benchmark datasets. In light of this, we are dedicated to fulfilling the unprecedented intelligence (compression) capacity of the LLM for lossless image compression tasks, thereby bridging the gap between theoretical and practical compression performance. Specifically, we propose P -LLM, a next-pixel prediction-based LLM, which integrates various elaborated insights and methodologies, \textit{e.g.,} pixel-level priors, the in-context ability of LLM, and a pixel-level semantic preservation strategy, to enhance the understanding capacity of pixel sequences for better next-pixel predictions. Extensive experiments on benchmark datasets demonstrate that P-LLM can beat SOTA classical and learned codecs. Kecheng Chen, Hui Liu 0036, Jie Liu 0044, Yibing Liu, Shiqi Wang 0001, Hong Yan 0001, Haoliang Li |
NeurIPS | 9 |
| 2025 | MedChain: Bridging the Gap Between LLM Agents and Clinical Practice with Interactive SequenceabstractClinical decision making (CDM) is a complex, dynamic process crucial to healthcare delivery, yet it remains a significant challenge for artificial intelligence systems. While Large Language Model (LLM)-based agents have been tested on general medical knowledge using licensing exams and knowledge question-answering tasks, their performance in the CDM in real-world scenarios is limited due to the lack of comprehensive benchmark that mirror actual medical practice. To address this gap, we present MedChain, a dataset of 12,163 clinical cases that covers five key stages of clinical workflow. MedChain distinguishes itself from existing benchmarks with three key features of real-world clinical practice: personalization, interactivity, and sequentiality. Further, to tackle real-world CDM challenges, we also propose MedChain-Agent, an AI system that integrates a feedback mechanism and a MedCase-RAG module to learn from previous cases and adapt its responses. MedChain-Agent demonstrates remarkable adaptability in gathering information dynamically and handling sequential clinical tasks, significantly outperforming existing approaches. The relevant dataset and code will be released upon acceptance of this paper. Jie Liu 0044, Wenxuan Wang 0001, Zizhan Ma, Guolin Huang, Yihang Su, Kao-Jung Chang, Haoliang Li, LinLin Shen, Michael R. Lyu, Wenting Chen |
NeurIPS | 7 |
| 2025 | SPACE: SPike-Aware Consistency Enhancement for Test-Time Adaptation in Spiking Neural NetworksabstractSpiking Neural Networks (SNNs), as a biologically plausible alternative to Artificial Neural Networks (ANNs), have demonstrated advantages in terms of energy efficiency, temporal processing, and biological plausibility. However, SNNs are highly sensitive to distribution shifts, which can significantly degrade their performance in real-world scenarios. Traditional test-time adaptation (TTA) methods designed for ANNs often fail to address the unique computational dynamics of SNNs, such as sparsity and temporal spiking behavior. To address these challenges, we propose SPike-Aware Consistency Enhancement (SPACE), the first source-free and single-instance TTA method specifically designed for SNNs. SPACE leverages the inherent spike dynamics of SNNs to maximize the consistency of spike-behavior-based local feature maps across augmented versions of a single test sample, enabling robust adaptation without requiring source data. We evaluate SPACE on multiple datasets. Furthermore, SPACE exhibits robust generalization across diverse network architectures, consistently enhancing the performance of SNNs on CNNs, Transformer, and ConvLSTM architectures. Experimental results show that SPACE outperforms state-of-the-art ANN methods while maintaining lower computational cost, highlighting its effectiveness and robustness for SNNs in real-world settings. The code will be available at https://github.com/ethanxyluo/SPACE. Xinyu Luo, Kecheng Chen, Pao-Sheng Sun, Chris Xing Tian, Arindam Basu, Haoliang Li |
NeurIPS | 6 |
| 2025 | Data-model interaction-driven transferable graph learning method for weak-shot onsite FTU health condition assessment
Jie Liu 0017, Haoliang Li, Ran Duan 0007, Zhongxu Hu, Tielin Shi |
Adv. Eng. Informatics | 3 |
| 2025 | Towards Data-Centric Face Anti-spoofing: Improving Cross-Domain Generalization via Physics-Based Data Synthesis
Rizhao Cai, Cecelia Soh, Zitong Yu, Haoliang Li, Wenhan Yang, Alex Chichung Kot |
Int. J. Comput. Vis. | 4 |
| 2025 | NPVForensics: Learning VA correlations in non-critical phoneme-viseme regions for deepfake detection
Yu Chen 0049, Yang Yu 0039, Haoliang Li, Wei Wang 0108, Yao Zhao 0001 |
Image Vis. Comput. | 4 |
| 2025 | Augmented Reality-Based Interactive Scheme for Robot-Assisted Percutaneous Renal Puncture NavigationabstractABSTRACT In this paper, we present an Augmented Reality (AR)‐based application combined with a robotic system for percutaneous renal puncture navigation interaction and demonstrate its technical feasibility. Our system provides an intuitive interaction scheme between the surgeon and the robot without the need for traditional external input devices, and applies an image‐target‐based 3D registration scheme to transform the coordinate system between Hololens2 and the robot without using additional tracking devices. Users can visualize the abdominal puncture phantom and obtain 3D depth information of the lesion site by wearing Hololens2 and control the robot directly using buttons or gestures. To investigate the accuracy and feasibility of the proposed interaction scheme, six subjects were recruited to complete 3D registration alignment accuracy experiments, and puncture positioning accuracy experiments using ultrasound unaided navigation, AR unaided navigation and AR robotic navigation. The results showed that the average alignment error of 3D registration was 3.61 ± 1.05 mm. The average positioning errors of ultrasound freehand navigation, AR freehand navigation and AR robotic navigation were 7.67 ± 2.00 mm, 6.13 ± 1.07 mm and 5.52 ± 0.37 mm, respectively; the average puncture times were 34.86 ± 1.67 s, 22.40 ± 2.07 s, and 29.41 ± 1.37 s. Yiwei Zhuang, Hua Xie, Wei Qing, Haoliang Li, Yuhan Shen, Yichun Shen |
Comput. Animat. Virtual Worlds | 5 |
| 2025 | Pixel-Inconsistency Modeling for Image Manipulation LocalizationabstractDigital image forensics plays a crucial role in image authentication and manipulation localization. Despite the progress powered by deep neural networks, existing forgery localization methodologies exhibit limitations when deployed to unseen datasets and perturbed images (i.e., lack of generalization and robustness to real-world applications). To circumvent these problems and aid image integrity, this paper presents a generalized and robust manipulation localization model through the analysis of pixel inconsistency artifacts. The rationale is grounded on the observation that most image signal processors (ISP) involve the demosaicing process, which introduces pixel correlations in pristine images. Moreover, manipulating operations, including splicing, copy-move, and inpainting, directly affect such pixel regularity. We, therefore, first split the input image into several blocks and design masked self-attention mechanisms to model the global pixel dependency in input images. Simultaneously, we optimize another local pixel dependency stream to mine local manipulation clues within input forgery images. In addition, we design novel Learning-to-Weight Modules (LWM) to combine features from the two streams, thereby enhancing the final forgery localization performance. To improve the training process, we propose a novel Pixel-Inconsistency Data Augmentation (PIDA) strategy, driving the model to focus on capturing inherent pixel-level artifacts instead of mining semantic forgery traces. This work establishes a comprehensive benchmark integrating 16 representative detection models across 12 datasets. Extensive experiments show that our method successfully extracts inherent pixel-inconsistency forgery fingerprints and achieve state-of-the-art generalization and robustness performances in image manipulation localization. Chenqi Kong, Anwei Luo, Shiqi Wang 0001, Haoliang Li, Anderson Rocha 0001, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | The NeRF Signature: Codebook-Aided Watermarking for Neural Radiance FieldsabstractNeural Radiance Fields (NeRF) have been gaining attention as a significant form of 3D content representation. With the proliferation of NeRF-based creations, the need for copyright protection has emerged as a critical issue. Although some approaches have been proposed to embed digital watermarks into NeRF, they often neglect essential model-level considerations and incur substantial time overheads, resulting in reduced imperceptibility and robustness, along with user inconvenience. In this paper, we extend the previous criteria for image watermarking to the model level and propose NeRF Signature, a novel watermarking method for NeRF. We employ a Codebook-aided Signature Embedding (CSE) that does not alter the model structure, thereby maintaining imperceptibility and enhancing robustness at the model level. Furthermore, after optimization, any desired signatures can be embedded through the CSE, and no fine-tuning is required when NeRF owners want to use new binary signatures. Then, we introduce a joint pose-patch encryption watermarking strategy to hide signatures into patches rendered from a specific viewpoint for higher robustness. In addition, we explore a Complexity-Aware Key Selection (CAKS) scheme to embed signatures in high visual complexity patches to enhance imperceptibility. The experimental results demonstrate that our method outperforms other baseline methods in terms of imperceptibility and robustness. Ziyuan Luo, Anderson Rocha 0001, Boxin Shi, Qing Guo 0005, Haoliang Li, Renjie Wan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Generalizing to New Dynamical Systems via Frequency Domain AdaptationabstractLearning the underlying dynamics from data with deep neural networks has shown remarkable potential in modeling various complex physical dynamics. However, current approaches are constrained in their ability to make reliable predictions in a specific domain and struggle with generalizing to unseen systems that are governed by the same general dynamics but differ in environmental characteristics. In this work, we formulate a parameter-efficient method, Fourier Neural Simulator for Dynamical Adaptation (FNSDA), that can readily generalize to new dynamics via adaptation in the Fourier space. Specifically, FNSDA identifies the shareable dynamics based on the known environments using an automatic partition in Fourier modes and learns to adjust the modes specific for each new environment by conditioning on low-dimensional latent systematic parameters for efficient generalization. We evaluate our approach on four representative families of dynamic systems, and the results show that FNSDA can achieve superior or competitive generalization performance compared to existing methods with a significantly reduced parameter cost. Our code is available at https://github.com/WonderSeven/FNSDA. Tiexin Qin, Hong Yan 0001, Haoliang Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Clean-Label Attack on Face Authentication Systems Through Rolling Shutter MechanismabstractWe introduce a novel clean-label black-box face presentation attack on face authentication systems, i.e., face recognition and verification systems, under mild conditions. Different from other clean-label attacks which require inserting complicated or intensity patterns after the image-capturing phase, our designed pattern can be automatically inserted during the exposure by utilizing the rolling shutter mechanism and modulating environment LEDs in a specialized waveform. This method provides a potential way to conduct backdoor attacks in the physical domain. Additionally, we propose an optimization strategy based on evolutionary computing to optimize the parameters of the stripe patterns, enhancing the attack success rate. The experimental results on several face recognition models and face verification services provided by the leading technology companies demonstrate the effectiveness of our attack method. Our study reveals a new attack applicable in the physical world, highlighting significant security concerns for existing face recognition, verification, and face anti-spoofing techniques. Yufei Wang 0006, Haoliang Li, Liepiao Zhang, Yongjian Hu, Alex Chichung Kot |
IEEE Signal Process. Lett. | 2 |
| 2025 | Generalized Time Series Classification via Component Decomposition and AlignmentabstractThe objective of domain generalization is to develop a model that can handle the domain shift problem without access to the target domain. In this paper, we propose a new domain generalization approach called Decomposition Framework with Dynamic Component Alignment (DFDCA), which employs signal decomposition on input data and conducts domain alignment on each component, providing another perspective on domain generalization for time series classification. Specifically, we first utilize a neural decomposition module to decompose the original time series data into several components, and design loss functions to guide the network to effectively perform signal decomposition for class-wise domain alignment on the decomposed components. The denoising attention mechanism is then introduced to enhance informative components while suppressing task-irrelevant components. Our proposed approach is evaluated on four publicly available datasets based on the cross-domain setting where the training and test samples are drawn from different distributions. The results demonstrate that it outperforms other baseline methods, achieving state-of-the-art performance. Yichuan Cheng, Darrick Lee, Harald Oberhauser, Haoliang Li |
IEEE Trans. Big Data | 4 |
| 2025 | Towards Extensible Detection of AI-Generated Images via Content-Agnostic Adapter-Based Category-Aware Incremental LearningabstractThe rapid evolution of image generation techniques has benefited several fields, but it has also given rise to security concerns. As countermeasures, a series of AI-generated image detection methods have been developed successfully. However, existing methods exhibit an inefficiency in handling the continual emergence of new generative models. To address this issue, we formulate the detection of AI-generated images in an extensible manner using an adapter-based domain incremental learning framework. Specifically, we first investigate the global consistency property of generation artifacts and design a content-agnostic adapter equipped on a vision transformer to extract common forensic features, where a token-level shuffling strategy is constructed for the dual-stream comparison to mitigate the fitting to specific image content. Then, motivated by the compactness of real images and the diversity of fake images due to their inherent generation processes, an asymmetric category-aware domain alignment method is designed to reduce the domain shift arisen from different generators. Finally, a multi-view knowledge distillation module, considering both point-to-point and structure-to-structure forensic knowledge, is devised to alleviate catastrophic forgetting. Experiments are conducted on several protocols using various image generators, and experimental results verify the superiority of our method compared to state-of-the-art methods for extensible detection. Shuai Tang 0001, Peisong He, Haoliang Li, Wei Wang 0108, Xinghao Jiang, Yao Zhao 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Image Provenance Analysis via Graph Encoding With Vision TransformerabstractRecent advances in AI-powered image editing tools have significantly lowered the barrier to image modification, raising pressing security concerns those related to spreading misinformation and disinformation on social platforms. Image provenance analysis is crucial in this context, as it identifies relevant images within a database and constructs a relationship graph by mining hidden manipulation and transformation cues, thereby providing concrete evidence chains. This paper introduces a novel end-to-end deep learning framework designed to explore the structural information of provenance graphs. Our proposed method distinguishes from previous approaches in two main ways. First, unlike earlier methods that rely on prior knowledge and have limited generalizability, our framework relies upon a patch attention mechanism to capture image provenance clues for local manipulations and global transformations, thereby enhancing graph construction performance. Second, while previous methods primarily focus on identifying tampering traces only between image pairs, they often overlook the hidden information embedded in the topology of the provenance graph. Our approach aligns the model training objectives with the final graph construction task, incorporating the overall structural information of the graph into the training process. We integrate graph structure information with the attention mechanism, enabling precise determination of the direction of transformation. Experimental results show the superiority of the proposed method over previous approaches, underscoring its effectiveness in addressing the challenges of image provenance analysis. Keyang Zhang, Chenqi Kong, Shiqi Wang 0001, Anderson Rocha 0001, Haoliang Li |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | A Novel AIoT-Based and User Behavior-Driven Dockless Bike-Sharing Management System for Chaotic Operations in a Condensed CityabstractRapid urbanization and the rising demand for sustainable mobility have accelerated the adoption of dockless bike-sharing systems. However, these systems frequently encounter challenges such as oversupply, inefficient resource allocation, and limited responsiveness to localized demand. To address these issues, this paper proposes the novel AIoT-based and user behavior-driven management system (AUMS)—a comprehensive AIoT-based framework that integrates demand prediction, dynamic clustering, and rebalancing algorithms to optimize the management of dockless bike-sharing services. Unlike existing approaches that focus narrowly on demand prediction, AUMS combines real-time IoT sensor data with user behavior analytics to support continuous, data-driven decision-making. The system employs a closed-loop control structure, enabling the dynamic reconfiguration of bike distribution in response to shifting urban mobility patterns. Its architecture includes three core modules: 1) demand prediction based on spatiotemporal behavioral data, 2) dynamic clustering to localize operational zones, and 3) intelligent rebalancing to minimize idle rates and unmet demand. The significance of this work lies in its holistic design and rigorous real-world validation. Over a 15-month period, AUMS was deployed in three escalating phases: a localized pilot, a district-level deployment in Tseung Kwan O, and a city-wide experiment across 13 districts in Hong Kong. Field results demonstrate notable performance improvements, including a 10.8% increase in bike utilization, 11.6% growth in trip frequency, and a 29.2% reduction in unmet demand. Ken Chun Ho Ching, Chaoqiang Jiang, Hassan Chun Wai Ching, Steve Kwan Po Ng, Ray C. C. Cheung, Haoliang Li, Alan H. F. Lam |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Critical Contour Prior-Guided Graph Learning With Pose Calibration for Identity-Aware Deepfake DetectionabstractDeepfake has recently raised severe public concerns about security issues, such as creating fake news of celebrities. As countermeasures, identity-aware detection methods leverage identity information to expose forged videos by measuring identity consistency between the suspicious input and its reference samples. However, the performance of existing methods suffers from notable degradation due to undesired variations of head poses and capturing environments. In this work, we first conduct a statistical analysis to illustrate the influence of different facial regions for forensic purposes, which infers more reliable identity information is located in critical face regions. Motivated by this analysis, we propose a graph learning-based identity-aware deepfake detection framework considering critical contour prior as guidance. First, feature sampling based on contour landmarks is applied to construct the graph data as the input of our critical contour prior-guided graph attention network (CP-GAT), where a node position prediction task is constructed as auxiliary supervision to explore rich relationships between nodes. To enhance pose-invariant ability, a rotation compensation block is integrated into CP-GAT and trained using a pose-calibrated contrastive learning to extract identity features, which takes high-quality front faces as the calibration goal with a progressively updating selection. Besides, an adversarial node masking-based training strategy is proposed as feature augmentation to further enhance the reliability. During the inference stage, the similarity between identity features of the input sample and its reference samples extracted by the trained CP-GAT is used to obtain the detection result. Extensive experiments are conducted on various face forgery datasets and state-of-the-art methods are compared to verify the superiority of the proposed method in terms of detection capability and robustness. Liyue Ming, Peisong He, Haoliang Li, Shiqi Wang 0001, Xinghao Jiang |
IEEE Trans. Multim. | 3 |
| 2025 | Unsupervised Domain Adaptation for Low-Dose CT Reconstruction via Bayesian Uncertainty AlignmentabstractLow-dose computed tomography (LDCT) image reconstruction techniques can reduce patient radiation exposure while maintaining acceptable imaging quality. Deep learning (DL) is widely used in this problem, but the performance of testing data (also known as target domain) is often degraded in clinical scenarios due to the variations that were not encountered in training data (also known as source domain). Unsupervised domain adaptation (UDA) of LDCT reconstruction has been proposed to solve this problem through distribution alignment. However, existing UDA methods fail to explore the usage of uncertainty quantification, which is crucial for reliable intelligent medical systems in clinical scenarios with unexpected variations. Moreover, existing direct alignment for different patients would lead to content mismatch issues. To address these issues, we propose to leverage a probabilistic reconstruction framework to conduct a joint discrepancy minimization between source and target domains in both the latent and image spaces. In the latent space, we devise a Bayesian uncertainty alignment to reduce the epistemic gap between the two domains. This approach reduces the uncertainty level of target domain data, making it more likely to render well-reconstructed results on target domains. In the image space, we propose a sharpness-aware distribution alignment (SDA) to achieve a match of second-order information, which can ensure that the reconstructed images from the target domain have similar sharpness to normal-dose CT (NDCT) images from the source domain. Experimental results on two simulated datasets and one clinical low-dose imaging dataset show that our proposed method outperforms other methods in quantitative and visualized performance. Kecheng Chen, Jie Liu 0044, Renjie Wan, Victor Ho-fun Lee, Varut Vardhanabhuti, Hong Yan 0001, Haoliang Li |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | Imaging Interiors: An Implicit Solution to Electromagnetic Inverse Scattering Problems
Ziyuan Luo, Boxin Shi, Haoliang Li, Renjie Wan |
ECCV (7) | 3 |
| 2024 | Event Trojan: Asynchronous Event-Based Backdoor Attacks
Ruofei Wang, Qing Guo 0005, Haoliang Li, Renjie Wan |
ECCV (7) | 3 |
| 2024 | Target-agnostic Source-free Domain Adaptation for Regression TasksabstractUnsupervised domain adaptation (UDA) seeks to bridge the domain gap between the target and source using unlabeled target data. Source-free UDA removes the requirement for labeled source data at the target to preserve data privacy and storage. However, previous works on source-free UDA assume knowledge of domain gap, and hence is limited to either target-aware or classification task. To overcome it, we propose TASFAR, a novel target-agnostic source-free domain adaptation approach for regression tasks. Using prediction confidence, TASFAR estimates a label density map as the target label distribution, which is then used to calibrate the source model on the target domain. We have conducted extensive experiments on four regression tasks with various domain gaps, namely, pedestrian dead reckoning for different users, image-based people counting in different scenes, housing-price prediction at different districts, and taxi-trip duration prediction from different departure points. TASFAR demonstrates significant superiority over state-of-the-art source-free UDA approaches, achieving an average error reduction of 22 % across the four tasks and comparable accuracy to source-based UDA, all without relying on source data. Tianlang He, Jierun Chen, Haoliang Li, Shueng-Han Gary Chan |
ICDE | 4 |
| 2024 | Neuron Activation Coverage: Rethinking Out-of-distribution Detection and GeneralizationabstractThe out-of-distribution (OOD) problem generally arises when neural networks encounter data that significantly deviates from the training data distribution, i.e., in-distribution (InD). In this paper, we study the OOD problem from a neuron activation view. We first formulate neuron activation states by considering both the neuron output and its influence on model decisions. Then, to characterize the relationship between neurons and OOD issues, we introduce the *neuron activation coverage* (NAC) -- a simple measure for neuron behaviors under InD data. Leveraging our NAC, we show that 1) InD and OOD inputs can be largely separated based on the neuron behavior, which significantly eases the OOD detection problem and beats the 21 previous methods over three benchmarks (CIFAR-10, CIFAR-100, and ImageNet-1K). 2) a positive correlation between NAC and model generalization ability consistently holds across architectures and datasets, which enables a NAC-based criterion for evaluating model robustness. Compared to prevalent InD validation criteria, we show that NAC not only can select more robust models, but also has a stronger correlation with OOD test performance. Yibing Liu, Chris Xing Tian, Haoliang Li, Lei Ma 0003, Shiqi Wang 0001 |
ICLR | 3 |
| 2024 | Log Neural Controlled Differential Equations: The Lie Brackets Make A DifferenceabstractThe vector field of a controlled differential equation (CDE) describes the relationship between a control path and the evolution of a solution path. Neural CDEs (NCDEs) treat time series data as observations from a control path, parameterise a CDE’s vector field using a neural network, and use the solution path as a continuously evolving hidden state. As their formulation makes them robust to irregular sampling rates, NCDEs are a powerful approach for modelling real-world data. Building on neural rough differential equations (NRDEs), we introduce Log-NCDEs, a novel, effective, and efficient method for training NCDEs. The core component of Log-NCDEs is the Log-ODE method, a tool from the study of rough paths for approximating a CDE’s solution. Log-NCDEs are shown to outperform NCDEs, NRDEs, the linear recurrent unit, S5, and MAMBA on a range of multivariate time series datasets with up to $50{,}000$ observations. Benjamin Walker 0001, Andrew D. McLeod, Tiexin Qin, Yichuan Cheng, Haoliang Li, Terry J. Lyons |
ICML | 5 |
| 2024 | Scenedoor: An Environmental Backdoor Attack for Face RecognitionabstractFace recognition is often used for biometric validation, which has become a significant technique in our society. Due to its sensitive applications, security vulnerabilities posed by backdoor attacks have attracted considerable focus. Current backdoor attack methods use digital perturbations or physical objects as triggers, while these additional requirements make existing backdoor attacks less viable in real-world applications. To address this issue, we propose a novel backdoor attack method named Scene Backdoor (Scenedoor), which injects a 3D scene as the trigger that effectively simplifies the backdoor activation. Any person who appears in this scene will be attacked as the attacker-desired identity. Specifically, we reconstruct a 3D scene from several 2D images and then blend the facial part extracted from the input sample with the reconstructed scene to generate the poisoned image. Extensive experiments are conducted on CelenDF (v2), CelebA-HQ, and PinsFace datasets, demonstrating that Scenedoor overtakes five state-of-the-art methods in terms of effectiveness, stealthiness, and robustness. Ruofei Wang, Ziyuan Luo, Haoliang Li, Renjie Wan |
VCIP | 3 |
| 2024 | Machine learning-driven high-fidelity ensemble surrogate modeling of Francis turbine unit based on data-model interactive simulation
Jie Liu 0017, Yanglong Lu, Haoliang Li |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | Data-model-interactive enhancement-based Francis turbine unit health condition assessment using graph driven health benchmark model
Jie Liu 0017, Haoliang Li, Xingxing Jiang |
Expert Syst. Appl. | 4 |
| 2024 | Domain Generalization with Small DataabstractAbstract In this work, we propose to tackle the problem of domain generalization in the context of insufficient samples. Instead of extracting latent feature embeddings based on deterministic models, we propose to learn a domain-invariant representation based on the probabilistic framework by mapping each data point into probabilistic embeddings. Specifically, we first extend empirical maximum mean discrepancy (MMD) to a novel probabilistic MMD that can measure the discrepancy between mixture distributions (i.e., source domains) consisting of a series of latent distributions rather than latent points. Moreover, instead of imposing the contrastive semantic alignment (CSA) loss based on pairs of latent points, a novel probabilistic CSA loss encourages positive probabilistic embedding pairs to be closer while pulling other negative ones apart. Benefiting from the learned representation captured by probabilistic models, our proposed method can marriage the measurement on the distribution over distributions (i.e., the global perspective alignment) and the distribution-based contrastive semantic alignment (i.e., the local perspective alignment). Extensive experimental results on three challenging medical datasets show the effectiveness of our proposed method in the context of insufficient data compared with state-of-the-art methods. Kecheng Chen, Elena Gal, Hong Yan 0001, Haoliang Li |
Int. J. Comput. Vis. | 4 |
| 2024 | Multi-teacher Universal Distillation Based on Information Hiding for Defense Against Facial Manipulation
Xin Li 0122, Yao Zhao 0001, Yu Ni, Haoliang Li |
Int. J. Comput. Vis. | 5 |
| 2024 | Universal and extensible language-vision models for organ segmentation and tumor detection from abdominal computed tomography
Jie Liu 0044, Yixiao Zhang 0001, Kang Wang 0016, Mehmet Can Yavuz, Xiaoxi Chen, Yixuan Yuan, Haoliang Li, Yang Yang 0009, Alan L. Yuille, Yucheng Tang, Zongwei Zhou |
Medical Image Anal. | 7 |
| 2024 | Temporal Diversified Self-Contrastive Learning for Generalized Face Forgery DetectionabstractFace forgery detection receives widespread attention due to the great security threats arising from the development of face forgery technologies. Most existing works define it as a binary classification problem by modeling the spatial and temporal artifacts to distinguish real and fake videos. However, the detector tends to heavily rely on the binary labels and overfit method-specific forgery patterns of the training set, resulting in limited generalization ability. To mitigate this issue, we propose a Temporal Diversified Self-Contrastive Learning (TDSCL) framework, which guides the model to exploit generalized temporal inconsistencies for face forgery detection. Firstly, a Temporally Diversified Transformation (TDT) strategy is designed to create diverse training samples with multiple temporal scales. Subsequently, Short-term Self-contrastive Learning (STSC) and Long-term Self-contrastive Learning (LTSC) are proposed to perform temporal representations of the video at different temporal granularities to capture intrinsic and generalized forensics clues to expose fake videos, which can serve as auxiliary supervisions equipped with different backbones flexibly. Moreover, a Similarity-Guided Adaptive Fusion (SGAF) module is designed to adaptively reinforce the temporal inconsistencies for reliable classification. Extensive experiments verify that the proposed method achieves superior generalization ability over various state-of-the-art methods in different benchmark datasets. Rongchuan Zhang, Peisong He, Haoliang Li, Shiqi Wang 0001, Yun Cao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | M$^{3}$3FAS: An Accurate and Robust MultiModal Mobile Face Anti-Spoofing SystemabstractFace presentation attacks (FPA), also known as face spoofing, have brought increasing concerns to the public through various malicious applications, such as financial fraud and privacy leakage. Therefore, safeguarding face recognition systems against FPA is of utmost importance. Although existing learning-based face anti-spoofing (FAS) models can achieve outstanding detection performance, they lack generalization capability and suffer significant performance drops in unforeseen environments. Many methodologies seek to use auxiliary modality data (e.g., depth and infrared maps) during the presentation attack detection (PAD) to address this limitation. However, these methods can be limited since (1) they require specific sensors such as depth and infrared cameras for data capture, which are rarely available on commodity mobile devices, and (2) they cannot work properly in practical scenarios when either modality is missing or of poor quality. In this paper, we devise an accurate and robustMultiModalMobileFaceAnti-Spoofing system namedM$^{3}$FASto overcome the issues above. The primary innovation of this work lies in the following aspects: (1) To achieve robust PAD, our system combines visual and auditory modalities using three commonly available sensors: camera, speaker, and microphone; (2) We design a novel two-branch neural network with three hierarchical feature aggregation modules to perform cross-modal feature fusion; (3). We propose a multi-head training strategy, allowing the model to output predictions from the vision, acoustic, and fusion heads, resulting in a more flexible PAD. Extensive experiments have demonstrated the accuracy, robustness, and flexibility of M$^{3}$FAS under various challenging experimental settings. The source code and dataset are available at:https://github.com/ChenqiKONG/M3FAS/. Chenqi Kong, Kexin Zheng, Yibing Liu, Shiqi Wang 0001, Anderson Rocha 0001, Haoliang Li |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2024 | PolSAR Ship Characterization and Robust Detection at Different Grazing Angles With Polarimetric Roll-Invariant FeaturesabstractPolarimetric synthetic aperture radar (PolSAR) plays an important role in remote sensing. As a valuable application, PolSAR ship detection receives great attention and obtains fruitful achievements recently. Since ship targets’ scattering responses are highly sensitive to the radar grazing angles, robust PolSAR ship detection still faces challenges especially the very low target to clutter ratio (TCR) phenomenon at large grazing angles. How to find stable polarimetric features for robust ship detection at different grazing angles becomes the key scientific problem. This work dedicates to this issue and the main idea is to explore the potentials of polarimetric roll-invariant features which are relatively independent to radar looking directions. The main contributions are threefold. First, quantitative investigations are conducted to disclose the variation law of ship targets and sea clutters’ polarimetric scattering mechanisms in terms of different grazing angles with electromagnetic computation data and PolSAR datasets. Second, the performances of polarimetric roll-invariant features are significantly examined with the TCR and the absolute difference (AD) indexes. Five optimal polarimetric roll-invariant features with stable TCRs over the wide grazing angle range are founded, which can clearly enhance the contrast between ship targets and sea clutters at large grazing angles. Finally, robust PolSAR ship detection approaches are established with the selected polarimetric roll-invariant features. Comparison studies with Gaofen-3 and Radarsat-2 PolSAR datasets of different grazing angles are carried out. Compared with traditional polarimetric features, the experimental results demonstrate that the selected polarimetric roll-invariant features exhibit superior detection performances in terms of both detection accuracy and detection robustness. Haoliang Li, Shen-Wen Liu, Si-Wei Chen 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | S-Adapter: Generalizing Vision Transformer for Face Anti-Spoofing With Statistical TokensabstractFace Anti-Spoofing (FAS) aims to detect malicious attempts to invade a face recognition system by presenting spoofed faces. State-of-the-art FAS techniques predominantly rely on deep learning models but their cross-domain generalization capabilities are often hindered by the domain shift problem, which arises due to different distributions between training and testing data. In this study, we develop a generalized FAS method under the Efficient Parameter Transfer Learning (EPTL) paradigm, where we adapt the pre-trained Vision Transformer models for the FAS task. During training, the adapter modules are inserted into the pre-trained ViT model, and the adapters are updated while other pre-trained parameters remain fixed. We find the limitations of previous vanilla adapters in that they are based on linear layers, which lack a spoofing-aware inductive bias and thus restrict the cross-domain generalization. To address this limitation and achieve cross-domain generalized FAS, we propose a novel Statistical Adapter (S-Adapter) that gathers local discriminative and statistical information from localized token histograms. To further improve the generalization of the statistical tokens, we propose a novel Token Style Regularization (TSR), which aims to reduce domain style variance by regularizing Gram matrices extracted from tokens across different domains. Our experimental results demonstrate that our proposed S-Adapter and TSR provide significant benefits in both zero-shot and few-shot cross-domain testing, outperforming state-of-the-art methods on several benchmark tests. We will release the source code upon acceptance. Rizhao Cai, Zitong Yu, Chenqi Kong, Haoliang Li, Changsheng Chen 0001, Yongjian Hu, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Robust Domain Misinformation Detection via Multi-Modal Feature AlignmentabstractSocial media misinformation harms individuals and societies and is potentialized by fast-growing multi-modal content (i.e., texts and images), which accounts for higher “credibility” than text-only news pieces. Although existing supervised misinformation detection methods have obtained acceptable performances in key setups, they may require large amounts of labeled data from various events, which can be time-consuming and tedious. In turn, directly training a model by leveraging a publicly available dataset may fail to generalize due to domain shifts between the training data (a.k.a. source domains) and the data from target domains. Most prior work on domain shift focuses on a single modality (e.g., text modality) and ignores the scenario where sufficient unlabeled target domain data may not be readily available in an early stage. The lack of data often happens due to the dynamic propagation trend (i.e., the number of posts related to fake news increases slowly before catching the public attention). We propose a novel robust domain and cross-modal approach (RDCM) for multi-modal misinformation detection. It reduces the domain shift by aligning the joint distribution of textual and visual modalities through an inter-domain alignment module and bridges the semantic gap between both modalities through a cross-modality alignment module. We also propose a framework that simultaneously considers application scenarios of domain generalization (in which the target domain data is unavailable) and domain adaptation (in which unlabeled target domain data is available). Evaluation results on two public multi-modal misinformation detection datasets (Pheme and Twitter Datasets) evince the superiority of the proposed model. Hui Liu 0036, Wenya Wang 0001, Anderson Rocha 0001, Haoliang Li |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Generalization Beyond Feature Alignment: Concept Activation-Guided Contrastive LearningabstractLearning invariant representations via contrastive learning has seen state-of-the-art performance in domain generalization (DG). Despite such success, in this paper, we find that its core learning strategy - feature alignment - could heavily hinder model generalization. Drawing insights in neuron interpretability, we characterize this problem from a neuron activation view. Specifically, by treating feature elements as neuron activation states, we show that conventional alignment methods tend to deteriorate the diversity of learned invariant features, as they indiscriminately minimize all neuron activation differences. This instead ignores rich relations among neurons - many of them often identify the same visual concepts despite differing activation patterns. With this finding, we present a simple yet effective approach, Concept Contrast (CoCo), which relaxes element-wise feature alignments by contrasting high-level concepts encoded in neurons. Our CoCo performs in a plug-and-play fashion, thus it can be integrated into any contrastive method in DG. We evaluate CoCo over four canonical contrastive methods, showing that CoCo promotes the diversity of feature representations and consistently improves model generalization capability. By decoupling this success through neuron coverage analysis, we further find that CoCo potentially invokes more meaningful neurons during training, thereby improving model learning. Yibing Liu, Chris Xing Tian, Haoliang Li, Shiqi Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | Privacy-Preserving Constrained Domain Generalization Via Gradient AlignmentabstractDeep neural networks (DNN) have demonstrated unprecedented success for various applications. However, due to the issue of limited dataset availability and the strict legal and ethical requirements for data privacy protection, the broad applications of DNN (e.g., medical imaging classification) with large-scale training data have been largely hindered, greatly constraining the model generalization capability. In this paper, we aim to tackle this problem by developing the privacy-preserving constrained domain generalization method, aiming to improve the generalization capability under the privacy-preserving condition. In particular, we propose to improve the information aggregation process on the centralized server side with a novel gradient alignment loss, expecting that the trained model can be better generalized to the “unseen” but related data. The rationale and effectiveness of our proposed method can be explained by connecting our proposed method with the Maximum Mean Discrepancy (MMD) which has been widely adopted as the distribution distance measure. Experimental results on three domain generalization benchmark datasets indicate that our method can achieve better cross-domain generalization capability compared to the state-of-the-art federated learning methods. Chris Xing Tian, Haoliang Li, Yufei Wang 0006, Shiqi Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2024 | Learning Robust Shape Regularization for Generalizable Medical Image SegmentationabstractGeneralizable medical image segmentation enables models to generalize to unseen target domains under domain shift issues. Recent progress demonstrates that the shape of the segmentation objective, with its high consistency and robustness across domains, can serve as a reliable regularization to aid the model for better cross-domain performance, where existing methods typically seek a shared framework to render segmentation maps and shape prior concurrently. However, due to the inherent texture and style preference of modern deep neural networks, the edge or silhouette of the extracted shape will inevitably be undermined by those domain-specific texture and style interferences of medical images under domain shifts. To address this limitation, we devise a novel framework with a separation between the shape regularization and the segmentation map. Specifically, we first customize a novel whitening transform-based probabilistic shape regularization extractor namely WT-PSE to suppress undesirable domain-specific texture and style interferences, leading to more robust and high-quality shape representations. Second, we deliver a Wasserstein distance-guided knowledge distillation scheme to help the WT-PSE to achieve more flexible shape extraction during the inference phase. Finally, by incorporating domain knowledge of medical images, we propose a novel instance-domain whitening transform method to facilitate a more stable training process with improved performance. Experiments demonstrate the performance of our proposed method on both multi-domain and single-domain generalization. Kecheng Chen, Tiexin Qin, Victor Ho-fun Lee, Hong Yan 0001, Haoliang Li |
IEEE Trans. Medical Imaging | 5 |
| 2024 | Disentangled Feature Representation for Few-Shot Image ClassificationabstractLearning the generalizable feature representation is critical to few-shot image classification. While recent works exploited task-specific feature embedding using meta-tasks for few-shot learning, they are limited in many challenging tasks as being distracted by the excursive features such as the background, domain, and style of the image samples. In this work, we propose a novel disentangled feature representation (DFR) framework, dubbed DFR, for few-shot learning applications. DFR can adaptively decouple the discriminative features that are modeled by the classification branch, from the class-irrelevant component of the variation branch. In general, most of the popular deep few-shot learning methods can be plugged in as the classification branch, thus DFR can boost their performance on various few-shot tasks. Furthermore, we propose a novel FS-DomainNet dataset based on DomainNet, for benchmarking the few-shot domain generalization (DG) tasks. We conducted extensive experiments to evaluate the proposed DFR on general, fine-grained, and cross-domain few-shot classification, as well as few-shot DG, using the corresponding four benchmarks, i.e., mini-ImageNet, tiered-ImageNet, Caltech-UCSD Birds 200-2011 (CUB), and the proposed FS-DomainNet. Thanks to the effective feature disentangling, the DFR-based few-shot classifiers achieved state-of-the-art results on all datasets. Hao Cheng 0016, Yufei Wang 0006, Haoliang Li, Alex Chichung Kot, Bihan Wen |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Rethinking Sensors Modeling: Hierarchical Information Enhanced Traffic ForecastingabstractWith the acceleration of urbanization, traffic forecasting has become an essential role in smart city construction. In the context of spatio-temporal prediction, the key lies in how to model the dependencies of sensors. However, existing works basically only consider the micro relationships between sensors, where the sensors are treated equally, and their macroscopic dependencies are neglected. In this paper, we argue to rethink the sensor's dependency modeling from two hierarchies: regional and global perspectives. Particularly, we merge original sensors with high intra-region correlation as a region node to preserve the inter-region dependency. Then, we generate representative and common spatio-temporal patterns as global nodes to reflect a global dependency between sensors and provide auxiliary information for spatio-temporal dependency learning. In pursuit of the generality and reality of node representations, we incorporate a Meta GCN to calibrate the regional and global nodes in the physical data space. Furthermore, we devise the cross-hierarchy graph convolution to propagate information from different hierarchies. In a nutshell, we propose a Hierarchical Information Enhanced Spatio-Temporal prediction method, HIEST, to create and utilize the regional dependency and common spatio-temporal patterns. Extensive experiments have verified the leading performance of our HIEST against state-of-the-art baselines. We publicize the code to ease reproducibility1. Qian Ma 0012, Zijian Zhang 0009, Xiangyu Zhao 0001, Haoliang Li, Yiqi Wang 0001, Zitao Liu 0001 |
CIKM | 4 |
| 2023 | Cross-Domain Object Classification Via Successive Subspace AlignmentabstractRecently, successive subspace learning (SSL)-based methods have shown to be effective for the task of visual object classification with mild data desire and mathematically transparent interpretable capability. However, existing SSL-based methods rely heavily on the data-centric subspace representations, leading to potential performance degradation problem in case of the domain shift between the training (a.k.a., source domain) and testing (a.k.a., target domain) data. To address this limitation, we propose an effective successive subspace learning method based on existing SSL-based methods. Specifically, we introduce a novel linear transformation layer to align eigenvectors in SSL module between source and target domains, as such, the discrepancy between source and target domains will be reduced, resulting in better cross-domain performance. The effectiveness of our proposed method is demonstrated on the Office-Caltech-10 and Office-31 benchmark datasets by using features extracted from pre-trained deep neural networks as input. Kecheng Chen, Haoliang Li, Hong Yan 0001 |
ICASSP | 2 |
| 2023 | Two-Branch Multi-Scale Deep Neural Network for Generalized Document Recapture Attack DetectionabstractThe image recapture attack is an effective image manipulation method to erase certain forensic traces, and when targeting on personal document images, it poses a great threat to the security of e-commerce and other web applications. Considering the current learning-based methods suffer from serious over-fitting problem, in this paper, we propose a novel two-branch deep neural network by mining better generalized recapture artifacts with a designed frequency filter bank and multi-scale cross-attention fusion module. In the extensive experiment, we show that our method can achieve better generalization capability compared with state-of-the-art techniques on different scenarios. Chenqi Kong, Shiqi Wang 0001, Haoliang Li |
ICASSP | 4 |
| 2023 | Rehearsal-Free Domain Continual Face Anti-Spoofing: Generalize More and Forget LessabstractFace Anti-Spoofing (FAS) is recently studied under the continual learning setting, where the FAS models are expected to evolve after encountering data from new domains. However, existing methods need extra replay buffers to store previous data for rehearsal, which becomes infeasible when previous data is unavailable because of privacy issues. In this paper, we propose the first rehearsal-free method for Domain Continual Learning (DCL) of FAS, which deals with catastrophic forgetting and unseen domain generalization problems simultaneously. For better generalization to unseen domains, we design the Dynamic Central Difference Convolutional Adapter (DCDCA) to adapt Vision Transformer (ViT) models during the continual learning sessions. To alleviate the forgetting of previous domains without using previous data, we propose the Proxy Prototype Contrastive Regularization (PPCR) to constrain the continual learning with previous domain knowledge from the proxy prototypes. Simulating practical DCL scenarios, we devise two new protocols which evaluate both generalization and anti-forgetting performance. Extensive experimental results show that our proposed method can improve the generalization performance in unseen domains and alleviate the catastrophic forgetting of previous knowledge. The code and protocol files are released on https://github.com/RizhaoCai/DCL-FAS-ICCV2023. Rizhao Cai, Yawen Cui, Zitong Yu, Haoliang Li, Yongjian Hu, Alex Chichung Kot |
ICCV | 5 |
| 2023 | Temporal Coherent Test Time Optimization for Robust Video Classification
Chenyu Yi, Siyuan Yang 0001, Yufei Wang 0006, Haoliang Li, Yap-Peng Tan, Alex Chichung Kot |
ICLR | 4 |
| 2023 | Trustworthy Machine Learning: Robustness, Generalization, and InterpretabilityabstractMachine learning is becoming increasingly important in today's world. Beyond its powerful performances, there has been an emerging concern about the trustworthiness of machine learning, including but not limited to: robustness to malicious attacks, generalization to unseen datasets, and interpretability to explain its outputs. Such concerns are even more urgent in some safety-critical applications such as medical diagnosis and autonomous driving. Trustworthy machine learning (TrustML) aims to tackle these challenges from the perspectives of theory, algorithm, and applications. In this tutorial, we will give a comprehensive introduction to the recent advance of trustworthy machine learning in robustness, generalization, and interpretability. We will cover their problem formulation, related research, popular algorithms, and successful applications. Additionally, we will also introduce some potential challenges for future research. We do hope that this tutorial will not only serve as a platform to understand TrustML, but also raise the awareness of everyone for more trustworthy applications. Jindong Wang 0001, Haoliang Li, Haohan Wang, Sinno Jialin Pan, Xing Xie 0001 |
KDD | 2 |
| 2023 | A Tutorial on Domain GeneralizationabstractWith the availability of massive labeled training data, powerful machine learning models can be trained. However, the traditional I.I.D. assumption that the training and testing data should follow the same distribution is often violated in reality. While existing domain adaptation approaches can tackle domain shift, it relies on the target samples for training. Domain generalization is a promising technology that aims to train models with good generalization ability to unseen distributions. In this tutorial, we will present the recent advance of domain generalization. Specifically, we introduce the background, formulation, and theory behind this topic. Our primary focus is on the methodology, evaluation, and applications. We hope this tutorial can draw interest of the community and provide a thorough review of this area. Eventually, more robust systems can be built for responsible AI. All tutorial materials and updates can be found online at https://dgresearch.github.io/. Jindong Wang 0001, Haoliang Li, Sinno Jialin Pan, Xing Xie 0001 |
WSDM | 2 |
| 2023 | Evolving Domain Generalization via Latent Structure-Aware Sequential AutoencoderabstractDomain generalization (DG) refers to the problem of generalizing machine learning systems to out-of-distribution (OOD) data with knowledge learned from several provided source domains. Most prior works confine themselves to stationary and discrete environments to tackle such generalization issue arising from OOD data. However, in practice, many tasks in non-stationary environments (e.g., autonomous-driving car system, sensor measurement) involve more complex and continuously evolving domain drift, emerging new challenges for model deployment. In this paper, we first formulate this setting as the problem of evolving domain generalization. To deal with the continuously changing domains, we propose MMD-LSAE, a novel framework that learns to capture the evolving patterns among domains for better generalization. Specifically, MMD-LSAE characterizes OOD data in non-stationary environments with two types of distribution shifts: covariate shift and concept shift, and employs deep autoencoder modules to infer their dynamics in latent space separately. In these modules, the inferred posterior distributions of latent codes are optimized to align with their corresponding prior distributions via minimizing maximum mean discrepancy (MMD). We theoretically verify that MMD-LSAE has the inherent capability to implicitly facilitate mutual information maximization, which can promote superior representation learning and improved generalization of the model. Furthermore, the experimental results on both synthetic and real-world datasets show that our proposed approach can consistently achieve favorable performance based on the evolving domain generalization setting. Tiexin Qin, Shiqi Wang 0001, Haoliang Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Neuron Coverage-Guided Domain GeneralizationabstractThis paper focuses on the domain generalization task where domain knowledge is unavailable, and even worse, only samples from a single domain can be utilized during training. Our motivation originates from the recent progresses in deep neural network (DNN) testing, which has shown that maximizing neuron coverage of DNN can help to explore possible defects of DNN (i.e., misclassification). More specifically, by treating the DNN as a program and each neuron as a functional point of the code, during the network training we aim to improve the generalization capability by maximizing the neuron coverage of DNN with the gradient similarity regularization between the original and augmented samples. As such, the decision behavior of the DNN is optimized, avoiding the arbitrary neurons that are deleterious for the unseen samples, and leading to the trained DNN that can be better generalized to out-of-distribution samples. Extensive studies on various domain generalization tasks based on both single and multiple domain(s) setting demonstrate the effectiveness of our proposed approach compared with state-of-the-art baseline methods. We also analyze our method by conducting visualization based on network dissection. The results further provide useful evidence on the rationality and effectiveness of our approach. Chris Xing Tian, Haoliang Li, Xiaofei Xie, Yang Liu 0003, Shiqi Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Benchmarking Single-Image Reflection Removal AlgorithmsabstractReflection removal has been discussed for more than decades. This paper aims to provide the analysis for different reflection properties and factors that influence image formation, an up-to-date taxonomy for existing methods, a benchmark dataset, and the unified benchmarking evaluations for state-of-the-art (especially learning-based) methods. Specifically, this paper presents a SIngle-image Reflection Removal Plus dataset “SIR$^{2+}$” with the new consideration for in-the-wild scenarios and glass with diverse color and unplanar shapes. We further perform quantitative and visual quality comparisons for state-of-the-art single-image reflection removal algorithms. Open problems for improving reflection removal algorithms are discussed at the end. Our dataset and follow-up update can be found athttps://reflectionremoval.github.io/sir2data/. Renjie Wan, Boxin Shi, Haoliang Li, Yuchen Hong, Ling-Yu Duan, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Dynamic Expert-Knowledge Ensemble for Generalizable Video Quality AssessmentabstractDespite the impressive progress of supervised methods in quality assessment for in- the-wild videos, models trained on one domain often fail to generalize well to others due to the domain shifts caused by distortion diversity and content variation. Domain generalizable video quality assessment (VQA) methods that can work across domains remain an open research challenge. Although combining more data following the mixed-domain training strategy can improve the generalization performance to a certain extent, the specific knowledge from each source domain, which could potentially be useful for improving unseen domain generalization, is ignored in this principle. Motivated by this, we propose a domain generalizable VQA method named Dynamic Ensemble of Expert-Knowledge (DEEK), a novel framework that dynamically exploits the expert-knowledge from each source domain to achieve a generalizable ensemble prediction. Specifically, based on the multiple experts each trained to specialize in a particular source domain, we aim to exploit complementary information provided by the expert-knowledge. We effectively train an ensemble model by proposing a quality-sensitive InfoNCE loss to regularize the collaborative training of all experts in the contrastive learning formulation, aiming to exploit complementary information provided by the expert-knowledge when forming the ensemble. By dynamically integrating the experts according to their relevances to the target data, these expert-knowledge could be leveraged for better generalization. Experiments on five VQA datasets verify that our approach outperforms the state-of-the-arts by large margins. Pengfei Chen 0003, Leida Li, Haoliang Li, Jinjian Wu, Weisheng Dong, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Asymmetric Modality Translation for Face Presentation Attack DetectionabstractFace presentation attack detection (PAD) is an essentialmeasure to protect face recognition systems from being spoofed by malicious users and has attracted great attention from both academia and industry. Although most of the existing methods can achieve desired performance to some extent, the generalization issue of face presentation attack detection under cross-domain settings (e.g., the setting of unseen attacks and varying illumination) remains to be solved. In this paper, we propose a novel framework based on asymmetric modality translation for face presentation attack detection in bi-modality scenarios. Under the framework, we establish connections between two modality images of genuine faces. Specifically, a novel modality fusion scheme is presented that the image of one modality is translated to the other one through an asymmetric modality translator, then fused with its corresponding paired image. The fusion result is fed as the input to a discriminator for inference. The training of the translator is supervised by an asymmetric modality translation loss. Besides, an illumination normalization module based on Pattern of Local Gravitational Force (PLGF) representation is used to reduce the impact of illumination variation. We conduct extensive experiments on three public datasets, which validate that our method is effective in detecting various types of attacks and achieves state-of-the-art performance under different evaluation protocols. Zhi Li 0054, Haoliang Li, Yongjian Hu, Kwok-Yan Lam, Alex Chichung Kot |
IEEE Trans. Multim. | 2 |
| 2022 | Low-Light Image Enhancement with Normalizing FlowabstractTo enhance low-light images to normally-exposed ones is highly ill-posed, namely that the mapping relationship between them is one-to-many. Previous works based on the pixel-wise reconstruction losses and deterministic processes fail to capture the complex conditional distribution of normally exposed images, which results in improper brightness, residual noise, and artifacts. In this paper, we investigate to model this one-to-many relationship via a proposed normalizing flow model. An invertible network that takes the low-light images/features as the condition and learns to map the distribution of normally exposed images into a Gaussian distribution. In this way, the conditional distribution of the normally exposed images can be well modeled, and the enhancement process, i.e., the other inference direction of the invertible network, is equivalent to being constrained by a loss function that better describes the manifold structure of natural images during the training. The experimental results on the existing benchmark datasets show our method achieves better quantitative and qualitative results, obtaining better-exposed illumination, less noise and artifact, and richer colors. Yufei Wang 0006, Renjie Wan, Wenhan Yang, Haoliang Li, Lap-Pui Chau, Alex Chichung Kot |
AAAI | 4 |
| 2022 | Towards Multi-Modal Sarcasm Detection via Hierarchical Congruity Modeling with Knowledge EnhancementabstractSarcasm is a linguistic phenomenon indicating a discrepancy between literal meanings and implied intentions.Due to its sophisticated nature, it is usually challenging to be detected from the text itself.As a result, multi-modal sarcasm detection has received more attention in both academia and industries.However, most existing techniques only modeled the atomic-level inconsistencies between the text input and its accompanying image, ignoring more complex compositions for both modalities.Moreover, they neglected the rich information contained in external knowledge, e.g., image captions.In this paper, we propose a novel hierarchical framework for sarcasm detection by exploring both the atomic-level congruity based on multi-head cross attention mechanism and the composition-level congruity based on graph neural networks, where a post with low congruity can be identified as sarcasm.In addition, we exploit the effect of various knowledge resources for sarcasm detection.Evaluation results on a public multi-modal sarcasm detection dataset based on Twitter demonstrate the superiority of our proposed model. Hui Liu 0036, Wenya Wang 0001, Haoliang Li |
EMNLP | 3 |
| 2022 | Rethinking Attention-Model Explainability through Faithfulness Violation TestabstractAttention mechanisms are dominating the explainability of deep models. They produce probability distributions over the input, which are widely deemed as feature-importance indicators. However, in this paper, we find one critical limitation in attention explanations: weakness in identifying the polarity of feature impact. This would be somehow misleading – features with higher attention weights may not faithfully contribute to model predictions; instead, they can impose suppression effects. With this finding, we reflect on the explainability of current attention-based techniques, such as Attention $\bigodot$ Gradient and LRP-based attention explanations. We first propose an actionable diagnostic methodology (henceforth faithfulness violation test) to measure the consistency between explanation weights and the impact polarity. Through the extensive experiments, we then show that most tested explanation methods are unexpectedly hindered by the faithfulness violation issue, especially the raw attention. Empirical analyses on the factors affecting violation issues further provide useful observations for adopting explanation methods in attention models. Yibing Liu, Haoliang Li, Chenqi Kong, Jing Li 0049, Shiqi Wang 0001 |
ICML | 2 |
| 2022 | Generalizing to Evolving Domains with Latent Structure-Aware Sequential AutoencoderabstractDomain generalization aims to improve the generalization capability of machine learning systems to out-of-distribution (OOD) data. Existing domain generalization techniques embark upon stationary and discrete environments to tackle the generalization issue caused by OOD data. However, many real-world tasks in non-stationary environments (e.g., self-driven car system, sensor measures) involve more complex and continuously evolving domain drift, which raises new challenges for the problem of domain generalization. In this paper, we formulate the aforementioned setting as the problem of evolving domain generalization. Specifically, we propose to introduce a probabilistic framework called Latent Structure-aware Sequential Autoencoder (LSSAE) to tackle the problem of evolving domain generalization via exploring the underlying continuous structure in the latent space of deep neural networks, where we aim to identify two major factors namely covariate shift and concept shift accounting for distribution shift in non-stationary environments. Experimental results on both synthetic and real-world datasets show that LSSAE can lead to superior performances based on the evolving domain generalization setting. Tiexin Qin, Shiqi Wang 0001, Haoliang Li |
ICML | 3 |
| 2022 | Hybrid RSS-TDOA Measurements Based Directional Target Localization in NLOS EnvironmentsabstractLocalization constitutes the basic requirement of many applications in wireless sensor networks. Directional transmission has become very common with the development of electronic and antenna technologies. This paper aims to estimate both the position and transmission orientation of the directional target on the basis of combined received-signal-strength (RSS) and time-difference-of-arrival (TDOA) measurements in the non-line-of-sight (NLOS) environments. The localization problem is first formulized by the maximum likelihood estimation. Considering the problem is nonconvex, this paper then converts it into three relatively simple and intuitive sub-problems. Through converting them into generalized trust region subproblem (GTRS) and using bisection as well as multi-resolution search methods, the estimates is finally obtained. Simulation results show that the proposed method can estimate the orientation and position of the target effectively, its localization accuracy also outperforms the comparison methods significantly. Peiliang Zuo, Haoliang Li |
VTC Spring | 3 |
| 2022 | Man-Made Target Structure Recognition With Polarimetric Correlation Pattern and Roll-Invariant Feature CodingabstractMan-made target recognition is of great significance for many applications within microwave remote sensing. The scattering diversity of various man-made target structures makes radar target identification a difficult task. This work aims at mitigating this issue by mining and utilization of man-made target scattering diversity in polarimetric rotation domain with the interpretation tool of polarimetric correlation pattern. The optimal polarimetric roll-invariant feature set is collected from polarimetric correlation pattern. Then, a polarimetric roll-invariant feature coding scheme is developed for man-made target structure recognition. Moreover, polarimetric radar measurement errors in terms of channel coupling and imbalance are also considered. Experimental studies with electromagnetic computation datasets including canonical structures and an unmanned aerial vehicle (UAV) target and real spaceborne polarimetric synthetic aperture radar (PolSAR) data of a ship target are carried out. Compared with the Cameron decomposition, the proposed method exhibits better recognition performance and stronger robustness, especially for oriented man-made structures. Haoliang Li, Ming-Dian Li, Xing-Chao Cui, Si-Wei Chen 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | GMFAD: Towards Generalized Visual Recognition via Multilayer Feature Alignment and DisentanglementabstractThe deep learning based approaches which have been repeatedly proven to bring benefits to visual recognition tasks usually make a strong assumption that the training and test data are drawn from similar feature spaces and distributions. However, such an assumption may not always hold in various practical application scenarios on visual recognition tasks. Inspired by the hierarchical organization of deep feature representation that progressively leads to more abstract features at higher layers of representations, we propose to tackle this problem with a novel feature learning framework, which is called GMFAD, with better generalization capability in a multilayer perceptron manner. We first learn feature representations at the shallow layer where shareable underlying factors among domains (e.g., a subset of which could be relevant for each particular domain) can be explored. In particular, we propose to align the domain divergence between domain pair(s) by considering both inter-dimension and inter-sample correlations, which have been largely ignored by many cross-domain visual recognition methods. Subsequently, to learn more abstract information which could further benefit transferability, we propose to conduct feature disentanglement at the deep feature layer. Extensive experiments based on different visual recognition tasks demonstrate that our proposed framework can learn better transferable feature representation compared with state-of-the-art baselines. Haoliang Li, Shiqi Wang 0001, Renjie Wan, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Detection of GAN-Generated Images by Estimating Artifact SimilarityabstractRecently, researchers have been dedicated to discovering Generative Adversarial Network (GAN) artifacts and using them to identify generated images. However, current approaches exhibit restricted performance when testing against unseen GAN models, which is also known as the cross-domain scenario. To overcome this limitation, we propose a novel GAN-generated image detection framework by estimating artifact similarity, which is inspired by relation network. The proposed method consists of two stages, including representation learning and representation comparison. For representation learning, ResNet-50 equipped with Instance Normalization in the Shallow layers (ResNet-INS) is constructed as the embedding network to extract generalized features. For representation comparison, Category and Domain-Aware loss function (CDA loss) is designed by leveraging both category and domain information efficiently, which can enlarge inter-class discrepancy of different categories (GAN-generated or pristine images) and improve intra-class compactness from different domains (source attributions) in the same category. Extensive experiments are conducted which consider various cross-domain scenarios to verify the generalization of the proposed method. Besides, our method exhibits satisfying robustness against common post-processings, even when data augmentation is not considered during the training stage. Weichuang Li, Peisong He, Haoliang Li, Hongxia Wang 0001, Ruimei Zhang |
IEEE Signal Process. Lett. | 3 |
| 2022 | Appearance Matters, So Does Audio: Revealing the Hidden Face via Cross-Modality TransferabstractRecently, there has been an exponential increase in the security concerns raised by faking face (e.g., deepfake), which automatically changes the identity with a specifically learned deep generative model. With numerous approaches proposed to identify the fake content, much less work has been dedicated to automatically revealing the authentic one that is originally acquired. Here, we propose a new paradigm that seeks to reveal the authentic face hidden behind the fake one by leveraging the joint information of face and audio. More specifically, given the fake face as well as the audio segment, the cross-modality transferable capability is exploited by learning to generate the feature of the authentic face, based on the underlying clues from the audio as well as the fake face appearance. The effectiveness of the proposed scheme is validated through a series of evaluations, and experimental results show that the proposed model achieves promising face reconstruction performance in revealing the hidden faces, in terms of reconstruction quality, as well as identity and face attribute inference accuracy. Chenqi Kong, Baoliang Chen, Wenhan Yang, Haoliang Li, Peilin Chen 0001, Shiqi Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Learning Meta Pattern for Face Anti-SpoofingabstractFace Anti-Spoofing (FAS) is essential to secure face recognition systems and has been extensively studied in recent years. Although deep neural networks (DNNs) for the FAS task have achieved promising results in intra-dataset experiments with similar distributions of training and testing data, the DNNs’ generalization ability is limited under the cross-domain scenarios with different distributions of training and testing data. To improve the generalization ability, recent hybrid methods have been explored to extract task-aware handcrafted features (e.g., Local Binary Pattern) as discriminative information for the input of DNNs. However, the handcrafted feature extraction relies on experts’ domain knowledge, and how to choose appropriate handcrafted features is underexplored. To this end, we propose a learnable network to extract Meta Pattern (MP) in our learning-to-learn framework. By replacing handcrafted features with the MP, the discriminative information from MP is capable of learning a more generalized model. Moreover, we devise a two-stream network to hierarchically fuse the input RGB image and the extracted MP by using our proposed Hierarchical Fusion Module (HFM). We conduct comprehensive experiments and show that our MP outperforms the compared handcrafted features. Also, our proposed method with HFM and the MP can achieve state-of-the-art performance on two different domain generalization evaluation benchmarks. Rizhao Cai, Zhi Li 0054, Renjie Wan, Haoliang Li, Yongjian Hu, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2022 | Detect and Locate: Exposing Face Manipulation by Semantic- and Noise-Level TelltalesabstractThe technological advancements of deep learning have enabled sophisticated face manipulation schemes, raising severe trust issues and security concerns in modern society. Generally speaking, detecting manipulated faces and locating the potentially altered regions are challenging tasks. Herein, we propose a conceptually simple but effective method to efficiently detect forged faces in an image while simultaneously locating the manipulated regions. The proposed scheme relies on a segmentation map that delivers meaningful high-level semantic information clues about the image. Furthermore, a noise map is estimated, playing a complementary role in capturing low-level clues and subsequently empowering decision-making. Finally, the features from these two modules are combined to distinguish fake faces. Extensive experiments show that the proposed model achieves state-of-the-art detection accuracy and remarkable localization performance. Chenqi Kong, Baoliang Chen, Haoliang Li, Shiqi Wang 0001, Anderson Rocha 0001, Sam Kwong |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | Beyond the Pixel World: A Novel Acoustic-Based Face Anti-Spoofing System for Smartphonesabstract2D face presentation attacks are one of the most notorious and pervasive face spoofing types, which have caused pressing security issues to facial authentication systems. While RGB-based face anti-spoofing (FAS) models have proven to counter the face spoofing attack effectively, most existing FAS models suffer from the overfitting problem (i.e., lack generalization capability to data collected from an unseen environment). Recently, many models have been devoted to capturing auxiliary information (e.g., depth and infrared maps) to achieve a more robust face liveness detection performance. However, these methods require expensive sensors and cost extra hardware to capture the specific modality information, limiting their applications in practical scenarios. To tackle these problems, we devise a novel and cost-effective FAS system based on the acoustic modality, named Echo-FAS, which employs the crafted acoustic signal as the probe to perform face liveness detection. We first propose to build a large-scale, high-diversity, and acoustic-based FAS database, Echo-Spoof. Then, based upon Echo-Spoof, we propose designing a novel two-branch framework that combines the global and local frequency clues of input signals to distinguish inputs, live vs. spoofing faces accurately. The devised Echo-FAS comprises the following three merits: (1) It only needs one available speaker and microphone as sensors while not requiring any expensive hardware; (2) It can successfully capture the 3D geometrical information of input queries and achieve a remarkable face anti-spoofing performance; and (3) It can be handily allied with other RGB-based FAS models to mitigate the overfitting problem in the RGB modality and make the FAS model more accurate and robust. Our proposed Echo-FAS provides new insights regarding the development of FAS systems for mobile devices. Chenqi Kong, Kexin Zheng, Shiqi Wang 0001, Anderson Rocha 0001, Haoliang Li |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2022 | One-Class Knowledge Distillation for Face Presentation Attack DetectionabstractFace presentation attack detection (PAD) has been extensively studied by research communities to enhance the security of face recognition systems. Although existing methods have achieved good performance on testing data with similar distribution as the training data, their performance degrades severely in application scenarios with data of unseen distributions. In situations where the training and testing data are drawn from different domains, a typical approach is to apply domain adaptation techniques to improve face PAD performance with the help of target domain data. However, it has always been a non-trivial challenge to collect sufficient data samples in the target domain, especially for attack samples. This paper introduces a teacher-student framework to improve the cross-domain performance of face PAD with one-class domain adaptation. In addition to the source domain data, the framework utilizes only a few genuine face samples of the target domain. Under this framework, a teacher network is trained with source domain samples to provide discriminative feature representations for face PAD. Student networks are trained to mimic the teacher network and learn similar representations for genuine face samples of the target domain. In the test phase, the similarity score between the representations of the teacher and student networks is used to distinguish attacks from genuine ones. To evaluate the proposed framework under one-class domain adaptation settings, we devised two new protocols and conducted extensive experiments. The experimental results show that our method outperforms baselines under one-class domain adaptation settings and even state-of-the-art methods with unsupervised domain adaptation. Zhi Li 0054, Rizhao Cai, Haoliang Li, Kwok-Yan Lam, Yongjian Hu, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | Man-Made Targets Characterization with Polarimetric Correlation Pattern Interpretation ToolabstractThe interpretation and recognition of man-made target is an important application of polarimetric radar. The scattering diversity of radar target contains rich information. The polarization interpretation theory in the rotation domain has been proposed to mine and characterize target hidden information. Recently, this interpretation technique has received plenty of attention and achieved successful applications. Based on the study, this work further utilizes the polarimetric correlation pattern interpretation tool for man-made target characterization and recognition. Its potential and advantages are verified by using canonical structures, the Slicy model with electromagnetic computation data and ship targets with polarimetric synthetic aperture radar (PolSAR) data. The experimental studies show that polarimetric correlation pattern has great potential for man-made targets recognition. Haoliang Li, Ming-Dian Li, Si-Wei Chen 0001 |
IGARSS | 1 |
| 2021 | Embracing the Dark Knowledge: Domain Generalization Using Regularized Knowledge DistillationabstractThough convolutional neural networks are widely used in different tasks, lack of generalization capability in the absence of sufficient and representative data is one of the challenges that hinders their practical application. In this paper, we propose a simple, effective, and plug-and-play training strategy named Knowledge Distillation for Domain Generalization (KDDG) which is built upon a knowledge distillation framework with the gradient filter as a novel regularization term. We find that both the "richer dark knowledge" from the teacher network, as well as the gradient filter we proposed, can reduce the difficulty of learning the mapping which further improves the generalization ability of the model. We also conduct experiments extensively to show that our framework can significantly improve the generalization capability of deep neural networks in different tasks including image classification, segmentation, reinforcement learning by comparing our method with existing state-of-the-art domain generalization techniques. Last but not the least, we propose to adopt two metrics to analyze our proposed method in order to better understand how our proposed method benefits the generalization capability of deep neural networks. Yufei Wang 0006, Haoliang Li, Lap-Pui Chau, Alex Chichung Kot |
ACM Multimedia | 2 |
| 2021 | A Residual Correction Approach for Semi-supervised Semantic Segmentation
Haoliang Li, Huicheng Zheng |
PRCV (4) | 1 |
| 2021 | Unsupervised Domain Adaptation in the Wild via Disentangling Representation Learning
Haoliang Li, Renjie Wan, Shiqi Wang 0001, Alex Chichung Kot |
Int. J. Comput. Vis. | 1 |
| 2021 | Face Image Reflection Removal
Renjie Wan, Boxin Shi, Haoliang Li, Ling-Yu Duan, Alex Chichung Kot |
Int. J. Comput. Vis. | 3 |
| 2021 | Special issue on low complexity methods for multimedia security
Guorui Feng, Sheng Li 0006, Haoliang Li, Shujun Li 0001 |
Multim. Syst. | 3 |
| 2021 | Detection of Spoofing Medium Contours for Face Anti-SpoofingabstractFace anti-spoofing is an important step for secure face recognition. In this paper, we target on building a general classifier to detect the face images with spoofing medium contours (termed as SMCs for simplicity). To this end, we consider the task of face anti-spoofing as the detection of SMCs from the image. We propose and train a Contour Enhanced Mask R-CNN (CEM-RCNN) model for the detection. This model detects the existence of the SMCs by incorporating the contour objectness which measures how likely an object contains the SMCs. The experimental results demonstrate the generality of the CEM-RCNN for identifying the face images with SMCs, which performs significantly better than the state-of-the-art on the cross-database scenario. Sheng Li 0006, Xinpeng Zhang 0001, Haoliang Li, Alex Chichung Kot |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | DRL-FAS: A Novel Framework Based on Deep Reinforcement Learning for Face Anti-SpoofingabstractInspired by the philosophy employed by human beings to determine whether a presented face example is genuine or not, i.e., to glance at the example globally first and then carefully observe the local regions to gain more discriminative information, for the face anti-spoofing problem, we propose a novel framework based on the Convolutional Neural Network (CNN) and the Recurrent Neural Network (RNN). In particular, we model the behavior of exploring face-spoofing-related information from image sub-patches by leveraging deep reinforcement learning. We further introduce a recurrent mechanism to learn representations of local information sequentially from the explored sub-patches with an RNN. Finally, for the classification purpose, we fuse the local information with the global one, which can be learned from the original input image through a CNN. Moreover, we conduct extensive experiments, including ablation study and visualization analysis, to evaluate our proposed framework on various public databases. The experiment results show that our method can generally achieve state-of-the-art performance among all scenarios, demonstrating its effectiveness. Rizhao Cai, Haoliang Li, Shiqi Wang 0001, Changsheng Chen 0001, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | Camera Invariant Feature Learning for Generalized Face Anti-SpoofingabstractThere has been an increasing consensus in learning based face anti-spoofing that the divergence in terms of camera models is causing a large domain gap in real application scenarios. We describe a framework that eliminates the influence of inherent variance from acquisition cameras at the feature level, leading to the generalized face spoofing detection model that could be highly adaptive to different acquisition devices. In particular, the framework is composed of two branches. The first branch aims to learn the camera invariant spoofing features via feature level decomposition in the high frequency domain. Motivated by the fact that the spoofing features exist not only in the high frequency domain, in the second branch the discrimination capability of extracted spoofing features is further boosted from the enhanced image based on the recomposition of the high-frequency and low-frequency information. Finally, the classification results of the two branches are fused together by a weighting strategy. Experiments show that the proposed method can achieve better performance in both intra-dataset and cross-dataset settings, demonstrating the high generalization capability in various application scenarios. Baoliang Chen, Wenhan Yang, Haoliang Li, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | No-Reference Screen Content Image Quality Assessment With Unsupervised Domain AdaptationabstractIn this paper, we quest the capability of transferring the quality of natural scene images to the images that are not acquired by optical cameras (e.g., screen content images, SCIs), rooted in the widely accepted view that the human visual system has adapted and evolved through the perception of natural environment. Here, we develop the first unsupervised domain adaptation based no reference quality assessment method for SCIs, leveraging rich subjective ratings of the natural images (NIs). In general, it is a non-trivial task to directly transfer the quality prediction model from NIs to a new type of content (i.e., SCIs) that holds dramatically different statistical characteristics. Inspired by the transferability of pair-wise relationship, the proposed quality measure operates based on the philosophy of improving the transferability and discriminability simultaneously. In particular, we introduce three types of losses which complementarily and explicitly regularize the feature space of ranking in a progressive manner. Regarding feature discriminatory capability enhancement, we propose a center based loss to rectify the classifier and improve its prediction capability not only for source domain (NI) but also the target domain (SCI). For feature discrepancy minimization, the maximum mean discrepancy (MMD) is imposed on the extracted ranking features of NIs and SCIs. Furthermore, to further enhance the feature diversity, we introduce the correlation penalization between different feature dimensions, leading to the features with lower rank and higher diversity. Experiments show that our method can achieve higher performance on different source-target settings based on a light-weight convolution neural network. The proposed method also sheds light on learning quality assessment measures for unseen application-specific content without the cumbersome and costing subjective evaluations. Baoliang Chen, Haoliang Li, Hongfei Fan, Shiqi Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Frame-Wise Detection of Double HEVC Compression by Learning Deep Spatio-Temporal Representations in Compression DomainabstractDetection of double compression is regarded as one primary step in analyzing the integrity of digital videos, which is of prominent importance in video forensics. However, current methods are vulnerable with the severe lossy quantization in the recompression process such that it is challenging to obtain reliable frame-wise detection results, especially for the high efficiency video coding (HEVC) standard. In view of these issues, in this paper, a hybrid neural network is proposed to reveal abnormal frames in HEVC videos with double compression by learning robust spatio-temporal representations from coding information in the compression domain. Based on the statistical analysis of Coding Units (CUs), it is interesting to find that HEVC video streams contain “rich” coding information that could be leveraged to identify abnormal traces caused by double compression. Two types of coding information maps, including CU Size Map (CSM) and CU Prediction mode Map (CPM), are exploited. In contrast with the conventional paradigm relying on pixel-level representations of decoded frames, CSMs and CPMs of a short-time video clip are treated as the input, aiming to achieve high robustness against recompression of low quality. In our hybrid neural network, an attention-based two-stream residual network is proposed to learn hierarchical representations from CSM and CPM, which are then jointly optimized by the attention-based fusion module. Finally, the temporal variation is modeled by Long Short-Term Memory (LSTM) to obtain frame-wise detection results. We have conducted extensive experiments considering various video content and coding parameters, such as bitrates and sizes of Group of Picture. Experimental results show that our approach can obtain state-of-the-art performance compared with conventional methods, especially when videos are recompressed in the low bitrate coding scenarios. Peisong He, Haoliang Li, Hongxia Wang 0001, Shiqi Wang 0001, Xinghao Jiang, Ruimei Zhang |
IEEE Trans. Multim. | 2 |
| 2020 | Reflection Scene Separation From a Single ImageabstractFor images taken through glass, existing methods focus on the restoration of the background scene by regarding the reflection components as noise. However, the scene reflected by glass surface also contains important information to be recovered, especially for the surveillance or criminal investigations. In this paper, instead of removing reflection components from the mixture image, we aim at recovering reflection scenes from the mixture image. We first propose a strategy to obtain such ground truth and its corresponding input images. Then, we propose a two-stage framework to obtain the visible reflection scene from the mixture image. Specifically, we train the network with a shift-invariant loss which is robust to misalignment between the input and output images. The experimental results show that our proposed method achieves promising results. Renjie Wan, Boxin Shi, Haoliang Li, Ling-Yu Duan, Alex Chichung Kot |
CVPR | 3 |
| 2020 | Unseen Face Presentation Attack Detection with Hypersphere LossabstractPresentation attack is one of the main threats to face verification systems and attracts great attention of research community. Recent methods achieve great success in intra-database test. However, the problem is more complex in practical scenario as the type of attack could be unseen to system designers. In this paper, we formulate the face presentation attack detection task under an open-set setting and address with our proposed deep anomaly detection based method. The training process is end-to-end supervised by a novel hypersphere loss function and the decision making is directly based on the learned feature representation. We conduct extensive experiments on multiple prevailing databases and evaluate our implemented models by using various metrics. The results show our proposed method is effective against unseen types of attacks and superior to latest state-of-the-art. Zhi Li 0054, Haoliang Li, Kwok-Yan Lam, Alex Chichung Kot |
ICASSP | 2 |
| 2020 | Heterogeneous Domain Generalization Via Domain MixupabstractOne of the main drawbacks of deep Convolutional Neural Networks (DCNN) is that they lack generalization capability. In this work, we focus on the problem of heterogeneous domain generalization which aims to improve the generalization capability across different tasks, which is, how to learn a DCNN model with multiple domain data such that the trained feature extractor can be generalized to supporting recognition of novel categories in a novel target domain. To solve this problem, we propose a novel heterogeneous domain generalization method by mixing up samples across multiple source domains with two different sampling strategies. Our experimental results based on the Visual Decathlon benchmark demonstrates the effectiveness of our proposed method. Yufei Wang 0006, Haoliang Li, Alex Chichung Kot |
ICASSP | 2 |
| 2020 | Low-Dose CT Image Blind Denoising with Graph Convolutional Networks
Kecheng Chen, Xiaorong Pu, Yazhou Ren 0001, Hang Qiu 0002, Haoliang Li |
ICONIP (1) | 5 |
| 2020 | Domain Generalization for Medical Imaging Classification with Linear-Dependency RegularizationabstractRecently, we have witnessed great progress in the field of medical imaging classification by adopting deep neural networks. However, the recent advanced models still require accessing sufficiently large and representative datasets for training, which is often unfeasible in clinically realistic environments. When trained on limited datasets, the deep neural network is lack of generalization capability, as the trained deep neural network on data within a certain distribution (e.g. the data captured by a certain device vendor or patient population) may not be able to generalize to the data with another distribution. In this paper, we introduce a simple but effective approach to improve the generalization capability of deep neural networks in the field of medical imaging classification. Motivated by the observation that the domain variability of the medical images is to some extent compact, we propose to learn a representative feature space through variational encoding with a novel linear-dependency regularization term to capture the shareable information among medical data collected from different domains. As a result, the trained neural network is expected to equip with better generalization capability to the ``unseen" medical data. Experimental results on two challenging medical imaging classification tasks indicate that our method can achieve better cross-domain generalization capability compared with state-of-the-art baselines. Haoliang Li, Yufei Wang 0006, Renjie Wan, Shiqi Wang 0001, Tie-Qiang Li, Alex Chichung Kot |
NeurIPS | 1 |
| 2020 | Improving Robustness of DNNs against Common Corruptions via Gaussian Adversarial TrainingabstractDeep neural networks have demonstrated tremendous success in image classification, but their performance sharply degrades when evaluated on slightly different test data (e.g., data with corruptions). To address these issues, we propose a minimax approach to improve common corruption robustness of deep neural networks via Gaussian Adversarial Training. To be specific, we propose to train neural networks with adversarial examples where the perturbations are Gaussian-distributed. Our experiments show that our proposed GAT can improve neural networks' robustness to noise corruptions more than other baseline methods. It also outperforms the state-of-the-art method in improving the overall robustness to common corruptions. Chenyu Yi, Haoliang Li, Renjie Wan, Alex Chichung Kot |
VCIP | 2 |
| 2020 | The Enhancement of Underexposed Images with Blurred ReflectanceabstractThe images captured in the low-light conditions always suffer from low visibility. Enhancing the visibility of the low-light image is of broad application to various computer vision tasks. Based on the classical Retinex model, previous methods assume the reflectance components as a well-exposed image. In this paper, we introduce the blurring distortion into the Retinex model to cover more general and challenging scenarios. We further propose a two-stage framework to extract the reflectance images and remove the blurring distortion separately. Specifically, we optimize the whole network by embedding a mechanism robust to the pixel misalignment in the training dataset. The experimental results show that our proposed method achieves promising results. Jinchao Zhou, Renjie Wan, Haoliang Li, Alex Chichung Kot |
VCIP | 3 |
| 2020 | CoRRN: Cooperative Reflection Removal NetworkabstractRemoving the undesired reflections from images taken through the glass is of broad application to various computer vision tasks. Non-learning based methods utilize different handcrafted priors such as the separable sparse gradients caused by different levels of blurs, which often fail due to their limited description capability to the properties of real-world reflections. In this paper, we propose a network with the feature-sharing strategy to tackle this problem in a cooperative and unified framework, by integrating image context information and the multi-scale gradient information. To remove the strong reflections existed in some local regions, we propose a statistic loss by considering the gradient level statistics between the background and reflections. Our network is trained on a new dataset with 3250 reflection images taken under diverse real-world scenes. Experiments on a public benchmark dataset show that the proposed method performs favorably against state-of-the-art methods. Renjie Wan, Boxin Shi, Haoliang Li, Ling-Yu Duan, Ah-Hwee Tan, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Discovering and incorporating latent target-domains for domain adaptation
Haoliang Li, Wen Li 0001, Shiqi Wang 0001 |
Pattern Recognit. | 1 |
| 2020 | Exposing Fake Bitrate Videos Using Hybrid Deep-Learning Network From Recompression ErrorabstractBitrate is generally regarded as an important criterion of video quality. However, with sophisticated video editing software, forgers can create fake bitrate videos by up-converting the bitrate of original videos with lower video quality to attract more viewers on video sharing websites. In this work, we first model the generation process of fake bitrate videos and analyze the dominant sources of information loss. It is found that the recompression error generated by the proposed one-step-further recompression operation is an efficient measurement to expose distinguishable quality variation tendencies between true and fake bitrate videos. Based on this analysis, we propose a detection method for fake bitrate videos using a hybrid deep-learning network from recompression error. For an input video, the patch-wise recompression errors are first calculated to increase the learning capability of the network. To learn robust representations of recompression errors in local regions with different degrees of predictability, a hybrid deep-learning network that contains two branches with heterogeneous structures is designed. For noise-like recompression errors, the first branch has a shallow CNN structure initialized with an Inception-like module using multisize convolutional kernels. For zero-element clustered recompression errors, the second branch has a multi-layer perceptron structure equipped with a unique layer that extracts the histogram of zero-element clustered square regions. The output vectors of different branches are concatenated and then jointly optimized to obtain the patch-wise detection results. Finally, the majority voting (local-to-global) strategy is applied to obtain the final detection result. Extensive experiments are conducted to evaluate the detection performance under various coding parameter settings, such as different bitrates, rate-distortion optimization strategies and so on. The experimental results demonstrate the superiority of the proposed method compared with several state-of-the-art methods to provide more fine-grained forensic clues. Peisong He, Haoliang Li, Bin Li 0011, Hongxia Wang 0001, Liang Liu 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Heterogeneous Domain Adaptation via Nonlinear Matrix FactorizationabstractHeterogeneous domain adaptation (HDA) aims to solve the learning problems where the source- and the target-domain data are represented by heterogeneous types of features. The existing HDA approaches based on matrix completion or matrix factorization have proven to be effective to capture shareable information between heterogeneous domains. However, there are two limitations in the existing methods. First, a large number of corresponding data instances between the source domain and the target domain are required to bridge the gap between different domains for performing matrix completion. These corresponding data instances may be difficult to collect in real-world applications due to the limited size of data in the target domain. Second, most existing methods can only capture linear correlations between features and data instances while performing matrix completion for HDA. In this paper, we address these two issues by proposing a new matrix-factorization-based HDA method in a semisupervised manner, where only a few labeled data are required in the target domain without requiring any corresponding data instances between domains. Such labeled data are more practical to obtain compared with cross-domain corresponding data instances. Our proposed algorithm is based on matrix factorization in an approximated reproducing kernel Hilbert space (RKHS), where nonlinear correlations between features and data instances can be exploited to learn heterogeneous features for both the source and the target domains. Extensive experiments are conducted on cross-domain text classification and object recognition, and experimental results demonstrate the superiority of our proposed method compared with the state-of-the-art HDA approaches. Haoliang Li, Sinno Jialin Pan, Shiqi Wang 0001, Alex Chichung Kot |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | Heterogeneous Transfer Learning via Deep Matrix Completion with Adversarial Kernel EmbeddingabstractHeterogeneous Transfer Learning (HTL) aims to solve transfer learning problems where a source domain and a target domain are of heterogeneous types of features. Most existing HTL approaches either explicitly learn feature mappings between the heterogeneous domains or implicitly reconstruct heterogeneous cross-domain features based on matrix completion techniques. In this paper, we propose a new HTL method based on a deep matrix completion framework, where kernel embedding of distributions is trained in an adversarial manner for learning heterogeneous features across domains. We conduct extensive experiments on two different vision tasks to demonstrate the effectiveness of our proposed method compared with a number of baseline methods. Haoliang Li, Sinno Jialin Pan, Renjie Wan, Alex Chichung Kot |
AAAI | 1 |
| 2019 | Detection of Fake Images Via The Ensemble of Deep Representations from Multi Color SpacesabstractRecently, the success of generating fake images by Generative Adversarial Network (GAN) has threatened the authentication of digital images. To address this issue, several automated fake image detectors have been proposed. However, current methods remain vulnerable when testing samples undergo post-processing attacks. In this work, we employed residual signals of chrominance components from multi color spaces, including YCbCr, HSV and Lab, to learn robust deep representations via the well-designed shallow convolutional neural network (CNN). Then, the learned deep representations from different color spaces are concatenated and then fed into the Random Forest (RF), which is the widely used ensemble classifier, to obtain final detection results. Extensive experiments are conducted on the fake image dataset generated by the advanced GAN technique. Experimental results demonstrate the proposed scheme outperforms state-of-the-art methods and achieves the promising average detection accuracy (above 99%) under several post-processing attacks, such as Gaussian blurring and so on. Peisong He, Haoliang Li, Hongxia Wang 0001 |
ICIP | 2 |
| 2018 | Multi-task Mid-level Feature Alignment Network for Unsupervised Cross-Dataset Person Re-Identification
Haoliang Li, Chang-Tsun Li, Alex Chichung Kot |
BMVC | 2 |
| 2018 | Domain Generalization With Adversarial Feature LearningabstractIn this paper, we tackle the problem of domain generalization: how to learn a generalized feature representation for an "unseen" target domain by taking the advantage of multiple seen source-domain data. We present a novel framework based on adversarial autoencoders to learn a generalized latent feature representation across domains for domain generalization. To be specific, we extend adversarial autoencoders by imposing the Maximum Mean Discrepancy (MMD) measure to align the distributions among different domains, and matching the aligned distribution to an arbitrary prior distribution via adversarial feature learning. In this way, the learned feature representation is supposed to be universal to the seen source domains because of the MMD regularization, and is expected to generalize well on the target domain because of the introduction of the prior distribution. We proposed an algorithm to jointly train different components of our proposed framework. Extensive experiments on various vision tasks demonstrate that our proposed framework can learn better generalized features for the unseen target domain compared with state-of-the-art domain generalization methods. Haoliang Li, Sinno Jialin Pan, Shiqi Wang 0001, Alex Chichung Kot |
CVPR | 1 |
| 2018 | Computer Graphics Identification Combining Convolutional and Recurrent Neural NetworksabstractIn this letter, a deep-learning-based pipeline is proposed to distinguish photographics (PGs) from computer-graphics (CGs) combining convolutional neural network (CNN) and recurrent neural network (RNN). In the preprocessing stage, the color space transformation and the Schmid filter bank are utilized to extract chrominance and luminance components, which suppress the irrelevant information of various image contents for the CG identification task. Then, a dual-path CNN architecture is designed to learn joint feature representations of local patches for exploiting their color and texture characteristics. To extract the global artifact, the directed acyclic graph RNN is applied to model the spatial dependence of local patterns. Finally, the output score of RNN is used to identify the input sample. The CG/PG dataset is constructed by collecting samples from the Internet. Experimental results show that the proposed framework can outperform state-of-the-art methods on identification ability of CGs, especially for images with low resolution. Peisong He, Xinghao Jiang, Tanfeng Sun, Haoliang Li |
IEEE Signal Process. Lett. | 4 |
| 2018 | Learning Generalized Deep Feature Representation for Face Anti-SpoofingabstractIn this paper, we propose a novel framework leveraging the advantages of the representational ability of deep learning and domain generalization for face spoofing detection. In particular, the generalized deep feature representation is achieved by taking both spatial and temporal information into consideration, and a 3D convolutional neural network architecture tailored for the spatial-temporal input is proposed. The network is first initialized by training with augmented facial samples based on cross-entropy loss and further enhanced with a specifically designed generalization loss, which coherently serves as the regularization term. The training samples from different domains can seamlessly work together for learning the generalized feature representation by manipulating their feature distribution distances. We evaluate the proposed framework with different experimental setups using various databases. Experimental results indicate that our method can learn more discriminative and generalized information compared with the state-of-the-art methods. Haoliang Li, Peisong He, Shiqi Wang 0001, Anderson Rocha 0001, Xinghao Jiang, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2018 | Unsupervised Domain Adaptation for Face Anti-SpoofingabstractFace anti-spoofing (a.k.a. presentation attack detection) has recently emerged as an active topic with great significance for both academia and industry due to the rapidly increasing demand in user authentication on mobile phones, PCs, tablets, and so on. Recently, numerous face spoofing detection schemes have been proposed based on the assumption that training and testing samples are in the same domain in terms of the feature space and marginal probability distribution. However, due to unlimited variations of the dominant conditions (illumination, facial appearance, camera quality, and so on) in face acquisition, such single domain methods lack generalization capability, which further prevents them from being applied in practical applications. In light of this, we introduce an unsupervised domain adaptation face anti-spoofing scheme to address the real-world scenario that learns the classifier for the target domain based on training samples in a different source domain. In particular, an embedding function is first imposed based on source and target domain data, which maps the data to a new space where the distribution similarity can be measured. Subsequently, the Maximum Mean Discrepancy between the latent features in source and target domains is minimized such that a more generalized classifier can be learned. State-of-the-art representations including both hand-crafted and deep neural network learned features are further adopted into the framework to quest the capability of them in domain adaptation. Moreover, we introduce a new database for face spoofing detection, which contains more than 4000 face samples with a large variety of spoofing types, capture devices, illuminations, and so on. Extensive experiments on existing benchmark databases and the new database verify that the proposed approach can gain significantly better generalization capability in cross-domain scenarios by providing consistently better anti-spoofing performance. Haoliang Li, Wen Li 0001, Shiqi Wang 0001, Feiyue Huang, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2017 | Memetic algorithm based location and topic aware recommender system
Shanfeng Wang, Maoguo Gong, Haoliang Li, Yue Wu 0004 |
Knowl. Based Syst. | 3 |
| 2016 | Color space identification from single imagesabstractIn this paper, we focus on the problem of RGB color space identification from a single image. At the moment, RGB color spaces are widely adopted in photography for image producing. The problems with respect to color space identification, such as to get the consistent printing or displaying quality on screen devices and software applications and prevention of multimedia unauthorized usage(shown or printed by other device via gamut mapping), need to be concerned. Current techniques are all relying on EXchangeable Image File Format (EXIF) to extract color space information. In this paper, we use a two-dimensional non-causal regressive model to explore the image demosaicing properties in order to extract discriminative features and train them on SVM classifier for image color space detection without relying on EXIF. In our experiment, images in three different color spaces (sRGB, adobeRGB and pro PhotoRGB) are generated for color space identification task. The experimental results show that the proposed technique has an good performance on image color space identification. Haoliang Li, Alex Chichung Kot, Leida Li |
ISCAS | 1 |
| 2016 | Multi-objective optimization for long tail recommendation
Shanfeng Wang, Maoguo Gong, Haoliang Li |
Knowl. Based Syst. | 3 |
| 2016 | Image Sharpness Assessment by Sparse RepresentationabstractRecent advances in sparse representation show that overcomplete dictionaries learned from natural images can capture high-level features for image analysis. Since atoms in the dictionaries are typically edge patterns and image blur is characterized by the spread of edges, an overcomplete dictionary can be used to measure the extent of blur. Motivated by this, this paper presents a no-reference sparse representation-based image sharpness index. An overcomplete dictionary is first learned using natural images. The blurred image is then represented using the dictionary in a block manner, and block energy is computed using the sparse coefficients. The sharpness score is defined as the variance-normalized energy over a set of selected high-variance blocks, which is achieved by normalizing the total block energy using the sum of block variances. The proposed method is not sensitive to training images, so a universal dictionary can be used to evaluate the sharpness of images. Experiments on six public image quality databases demonstrate the advantages of the proposed method. Leida Li, Jinjian Wu, Haoliang Li, Weisi Lin, Alex Chichung Kot |
IEEE Trans. Multim. | 4 |
| 2015 | GridSAR: Grid strength and regularity for robust evaluation of blocking artifacts in JPEG images
Leida Li, Yu Zhou 0009, Jinjian Wu, Weisi Lin, Haoliang Li |
J. Vis. Commun. Image Represent. | 5 |
| 2015 | Blurred Image Splicing Localization by Exposing Blur Type InconsistencyabstractIn a tampered blurred image generated by splicing, the spliced region and the original image may have different blur types. Splicing localization in this image is a challenging problem when a forger uses some postprocessing operations as antiforensics to remove the splicing traces anomalies by resizing the tampered image or blurring the spliced region boundary. Such operations remove the artifacts that make detection of splicing difficult. In this paper, we overcome this problem by proposing a novel framework for blurred image splicing localization based on the partial blur type inconsistency. In this framework, after the block-based image partitioning, a local blur type detection feature is extracted from the estimated local blur kernels. The image blocks are classified into out-of-focus or motion blur based on this feature to generate invariant blur type regions. Finally, a fine splicing localization is applied to increase the precision of regions boundary. We can use the blur type differences of the regions to trace the inconsistency for the splicing localization. Our experimental results show the efficiency of the proposed method in the detection and the classification of the out-of-focus and motion blur types. For splicing localization, the result demonstrates that our method works well in detecting the inconsistency in the partial blur types of the tampered images. However, our method can be applied to blurred images only. Khosro Bahrami, Alex Chichung Kot, Leida Li, Haoliang Li |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2013 | A novel low-power filter design via reduced-precision redundancy for voltage overscaling applicationsabstractIn this paper, we apply adaptive-signal processing in the reduced precision redundancy filter design for the voltage scaling (VOS) application to achieve high energy efficiency and high SNR performance. RPR technique can mitigate the soft error caused by VOS in the critical path, but the SNR performance of RPR is limited. Thus, we combined RPR and adaptive signal processing to achieve high SNR and low power consumption performance simultaneously for VOS applications. The adaptive signal processing is used to recover the middle significant bits in the result of the filter. From case studies, we find that the proposed method can obtain up to 64% energy reduction with much less SNR performance degradation losing than the traditional RPR and its optimization schemes. Haoliang Li, Jianhao Hu, Jienan Chen |
GLOBECOM | 1 |