EDBT 2026 Demo / reviewers in the wild / expert
Hong Liu 0009
dblp:29/5010-9
· DBLP profile ↗
45ranked-venue papers
12as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 10 first-author · 15 since 2021Artificial intelligence and machine learning · 30 · 9 first-author · 16 since 2021Security and privacy · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SimLabel: Similarity-Weighted Semi-supervision for Multi-annotator Learning with Missing LabelsabstractMulti-annotator learning (MAL) aims to model annotator-specific labeling patterns. However, existing methods face a critical challenge: they simply skip updating annotator-specific model parameters when encountering missing labels—a common scenario in real-world crowdsourced datasets where each annotator labels only small subsets of samples. This leads to inefficient data utilization and overfitting risks. To this end, we propose a novel similarity-weighted semi-supervised learning framework (SimLabel) that leverages inter-annotator similarities to generate weighted soft labels for missing annotations, enabling the utilization of unannotated samples rather than skipping them entirely. We further introduce a confidence-based iterative refinement mechanism that combines maximum probability with entropy-based uncertainty to prioritize predicted high-quality pseudo-labels to impute missing labels, jointly enhancing similarity estimation and model performance over time. For evaluation, we contribute a new multimodal multi-annotator dataset, AMER2, with high and more variable missing rates, reflecting real-world annotation sparsity and enabling evaluation across different sparsity levels. Extensive experiments validate the effectiveness of our method. Hong Liu 0009, Takanori Takebe, Yuta Nakashima |
AAAI | 3 |
| 2026 | QuMAB: Query-based Multi-annotator Behavior Pattern LearningabstractMulti-annotator learning traditionally aggregates diverse annotations to approximate a single “ground truth”, treating disagreements as noise. However, this paradigm faces fundamental challenges: subjective tasks often lack absolute ground truth, and sparse annotation coverage makes aggregation statistically unreliable. We introduce a paradigm shift from sample-wise aggregation to annotator-wise behavior modeling. By treating annotator disagreements as valuable information rather than noise, modeling annotator-specific behavior patterns can reconstruct unlabeled data to reduce annotation cost, enhance aggregation reliability, and explain annotator decision behavior. To this end, we propose QuMAB (Query-based Multi-Annotator Behavior Pattern Learning), which uses lightweight queries to model individual annotators while capturing inter-annotator correlations as implicit regularization, preventing overfitting to sparse individual data while maintaining individualization and improving generalization, with a visualization of annotator focus regions offering an explainable analysis of behavior understanding. We contribute two large-scale datasets with dense per-annotator labels: STREET (4,300 labels/annotator) and AMER (average 3,118 labels/annotator), the first multimodal multi-annotator dataset. Extensive experiments demonstrate the superiority of our QuMAB in modeling individual annotators’ behavior patterns, their utility for consensus prediction, and applicability under sparse annotations. Hong Liu 0009, Takanori Takebe, Shozo Nishii, Yuta Nakashima |
AAAI | 3 |
| 2026 | Quality-Agnostic Deepfake Detection With Saliency-Guided Restoration and Adaptive FusionabstractDeepfake technology poses a significant threat to the Internet of Things by enabling identity spoofing and the dissemination of misinformation. Although numerous face forgery detectors have been developed to counter the risks of facial deepfakes, detecting forgeries across varying quality levels, particularly under extreme degradations, remains a critical challenge. Current face forgery detection methods largely rely on identifying low-level artifacts, which are highly susceptible to distortions introduced by image degradation. Recognizing that restoration can recover critical forensic cues from severely degraded forgeries, this study proposes a quality-agnostic deepfake detection framework that leverages a blind face restoration model to enhance robustness. The framework incorporates an auxiliary restoration branch alongside the original detection pathway. While the original branch operates directly on the degraded input, the restoration branch performs detection on facial images restored by a blind face restoration model. In addition, a Saliency-Guided Restoration objective is introduced to enhance alignment between the restoration and detection tasks. A Restoration Similarity-Aware Fusion mechanism is further designed to adaptively integrate predictions from both branches based on input quality. To assess the robustness of our approach under extreme degradation, we establish a specialized, real-world-inspired benchmark that simulates diverse degradation scenarios. Comprehensive experiments on both the proposed and existing benchmarks demonstrate that our method consistently achieves superior robustness in various degradation scenarios. Xiaotian Si, Linghui Li 0001, Zhihao Tang 0002, Bingyu Li 0003, Kaiguo Yuan, Hong Liu 0009, Qi Tian 0001 |
IEEE Internet Things J. | 7 |
| 2026 | Rethinking Crowd Localization Evaluation via Optimal Transportation Cost
Jun-Xiu Li, Hong Liu 0009, Xiao Wu 0001, Yu-Pei Song, Zhenhua Zeng, Shin'ichi Satoh 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Deep Dual Internal Learning for Hyperspectral Image Super-Resolution
Yongqing Sun, Hong Liu 0009, Qiong Chang, Xianhua Han |
MMM (1) | 2 |
| 2025 | Paladin: Understanding Video Intentions in Political Advertisement VideosabstractIn this paper, we introduce a novel taskfor video understanding that focuses on detecting editing intentions in political advertisement videos. Political advertisement videos are edited with some intentions (e.g., “associating some candidates with negative emotions”) of making people unthinkingly believe the messages in the videos, potentially ending up with some irrational bias. Detecting such intentions is thus the primary step toward fairer decision-making based on the messages themselves. To this end, we classify such editing intentions into 10 categories (referred to as communication techniques) in consultation with a professional editor as well as based on communication techniques presented in the natural language processing community, and build a dataset of 12,526 political advertisement videos, each of which is annotated with several communication technique segments. We also explore the capability of existing video understanding models in detecting editing intentions over the dataset, which identifies new dimensions of challenges to be addressed. Hong Liu 0009, Yuta Nakashima, Noboru Babaguchi |
WACV | 1 |
| 2025 | Guest Editorial: Special Issue on Open-World Visual Recognition
Zhun Zhong, Hong Liu 0009, Yin Cui, Shin'ichi Satoh 0001, Nicu Sebe, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 2 |
| 2025 | Conditional Diffusion Models for Camouflaged and Salient Object DetectionabstractCamouflaged Object Detection (COD) poses a significant challenge in computer vision, playing a critical role in applications. Existing COD methods often exhibit challenges in accurately predicting nuanced boundaries with high-confidence predictions. In this work, we introduce CamoDiffusion, a new learning method that employs a conditional diffusion model to generate masks that progressively refine the boundaries of camouflaged objects. In particular, we first design an adaptive transformer conditional network, specifically designed for integration into a Denoising Network, which facilitates iterative refinement of the saliency masks. Second, based on the classical diffusion model training, we investigate a variance noise schedule and a structure corruption strategy, which aim to enhance the accuracy of our denoising model by effectively handling uncertain input. Third, we introduce a Consensus Time Ensemble technique, which integrates intermediate predictions using a sampling mechanism, thus reducing overconfidence and incorrect predictions. Finally, we conduct extensive experiments on three benchmark datasets that show that: 1) the efficacy and universality of our method is demonstrated in both camouflaged and salient object detection tasks. 2) compared to existing state-of-the-art methods, CamoDiffusion demonstrates superior performance 3) CamoDiffusion offers flexible enhancements, such as an accelerated version based on the VQ-VAE model and a skip approach. Ke Sun 0016, Zhongxi Chen, Xianming Lin, Xiaoshuai Sun, Hong Liu 0009, Rongrong Ji |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | CNNGen: A Generator and a Dataset for Energy-Aware Neural Architecture SearchabstractNeural Architecture Search (NAS) methods seek optimal networks by exploring thousands of variants of a reference architecture.Yet, optimality is typically related to prediction performance, overlooking the environmental impacts of training.Thus, NAS search spaces are unfit for performance and energy consumption trade-offs.We contribute to energy-aware NAS with (i) a grammar-based Convolutional Neural Network generator (CN-NGen) producing diverse architectures not based on a reference one; (ii) 1,300 available architectures obtained via CNNGen with their implementation, energy consumption and performance measurements; (iii) Three state-of-the-art predictors releasing the need for trained models for performance and energy estimation. CNN Generator (CNNGen)CNNGen uses the Xtext context-free grammar framework [3] to generate CNN architectures.The sequence of grammar tokens describes the CNN's topology (i.e., the succession of layers).Our grammar captures the CNN domain knowledge to produce valid architectures.Thus, CNNGen differs from other NAS methods like NASBench [2] Indeed, CNNGen produces architectures from scratch and not as variants of existing ones.CNNGen also comes with an editor allowing to specify architectures.From a valid sequence of grammar tokens, 173 Antoine Gratia, Hong Liu 0009, Shin'ichi Satoh 0001, Paul Temple, Pierre-Yves Schobbens, Gilles Perrouin |
ESANN | 2 |
| 2024 | Mitigating robust overfitting via self-residual-calibration regularization (Abstract Reprint)
Hong Liu 0009, Zhun Zhong, Nicu Sebe, Shin'ichi Satoh 0001 |
IJCAI | 1 |
| 2024 | DiffusionFake: Enhancing Generalization in Deepfake Detection via Guided Stable DiffusionabstractThe rapid progress of Deepfake technology has made face swapping highly realistic, raising concerns about the malicious use of fabricated facial content. Existing methods often struggle to generalize to unseen domains due to the diverse nature of facial manipulations. In this paper, we revisit the generation process and identify a universal principle: Deepfake images inherently contain information from both source and target identities, while genuine faces maintain a consistent identity. Building upon this insight, we introduce DiffusionFake, a novel plug-and-play framework that reverses the generative process of face forgeries to enhance the generalization of detection models. DiffusionFake achieves this by injecting the features extracted by the detection model into a frozen pre-trained Stable Diffusion model, compelling it to reconstruct the corresponding target and source images. This guided reconstruction process constrains the detection network to capture the source and target related features to facilitate the reconstruction, thereby learning rich and disentangled representations that are more resilient to unseen forgeries. Extensive experiments demonstrate that DiffusionFake significantly improves cross-domain generalization of various detector architectures without introducing additional parameters during inference. The code are available in https://github.com/skJack/DiffusionFake.git. Ke Sun 0016, Shen Chen 0004, Taiping Yao, Hong Liu 0009, Xiaoshuai Sun, Shouhong Ding, Rongrong Ji |
NeurIPS | 4 |
| 2024 | Defense Against Adversarial Attacks Using Topology Aligning Adversarial TrainingabstractRecent works have indicated that deep neural networks (DNNs) are vulnerable to adversarial attacks, wherein an attacker perturbs an input example with human-imperceptible noise that can easily fool the DNNs, resulting in incorrect predictions. This severely limits the application of deep learning in security-critical scenarios, such as face authentication. Adversarial training (AT) is one of the most practical approaches to strengthening the robustness of DNNs. However, existing AT-based methods treat each training sample independently, thereby ignoring the underlying topological structure in the training data. To this end, in this paper, we take full advantage of the topology information and introduce a Topology Aligning Adversarial Training (TAAT) algorithm. TAAT aims to encourage the trained model to maintain consistency in the topological structure within the feature space of both natural and adversarial examples. To ensure the stability and efficiency of topology alignment, we further introduce a novel Knowledge-Guided (KG) training scheme. This scheme explicitly aligns local logit outputs with global topological structures, leveraging a robust auxiliary model to enhance the target model’s performance. To verify the effectiveness of the proposed method, we conduct extensive experiments on popular benchmark datasets (e.g., CIFAR and ImageNet) and evaluate the robustness against state-of-the-art adversarial attacks (e.g., PGD-attack and AutoAttack). The experimental results demonstrate that the proposed method has superior robustness over the previous state-of-the-art methods. Our code and pre-trained models are available at https://github.com/SkyKuang/TAAT. Huafeng Kuang, Hong Liu 0009, Xianming Lin, Rongrong Ji |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Deep Counterfactual Representation Learning for Visual Recognition Against Weather CorruptionsabstractDeep learning has been widely studied for processing and understanding multimedia data, and it does help improve performance. Recent research has shown that deep models are vulnerable to images containing adverse weather corruptions, leading to a safety risk for numerous safety-critical systems (e.g., autonomous driving systems). There are two problems with the current situation. First, collecting data under different weather scenarios is highly difficult in practice. Second, the performance degrades significantly when the training and test data are from different distributions, as exemplified by the weather corrupted test data. As a result, it is challenging to train a model without access to the images containing variations of various weather conditions, and it is difficult to make trained model generalized to unknown data under different weather conditions. In this paper, we introduce aCounterfactual Representation Learning(CRL) method to address these problems. Without access to training data including weather condition variations, our CRL makes the model resistant to unseen test data that has been corrupted by weather condition variations. Our basic idea is inspired by the perspective of counterfactual regularization. We build a causal model that introduces a counterfactual variable to eliminate the unobserved characteristics brought about by weather conditions. In particular, such a counterfactual variable is approximated by randomly shuffled features, echoing the previous empirical observation that the shuffling technique can perturb the shape details while preserving the local textures. We use information theoretic representation learning to encourage the neural networks to learn more powerful and robust features, which consist of two components. We conduct experiments on five benchmark datasets, namely, CIFAR-100-C, ImageNet-C, KITTI-C, BDD100 k, and CityScapes-C, all of which contain weather corruption. The results of our experiments show that our proposed method can not only be a plug-and-play technique but also work nicely for both object recognition and detection. Hong Liu 0009, Yongqing Sun, Yukihiro Bandoh, Masaki Kitahara, Shin'ichi Satoh 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | STAR Loss: Reducing Semantic Ambiguity in Facial Landmark DetectionabstractRecently, deep learning-based facial landmark detection has achieved significant improvement. However, the semantic ambiguity problem degrades detection performance. Specifically, the semantic ambiguity causes inconsistent annotation and negatively affects the model's convergence, leading to worse accuracy and instability prediction. To solve this problem, we propose a Self-adapTive Ambiguity Reduction (STAR) loss by exploiting the properties of se-mantic ambiguity. We find that semantic ambiguity results in the anisotropic predicted distribution, which inspires us to use predicted distribution to represent semantic ambiguity. Based on this, we design the STAR loss that measures the anisotropism of the predicted distribution. Compared with the standard regression loss, STAR loss is encouraged to be small when the predicted distribution is anisotropic and thus adaptively mitigates the impact of semantic ambiguity. Moreover, we propose two kinds of eigen-value restriction methods that could avoid both distribution's abnormal change and the model's premature convergence. Finally, the comprehensive experiments demonstrate that STAR loss outperforms the state-of-the-art methods on three benchmarks, i.e., COFW, 300W, and WFLW, with negligible computation overhead. Code is at https://github.com/ZhenglinZhou/STAR Zhenglin Zhou, Huaxia Li, Hong Liu 0009, Nanyang Wang, Gang Yu 0002, Rongrong Ji |
CVPR | 3 |
| 2023 | Improving Adversarial Robustness via Information Bottleneck DistillationabstractPrevious studies have shown that optimizing the information bottleneck can significantly improve the robustness of deep neural networks. Our study closely examines the information bottleneck principle and proposes an Information Bottleneck Distillation approach. This specially designed, robust distillation technique utilizes prior knowledge obtained from a robust pre-trained model to boost information bottlenecks. Specifically, we propose two distillation strategies that align with the two optimization processes of the information bottleneck. Firstly, we use a robust soft-label distillation method to increase the mutual information between latent features and output prediction. Secondly, we introduce an adaptive feature distillation method that automatically transfers relevant knowledge from the teacher model to the student model, thereby reducing the mutual information between the input and latent features. We conduct extensive experiments to evaluate our approach's robustness against state-of-the-art adversarial attackers such as PGD-attack and AutoAttack. Our experimental results demonstrate the effectiveness of our approach in significantly improving adversarial robustness. Our code is available at https://github.com/SkyKuang/IBD. Huafeng Kuang, Hong Liu 0009, Yongjian Wu 0001, Shin'ichi Satoh 0001, Rongrong Ji |
NeurIPS | 2 |
| 2023 | Mitigating robust overfitting via self-residual-calibration regularization
Hong Liu 0009, Zhun Zhong, Nicu Sebe, Shin'ichi Satoh 0001 |
Artif. Intell. | 1 |
| 2023 | Towards Robust Person Re-Identification by Defending Against Universal AttackersabstractRecent studies show that deep person re-identification (re-ID) models are vulnerable to adversarial examples, so it is critical to improving the robustness of re-ID models against attacks. To achieve this goal, we explore the strengths and weaknesses of existing re-ID models, i.e., designing learning-based attacks and training robust models by defending against the learned attacks. The contributions of this paper are three-fold: First, we build a holistic attack-defense framework to study the relationship between the attack and defense for person re-ID. Second, we introduce a combinatorial adversarial attack that is adaptive to unseen domains and unseen model types. It consists of distortions in pixel and color space (i.e., mimicking camera shifts). Third, we propose a novel virtual-guided meta-learning algorithm for our attack-defense system. We leverage a virtual dataset to conduct experiments under our meta-learning framework, which can explore the cross-domain constraints for enhancing the generalization of the attack and the robustness of the re-ID model. Comprehensive experiments on three large-scale re-ID benchmarks demonstrate that: 1) Our combinatorial attack is effective and highly universal in cross-model and cross-dataset scenarios; 2) Our meta-learning algorithm can be readily applied to different attack and defense approaches, which can reach consistent improvement; 3) The defense model trained on the learning-to-learn framework is robust to recent SOTA attacks that are not even used during training. Fengxiang Yang, Juanjuan Weng, Zhun Zhong, Hong Liu 0009, Zheng Wang 0007, Zhiming Luo, Donglin Cao, Shaozi Li, Shin'ichi Satoh 0001, Nicu Sebe |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Semantically Consistent Visual Representation for Adversarial RobustnessabstractDeep neural networks have been widely used in various domains owing to the success of deep learning. However, recent studies have shown that these models are vulnerable to adversarial examples, leading to inaccurate predictions. In this paper, we focus on the issue of adversarial robustness by examining it through the lens of semantic information, which drives us to propose a new perspective,i.e., adversarial attacks destroy the correlation between visual representations and semantic word vectors, while adversarial training restores it. Additionally, we discover that the correlation among robust representations of different categories aligns with the correlation among the corresponding semantic word vectors. Based on these empirical observations, we incorporate the semantic information into the model training process and propose Semantic Constraint Adversarial Robust Learning (SCARL). Firstly, inspired by the information-theoretical perspective, we maximize mutual information to bridge the information gap between the visual representations and the corresponding semantic word vectors in the embedding space. We further provide a differentiable lower bound to optimize such mutual information efficiently. Secondly, we introduce a novel semantic structure constraint that maintains the structure of visual representations consistent with that of semantic word vectors. Finally, we integrate these techniques with adversarial training to learn robust visual representations. We conduct extensive experiments on several datasets (such as CIFAR and TinyImageNet) and evaluate the robustness against various adversarial attacks (such as PGD-attack and AutoAttack), demonstrating the benefits of incorporating semantic information for improving model robustness. Our code is available at https://github.com/SkyKuang/SCARL. Huafeng Kuang, Hong Liu 0009, Yongjian Wu 0001, Rongrong Ji |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | Win-Win by Competition: Auxiliary-Free Cloth-Changing Person Re-IdentificationabstractRecent person Re-IDentification (ReID) systems have been challenged by changes in personnel clothing, leading to the study of Cloth-Changing person ReID (CC-ReID). Commonly used techniques involve incorporating auxiliary information (e.g., body masks, gait, skeleton, and keypoints) to accurately identify the target pedestrian. However, the effectiveness of these methods heavily relies on the quality of auxiliary information and comes at the cost of additional computational resources, ultimately increasing system complexity. This paper focuses on achieving CC-ReID by effectively leveraging the information concealed within the image. To this end, we introduce an Auxiliary-free Competitive IDentification (ACID) model. It achieves a win-win situation by enriching the identity (ID)-preserving information conveyed by the appearance and structure features while maintaining holistic efficiency. In detail, we build a hierarchical competitive strategy that progressively accumulates meticulous ID cues with discriminating feature extraction at the global, channel, and pixel levels during model inference. After mining the hierarchical discriminative clues for appearance and structure features, these enhanced ID-relevant features are crosswise integrated to reconstruct images for reducing intra-class variations. Finally, by combing with self- and cross-ID penalties, the ACID is trained under a generative adversarial learning framework to effectively minimize the distribution discrepancy between the generated data and real-world data. Experimental results on four public cloth-changing datasets (i.e., PRCC-ReID, VC-Cloth, LTCC-ReID, and Celeb-ReID) demonstrate the proposed ACID can achieve superior performance over state-of-the-art methods. The code is available soon at: https://github.com/BoomShakaY/Win-CCReID. Zhengwei Yang 0001, Xian Zhong, Zhun Zhong, Hong Liu 0009, Zheng Wang 0007, Shin'ichi Satoh 0001 |
IEEE Trans. Image Process. | 4 |
| 2022 | An Information Theoretic Approach for Attention-Driven Face Forgery Detection
Ke Sun 0016, Hong Liu 0009, Taiping Yao, Xiaoshuai Sun, Shen Chen 0004, Shouhong Ding, Rongrong Ji |
ECCV (14) | 2 |
| 2022 | Black-Box Dissector: Towards Erasing-Based Hard-Label Model Stealing Attack
Yixu Wang, Jie Li 0052, Hong Liu 0009, Yan Wang 0059, Yongjian Wu 0001, Feiyue Huang, Rongrong Ji |
ECCV (5) | 3 |
| 2022 | Attentive Decoupling Network for Cloth-Changing Re-IdentificationabstractRecently, Cloth-Changing person Re-IDentification (CC-ReID) plays a vital role in the public security system and social livelihood, and suffers the problem of considerable intra-class variation. This paper demonstrates that coarse-grained appearance and body shape features are helpful for CC-ReID. We propose an Attentive DeCoupling (ADC) Network for CC-ReID without auxiliary information. The proposed network is built on two core designs. First, a joint identification structure is proposed to retain ID-relevant information at appearance and shape levels. Second, Competitive Attention (CA) is adopted, where the model progressively updates attention to accumulate sound cues for discriminating identity (ID). The proposed decoupling process is continuously improved through constant self-defeating competition of the network. Experimental results on the public cloth-changing dataset show the proposed method's effectiveness and generalizability. Zhengwei Yang 0001, Xian Zhong, Hong Liu 0009, Zhun Zhong, Zheng Wang 0007 |
ICME | 3 |
| 2021 | Domain General Face Forgery Detection by Learning to WeightabstractIn this paper, we propose a domain-general model, termed learning-to-weight (LTW), that guarantees face detection performance across multiple domains, particularly the target domains that are never seen before. However, various face forgery methods cause complex and biased data distributions, making it challenging to detect fake faces in unseen domains. We argue that different faces contribute differently to a detection model trained on multiple domains, making the model likely to fit domain-specific biases. As such, we propose the LTW approach based on the meta-weight learning algorithm, which configures different weights for face images from different domains. The LTW network can balance the model's generalizability across multiple domains. Then, the meta-optimization calibrates the source domain's gradient enabling more discriminative features to be learned. The detection ability of the network is further improved by introducing an intra-class compact loss. Extensive experiments on several commonly used deepfake datasets to demonstrate the effectiveness of our method in detecting synthetic faces. Code and supplemental material are available at https://github.com/skJack/LTW. Ke Sun 0016, Hong Liu 0009, Qixiang Ye, Yue Gao 0002, Jianzhuang Liu, Ling Shao 0001, Rongrong Ji |
AAAI | 2 |
| 2021 | Learning to Attack Real-World Models for Person Re-identification via Virtual-Guided Meta-LearningabstractRecent advances in person re-identification (re-ID) have led to impressive retrieval accuracy. However, existing re-ID models are challenged by the adversarial examples crafted by adding quasi-imperceptible perturbations. Moreover, re-ID systems face the domain shift issue that training and testing domains are not consistent. In this study, we argue that learning powerful attackers with high universality that works well on unseen domains is an important step in promoting the robustness of re-ID systems. Therefore, we introduce a novel universal attack algorithm called ``MetaAttack'' for person re-ID. MetaAttack can mislead re-ID models on unseen domains by a universal adversarial perturbation. Specifically, to capture common patterns across different domains, we propose a meta-learning scheme to seek the universal perturbation via the gradient interaction between meta-train and meta-test formed by two datasets. We also take advantage of a virtual dataset (PersonX), instead of real ones, to conduct meta-test. This scheme not only enables us to learn with more comprehensive variation factors but also mitigates the negative effects caused by biased factors of real datasets. Experiments on three large-scale re-ID datasets demonstrate the effectiveness of our method in attacking re-ID models on unseen domains. Our final visualization results reveal some new properties of existing re-ID systems, which can guide us in designing a more robust re-ID model. Code and supplemental material are available at \url{https://github.com/FlyingRoastDuck/MetaAttack_AAAI21}. Fengxiang Yang, Zhun Zhong, Hong Liu 0009, Zheng Wang 0007, Zhiming Luo, Shaozi Li, Nicu Sebe, Shin'ichi Satoh 0001 |
AAAI | 3 |
| 2021 | Towards Robustness Against Natural Language Word Substitutions
Xinshuai Dong, Anh Tuan Luu, Rongrong Ji, Hong Liu 0009 |
ICLR | 4 |
| 2021 | Improving Camouflaged Object Detection with the Uncertainty of Pseudo-edge LabelsabstractThis paper focuses on camouflaged object detection (COD), which is a task to detect objects hidden in the background. Most of the current COD models aim to highlight the target object directly while outputting ambiguous camouflaged boundaries. On the other hand, the performance of the models considering edge information is not yet satisfactory. To this end, we propose a new framework that makes full use of multiple visual cues, i.e., saliency as well as edges, to refine the predicted camouflaged map. This framework consists of three key components, i.e., a pseudo-edge generator, a pseudo-map generator, and an uncertainty-aware refinement module. In particular, the pseudo-edge generator estimates the boundary that outputs the pseudo-edge label, and the conventional COD method serves as the pseudo-map generator that outputs the pseudo-map label. Then, we propose an uncertainty-based module to reduce the uncertainty and noise of such two pseudo labels, which takes both pseudo labels as input and outputs an edge-accurate camouflaged map. Experiments on various COD datasets demonstrate the effectiveness of our method with superior performance to the existing state-of-the-art methods. Nobukatsu Kajiura, Hong Liu 0009, Shin'ichi Satoh 0001 |
MMAsia | 2 |
| 2020 | Projection & Probability-Driven Black-Box AttackabstractGenerating adversarial examples in a black-box setting retains a significant challenge with vast practical application prospects. In particular, existing black-box attacks suffer from the need for excessive queries, as it is non-trivial to find an appropriate direction to optimize in the high-dimensional space. In this paper, we propose Projection & Probability-driven Black-box Attack (PPBA) to tackle this problem by reducing the solution space and providing better optimization. For reducing the solution space, we first model the adversarial perturbation optimization problem as a process of recovering frequency-sparse perturbations with compressed sensing, under the setting that random noise in the low-frequency space is more likely to be adversarial. We then propose a simple method to construct a low-frequency constrained sensing matrix, which works as a plug-and-play projection matrix to reduce the dimensionality. Such a sensing matrix is shown to be flexible enough to be integrated into existing methods like NES and BanditsTD. For better optimization, we perform a random walk with a probability-driven strategy, which utilizes all queries over the whole progress to make full use of the sensing matrix for a less query budget. Extensive experiments show that our method requires at most 24% fewer queries with a higher attack success rate compared with state-of-the-art approaches. Finally, the attack method is evaluated on the real-world online service, i.e., Google Cloud Vision API, which further demonstrates our practical potentials. Jie Li 0052, Rongrong Ji, Hong Liu 0009, Jianzhuang Liu, Bineng Zhong 0001, Cheng Deng 0002, Qi Tian 0001 |
CVPR | 3 |
| 2020 | Anti-bandit Neural Architecture Search for Model Defense
Baochang Zhang 0001, Hong Liu 0009, Rongrong Ji, David S. Doermann |
ECCV (13) | 5 |
| 2020 | API-Net: Robust Generative Classifier via a Single Discriminator
Xinshuai Dong, Hong Liu 0009, Rongrong Ji, Liujuan Cao, Qixiang Ye, Jianzhuang Liu, Qi Tian 0001 |
ECCV (13) | 2 |
| 2020 | Hadamard Matrix Guided Online Hashing
Mingbao Lin, Rongrong Ji, Hong Liu 0009, Xiaoshuai Sun, Shen Chen 0004, Qi Tian 0001 |
Int. J. Comput. Vis. | 3 |
| 2019 | Towards Optimal Discrete Online Hashing with Balanced SimilarityabstractWhen facing large-scale image datasets, online hashing serves as a promising solution for online retrieval and prediction tasks. It encodes the online streaming data into compact binary codes, and simultaneously updates the hash functions to renew codes of the existing dataset. To this end, the existing methods update hash functions solely based on the new data batch, without investigating the correlation between such new data and the existing dataset. In addition, existing works update the hash functions using a relaxation process in its corresponding approximated continuous space. And it remains as an open problem to directly apply discrete optimizations in online hashing. In this paper, we propose a novel supervised online hashing method, termed Balanced Similarity for Online Discrete Hashing (BSODH), to solve the above problems in a unified framework. BSODH employs a well-designed hashing algorithm to preserve the similarity between the streaming data and the existing dataset via an asymmetric graph regularization. We further identify the “data-imbalance” problem brought by the constructed asymmetric graph, which restricts the application of discrete optimization in our problem. Therefore, a novel balanced similarity is further proposed, which uses two equilibrium factors to balance the similar and dissimilar weights and eventually enables the usage of discrete optimizations. Extensive experiments conducted on three widely-used benchmarks demonstrate the advantages of the proposed method over the stateof-the-art methods. Mingbao Lin, Rongrong Ji, Hong Liu 0009, Xiaoshuai Sun, Yongjian Wu 0001, Yunsheng Wu |
AAAI | 3 |
| 2019 | Learning Neural Bag-of-Matrix-Summarization with Riemannian NetworkabstractSymmetric positive defined (SPD) matrix has attracted increasing research focus in image/video analysis, which merits in capturing the Riemannian geometry in its structured 2D feature representation. However, computation in the vector space on SPD matrices cannot capture the geometric properties, which corrupts the classification performance. To this end, Riemannian based deep network has become a promising solution for SPD matrix classification, because of its excellence in performing non-linear learning over SPD matrix. Besides, Riemannian metric learning typically adopts a kNN classifier that cannot be extended to large-scale datasets, which limits its application in many time-efficient scenarios. In this paper, we propose a Bag-of-Matrix-Summarization (BoMS) method to be combined with Riemannian network, which handles the above issues towards highly efficient and scalable SPD feature representation. Our key innovation lies in the idea of summarizing data in a Riemannian geometric space instead of the vector space. First, the whole training set is compressed with a small number of matrix features to ensure high scalability. Second, given such a compressed set, a constant-length vector representation is extracted by efficiently measuring the distribution variations between the summarized data and the latent feature of the Riemannian network. Finally, the proposed BoMS descriptor is integrated into the Riemannian network, upon which the whole framework is end-to-end trained via matrix back-propagation. Experiments on four different classification tasks demonstrate the superior performance of the proposed method over the state-of-the-art methods. Hong Liu 0009, Jie Li 0052, Yongjian Wu 0001, Rongrong Ji |
AAAI | 1 |
| 2019 | Towards Visual Feature TranslationabstractMost existing visual search systems are deployed based upon fixed kinds of visual features, which prohibits the feature reusing across different systems or when upgrading systems with a new type of feature. Such a setting is obviously inflexible and time/memory consuming, which is indeed mendable if visual features can be ``translated" across systems. In this paper, we make the first attempt towards visual feature translation to break through the barrier of using features across different visual search systems. To this end, we propose a Hybrid Auto-Encoder (HAE) to translate visual features, which learns a mapping by minimizing the translation and reconstruction errors. Based upon HAE, an Undirected Affinity Measurement (UAM) is further designed to quantify the affinity among different types of visual features. Extensive experiments have been conducted on several public datasets with sixteen different types of widely-used features in visual search systems. Quantitative results show the encouraging possibilities of feature translation. For the first time, the affinity among widely-used features like SIFT and DELF is reported. Jie Hu 0018, Rongrong Ji, Hong Liu 0009, Shengchuan Zhang, Cheng Deng 0002, Qi Tian 0001 |
CVPR | 3 |
| 2019 | Universal Adversarial Perturbation via Prior Driven Uncertainty ApproximationabstractDeep learning models have shown their vulnerabilities to universal adversarial perturbations (UAP), which are quasi-imperceptible. Compared to the conventional supervised UAPs that suffer from the knowledge of training data, the data-independent unsupervised UAPs are more applicable. Existing unsupervised methods fail to take advantage of the model uncertainty to produce robust perturbations. In this paper, we propose a new unsupervised universal adversarial perturbation method, termed as Prior Driven Uncertainty Approximation (PD-UA), to generate a robust UAP by fully exploiting the model uncertainty at each network layer. Specifically, a Monte Carlo sampling method is deployed to activate more neurons to increase the model uncertainty for a better adversarial perturbation. Thereafter, a textural bias prior to revealing a statistical uncertainty is proposed, which helps to improve the attacking performance. The UAP is crafted by the stochastic gradient descent algorithm with a boosted momentum optimizer, and a Laplacian pyramid frequency model is finally used to maintain the statistical uncertainty. Extensive experiments demonstrate that our method achieves well attacking performances on the ImageNet validation set, and significantly improves the fooling rate compared with the state-of-the-art methods. Hong Liu 0009, Rongrong Ji, Jie Li 0052, Baochang Zhang 0001, Yue Gao 0002, Yongjian Wu 0001, Feiyue Huang |
ICCV | 1 |
| 2019 | Universal Perturbation Attack Against Image RetrievalabstractUniversal adversarial perturbations (UAPs), a.k.a. input-agnostic perturbations, has been proved to exist and be able to fool cutting-edge deep learning models on most of the data samples. Existing UAP methods mainly focus on attacking image classification models. Nevertheless, little attention has been paid to attacking image retrieval systems. In this paper, we make the first attempt in attacking image retrieval systems. Concretely, image retrieval attack is to make the retrieval system return irrelevant images to the query at the top ranking list. It plays an important role to corrupt the neighbourhood relationships among features in image retrieval attack. To this end, we propose a novel method to generate retrieval-against UAP to break the neighbourhood relationships of image features via degrading the corresponding ranking metric. To expand the attack method to scenarios with varying input sizes or untouchable network parameters, a multi-scale random resizing scheme and a ranking distillation strategy are proposed. We evaluate the proposed method on four widely-used image retrieval datasets, and report a significant performance drop in terms of different metrics, such as mAP and mP@10. Finally, we test our attack methods on the real-world visual search engine, i.e., Google Images, which demonstrates the practical potentials of our methods. Jie Li 0052, Rongrong Ji, Hong Liu 0009, Xiaopeng Hong, Yue Gao 0002, Qi Tian 0001 |
ICCV | 3 |
| 2019 | Multi-modal Multi-layer Fusion Network with Average Binary Center Loss for Face Anti-spoofingabstractFace anti-spoofing detection is critical to guarantee the security of biometric face recognition systems. Despite extensive advances in facial anti-spoofing based on single-model image, little work has been devoted to multi-modal anti-spoofing, which is however widely encountered in real-world scenarios. Following the recent progress, this paper mainly focuses on multi-modal face anti-spoofing and aims to solve the following two challenges: (1) how to effectively fuse multi-modal information; and (2) how to effectively learn distinguishable features despite single cross-entropy loss. We propose a novel Multi-modal Multi-layer Fusion Convolutional Neural Network (mmfCNN), which targets at finding a discriminative model for recognizing the subtle differences between live and spoof faces. The mmfCNN can fully use different information provided by diverse modalities, which is based on a weight-adaptation aggregation approach. Specifically, we utilize a multi-layer fusion model to further aggregate the features from different layers, which fuses the low-, mid- and high-level information from different modalities in a unified framework. Moreover, a novel Average Binary Center (ABC) loss is proposed to maximize the dissimilarity between the features of live and spoof faces, which helps to stabilize the training to generate a robust and discriminative model. Extensive experiments conducted on the CISIA-SURF and 3DMAD datasets verify the significance and generalization capability of the proposed method for the face anti-spoofing task. Code is available at: https://github.com/SkyKuang/Face-anti-spoofing. Huafeng Kuang, Rongrong Ji, Hong Liu 0009, Shengchuan Zhang, Xiaoshuai Sun, Feiyue Huang, Baochang Zhang 0001 |
ACM Multimedia | 3 |
| 2019 | Ordinal Constraint Binary Coding for Approximate Nearest Neighbor SearchabstractBinary code learning,a.k.a.hashing, has been successfully applied to the approximate nearest neighbor search in large-scale image collections. The key challenge lies in reducing the quantization error from the original real-valued feature space to a discrete Hamming space. Recent advances in unsupervised hashing advocate the preservation of ranking information, which is achieved by constraining the binary code learning to be correlated with pairwise similarity. However, few unsupervised methods consider the preservation ofordinal relationsin the learning process, which serves as a more basic cue to learn optimal binary codes. In this paper, we propose a novel hashing scheme, termed Ordinal Constraint Hashing (OCH), which embeds the ordinal relation among data points to preserve ranking into binary codes. The core idea is to construct an ordinal graph via tensor product, and then train the hash function over this graph to preserve the permutation relations among data points in the Hamming space. Subsequently, an in-depth acceleration scheme, termed Ordinal Constraint Projection (OCP), is introduced, which approximates the$n$-pair ordinal graph by$L$-pair anchor-based ordinal graph, and reduce the corresponding complexity from$O(n^4)$to$O(L^3)$($L\ll n$). Finally, to make the optimization tractable, we further relax the discrete constrains and design a customized stochastic gradient decent algorithm on the Stiefel manifold. Experimental results on serval large-scale benchmarks demonstrate that the proposed OCH method can achieve superior performance over the state-of-the-art approaches. Hong Liu 0009, Rongrong Ji, Jingdong Wang 0001, Chunhua Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Towards Compact Visual Descriptor via Deep Fisher Network with Binary EmbeddingabstractFisher Vector (FV) has been widely used to aggregate the local descriptors of an image into a global representation in large-scale image retrieval. However, FV has limited learning capability and its parameters are mostly fixed after constructing the codebook, which is inflexible and cannot be trained jointly with deep networks. Moreover, the high dimension of FV makes it difficult to be applied in scenarios compact descriptors are needed. In this paper, we propose a novel compact image description scheme based on Fisher network with binary embedding to solve the large-scale image retrieval problem, which consists of two components: a Fisher encoder component and a binary embedding component. Concretely, the Fisher encoder is a trainable neural network functions as the traditional FV, which aggregates the local descriptors into a global representation. And the binary encoder embeds the high-dimensional FV to a binary vector, which outputs the compact global binary descriptor. To learn such a descriptor, we further introduce a novel and effective loss function, in which maximum margin criterion is exploited to minimize the distances of positive pairs, as well as maximizing the distances of negative pairs. Extensive experiments performed on MPEG-7 CDVS benchmarks and ILSVR2010 demonstrate that the proposed framework can achieve very superior performance over the state-of-the-art methods. Jianqiang Qian, Xianming Lin, Hong Liu 0009, Youming Deng, Rongrong Ji |
ICME | 3 |
| 2018 | Supervised Online Hashing via Hadamard Codebook LearningabstractIn recent years, binary code learning, a.k.a. hashing, has received extensive attention in large-scale multimedia retrieval. It aims to encode high-dimensional data points into binary codes, hence the original high-dimensional metric space can be efficiently approximated via Hamming space. However, most existing hashing methods adopted offline batch learning, which is not suitable to handle incremental datasets with streaming data or new instances. In contrast, the robustness of the existing online hashing remains as an open problem, while the embedding of supervised/semantic information hardly boosts the performance of the online hashing, mainly due to the defect of unknown category numbers in supervised learning. In this paper, we propose an online hashing scheme, termed Hadamard Codebook based Online Hashing (HCOH), which aims to solving the above problems towards robust and supervised online hashing. In particular, we first assign an appropriate high-dimensional binary codes to each class label, which is generated randomly by Hadamard codes. Subsequently, LSH is adopted to reduce the length of such Hadamard codes in accordance with the hash bits, which can adapt the predefined binary codes online, and theoretically guarantee the semantic similarity. Finally, we consider the setting of stochastic data acquisition, which facilitates our method to efficiently learn the corresponding hashing functions via stochastic gradient descend (SGD) online. Notably, the proposed HCOH can be embedded with supervised labels and is not limited to a predefined category number. Extensive experiments on three widely-used benchmarks demonstrate the merits of the proposed scheme over the state-of-the-art methods. Mingbao Lin, Rongrong Ji, Hong Liu 0009, Yongjian Wu 0001 |
ACM Multimedia | 3 |
| 2018 | Dense Auto-Encoder Hashing for Robust Cross-Modality RetrievalabstractCross-modality retrieval has been widely studied, which aims to search images as response to text queries or vice versa. When faced with large-scale dataset, cross-modality hashing serves as an efficient and effective solution, which learns binary codes to approximate the cross-modality similarity in the Hamming space. Most recent cross-modality hashing schemes focus on learning the hash functions from data instances with fully modalities. However, how to learn robust binary codes when facing incomplete modality (i.e., with one modality missed or partially observed), is left unexploited, which however widely occurs in real-world applications. In this paper, we propose a novel cross-modality hashing, termed Dense Auto-encoder Hashing (DAH), which can explicitly impute the missed modality and produce robust binary codes by leveraging the relatedness among different modalities. To that effect, we propose a novel Dense Auto-encoder Network (DAN) to impute the missing modalities, which densely connects each layer to every other layer in a feed-forward fashion. For each layer, a noisy auto-encoder block is designed to calculate the residue between the current prediction and original data. Finally, a hash-layer is added to the end of DAN, which serves as a special binary encoder model to deal with the incomplete modality input. Quantitative experiments on three cross-modality visual search benchmarks, i.e., the Wiki, NUS-WIDE, and FLICKR-25K, have shown that the proposed DAH has superior performance over the state-of-the-art approaches. Hong Liu 0009, Mingbao Lin, Shengchuan Zhang, Yongjian Wu 0001, Feiyue Huang, Rongrong Ji |
ACM Multimedia | 1 |
| 2017 | Ordinal Constrained Binary Code Learning for Nearest Neighbor SearchabstractRecent years have witnessed extensive attention in binary code learning, a.k.a. hashing, for nearest neighbor search problems. It has been seen that high-dimensional data points can quantize into binary codes to give an efficient similarity approximation via Hamming distance. Among the existing schemes, ranking-based hashing is recent promising that targets at preserving ordinal relations of ranking in the Hamming space to minimize retrieval loss. However, the size of the ranking tuples that show the ordinal relations, is quadratic or cubic to the size of training samples. It is so very expensive to embed such ranking tuples in binary code learning, especially given a large-scale training data set. Besides, it remains difficult to build ranking tuples efficiently for most ranking-preserving hashing, which are deployed over an ordinal graph-based setting. To handle these problems, we propose a novel ranking-preserving hashing method, dubbed Ordinal Constraint Hashing (OCH), which efficiently learns the optimal hashing functions with a graph-based approximation to embed the ordinal relations. The core idea is to reduce the size of ordinal graph with ordinal constraint projection, which preserves the ordinal relations through a small data set (such as clusters or random samples). In particular, to learn such hash functions effectively, we further relax the discrete constraints and design a specific stochastic gradient decent algorithm for optimization. Experimental results on three large-scale visual search benchmark datasets, i.e. LabelMe, Tiny100K and GIST1M, show that the proposed OCH method can achieve superior performance over the state-of-the-arts approaches. Hong Liu 0009, Rongrong Ji, Yongjian Wu 0001, Feiyue Huang |
AAAI | 1 |
| 2017 | Cross-Modality Binary Code Learning via Fusion Similarity HashingabstractBinary code learning has been emerging topic in large-scale cross-modality retrieval recently. It aims to map features from multiple modalities into a common Hamming space, where the cross-modality similarity can be approximated efficiently via Hamming distance. To this end, most existing works learn binary codes directly from data instances in multiple modalities, which preserve both intra-and inter-modal similarities respectively. Few methods consider to preserve the fusion similarity among multi-modal instances instead, which can explicitly capture their heterogeneous correlation in cross-modality retrieval. In this paper, we propose a hashing scheme, termed Fusion Similarity Hashing (FSH), which explicitly embeds the graph-based fusion similarity across modalities into a common Hamming space. Inspired by the fusion by diffusion, our core idea is to construct an undirected asymmetric graph to model the fusion similarity among different modalities, upon which a graph hashing scheme with alternating optimization is introduced to learn binary codes that embeds such fusion similarity. Quantitative evaluations on three widely used benchmarks, i.e., UCI Handwritten Digit, MIR-Flickr25K and NUS-WIDE, demonstrate that the proposed FSH approach can achieve superior performance over the state-of-the-art methods. Hong Liu 0009, Rongrong Ji, Yongjian Wu 0001, Feiyue Huang, Baochang Zhang 0001 |
CVPR | 1 |
| 2017 | Toward Optimal Manifold Hashing via Discrete Locally Linear EmbeddingabstractBinary code learning, also known as hashing, has received increasing attention in large-scale visual search. By transforming high-dimensional features to binary codes, the original Euclidean distance is approximated via Hamming distance. More recently, it is advocated that it is the manifold distance, rather than the Euclidean distance, that should be preserved in the Hamming space. However, it retains as an open problem to directly preserve the manifold structure by hashing. In particular, it first needs to build the local linear embedding in the original feature space, and then quantize such embedding to binary codes. Such a two-step coding is problematic and less optimized. Besides, the off-line learning is extremely time and memory consuming, which needs to calculate the similarity matrix of the original data. In this paper, we propose a novel hashing algorithm, termed discrete locality linear embedding hashing (DLLH), which well addresses the above challenges. The DLLH directly reconstructs the manifold structure in the Hamming space, which learns optimal hash codes to maintain the local linear relationship of data points. To learn discrete locally linear embeddingcodes, we further propose a discrete optimization algorithm with an iterative parameters updating scheme. Moreover, an anchor-based acceleration scheme, termed Anchor-DLLH, is further introduced, which approximates the large similarity matrix by the product of two low-rank matrices. Experimental results on three widely used benchmark data sets, i.e., CIFAR10, NUS-WIDE, and YouTube Face, have shown superior performance of the proposed DLLH over the state-of-the-art approaches. Rongrong Ji, Hong Liu 0009, Liujuan Cao, Yongjian Wu 0001, Feiyue Huang |
IEEE Trans. Image Process. | 2 |
| 2016 | Towards Optimal Binary Code Learning via Ordinal EmbeddingabstractBinary code learning, a.k.a., hashing, has been recently popular due to its high efficiency in large-scale similarity search and recognition. It typically maps high-dimensional data points to binary codes, where data similarity can be efficiently computed via rapid Hamming distance. Most existing unsupervised hashing schemes pursue binary codes by reducing the quantization error from an original real-valued data space to a resulting Hamming space. On the other hand, most existing supervised hashing schemes constrain binary code learning to correlate with pairwise similarity labels. However, few methods consider ordinal relations in the binary code learning process, which serve as a very significant cue to learn the optimal binary codes for similarity search. In this paper, we propose a novel hashing scheme, dubbed Ordinal Embedding Hashing (OEH), which embeds given ordinal relations among data points to learn the ranking-preserving binary codes. The core idea is to construct a directed unweighted graph to capture the ordinal relations, and then train the hash functions using this ordinal graph to preserve the permutation relations in the Hamming space. To learn such hash functions effectively, we further relax the discrete constraints and design a stochastic gradient decent algorithm to obtain the optimal solution. Experimental results on two large-scale benchmark datasets demonstrate that the proposed OEH method can achieve superior performance over the state-of-the-arts approaches.At last, the evaluation on query by humming dataset demonstrates the OEH also has good performance for music retrieval by using user's humming or singing. Hong Liu 0009, Rongrong Ji, Yongjian Wu 0001, Wei Liu 0005 |
AAAI | 1 |
| 2016 | Supervised Matrix Factorization for Cross-Modality Hashing
Hong Liu 0009, Rongrong Ji, Yongjian Wu 0001, Gang Hua 0001 |
IJCAI | 1 |