VLDB 2026 Research / reviewers in the wild / expert
Hui Xue 0001
dblp:27/3541-1 · also Hui Xue'
· DBLP profile ↗
73ranked-venue papers
0as first author
63since 2021 · last 2026
0000-0002-2093-2839ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 38 since 2021Graphics, computer vision, multimedia, augmented reality and games · 39 · 30 since 2021Databases, data management, data science and information retrieval · 7 · 7 since 2021Security and privacy · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-ExpertsabstractHaolei Xu, Haiwen Hong, Hongxing Li, Rui Zhou, Yang Zhang, Longtao Huang, Hui Xue, Yongliang Shen, Weiming Lu, Yueting Zhuang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Haolei Xu, Haiwen Hong, Longtao Huang, Hui Xue 0001, Yongliang Shen 0001, Weiming Lu 0001, Yueting Zhuang |
ACL (1) | 7 |
| 2026 | Why Steering Works: Toward a Unified View of Language Model Parameter DynamicsabstractZiwen Xu, Chenyan WU, Hengyu Sun, Haiwen Hong, Mengru Wang, Yunzhi Yao, Longtao Huang, Hui Xue, Shumin Deng, Zhixuan Chu, Huajun Chen, Ningyu Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ziwen Xu, Chenyan Wu, Hengyu Sun, Haiwen Hong, Yunzhi Yao, Longtao Huang, Hui Xue 0001, Shumin Deng, Zhixuan Chu, Huajun Chen, Ningyu Zhang 0001 |
ACL (1) | 8 |
| 2026 | How Controllable Are Large Language Models? A Unified Evaluation across Behavioral GranularitiesabstractZiwen Xu, Kewei Xu, Haoming Xu, Haiwen Hong, Longtao Huang, Hui Xue, Ningyu Zhang, Yongliang Shen, Guozhou Zheng, Huajun Chen, Shumin Deng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ziwen Xu, Kewei Xu, Haiwen Hong, Longtao Huang, Hui Xue 0001, Ningyu Zhang 0001, Yongliang Shen 0001, Guozhou Zheng, Huajun Chen, Shumin Deng |
ACL (1) | 6 |
| 2026 | A Causal Perspective for Enhancing Jailbreak Attack and Defense
Licheng Pan, Yunsheng Lu, Jiexi Liu 0005, Jialing Tao, Haozhe Feng, Hui Xue 0001, Zhixuan Chu, Kui Ren 0001 |
NDSS | 6 |
| 2025 | Learning from Graph: Mitigating Label Noise on Graph through Topological Feature ReconstructionabstractGraph Neural Networks (GNNs) have shown remarkable performance in modeling graph data. However, Labeling graph data typically relies on unreliable information, leading to noisy node labels. Existing approaches for GNNs under Label Noise (GLN) employ supervision signals beyond noisy labels for robust learning. While empirically effective, they tend to over-reliance on supervision signals built upon external assumptions, leading to restricted applicability. In this work, we shift the focus to exploring how to extract useful information and learn from the graph itself, thus achieving robust graph learning. From an information theory perspective, we theoretically and empirically demonstrate that the graph itself contains reliable information for graph learning under label noise. Based on these insights, we propose the Topological Feature Reconstruction (TFR) method. Specifically, TFR leverages the fact that the pattern of clean labels can more accurately reconstruct graph features through topology, while noisy labels cannot. TFR is a simple and theoretically guaranteed model for robust graph learning under label noise. We conduct extensive experiments across datasets with varying properties. The results demonstrate the robustness and broad applicability of our proposed TFR compared to state-of-the-art baselines. Codes are available at https://github.com/eaglelab-zju/TFR. Zhonghao Wang 0002, Yuanchen Bei, Sheng Zhou 0004, Zhiyao Zhou, Jiapei Fan, Hui Xue 0001, Haishuai Wang, Jiajun Bu |
CIKM | 6 |
| 2025 | A Comprehensive Toolkit for Generalized Robust VisionabstractWhile deep neural networks (DNNs) excel in computer vision tasks, their real-world deployment is hindered by robustness limitations compared to human perception. Adversarial attacks and data distribution shifts remain critical vulnerabilities, degrading model performance under practical conditions. To address these challenges and advance robustness research, we introduce a comprehensive, user-friendly toolkit for training, evaluating, and analyzing robust vision models. It targets two key dimensions of robustness: 1) Adversarial robustness-defending against malicious worst-case perturbations (adversarial examples); 2) Natural robustness-maintaining performance under real-world corruptions and distribution shifts. Through extensive image classification benchmarks, our toolkit enables precise model assessment. We envision this toolkit accelerating the development of practically robust models and bridging the gap between machine and human vision capabilities. Zhao Li 0007, Yuefeng Chen, Hui Xue 0001, Xiaofeng Mao |
CIKM | 3 |
| 2025 | A New Model for Prototype-based Continual Learning in Hyperspherical SpaceabstractThe continuous emergence of new objects in the visual world poses a serious challenge to deep object recognition methods, which sparks the increasing study on continual or incremental learning. However, learning new tasks faces the tough catastrophic forgetting problem, i.e., dramatic performance degradation on old tasks. A good continual learning model should be robustly adapted to the upcoming tasks while effectively handling catastrophic forgetting. In this paper, we focus on the class-incremental learning (CIL) task, and propose a novel prototype-based continual learning model C-HPN that projects the visual features into a hypersphere geometric space, where continual learning is conducted. C-HPN features two-fold contributions. On the one hand, instead of using the popular cross-entropy loss, we develop an instance-prototype compact loss to obtain well-clustered hyperspherical embeddings and a prototype-prototype separability loss to boost the model’s generalization by introducing large angle distance inductive bias between prototypes in the hyperspherical space. On the other hand, prototype construction and adaptation strategies are designed for effectively adapting new classes, and an instance-prototype relationship preservation distillation mechanism is introduced to overcome catastrophic forgetting. Extensive experiments on several image datasets validate the effectiveness of the proposed method. Yixin Ren, Yewei Xia, Longtao Huang, Hui Xue 0001, Shuigeng Zhou |
ICASSP | 5 |
| 2025 | Jailbreaking Multimodal Large Language Models via Shuffle InconsistencyabstractMultimodal Large Language Models (MLLMs) have achieved impressive performance and have been put into practical use in commercial applications, but they still have potential safety mechanism vulnerabilities. Jailbreak attacks are red teaming methods that aim to bypass safety mechanisms and discover MLLMs' potential risks. Existing MLLMs' jailbreak methods often bypass the model's safety mechanism through complex optimization methods or carefully designed image and text prompts. Despite achieving some progress, they have a low attack success rate on commercial closed-source MLLMs. Unlike previous research, we empirically find that there exists a Shuffle Inconsistency between MLLMs' comprehension ability and safety ability for the shuffled harmful instruction. That is, from the perspective of comprehension ability, MLLMs can understand the shuffled harmful text-image instructions well. However, they can be easily bypassed by the shuffled harmful instructions from the perspective of safety ability, leading to harmful responses. Then we innovatively propose a text-image jailbreak attack named SI-Attack. Specifically, to fully utilize the Shuffle Inconsistency and overcome the shuffle randomness, we apply a query-based black-box optimization method to select the most harmful shuffled inputs based on the feedback of the toxic judge model. A series of experiments show that SI-Attack can improve the attack's performance on three benchmarks. In particular, SI-Attack can obviously improve the attack success rate for commercial MLLMs such as GPT-4o or Claude-3.5-Sonnet. Ranjie Duan, Caixin Kang, Shouwei Ruan, Jialing Tao, Yuefeng Chen, Hui Xue 0001, Xingxing Wei 0001 |
ICCV | 9 |
| 2025 | Modularized Self-Reflected Video Reasoner for Multimodal LLM with Application to Video Question AnsweringabstractMultimodal Large Language Models (Multimodal LLMs) have shown their strength in Video Question Answering (VideoQA). However, due to the black-box nature of end-to-end training strategies, existing approaches based on Multimodal LLMs suffer from the lack of interpretability for VideoQA: they can neither present reasoning paths nor indicate where the answers are derived from the video. To address this issue, we propose **MSR-ViR** (**M**odularized **S**elf-**R**eflected **Vi**deo **R**easoner), which for the first time integrates modular networks to Multimodal LLMs, capable of providing VideoQA with explicit reasoning paths for more interpretability. Specifically, a **MoST-Grounding** (Modularized Spatial-Temporal Grounding) network is proposed to decompose complex questions via tree-structured policies, localizing relevant temporal and spatial segments within videos through step-by-step reasoning. The proposed MoST-Grounding network provides explicit visually grounded information for Multimodal LLMs with clear reasoning paths, thus enhancing interpretability for the predicted answers. To further improve the reasoning quality, we design an **Alternate Self-reflection Training Strategy** to jointly optimize policy generation and Multimodal LLMs. Experiments on real-world datasets demonstrate the superiority of our proposed MSR-ViR framework in video understanding, reasoning transparency, and providing explicit localization evidence for answers. Zihan Song 0003, Xin Wang 0019, Zi Qian, Hong Chen 0011, Longtao Huang, Hui Xue 0001, Wenwu Zhu 0001 |
ICML | 6 |
| 2025 | Dynamic Mixture of Curriculum LoRA Experts for Continual Multimodal Instruction TuningabstractContinual multimodal instruction tuning is crucial for adapting Multimodal Large Language Models (MLLMs) to evolving tasks. However, most existing methods adopt a fixed architecture, struggling with adapting to new tasks due to static model capacity. We propose to evolve the architecture under parameter budgets for dynamic task adaptation, which remains unexplored and imposes two challenges: 1) task architecture conflict, where different tasks require varying layer-wise adaptations, and 2) modality imbalance, where different tasks rely unevenly on modalities, leading to unbalanced updates. To address these challenges, we propose a novel Dynamic Mixture of Curriculum LoRA Experts (D-MoLE) method, which automatically evolves MLLM's architecture with controlled parameter budgets to continually adapt to new tasks while retaining previously learned knowledge. Specifically, we propose a dynamic layer-wise expert allocator, which automatically allocates LoRA experts across layers to resolve architecture conflicts, and routes instructions layer-wisely to facilitate knowledge sharing among experts. Then, we propose a gradient-based inter-modal continual curriculum, which adjusts the update ratio of each module in MLLM based on the difficulty of each modality within the task to alleviate the modality imbalance problem. Extensive experiments show that D-MoLE significantly outperforms state-of-the-art baselines, achieving a 15 percent average improvement over the best baseline. To the best of our knowledge, this is the first study of continual learning for MLLMs from an architectural perspective. Chendi Ge, Xin Wang 0019, Zeyang Zhang 0001, Hong Chen 0011, Jiapei Fan, Longtao Huang, Hui Xue 0001, Wenwu Zhu 0001 |
ICML | 7 |
| 2025 | Score-based Generative Modeling for Conditional Independence TestingabstractDetermining conditional independence (CI) relationships between random variables is a fundamental yet challenging task in machine learning and statistics, especially in high-dimensional settings. Existing generative model-based CI testing methods, such as those utilizing generative adversarial networks (GANs), often struggle with undesirable modeling of conditional distributions and training instability, resulting in subpar performance. To address these issues, we propose a novel CI testing method via score-based generative modeling, which achieves precise Type I error control and strong testing power. Concretely, we first employ a sliced conditional score matching scheme to accurately estimate conditional score and use Langevin dynamics conditional sampling to generate null hypothesis samples, ensuring precise Type I error control. Then, we incorporate a goodness-of-fit stage into the method to verify generated samples and enhance interpretability in practice. We theoretically establish the error bound of conditional distributions modeled by score-based generative models and prove the validity of our CI tests. Extensive experiments on both synthetic and real-world datasets show that our method significantly outperforms existing state-of-the-art methods, providing a promising way to revitalize generative model-based CI testing. Yixin Ren, Chenghou Jin, Yewei Xia, Longtao Huang, Hui Xue 0001, Hao Zhang 0079, Jihong Guan, Shuigeng Zhou |
KDD (2) | 6 |
| 2025 | Towards the Resistance of Neural Network Fingerprinting to Fine-tuningabstractThis paper proves a new fingerprinting method to embed the ownership information into a deep neural network (DNN) with theoretically guaranteed robustness to fine-tuning. Specifically, we prove that when the input feature of a convolutional layer only contains low-frequency components, specific frequency components of the convolutional filter will not be changed by gradient descent during the fine-tuning process, where we propose a revised Fourier transform to extract frequency components from the convolutional filter. Additionally, we also prove that these frequency components are equivariant to weight scaling and weight permutations. In this way, we design a fingerprint module to embed the fingerprint information into specific frequency components of convolutional filters. Preliminary experiments demonstrate the effectiveness of our method. Ling Tang 0002, Yuefeng Chen, Hui Xue 0001, Quanshi Zhang |
NeurIPS | 3 |
| 2025 | Mirage in the Eyes: Hallucination Attack on Multi-modal Large Language Models with Only Attention Sink
Mi Zhang 0001, Junjie Sun, Chenyue Wang, Min Yang 0002, Hui Xue 0001, Jialing Tao, Ranjie Duan, Jiexi Liu 0005 |
USENIX Security Symposium | 6 |
| 2025 | A Comprehensive Study on Robustness of Image Classification Models: Benchmarking and Rethinking
Chang Liu 0077, Yinpeng Dong, Wenzhao Xiang 0001, Xiao Yang 0028, Hang Su 0006, Jun Zhu 0001, Yuefeng Chen, Yuan He 0011, Hui Xue 0001, Shibao Zheng |
Int. J. Comput. Vis. | 9 |
| 2025 | Improving model generalization by on-manifold adversarial augmentation in the frequency domainabstractDeep Neural Networks (DNNs) often suffer from performance drops when training and test data distributions differ. Ensuring model generalization for Out-Of-Distribution (OOD) data is crucial, but current models still struggle with accuracy on such data. Recent studies have shown that regular or off-manifold adversarial examples as data augmentation improve OOD generalization. Building on this, we provide theoretical validation that on-manifold adversarial examples can enhance OOD generalization even more. However, generating these examples is challenging due to the complexity of real manifolds. To address this, we propose AdvWavAug, an on-manifold adversarial data augmentation method using a Wavelet module. This approach, based on the AdvProp training framework, leverages wavelet transformation to project an image into the wavelet domain and modifies it within the estimated data manifold. Experiments on various models and datasets, including ImageNet and its distorted versions, show that our method significantly improves model generalization, especially for OOD data. Chang Liu 0077, Wenzhao Xiang 0001, Yuan He 0011, Hui Xue 0001, Shibao Zheng, Hang Su 0006 |
J. Vis. Commun. Image Represent. | 4 |
| 2025 | ANF: Crafting Transferable Adversarial Point Clouds via Adversarial Noise FactorizationabstractTransfer-based adversarial attacks involve generating adversarial point clouds in surrogate models and transferring them to other models to assess 3D model robustness. However, current methods rely too much on surrogate model parameters, limiting transferability. In this work, we use Shapley value to identify positive and negative features, guiding optimization of adversarial noise in feature space. To effectively mislead the 3D classifier, we factorize the adversarial noise into positive and negative noise, with the former keeping the features of the adversarial point cloud close to the negative features, and the latter and the adversarial noise moving it away from the positive features. Finally, a novel adversarial point cloud attack method with Adversarial Noise Factorization is proposed, which is abbreviated asANF. ANF simultaneously optimizes the adversarial noise and its positive and negative noise in the feature space, only relying on partial network parameters, which significantly reduces the reliance on the surrogate model and improves the transferability of the adversarial point cloud. Experiments on well-recognized benchmark datasets show that the transferability of adversarial point clouds generated by ANF could be improved by more than 26.7$\%$on average over state-of-the-art transfer-based adversarial attack methods. Hai Chen, Shu Zhao 0005, Xiao Yang 0028, Huanqian Yan, Yuan He 0011, Hui Xue 0001, Fulan Qian, Hang Su 0006 |
IEEE Trans. Big Data | 6 |
| 2024 | UniPSDA: Unsupervised Pseudo Semantic Data Augmentation for Zero-Shot Cross-Lingual Natural Language UnderstandingabstractCross-lingual representation learning transfers knowledge from resource-rich data to resource-scarce ones to improve the semantic understanding abilities of different languages. However, previous works rely on shallow unsupervised data generated by token surface matching, regardless of the global context-aware semantics of the surrounding text tokens. In this paper, we propose an Unsupervised Pseudo Semantic Data Augmentation (UniPSDA) mechanism for cross-lingual natural language understanding to enrich the training data without human interventions. Specifically, to retrieve the tokens with similar meanings for the semantic data augmentation across different languages, we propose a sequential clustering process in 3 stages: within a single language, across multiple languages of a language family, and across languages from multiple language families. Meanwhile, considering the multi-lingual knowledge infusion with context-aware semantics while alleviating computation burden, we directly replace the key constituents of the sentences with the above-learned multi-lingual family knowledge, viewed as pseudo-semantic. The infusion process is further optimized via three de-biasing techniques without introducing any neural parameters. Extensive experiments demonstrate that our model consistently improves the performance on general zero-shot cross-lingual natural language understanding tasks, including sequence classification, information extraction, and question answering. Taolin Zhang 0001, Jiali Deng, Longtao Huang, Chengyu Wang 0001, Hui Xue 0001 |
LREC/COLING | 7 |
| 2024 | KEHRL: Learning Knowledge-Enhanced Language Representations with Hierarchical Reinforcement LearningabstractKnowledge-enhanced pre-trained language models (KEPLMs) leverage relation triples from knowledge graphs (KGs) and integrate these external data sources into language models via self-supervised learning. Previous works treat knowledge enhancement as two independent operations, i.e., knowledge injection and knowledge integration. In this paper, we propose to learn Knowledge-Enhanced language representations with Hierarchical Reinforcement Learning (KEHRL), which jointly addresses the problems of detecting positions for knowledge injection and integrating external knowledge into the model in order to avoid injecting inaccurate or irrelevant knowledge. Specifically, a high-level reinforcement learning (RL) agent utilizes both internal and prior knowledge to iteratively detect essential positions in texts for knowledge injection, which filters out less meaningful entities to avoid diverting the knowledge learning direction. Once the entity positions are selected, a relevant triple filtration module is triggered to perform low-level RL to dynamically refine the triples associated with polysemic entities through binary-valued actions. Experiments validate KEHRL’s effectiveness in probing factual knowledge and enhancing the model’s performance on various natural language understanding tasks. Taolin Zhang 0001, Longtao Huang, Chengyu Wang 0001, Hui Xue 0001 |
LREC/COLING | 6 |
| 2024 | TRELM: Towards Robust and Efficient Pre-training for Knowledge-Enhanced Language ModelsabstractKEPLMs are pre-trained models that utilize external knowledge to enhance language understanding. Previous language models facilitated knowledge acquisition by incorporating knowledge-related pre-training tasks learned from relation triples in knowledge graphs. However, these models do not prioritize learning embeddings for entity-related tokens. Updating all parameters in KEPLM is computationally demanding. This paper introduces TRELM, a Robust and Efficient Pre-training framework for Knowledge-Enhanced Language Models. We observe that text corpora contain entities that follow a long-tail distribution, where some are suboptimally optimized and hinder the pre-training process. To tackle this, we employ a robust approach to inject knowledge triples and employ a knowledge-augmented memory bank to capture valuable information. Moreover, updating a small subset of neurons in the feed-forward networks (FFNs) that store factual knowledge is both sufficient and efficient. Specifically, we utilize dynamic knowledge routing to identify knowledge paths in FFNs and selectively update parameters during pre-training. Experimental results show that TRELM achieves at least a 50% reduction in pre-training time and outperforms other KEPLMs in knowledge probing tasks and multiple knowledge-aware language understanding tasks. Chengyu Wang 0001, Taolin Zhang 0001, Jun Huang 0007, Longtao Huang, Hui Xue 0001 |
LREC/COLING | 8 |
| 2024 | One-dimensional Adapter to Rule Them All: Concepts, Diffusion Models and Erasing ApplicationsabstractThe prevalent use of commercial and open-source diffusion models (DMs) for text-to-image generation prompts risk mitigation to prevent undesired behaviors. Existing concept erasing methods in academia are all based on full parameter or specification-based fine-tuning, from which we observe the following issues: 1) Generation alteration towards erosion: Parameter drift during target elimination causes alterations and potential deformations across all generations, even eroding other concepts at varying degrees, which is more evident with multi-concept erased; 2) Transfer in-ability & deployment inefficiency: Previous model-specific erasure impedes the flexible combination of concepts and the training-free transfer towards other models, resulting in linear cost growth as the deployment scenarios increase. To achieve non-invasive, precise, customizable, and transferable elimination, we ground our erasing framework on one-dimensional adapters to erase multiple concepts from most DMs at once across versatile erasing applications. The concept-SemiPermeable structure is injected as a Membrane (SPM) into any DM to learn targeted erasing, and mean-time the alteration and erosion phenomenon is effectively mitigated via a novel Latent Anchoring fine-tuning strategy. Once obtained, SPMs can be flexibly combined and plug-and-play for other DMs without specific re-tuning, enabling timely and efficient adaptation to diverse scenarios. During generation, our Facilitated Transport mechanism dynamically regulates the permeability of each SPM to re-spond to different input prompts, further minimizing the impact on other concepts. Quantitative and qualitative results across ~40 concepts, 7 DMs and 4 erasing applications have demonstrated the superior erasing of SPM. Our code and pre-tuned SPMs are available on the project page https:/lyumengyao.github.io/projects/spm. Mengyao Lyu, Yuhong Yang 0008, Haiwen Hong, Hui Chen 0013, Xuan Jin, Yuan He 0011, Hui Xue 0001, Jungong Han, Guiguang Ding |
CVPR | 7 |
| 2024 | R4: Reinforced Retriever-Reorder-Responder for Retrieval-Augmented Large Language ModelsabstractRetrieval-augmented large language models (LLMs) leverage relevant content retrieved by information retrieval systems to generate correct responses, aiming to alleviate the hallucination problem. However, existing retriever-responder methods typically append relevant documents to the prompt of LLMs to perform text generation tasks without considering the interaction of fine-grained structural semantics between the retrieved documents and the LLMs. This issue is particularly important for accurate response generation as LLMs tend to “lose in the middle” when dealing with input prompts augmented with lengthy documents. In this work, we propose a new pipeline named “Reinforced Retriever-Reorder-Responder” (R4) to learn document orderings for retrieval-augmented LLMs, thereby further enhancing their generation abilities while the large numbers of parameters of LLMs remain frozen. The reordering learning process is divided into two steps according to the quality of the generated responses: document order adjustment and document representation enhancement. Specifically, document order adjustment aims to organize retrieved document orderings into beginning, middle, and end positions based on graph attention learning, which maximizes the reinforced reward of response quality. Document representation enhancement further refines the representations of retrieved documents for responses of poor quality via document-level gradient adversarial learning. Extensive experiments demonstrate that our proposed pipeline achieves better factual question-answering performance on knowledge-intensive tasks compared to strong baselines across various public datasets. The source codes and trained models will be released upon paper acceptance. Taolin Zhang 0001, Qizhou Chen, Chengyu Wang 0001, Longtao Huang, Hui Xue 0001, Jun Huang 0007 |
ECAI | 6 |
| 2024 | Lifelong Knowledge Editing for LLMs with Retrieval-Augmented Continuous Prompt LearningabstractModel editing aims to correct outdated or erroneous knowledge in large language models (LLMs) without the need for costly retraining.Lifelong model editing is the most challenging task that caters to the continuous editing requirements of LLMs.Prior works primarily focus on single or batch editing; nevertheless, these methods fall short in lifelong editing scenarios due to catastrophic knowledge forgetting and the degradation of model performance.Although retrieval-based methods alleviate these issues, they are impeded by slow and cumbersome processes of integrating the retrieved knowledge into the model.In this work, we introduce RECIPE, a RetriEval-augmented ContInuous Prompt lEarning method, to boost editing efficacy and inference efficiency in lifelong learning.RECIPE first converts knowledge statements into short and informative continuous prompts, prefixed to the LLM's input query embedding, to efficiently refine the response grounded on the knowledge.It further integrates the Knowledge Sentinel (KS) that acts as an intermediary to calculate a dynamic threshold, determining whether the retrieval repository contains relevant knowledge.Our retriever and prompt encoder are jointly trained to achieve editing properties, i.e., reliability, generality, and locality.In our experiments, RECIPE is assessed extensively across multiple LLMs and editing datasets, where it achieves superior editing performance.RECIPE also demonstrates its capability to maintain the overall performance of LLMs alongside showcasing fast editing and inference speed. Qizhou Chen, Taolin Zhang 0001, Chengyu Wang 0001, Longtao Huang, Hui Xue 0001 |
EMNLP | 7 |
| 2024 | Large Language Model with Curriculum Reasoning for Visual Concept RecognitionabstractVisual concept recognition aims to capture the basic attributes of an image and reason about the relationships among them to determine whether the image satisfies a certain concept, and has been widely used in various tasks such as human action recognition and image risk warning. Most existing works adopt deep neural networks for visual concept recognition, which are black-box and incomprehensible to humans, thus making them unacceptable for sensitive domains such as prohibited event detection and risk early warning etc. To address this issue, we propose to combine large language model (LLM) with explainable symbolic reasoning via curriculum reweighting to increase the interpretability and accuracy of visual concept recognition in this paper. However, realizing this goal is challenging given that i) the performance of symbolic representations are limited by the lack of annotated reasoning symbols and rules for most tasks, and ii) the LLMs may suffer from knowlege hallucination and dynamic open environment. To address these issues, in this paper, we propose CurLLM-Reasoner, a curriculum reasoning method based on symbolic reasoning and large language model for visual concept recognition. Specifically, we propose a novel rule enhancement module with a tool library, which fully leverage the reasoning capability of large language models and can generate human-understandable rules without any annotation. We further propose a curriculum data resampling methodology to help the large language model accurately extract from easy to complex rules at different reasoning stages. Extensive experiments on various datasets demonstrate that CurLLM-Reasoner can achieve the state-of-the-art visual concept recognition results with explainable rules while free of human annotations. Yipeng Zhang 0003, Xin Wang 0019, Hong Chen 0011, Jiapei Fan, Weigao Wen, Hui Xue 0001, Hong Mei 0001, Wenwu Zhu 0001 |
KDD | 6 |
| 2024 | Differential-Perceptive and Retrieval-Augmented MLLM for Change CaptioningabstractChange captioning involves describing the subtle changes between a pair of similar images. Although existing efforts have achieved compelling success, they overlook the potential of multimodal large language models (MLLMs) in tackling this challenging task. In this work, we aim to empower MLLMs with the capability to perceive subtle differences between paired images and enhance their performance in generating change captions. Specifically, we present a diFferentIal-perceptive aNd rEtRieval-augmented MLLM (FINER-MLLM) tailored for this task. In particular, FINER-MLLM leverages LoRA fine-tuned MLLM's image encoder to extract image patch features, enabling the capture of detailed image information. Subsequently, within MLLM's feature extraction, typically Q-Former, FINER-MLLM incorporates dual constraints: the intra-image feature independence constraint and the inter-image feature alignment constraint. These constraints ensure that the features can comprehensively extract subtle visual information within each image and that corresponding features across images align effectively. Last, we introduced the retrieval augmentation to first retrieve the relevant corpus to facilitate the MLLM's decoder i.e., LLM, in generating accurate change captions. Extensive experiments on three benchmark datasets, i.e., CLEVR-Change, Spot-the-Diff, and Image-Editing-Request, demonstrate the superiority of our proposed method. Haokun Wen, Jianlong Wu, Pengda Qin, Hui Xue 0001, Liqiang Nie |
ACM Multimedia | 5 |
| 2024 | Synthesizing Coherent Story with Auto-Regressive Latent Diffusion ModelsabstractConditioned diffusion models have demonstrated state-of-the-art text-to-image synthesis capacity. Recently, most works focus on synthesizing independent images; While for real-world applications, it is common and necessary to generate a series of coherent images for story-stelling. In this work, we mainly focus on story visualization and continuation tasks and propose AR-LDM, a latent diffusion model auto-regressively conditioned on history captions and generated images. Moreover, AR-LDM can generalize to new characters through adaptation. To our best knowledge, this is the first work successfully leveraging diffusion models for coherent visual story synthesizing. It also extends the text-conditioned method to multimodal conditioning. Quantitative results show that AR-LDM achieves SoTA FID scores on PororoSV, FlintstonesSV, and the adopted challenging dataset VIST containing natural images. Large-scale human evaluations show that AR-LDM has superior performance in terms of quality, relevance, and consistency. Code available at this https URL Xichen Pan, Pengda Qin, Hui Xue 0001, Wenhu Chen |
WACV | 4 |
| 2024 | Context-Aware Robust Fine-Tuning
Xiaofeng Mao, Xiaojun Jia, Rong Zhang 0006, Hui Xue 0001, Zhao Li 0007 |
Int. J. Comput. Vis. | 5 |
| 2024 | Rethinking Out-of-Distribution Detection From a Human-Centric Perspective
Yao Zhu 0003, Yuefeng Chen, Rong Zhang 0006, Hui Xue 0001, Xiang Tian 0002, Rongxin Jiang 0001, Bolun Zheng, Yaowu Chen |
Int. J. Comput. Vis. | 5 |
| 2024 | Revisiting and Exploring Efficient Fast Adversarial Training via LAW: Lipschitz Regularization and Auto Weight AveragingabstractFast Adversarial Training (FAT) not only improves the model robustness but also reduces the training cost of standard adversarial training. However, fast adversarial training often suffers from Catastrophic Overfitting (CO), which results in poor robustness performance. Catastrophic Overfitting describes the phenomenon of a sudden and significant decrease in robust accuracy during the training of fast adversarial training. Many effective techniques have been developed to prevent Catastrophic Overfitting and improve the model robustness from different perspectives. However, these techniques adopt inconsistent training settings and require different training costs, i.e., training time and memory costs, leading to unfair comparisons. In this paper, we conduct a comprehensive study of over 10 fast adversarial training methods in terms of adversarial robustness and training costs. We revisit the effectiveness and efficiency of fast adversarial training techniques in preventing Catastrophic Overfitting from the perspective of model local nonlinearity and propose an effective Lipschitz regularization method for fast adversarial training. Furthermore, we explore the effect of data augmentation and weight averaging in fast adversarial training and propose a simple yet effective auto weight averaging method to improve robustness further. By assembling these techniques, we propose an effective FGSM-based fast adversarial training method equipped with Lipschitz regularization and Auto Weight averaging, abbreviated as FGSM-LAW. Experimental evaluations on four benchmark databases demonstrate the superiority of our method over state-of-the-art fast adversarial training methods and the advanced standard adversarial training methods. Xiaojun Jia, Yuefeng Chen, Xiaofeng Mao, Ranjie Duan, Jindong Gu, Rong Zhang 0006, Hui Xue 0001, Yang Liu 0003, Xiaochun Cao |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2023 | Incremental Graph Classification by Class Prototype Construction and AugmentationabstractGraph neural networks (GNNs) are prone to catastrophic forgetting of past experience in continuous learning scenarios. In this work, we propose a novel method for class-incremental graph learning (CGL) by class prototype construction and augmentation, which can effectively overcome catastrophic forgetting and requires no storage of exemplars (i.e., data-free). Concretely, on the one hand, we construct class prototypes in the embedding space that contain rich topological information of nodes or graphs to represent past data, which are then used for future learning. On the other hand, to boost the adaptability of the model to new classes, we employ class prototype augmentation (PA) to create virtual classes by combining current prototypes. Theoretically, we show that PA can promote the model's adaptation to new data and reduce the inconsistency of old prototypes in the embedding space, therefore further mitigate catastrophic forgetting. Extensive experiments on both node and graph classification datasets show that our method significantly outperforms the existing methods in reducing catastrophic forgetting, and beats the existing methods in most cases in terms of classification accuracy. Yixin Ren, Dong Li 0037, Hui Xue 0001, Zhao Li 0007, Shuigeng Zhou |
CIKM | 4 |
| 2023 | ImageNet-E: Benchmarking Neural Network Robustness via Attribute EditingabstractRecent studies have shown that higher accuracy on ImageNet usually leads to better robustness against different corruptions. Therefore, in this paper, instead of following the traditional research paradigm that investigates new out-of-distribution corruptions or perturbations deep models may encounter, we conduct model debugging in in-distribution data to explore which object attributes a model may be sensitive to. To achieve this goal, we create a toolkit for object editing with controls of backgrounds, sizes, positions, and directions, and create a rigorous benchmark named ImageNet-E(diting) for evaluating the image classifier robustness in terms of object attributes. With our ImageNet-E, we evaluate the performance of current deep learning models, including both convolutional neural networks and vision transformers. We find that most models are quite sensitive to attribute changes. A small change in the background can lead to an average of 9.23% drop on top-1 accuracy. We also evaluate some robust models including both adversarially trained models and other robust trained models and find that some models show worse robustness against attribute changes than vanilla models. Based on these findings, we discover ways to enhance attribute robustness with preprocessing, architecture designs, and training strategies. We hope this work can provide some insights to the community and open up a new avenue for research in robust computer vision. The code and dataset are available at https://github.com/alibaba/easyrobust. Yuefeng Chen, Yao Zhu 0003, Shuhui Wang, Rong Zhang 0006, Hui Xue 0001 |
CVPR | 6 |
| 2023 | Improving Hyper-relational Knowledge Graph Representation with Multi-grained Encoding
Longtao Huang, Hui Xue 0001 |
DASFAA (2) | 3 |
| 2023 | Transaudio: Towards the Transferable Adversarial Audio Attack Via Learning Contextualized PerturbationsabstractIn a transfer-based attack against Automatic Speech Recognition (ASR) systems, attacks are unable to access the architecture and parameters of the target model. Existing attack methods are mostly investigated in voice assistant scenarios with restricted voice commands, prohibiting their applicability to more general ASR related applications. To tackle this challenge, we propose a novel contextualized attack with deletion, insertion, and substitution adversarial behaviors, namely TransAudio, which achieves arbitrary word-level attacks based on the proposed two-stage framework. To strengthen the attack transferability, we further introduce an audio score-matching optimization strategy to regularize the training process, which mitigates adversarial example over-fitting to the surrogate model. Extensive experiments and analysis demonstrate the effectiveness of TransAudio against open-source ASR models and commercial APIs. Gege Qi, Yuefeng Chen, Yao Zhu 0003, Binyuan Hui, Xiaofeng Mao, Rong Zhang 0006, Hui Xue 0001 |
ICASSP | 8 |
| 2023 | COCO-O: A Benchmark for Object Detectors under Natural Distribution ShiftsabstractPractical object detection application can lose its effectiveness on image inputs with natural distribution shifts. This problem leads the research community to pay more attention on the robustness of detectors under Out-Of-Distribution (OOD) inputs. Existing works construct datasets to benchmark the detector’s OOD robustness for a specific application scenario, e.g., Autonomous Driving. However, these datasets lack universality and are hard to benchmark general detectors built on common tasks such as COCO. To give a more comprehensive robustness assessment, we introduce COCO-O(ut-of-distribution), a test dataset based on COCO with 6 types of natural distribution shifts. COCO-O has a large distribution gap with training data and results in a significant 55.7% relative performance drop on a Faster R-CNN detector. We leverage COCO-O to conduct experiments on more than 100 modern object detectors to investigate if their improvements are credible or just over-fitting to the COCO test set. Unfortunately, most classic detectors in early years do not exhibit strong OOD generalization. We further study the robustness effect on recent breakthroughs of detector’s architecture design, augmentation and pre-training techniques. Some empirical findings are revealed: 1) Compared with detection head or neck, backbone is the most important part for robustness; 2) An end-to-end detection transformer design brings no enhancement, and may even reduce robustness; 3) Large-scale foundation models have made a great leap on robust object detection. We hope our COCO-O could provide a rich testbed for robustness study of object detection. The dataset will be available at https://github.com/alibaba/easyrobust/tree/main/benchmarks/coco_o. Xiaofeng Mao, Yuefeng Chen, Yao Zhu 0003, Da Chen 0003, Hang Su 0006, Rong Zhang 0006, Hui Xue 0001 |
ICCV | 7 |
| 2023 | Open-Vocabulary Object Detection With an Open CorpusabstractExisting open vocabulary object detection (OVD) works expand the object detector toward open categories by replacing the classifier with the category text embeddings and optimizing the region-text alignment on data of the base categories. However, both the class-agnostic proposal generator and the classifier are biased to the seen classes as demonstrated by the gaps of objectness and accuracy assessment between base and novel classes. In this paper, an open corpus, composed of a set of external object concepts and clustered to several centroids, is introduced to improve the generalization ability in the detector. We propose the generalized objectness assessment (GOAT) in the proposal generator based on the visual-text alignment, where the similarities of visual feature to the cluster centroids are summarized as the objectness. This simple heuristic evaluates objectness with concepts in open corpus and is thus generalized to open categories. We further propose category expanding (CE) with open corpus in two training tasks, which enables the detector to perceive more categories in the feature space and get more reasonable optimization direction. For the classification task, we introduce an open corpus classifier by reconstructing original classifier with similar words in text space. For the image-caption alignment task, the open corpus centroids are incorporated to enlarge the negative samples in the contrastive loss. Extensive experiments demonstrate the effectiveness of GOAT and CE, which greatly improve the performance on novel classes and get new state-of-the-art on the OVD benchmarks. Haiwen Hong, Xuan Jin, Yuan He 0011, Hui Xue 0001, Zhou Zhao 0001 |
ICCV | 6 |
| 2023 | Inequality phenomenon in l∞-adversarial training, and its unrealized threats
Ranjie Duan, Yuefeng Chen, Yao Zhu 0003, Xiaojun Jia, Rong Zhang 0006, Hui Xue 0001 |
ICLR | 6 |
| 2023 | Budgeted Training for Vision Transformer
Zhuofan Xia, Xuran Pan, Xuan Jin, Yuan He 0011, Hui Xue 0001, Shiji Song, Gao Huang 0001 |
ICLR | 5 |
| 2023 | DR-Net: A Multi-View Face Synthesis Network Driven by Dual RepresentationabstractWith the rapid growth of virtual human techniques and the popularity of 3D face modeling, multi-view face synthesis has received considerable attention in the computer vision community. Though learning-based methods have achieved photo-realistic effect, the lack of geometry guidance remains a great hurdle to their results in high fidelity of identity. To relieve it, we propose a novel network named Dual-Representation-Net (DR-Net) driven by both learning and geometry information. In addition, we utilize face parsing to lead the learning of regional attention, in this way, the two synthesizing representations can collaborate and compete dynamically under different face angles, which improves the stability of face synthesis by assimilating the benefits of both kinds of information. What is more, we want to emphasize that DR-Net can serve as a general framework for multi-view face synthesis by jointly exploiting learning and geometry information. That is, our geometry module and learning module are actually substitutable, which can be replaced by any other SOTA methods. This makes our method both flexible and scalable. Extensive experiments verify that our approach is superior to other counterparts in both synthesis quality and identity preservation. Moreover, with the geometric information, we can decouple the synthesis of face and background to obtain more robust results on in-the-wild data. Xianliang Huang, Yining Lang, Yuan He 0011, Hui Xue 0001, Shuigeng Zhou |
ICME | 5 |
| 2023 | Robust Automatic Speech Recognition via WavAugment Guided Phoneme Adversarial Training
Gege Qi, Yuefeng Chen, Xiaofeng Mao, Xiaojun Jia, Ranjie Duan, Rong Zhang 0006, Hui Xue 0001 |
INTERSPEECH | 7 |
| 2023 | FairRec: Fairness Testing for Deep Recommender SystemsabstractDeep learning-based recommender systems (DRSs) are increasingly and widely deployed in the industry, which brings significant convenience to people’s daily life in different ways. However, recommender systems are also shown to suffer from multiple issues, e.g., the echo chamber and the Matthew effect, of which the notation of “fairness” plays a core role. For instance, the system may be regarded as unfair to 1) a specific user, if the user gets worse recommendations than other users, or 2) an item (to recommend), if the item is much less likely to be exposed to the users than other items. While many fairness notations and corresponding fairness testing approaches have been developed for traditional deep classification models, they are essentially hardly applicable to DRSs. One major challenge is that there still lacks a systematic understanding and mapping between the existing fairness notations and the diverse testing requirements for deep recommender systems, not to mention further testing or debugging activities. To address the gap, we propose FairRec, a unified framework that supports fairness testing of DRSs from multiple customized perspectives, e.g., model utility, item diversity, item popularity, etc. We also propose a novel, efficient search-based testing approach to tackle the new challenge, i.e., double-ended discrete particle swarm optimization (DPSO) algorithm, to effectively search for hidden fairness issues in the form of certain disadvantaged groups from a vast number of candidate groups. Given the testing report, by adopting a simple re-ranking mitigation strategy on these identified disadvantaged groups, we show that the fairness of DRSs can be significantly improved. We conducted extensive experiments on multiple industry-level DRSs adopted by leading companies. The results confirm that FairRec is effective and efficient in identifying the deeply hidden fairness issues, e.g., achieving ∼95% testing accuracy with ∼half to 1/8 time. Huizhong Guo 0001, Jingyi Wang 0004, Dongxia Wang 0002, Zehong Hu, Rong Zhang 0006, Hui Xue 0001 |
ISSTA | 8 |
| 2023 | Spatio-Temporal Catcher: A Self-Supervised Transformer for Deepfake Video DetectionabstractAs deepfake technology has become increasingly sophisticated and accessible, making it easier for individuals with malicious intent to create convincing fake content, which has raised considerable concern in the multimedia and computer vision community. Despite significant advances in deepfake video detection, most existing methods mainly focused on model architecture and training processes with little focus on data perspectives. In this paper, we argue that data quality has become the main bottleneck of current research. To be specific, in the pre-training phase, the domain shift between pre-training and target datasets may lead to poor generalization ability. Meanwhile, in the training phase, the low fidelity of the existing datasets leads to detectors relying on specific low-level visual artifacts or inconsistency. To overcome the shortcomings, (1). In the pre-training phase, pre-train our model on high-quality facial videos by utilizing data-efficient reconstruction-based self-supervised learning to solve domain shift. (2). In the training phase, we develop a novel spatio-temporal generator that can synthesize various high-quality "fake" videos in large quantities at a low cost, which enables our model to learn more general spatio-temporal representations in a self-supervised manner. (3). Additinally, to take full advantage of synthetic "fake" videos, we adopt diversity losses at both frame and video levels to explore the diversity of clues in "fake" videos. Our proposed framework is data-efficient and does not require any real-world deepfake videos. Extensive experiments demonstrate that our method significantly improves the generalization capability. Particularly on the most challenging CDF and DFDC datasets, our method outperforms the baselines by 8.88% and 7.73% points, respectively. Maosen Li, Xurong Li, Cheng Deng 0002, Heng Huang 0001, Feng Mao, Hui Xue 0001, Minghao Li 0008 |
ACM Multimedia | 7 |
| 2023 | BiFPro: A Bidirectional Facial-data Protection Framework against DeepFakeabstractThe rapid progress of the DeepFake technique has caused severe privacy problems. Thus protecting facial data against DeepFake becomes an urgent requirement. Face protection can be regarded as a bidirectional process: Face-out-detection (FOD) and Face-in-forensics (FIF). For FOD, the detectability should be satisfied when using the protected face to replace other faces. For FIF, traceability should be guaranteed when the protected face is replaced by others. For this, we propose a Bidirectional Facial-data Protection Framework (BiFPro) to protect face data comprehensively. This framework is composed of three main parts: Watermarking embedding, Face-out-detection (FOD) and Face-in-forensics (FIF). For the FOD case, we ensure the vulnerability of the original face by embedding fragile watermarking. Once the protected facial image is used to replace other faces, the watermarking information will be corrupted in the synthesized face images which can be used to detect the authenticity of the protected facial images. As for the FIF case, we guarantee the traceability of the protected face image by embedding robust watermarking, with which the fake faces can be traced with the reserved watermarking even after the face is swapped. Experimental results demonstrate that our proposed BiFPro could generate the watermarking which is fragile to FOD and at the same time robust to FIF with an average watermark extraction success rate reaching more than 95% when defending against the four advanced DeepFake techniques. Finally, we hope this work can encourage more initiative countermeasures against DeepFake. Honggu Liu, Wenbo Zhou 0004, Han Fang 0004, Paolo Bestagini, Weiming Zhang 0001, Yuefeng Chen, Stefano Tubaro, Nenghai Yu, Yuan He 0011, Hui Xue 0001 |
ACM Multimedia | 11 |
| 2023 | Model Inversion Attack via Dynamic Memory LearningabstractModel Inversion (MI) attacks aim to recover the private training data from the target model, which has raised security concerns about the deployment of DNNs in practice. Recent advances in generative adversarial models have rendered them particularly effective in MI attacks, primarily due to their ability to generate high-fidelity and perceptually realistic images that closely resemble the target data. In this work, we propose a novel Dynamic Memory Model Inversion Attack (DMMIA) to leverage historically learned knowledge, which interacts with samples (during the training) to induce diverse generations. DMMIA constructs two types of prototypes to inject the information about historically learned knowledge: Intra-class Multicentric Representation (IMR) representing target-related concepts by multiple learnable prototypes, and Inter-class Discriminative Representation (IDR) characterizing the memorized samples as learned prototypes to capture more privacy-related information. As a result, our DMMIA has a more informative representation, which brings more diverse and discriminative generated results. Experiments on multiple benchmarks show that DMMIA performs better than state-of-the-art MI attack methods. Gege Qi, Yuefeng Chen, Xiaofeng Mao, Binyuan Hui, Rong Zhang 0006, Hui Xue 0001 |
ACM Multimedia | 7 |
| 2023 | Spectral Invariant Learning for Dynamic Graphs under Distribution ShiftsabstractDynamic graph neural networks (DyGNNs) currently struggle with handling distribution shifts that are inherent in dynamic graphs.
Existing work on DyGNNs with out-of-distribution settings only focuses on the time domain, failing to handle cases involving distribution shifts in the spectral domain. In this paper, we discover that there exist cases with distribution shifts unobservable in the time domain while observable in the spectral domain, and propose to study distribution shifts on dynamic graphs in the spectral domain for the first time.
However, this investigation poses two key challenges: i) it is non-trivial to capture different graph patterns that are driven by various frequency components entangled in the spectral domain; and ii) it remains unclear how to handle distribution shifts with the discovered spectral patterns. To address these challenges, we propose Spectral Invariant Learning for Dynamic Graphs under Distribution Shifts (SILD), which can handle distribution shifts on dynamic graphs by capturing and utilizing invariant and variant spectral patterns. Specifically, we first design a DyGNN with Fourier transform to obtain the ego-graph trajectory spectrums, allowing the mixed dynamic graph patterns to be transformed into separate frequency components. We then develop a disentangled spectrum mask to filter graph dynamics from various frequency components and discover the invariant and variant spectral patterns. Finally, we propose invariant spectral filtering, which encourages the model to rely on invariant patterns for generalization under distribution shifts. Experimental results on synthetic and real-world dynamic graph datasets demonstrate the superiority of our method for both node classification and link prediction tasks under distribution shifts. Zeyang Zhang 0001, Xin Wang 0019, Ziwei Zhang 0001, Zhou Qin 0002, Weigao Wen, Hui Xue 0001, Haoyang Li 0001, Wenwu Zhu 0001 |
NeurIPS | 6 |
| 2023 | Continual Few-shot Learning with Transformer Adaptation and Knowledge RegularizationabstractContinual few-shot learning, as a paradigm that simultaneously solves continual learning and few-shot learning, has become a challenging problem in machine learning. An eligible continual few-shot learning model is expected to distinguish all seen classes upon new categories arriving, where each category only includes very few labeled data. However, existing continual few-shot learning methods only consider the visual modality, where the distributions of new categories often indistinguishably overlap with old categories, thus resulting in the severe catastrophic forgetting problem. To tackle this problem, in this paper we study continual few-shot learning with the assistance of semantic knowledge by simultaneously taking both visual modality and semantic concepts of categories into account. We propose a Continual few-shot learning algorithm with Semantic knowledge Regularization (CoSR) for adapting to the distribution changes of visual prototypes through a Transformer-based prototype adaptation mechanism. Specifically, the original visual prototypes from the backbone are fed into the well-designed Transformer with corresponding semantic concepts, where the semantic concepts are extracted from all categories. The semantic-level regularization forces the categories with similar semantics to be closely distributed, while the opposite ones are constrained to be far away from each other. The semantic regularization improves the model’s ability to distinguish between new and old categories, thus significantly mitigating the catastrophic forgetting problem in continual few-shot learning. Extensive experiments on CIFAR100, miniImageNet, CUB200 and an industrial dataset with long-tail distribution demonstrate the advantages of our CoSR model compared with state-of-the-art methods. Xin Wang 0019, Yue Liu 0025, Jiapei Fan, Weigao Wen, Hui Xue 0001, Wenwu Zhu 0001 |
WWW | 5 |
| 2023 | To make yourself invisible with Adversarial Semantic Contours
Yichi Zhang 0012, Hang Su 0006, Jun Zhu 0001, Shibao Zheng, Yuan He 0011, Hui Xue 0001 |
Comput. Vis. Image Underst. | 7 |
| 2022 | Towards Robust Vision TransformerabstractRecent advances on Vision Transformer (ViT) and its improved variants have shown that self-attention-based networks surpass traditional Convolutional Neural Networks (CNNs) in most vision tasks. However, existing ViTs focus on the standard accuracy and computation cost, lacking the investigation of the intrinsic influence on model robustness and generalization. In this work, we conduct systematic evaluation on components of ViTs in terms of their impact on robustness to adversarial examples, common corruptions and distribution shifts. We find some components can be harmful to robustness. By leveraging robust components as building blocks of ViTs, we propose Robust Vision Transformer (RVT), which is a new vision transformer and has superior performance with strong robustness. Inspired by the findings during the evaluation, we further propose two new plug-and-play techniques called position-aware attention scaling and patch-wise augmentation to augment our RVT, which we abbreviate as RVT*. The experimental results of RVT on ImageNet and six robustness benchmarks demonstrate its advanced robustness and generalization ability compared with previous ViTs and state-of-the-art CNNs. Furthermore, RVT-S* achieves Top-1 rank on multiple robustness leaderboards including ImageNet-C, ImageNet-Sketch and ImageNet-R. Xiaofeng Mao, Gege Qi, Yuefeng Chen, Ranjie Duan, Shaokai Ye, Yuan He 0011, Hui Xue 0001 |
CVPR | 8 |
| 2022 | Supervised Prototypical Contrastive Learning for Emotion Recognition in ConversationabstractCapturing emotions within a conversation plays an essential role in modern dialogue systems.However, the weak correlation between emotions and semantics brings many challenges to emotion recognition in conversation (ERC).Even semantically similar utterances, the emotion may vary drastically depending on contexts or speakers.In this paper, we propose a Supervised Prototypical Contrastive Learning (SPCL) loss for the ERC task.Leveraging the Prototypical Network, the SPCL targets at solving the imbalanced classification problem through contrastive learning and does not require a large batch size.Meanwhile, we design a difficulty measure function based on the distance between classes and introduce curriculum learning to alleviate the impact of extreme samples.We achieve state-of-the-art results on three widely used benchmarks.Further, we conduct analytical experiments to demonstrate the effectiveness of our proposed SPCL and curriculum learning strategy.We release the code at https://github.com/caskcsg/SPCL.⋆ Longtao Huang, Hui Xue 0001, Songlin Hu 0001 |
EMNLP | 3 |
| 2022 | MAKD: MULTIPLE Auxiliary Knowledge DistillationabstractKnowledge distillation aims to learn a small student model by leveraging knowledge from a larger teacher model. The gap between these heterogeneous models hinder their knowledge transfer and it would be more challenging when the teacher model is from another task. Previous methods view the teacher model as a perfect feature extractor and train the student model to mimic it. While we notice that the teacher model has defect in extracting features of another task samples. To improve knowledge distillation under such situation, we propose Multiple Auxiliary Subspaces (MAS). Most previous methods improve distillation performance by representation alignment, while we resort to the promotion of the teacher model which is more suitable for cross-task distillation. The MAS distills the knowledge in a mutual learning way based on an auxiliary network. Along with the training procedure, the teacher model is improved by the auxiliary network which works as a trainable part of the teacher model and learn the features distribution of target samples from the student model. And this promotion of the teacher model will benefit the student model via the following distillation procedure. We adopt the representation alignment technique, multiple auxiliary networks to further enhance the proposed method. The MAS works well with limited or sufficient labeled target data. If the source data is available, based on it the MAS can construct another auxiliary network to further improve the distillation. Experiments are conducted to validate that the MAS outperforms baseline methods and achieves state-of-the-art results on several standard benchmarks. Zehan Chen, Xuan Jin, Yuan He 0011, Hui Xue 0001 |
ICASSP | 4 |
| 2022 | LightPose: A Lightweight and Efficient Model with Transformer for Human Pose EstimationabstractThe prediction of keypoints by generating high-resolution heatmaps has become a popular solution in human pose estimation. While this kind of method requires up-sampling or deconvolution operations, which would bring a great challenge to the acceleration of model inference. If performing keypoint prediction on low-resolution heatmaps, the performance is unsatisfied due to serious quantization errors. To solve this contradiction, we propose to perform joint training of the heatmap and center offset on low-resolution heatmaps to reduce quantization errors, which could achieve the comparable performance to the high-resolution heatmap and reduce the computational complexity. In addition, we utilize transformer to enhance the representation ability of low-resolution features, instead of increasing the network layers or the convolution kernel size. The transformer could bring the significant improvement of the performance with little computation cost. Combining the above two modules, we design a new lightweight pose estimation model, named LightPose. Experimental results have shown that, compared with HRNet, our method could achieve the state-of-the-art performance on COCO and MPII datasets with a massive reduction of the parameters by 86% and GFLOPs by 67%. Ding Ni, Yan Wang 0002, Hui Xue 0001 |
ICASSP | 5 |
| 2022 | Beyond ImageNet Attack: Towards Crafting Adversarial Examples for Black-box Domains
Yuefeng Chen, Jingkuan Song, Lianli Gao, Yuan He 0011, Hui Xue 0001 |
ICLR | 7 |
| 2022 | Diverse Instance Discovery: Vision-Transformer for Instance-Aware Multi-Label Image RecognitionabstractPrevious works on multi-label image recognition (MLIR) usually use CNNs as a starting point for research. In this paper, we take pure Vision Transformer (ViT) as the research base and make full use of the advantages of Transformer with long-range dependency modeling to circumvent the disadvantages of CNNs limited to local receptive field. However, for multi-label images containing multiple objects from different categories, scales, and spatial relations, it is not optimal to use global information alone. Our goal is to leverage ViT's patch tokens and self-attention mechanism to mine rich instances in multi-label images, named diverse instance discovery (DiD). To this end, we propose a semantic category-aware module and a spatial relationship-aware module, respectively, and then combine the two by a re-constraint strategy to obtain instance-aware attention maps. Finally, we propose a weakly supervised object localization-based approach to extract multi-scale local features, to form a multi-view pipeline. Our method requires only weakly supervised information at the label level, no additional knowledge injection or other strongly supervised information is required. Experiments on three benchmark datasets show that our method significantly outperforms previous works and achieves state-of-the-art results under fair experimental comparisons. Yunqing Hu, Xuan Jin, Yin Zhang 0006, Haiwen Hong, Jingfeng Zhang, Feihu Yan, Yuan He 0011, Hui Xue 0001 |
ICME | 8 |
| 2022 | Enhance the Visual Representation via Discrete Adversarial TrainingabstractAdversarial Training (AT), which is commonly accepted as one of the most effective approaches defending against adversarial examples, can largely harm the standard performance, thus has limited usefulness on industrial-scale production and applications. Surprisingly, this phenomenon is totally opposite in Natural Language Processing (NLP) task, where AT can even benefit for generalization. We notice the merit of AT in NLP tasks could derive from the discrete and symbolic input space. For borrowing the advantage from NLP-style AT, we propose Discrete Adversarial Training (DAT). DAT leverages VQGAN to reform the image data to discrete text-like inputs, i.e. visual words. Then it minimizes the maximal risk on such discrete images with symbolic adversarial perturbations. We further give an explanation from the perspective of distribution to demonstrate the effectiveness of DAT. As a plug-and-play technique for enhancing the visual representation, DAT achieves significant improvement on multiple tasks including image classification, object detection and self-supervised learning. Especially, the model pre-trained with Masked Auto-Encoding (MAE) and fine-tuned by our DAT without extra data can get 31.40 mCE on ImageNet-C and 32.77% top-1 accuracy on Stylized-ImageNet, building the new state-of-the-art. The code will be available at https://github.com/alibaba/easyrobust. Xiaofeng Mao, Yuefeng Chen, Ranjie Duan, Yao Zhu 0003, Gege Qi, Shaokai Ye, Rong Zhang 0006, Hui Xue 0001 |
NeurIPS | 9 |
| 2022 | Boosting Out-of-distribution Detection with Typical FeaturesabstractOut-of-distribution (OOD) detection is a critical task for ensuring the reliability and safety of deep neural networks in real-world scenarios. Different from most previous OOD detection methods that focus on designing OOD scores or introducing diverse outlier examples to retrain the model, we delve into the obstacle factors in OOD detection from the perspective of typicality and regard the feature's high-probability region of the deep model as the feature's typical set. We propose to rectify the feature into its typical set and calculate the OOD score with the typical features to achieve reliable uncertainty estimation. The feature rectification can be conducted as a plug-and-play module with various OOD scores. We evaluate the superiority of our method on both the commonly used benchmark (CIFAR) and the more challenging high-resolution benchmark with large label space (ImageNet). Notably, our approach outperforms state-of-the-art methods by up to 5.11% in the average FPR95 on the ImageNet benchmark. Yao Zhu 0003, Yuefeng Chen, Chuanlong Xie, Rong Zhang 0006, Hui Xue 0001, Xiang Tian 0002, Bolun Zheng, Yaowu Chen |
NeurIPS | 6 |
| 2022 | HiSA: Hierarchically Semantic Associating for Video Temporal GroundingabstractVideo Temporal Grounding (VTG) aims to locate the time interval in a video that is semantically relevant to a language query. Existing VTG methods interact the query with entangled video features and treat the instances in a dataset independently. However, intra-video entanglement and inter-video connection are rarely considered in these methods, leading to mismatches between the video and language. To this end, we propose a novel method, dubbed Hierarchically Semantic Associating (HiSA), which aims to precisely align the video with language and obtain discriminative representation for further location regression. Specifically, the action factors and background factors are disentangled from adjacent video segments, enforcing precise multimodal interaction and alleviating the intra-video entanglement. In addition, cross-guided contrast is elaborately framed to capture the inter-video connection, which benefits the multimodal understanding to locate the time interval. Extensive experiments on three benchmark datasets demonstrate that our approach significantly outperforms the state-of-the-art methods. The project page is available at: https://github.com/zhexu1997/HiSA. Zhe Xu 0009, Da Chen 0003, Cheng Deng 0002, Hui Xue 0001 |
IEEE Trans. Image Process. | 5 |
| 2021 | Composite Adversarial AttacksabstractAdversarial attack is a technique for deceiving Machine Learning (ML) models, which provides a way to evaluate the adversarial robustness. In practice, attack algorithms are artificially selected and tuned by human experts to break a ML system. However, manual selection of attackers tends to be sub-optimal, leading to a mistakenly assessment of model security. In this paper, a new procedure called Composite Adversarial Attack (CAA) is proposed for automatically searching the best combination of attack algorithms and their hyper-parameters from a candidate pool of 32 base attackers. We design a search space where attack policy is represented as an attacking sequence, i.e., the output of the previous attacker is used as the initialization input for successors. Multi-objective NSGA-II genetic algorithm is adopted for finding the strongest attack policy with minimum complexity. The experimental result shows CAA beats 10 top attackers on 11 diverse defenses with less elapsed time (6 × faster than AutoAttack), and achieves the new state-of-the-art on linf, l2 and unrestricted adversarial attacks. Xiaofeng Mao, Yuefeng Chen, Shuhui Wang, Hang Su 0006, Yuan He 0011, Hui Xue 0001 |
AAAI | 6 |
| 2021 | QAIR: Practical Query-Efficient Black-Box Attacks for Image RetrievalabstractWe study the query-based attack against image retrieval to evaluate its robustness against adversarial examples under the black-box setting, where the adversary only has query access to the top-k ranked unlabeled images from the database. Compared with query attacks in image classification, which produce adversaries according to the returned labels or confidence score, the challenge becomes even more prominent due to the difficulty in quantifying the attack effectiveness on the partial retrieved list. In this paper, we make the first attempt in Query-based Attack against Image Retrieval (QAIR), to completely subvert the top-k retrieval results. Specifically, a new relevance-based loss is designed to quantify the attack effects by measuring the set similarity on the top-k retrieval results before and after attacks and guide the gradient optimization. To further boost the attack efficiency, a recursive model stealing method is proposed to acquire transferable priors on the target model and generate the prior-guided gradients. Comprehensive experiments show that the proposed attack achieves a high attack success rate with few queries against the image retrieval systems under the black-box setting. The attack evaluations on the real-world visual search engine show that it successfully deceives a commercial system such as Bing Visual Search with 98% attack success rate by only 33 queries on average. Yuefeng Chen, Shaokai Ye, Yuan He 0011, Shuhui Wang, Hang Su 0006, Hui Xue 0001 |
CVPR | 8 |
| 2021 | Spatial-Phase Shallow Learning: Rethinking Face Forgery Detection in Frequency DomainabstractThe remarkable success in face forgery techniques has received considerable attention in computer vision due to security concerns. We observe that up-sampling is a necessary step of most face forgery techniques, and cumulative up-sampling will result in obvious changes in the frequency domain, especially in the phase spectrum. According to the property of natural images, the phase spectrum preserves abundant frequency components that provide extra information and complement the loss of the amplitude spectrum. To this end, we present a novel Spatial-Phase Shallow Learning (SPSL) method, which combines spatial image and phase spectrum to capture the up-sampling artifacts of face forgery to improve the transferability, for face forgery detection. And we also theoretically analyze the validity of utilizing the phase spectrum. Moreover, we notice that local texture information is more crucial than high-level semantic information for the face forgery detection task. So we reduce the receptive fields by shallowing the network to suppress high-level features and focus on the local region. Extensive experiments show that SPSL can achieve the state-of-the-art performance on cross-datasets evaluation as well as multi-class classification and obtain comparable results on single dataset evaluation. Honggu Liu, Wenbo Zhou 0004, Yuefeng Chen, Yuan He 0011, Hui Xue 0001, Weiming Zhang 0001, Nenghai Yu |
CVPR | 6 |
| 2021 | Self-Supervised Learning for Few-Shot Image ClassificationabstractFew-shot image classification aims to classify unseen classes with limited labelled samples. Recent works benefit from the meta-learning process with episodic tasks and can fast adapt to class from training to testing. Due to the limited number of samples for each task, the initial embedding network for meta-learning becomes an essential component and can largely affect the performance in practice. To this end, most of the existing methods highly rely on the efficient embedding network. Due to the limited labelled data, the scale of embedding network is constrained under a supervised learning(SL) manner which becomes a bottleneck of the few-shot learning methods. In this paper, we proposed to train a more generalized embedding network with self-supervised learning (SSL) which can provide robust representation for downstream tasks by learning from the data itself. We evaluate our work by extensive comparisons with previous baseline methods on two few-shot classification datasets (i.e., MiniImageNet and CUB) and achieve better performance over baselines. Tests on four datasets in cross-domain few-shot learning classification show that the proposed method achieves state-of-the-art results and further prove the robustness of the proposed model. Our code is available at https://github.com/phecy/SSL-FEW-SHOT. Da Chen 0003, Yuefeng Chen, Feng Mao, Yuan He 0011, Hui Xue 0001 |
ICASSP | 6 |
| 2021 | Enhancing Model Robustness by Incorporating Adversarial Knowledge into Semantic RepresentationabstractDespite that deep neural networks (DNNs) have achieved enormous success in many domains like natural language processing (NLP), they have also been proven to be vulnerable to maliciously generated adversarial examples. Such inherent vulnerability has threatened various real-world deployed DNNs-based applications. To strength the model robustness, several countermeasures have been proposed in the English NLP domain and obtained satisfactory performance. However, due to the unique language properties of Chinese, it is not trivial to extend existing defenses to the Chinese domain. Therefore, we propose AdvGraph, a novel defense which enhances the robustness of Chinese-based NLP models by incorporating adversarial knowledge into the semantic representation of the input. Extensive experiments on two real-world tasks show that AdvGraph exhibits better performance compared with previous work: (i) effective – it significantly strengthens the model robustness even under the adaptive attacks setting without negative impact on model performance over legitimate input; (ii) generic – its key component, i.e., the representation of connotative adversarial knowledge is task-agnostic, which can be reused in any Chinese-based NLP models without retraining; and (iii) efficient – it is a light-weight defense with sub-linear computational complexity, which can guarantee the efficiency required in practical scenarios. Tianyu Du, Rong Zhang 0006, Hui Xue 0001, Shouling Ji |
ICASSP | 5 |
| 2021 | Towards Face Encryption by Generating Adversarial Identity MasksabstractAs billions of personal data being shared through social media and network, the data privacy and security have drawn an increasing attention. Several attempts have been made to alleviate the leakage of identity information from face photos, with the aid of, e.g., image obfuscation techniques. However, most of the present results are either perceptually unsatisfactory or ineffective against face recognition systems. Our goal in this paper is to develop a technique that can encrypt the personal photos such that they can protect users from unauthorized face recognition systems but remain visually identical to the original version for human beings. To achieve this, we propose a targeted identity-protection iterative method (TIP-IM) to generate adversarial identity masks which can be overlaid on facial images, such that the original identities can be concealed without sacrificing the visual quality. Extensive experiments demonstrate that TIP-IM provides 95%+ protection success rate against various state-of-the-art face recognition models under practical test scenarios. Besides, we also show the practical and effective applicability of our method on a commercial API service. Xiao Yang 0028, Yinpeng Dong, Tianyu Pang, Hang Su 0006, Jun Zhu 0001, Yuefeng Chen, Hui Xue 0001 |
ICCV | 7 |
| 2021 | DRDF: Determining the Importance of Different Multimodal Information with Dual-Router Dynamic FrameworkabstractIn multimodal tasks, the importance of text and image modal information often varies for different input cases. To model the difference of importance of different modal information, we propose a high-performance and highly general Dual-Router Dynamic Framework (DRDF), consisting of Dual-Router, MWF-Layer, experts and expert fusion unit. The text router and image router in Dual-Router take text modal information and image modal information respectively, and MWF-Layer is responsible to determine the importance of modal information. Based on the result of the determination, MWF-Layer generates fused weights for the subsequent experts fusion. Experts can adopt a variety of backbones that match the current multimodal or unimodal task. DRDF features high generality and modularity, and we test 12 backbones such as Visual BERT and their corresponding DRDF instances on the multimodal dataset Hateful memes, and unimodal datasets CIFAR10, CIFAR100, and TinyImagenet. Our DRDF instance outperforms those backbones. We also validate the effectiveness of components of DRDF by ablation studies, and discuss the reasons and ideas of DRDF design. Haiwen Hong, Xuan Jin, Yin Zhang 0006, Yunqing Hu, Jingfeng Zhang, Yuan He 0011, Hui Xue 0001 |
ACM Multimedia | 7 |
| 2021 | RAMS-Trans: Recurrent Attention Multi-scale Transformer for Fine-grained Image RecognitionabstractIn fine-grained image recognition (FGIR), the localization and amplification of region attention is an important factor, which has been explored extensively convolutional neural networks (CNNs) based approaches. The recently developed vision transformer (ViT) has achieved promising results in computer vision tasks. Compared with CNNs, Image sequentialization is a brand new manner. However, ViT is limited in its receptive field size and thus lacks local attention like CNNs due to the fixed size of its patches, and is unable to generate multi-scale features to learn discriminative region attention. To facilitate the learning of discriminative region attention without box/part annotations, we use the strength of the attention weights to measure the importance of the patch tokens corresponding to the raw images. We propose the recurrent attention multi-scale transformer (RAMS-Trans), which uses the transformer's self-attention to recursively learn discriminative region attention in a multi-scale manner. Specifically, at the core of our approach lies the dynamic patch proposal module (DPPM) responsible for guiding region amplification to complete the integration of multi-scale image patches. The DPPM starts with the full-size image patches and iteratively scales up the region attention to generate new patches from global to local by the intensity of the attention weights generated at each scale as an indicator. Our approach requires only the attention weights that come with ViT itself and can be easily trained end-to-end. Extensive experiments demonstrate that RAMS-Trans performs better than exising works, in addition to efficient CNN models, achieving state-of-the-art results on three benchmark datasets. Yunqing Hu, Xuan Jin, Yin Zhang 0006, Haiwen Hong, Jingfeng Zhang, Yuan He 0011, Hui Xue 0001 |
ACM Multimedia | 7 |
| 2021 | Pano-SfMLearner: Self-Supervised Multi-Task Learning of Depth and Semantics in Panoramic VideosabstractWith the advent of virtual reality and augment reality applications, omnidirectional imaging and$360^{\circ }$cameras become increasingly popular in many scenarios such as entertainment and autonomous systems. In this paper, we propose a self-supervised framework for multi-task learning on depth, camera motion and semantics from panoramic videos. Specifically, our method is based on differentiable warping of adjacent views to the target. Two improvements are provided. First, we introduce a view synthesis module based on equirectangular projection to enable direct optimization on panoramic images. Second, we introduce a self-supervised segmentation branch to involve the constraint of semantic consistency for further improvement. Extensive experiments on two$360^{\circ }$video and two$360^{\circ }$image datasets demonstrate that our method outperforms the state-of-the-art and achieves favorable cross-modality performance. Shuhui Wang, Yulan Guo, Yuan He 0011, Hui Xue 0001 |
IEEE Signal Process. Lett. | 5 |
| 2020 | Fine-Grained Fashion Similarity Learning by Attribute-Specific Embedding NetworkabstractThis paper strives to learn fine-grained fashion similarity. In this similarity paradigm, one should pay more attention to the similarity in terms of a specific design/attribute among fashion items, which has potential values in many fashion related applications such as fashion copyright protection. To this end, we propose an Attribute-Specific Embedding Network (ASEN) to jointly learn multiple attribute-specific embeddings in an end-to-end manner, thus measure the fine-grained similarity in the corresponding space. With two attention modules, i.e., Attribute-aware Spatial Attention and Attribute-aware Channel Attention, ASEN is able to locate the related regions and capture the essential patterns under the guidance of the specified attribute, thus make the learned attribute-specific embeddings better reflect the fine-grained similarity. Extensive experiments on four fashion-related datasets show the effectiveness of ASEN for fine-grained fashion similarity learning and its potential for fashion reranking. Code and data are available at https://github.com/Maryeon/asen. Zhe Ma 0002, Jianfeng Dong, Zhongzi Long, Yao Zhang 0019, Yuan He 0011, Hui Xue 0001, Shouling Ji |
AAAI | 6 |
| 2020 | SpanMlt: A Span-based Multi-Task Learning Framework for Pair-wise Aspect and Opinion Terms ExtractionabstractAspect terms extraction and opinion terms extraction are two key problems of fine-grained Aspect Based Sentiment Analysis (ABSA).The aspect-opinion pairs can provide a global profile about a product or service for consumers and opinion mining systems.However, traditional methods can not directly output aspect-opinion pairs without given aspect terms or opinion terms.Although some recent co-extraction methods have been proposed to extract both terms jointly, they fail to extract them as pairs.To this end, this paper proposes an end-to-end method to solve the task of Pair-wise Aspect and Opinion Terms Extraction (PAOTE).Furthermore, this paper treats the problem from a perspective of joint term and relation extraction rather than under the sequence tagging formulation performed in most prior works.We propose a multi-task learning framework based on shared spans, where the terms are extracted under the supervision of span boundaries.Meanwhile, the pair-wise relations are jointly identified using the span representations.Extensive experiments show that our model consistently outperforms stateof-the-art methods. Longtao Huang, Rong Zhang 0006, Hui Xue 0001 |
ACL | 5 |
| 2020 | Which Is Plagiarism: Fashion Image Retrieval Based on Regional Representation for Design ProtectionabstractWith the rapid growth of e-commerce and the popularity of online shopping, fashion retrieval has received considerable attention in the computer vision community. Different from the existing works that mainly focus on identical or similar fashion item retrieval, in this paper, we aim to study the plagiarized clothes retrieval which is somewhat ignored in the academic community while itself has great application value. One of the key challenges is that plagiarized clothes are usually modified in a certain region on the original design to escape the supervision by traditional retrieval methods. To relieve it, we propose a novel network named Plagiarized-Search-Net (PS-Net) based on regional representation, where we utilize the landmarks to guide the learning of regional representations and compare fashion items region by region. Besides, we propose a new dataset named Plagiarized Fashion for plagiarized clothes retrieval, which provides a meaningful complement to the existing fashion retrieval field. Experiments on Plagiarized Fashion dataset verify that our approach is superior to other instance-level counterparts for plagiarized clothes retrieval, showing a promising result for original design protection. Moreover, our PS-Net can also be adapted to traditional fashion retrieval and landmark estimation tasks and achieves the state-of-the-art performance on the DeepFashion and DeepFashion2 datasets. Yining Lang, Yuan He 0011, Jianfeng Dong, Hui Xue 0001 |
CVPR | 5 |
| 2020 | Self-Supervised Adversarial TrainingabstractRecent work has demonstrated that neural networks are vulnerable to adversarial examples. To escape from the predicament, many works try to harden the model in various ways, in which adversarial training is an effective way which learns robust feature representation so as to resist adversarial attacks. Meanwhile, the self-supervised learning aims to learn robust and semantic embedding from data itself. With these views, we introduce self-supervised learning to against adversarial examples in this paper. Specifically, the self-supervised representation coupled with k-Nearest Neighbour is proposed for classification. To further strengthen the defense ability, self-supervised adversarial training is proposed, which maximizes the mutual information between the representations of original examples and the corresponding adversarial examples. Experimental results show that the self-supervised representation outperforms its supervised version in respect of robustness and self-supervised adversarial training can further improve the defense ability efficiently. Kejiang Chen, Yuefeng Chen, Hang Zhou 0007, Xiaofeng Mao, Yuan He 0011, Hui Xue 0001, Weiming Zhang 0001, Nenghai Yu |
ICASSP | 7 |
| 2020 | Hierarchical Sequence Representation with Graph NetworkabstractVideo classification problem is a challenging task in computer vision. The performance of this task is highly relied on the scale of training data and the effectiveness of video embedding via a robust embedding network. Unsupervised solutions such as feature average pooling technique, as a simple label-independent and parameter-free based method, cannot efficiently represent the video sequences. While supervised methods, such as RNN, can improve the recognition accuracy. The performance of RNN based methods, however, is decreased with the increasing length of the videos and the hierarchical relationships between frames across events in the video. In this paper, we propose a novel video classification method based on a deep convolutional graph neural network (DCGN). The proposed method utilizes the characteristics of the hierarchical structure of the video, and performed multi-level embedding feature extraction on the video frame sequence through the graph network, and obtained a video representation which reflects the event semantics hierarchically. Experiments on YouTube-8M Large-Scale Video Understanding dataset show that our proposed model outperforms the commonly used RNN based models, verifying its effectiveness for video classification. Da Chen 0003, Xiang Wu 0004, Jianfeng Dong, Yuan He 0011, Hui Xue 0001, Feng Mao |
ICASSP | 5 |
| 2020 | The Open Brands Dataset: Unified Brand Detection and Recognition at ScaleabstractIntellectual property protection(IPP) have received more and more attention recently due to the development of the global e-commerce platforms. brand recognition plays a significant role in IPP. Recent studies for brand recognition and detection are based on small-scale datasets that are not comprehensive enough when exploring emerging deep learning techniques. Moreover, it is challenging to evaluate the true performance of brand detection methods in realistic and open scenes. In order to tackle these problems, we first define the special issues of brand detection and recognition compared with generic object detection. Second, a novel brands benchmark called "Open Brands" is established. The dataset contains 1,437,812 images which have brands and 50,000 images without any brand. The part with brands in Open Brands contains 3,113,828 instances annotated in 3 dimensions: 4 types, 559 brands and 1216 logos. To the best of our knowledge, it is the largest dataset for brand detection and recognition with rich annotations. We provide in-depth comprehensive statistics about the dataset, validate the quality of the annotations and study how the performance of many modern models evolves with an increasing amount of training data. Third, we design a network called "Brand Net" to handle brand recognition. Brand Net gets state-of-art mAP on Open Brand compared with existing detection methods. Xuan Jin, Rong Zhang 0006, Yuan He 0011, Hui Xue 0001 |
ICASSP | 5 |
| 2020 | Design-Gan: Cross-Category Fashion Translation Driven By Landmark AttentionabstractThe rise of generative adversarial networks has boosted a vast interest in the field of fashion image-to-image translation. However, previous methods do not perform well in cross-category translation tasks, e.g., translating jeans to skirts in fashion images. The translated skirts are easier to lose the detail texture of the jeans, and the generated legs or arms often look unnatural. In this paper, we propose a novel approach, called DesignGAN, that utilizes the landmark guided attention and a similarity constraint mechanism to achieve fashion cross-category translation. Moreover, we can achieve texture editing on any customized input, which can even be used as an effective way to empower fashion designers. Experiments on fashion datasets verify that DesignGAN is superior to other image-to-image translation methods. Yining Lang, Yuan He 0011, Jianfeng Dong, Hui Xue 0001 |
ICASSP | 5 |
| 2020 | Learning to Characterize Adversarial SubspacesabstractDeep Neural Networks (DNNs) are known to be vulnerable to the maliciously generated adversarial examples. To detect these adversarial examples, previous methods use artificially designed metrics to characterize the properties of adversarial subspaces where adversarial examples lie. However, we find these methods are not working in practical attack detection scenarios. Because the artificially defined features are lack of robustness and show limitation in discriminative power to detect strong attacks. To solve this problem, we propose a novel adversarial detection method which identifies adversaries by adaptively learning reasonable metrics to characterize adversarial subspaces. As auxiliary context information, k nearest neighbors are used to represent the surrounded subspace of the detected sample. We propose an innovative model called Neighbor Context Encoder (NCE) to learn from k neighbors context and infer if the detected sample is normal or adversarial. We conduct thorough experiment on CIFAR-10, CIFAR-100 and ImageNet dataset. The results demonstrate that our approach surpasses all existing methods under three settings: attack-aware black-box detection, attack-unaware black-box detection and white-box detection. Xiaofeng Mao, Yuefeng Chen, Yuan He 0011, Hui Xue 0001 |
ICASSP | 5 |
| 2020 | Sharp Multiple Instance Learning for DeepFake Video DetectionabstractWith the rapid development of facial manipulation techniques, face forgery has received considerable attention in multimedia and computer vision community due to security concerns. Existing methods are mostly designed for single-frame detection trained with precise image-level labels or for video-level prediction by only modeling the inter-frame inconsistency, leaving potential high risks for DeepFake attackers. In this paper, we introduce a new problem of partial face attack in DeepFake video, where only video-level labels are provided but not all the faces in the fake videos are manipulated. We address this problem by multiple instance learning framework, treating faces and input video as instances and bag respectively. A sharp MIL (S-MIL) is proposed which builds direct mapping from instance embeddings to bag prediction, rather than from instance embeddings to instance prediction and then to bag prediction in traditional MIL. Theoretical analysis proves that the gradient vanishing in traditional MIL is relieved in S-MIL. To generate instances that can accurately incorporate the partially manipulated faces, spatial-temporal encoded instance is designed to fully model the intra-frame and inter-frame inconsistency, which further helps to promote the detection performance. We also construct a new dataset FFPMS for partially attacked DeepFake video detection, which can benefit the evaluation of different methods at both frame and video levels. Experiments on FFPMS and the widely used DFDC dataset verify that S-MIL is superior to other counterparts for partially attacked DeepFake video detection. In addition, S-MIL can also be adapted to traditional DeepFake image detection tasks and achieve state-of-the-art performance on single-frame datasets. Yining Lang, Yuefeng Chen, Xiaofeng Mao, Yuan He 0011, Shuhui Wang, Hui Xue 0001 |
ACM Multimedia | 7 |
| 2019 | Bilinear Representation for Language-based Image Editing Using Conditional Generative Adversarial NetworksabstractThe task of Language-Based Image Editing (LBIE) aims at generating a target image by editing the source image based on the given language description. The main challenge of LBIE is to disentangle the semantics in image and text and then combine them to generate realistic images. Therefore, the editing performance is heavily dependent on the learned representation. In this work, conditional generative adversarial network (cGAN) is utilized for LBIE. We find that existing conditioning methods in cGAN lack of representation power as they cannot learn the second-order correlation between two conditioning vectors. To solve this problem, we propose an improved conditional layer named Bilinear Residual Layer (BRL) to learning more powerful representations for LBIE task. Qualitative and quantitative comparisons demonstrate that our method can generate images with higher quality when compared to previous LBIE techniques. Xiaofeng Mao, Yuefeng Chen, Yuan He 0011, Hui Xue 0001 |
ICASSP | 6 |