EDBT 2026 Demo / reviewers in the wild / expert
Xingxing Zhang 0001
dblp:59/9985-1
· DBLP profile ↗
43ranked-venue papers
8as first author
32since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 3 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 6 first-author · 18 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BiKT: Unleashing the Potential of GNNs via Bi-Directional Knowledge TransferabstractBased on the message-passing paradigm, there has been an amount of research proposing diverse and impressive feature propagation mechanisms to improve the performance of GNNs. However, less focus has been put on feature transformation, another major operation of the message-passing framework. In this paper, we first empirically investigate the performance of the feature transformation operation in several typical GNNs. Unexpectedly, we notice that GNNs do not completely free up the power of the inherent feature transformation operation. By this observation, we propose the Bi-directional Knowledge Transfer (BiKT), a plug-and-play approach to unleash the potential of the feature transformation operations without modifying the original architecture. Taking the feature transformation operation as a derived representation learning model that shares parameters with the original GNN, the direct prediction by this model provides a topological-agnostic knowledge feedback that can further instruct the learning of GNN and the feature transformations therein. On this basis, BiKT not only allows us to acquire knowledge from both the GNN and its derived model but also promotes each other by injecting the knowledge into the other. In addition, a theoretical analysis is further provided to demonstrate that BiKT improves the generalization bound of the GNNs from the perspective of domain adaptation. An extensive group of experiments on up to 7 datasets with 5 typical GNNs demonstrates that BiKT brings up to 0.5% - 4% performance gain over the original GNN, which means a boosted GNN is obtained. Meanwhile, the derived model also shows a powerful performance to compete with or even surpass the original GNN, enabling us to flexibly apply it independently to some other specific downstream tasks. Shuai Zheng 0005, Zhizhe Liu, Zhenfeng Zhu, Xingxing Zhang 0001, Jianxin Li 0002, Yao Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Rethinking Obscured Sub-Optimality in Analytic Learning for Exemplar-Free Class-Incremental LearningabstractExemplar-free Class-Incremental Learning (EFCIL) poses a significant challenge in mitigating catastrophic forgetting, due to the absence of exemplars. Recently, analytic learning-based methods propose a recursive alignment procedure to execute EFCIL in a phase-invariant manner and show state-of-the-art performance. However, they heavily rely on a frozen feature extractor trained with the initial dataset to avoid the misalignment between feature and label spaces, ignoring the importance of acquiring generalizable features across incremental tasks for performance improvement. To tackle this, we rethink the obscured sub-optimality of analytic learning-based methods, particularly through empirical reevaluation, and then introduce the Multi-head analytic learning (Muheal) approach. Muheal forms the multi-head model with a delicate feature extractor, thereby introducing a feature optimization procedure and a forgetting compensation module to balance the learning and forgetting. Specifically, within the feature optimization procedure, the feature extractor seeks to learn more generalizable features in a self-supervised manner using the fully-connected classification head. An analytic learning-based classification head follows to align the feature-label space. Additionally, we employ the compensation module to generate and align pseudo-features with a replicated analytic head, thus preventing overfitting and testing. Comprehensive experiments on several benchmark datasets have demonstrated that Muheal significantly outperforms existing state-of-the-art EFCIL methods and is comparable, if not superior, to methods that use replay techniques. Zijian Gao, Kele Xu, Xingxing Zhang 0001, Huiping Zhuang, Tianjiao Wan, Bo Ding 0001, Xinjun Mao, Huaimin Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Maintaining Fairness in Logit-based Knowledge Distillation for Class-Incremental LearningabstractLogit-based knowledge distillation (KD) is commonly used to mitigate catastrophic forgetting in class-incremental learning (CIL) caused by data distribution shifts. However, the strict match of logit values between student and teacher models conflicts with the cross-entropy (CE) loss objective of learning new classes, leading to significant recency bias (i.e. unfairness). To address this issue, we rethink the overlooked limitations of KD-based methods through empirical analysis. Inspired by our findings, we introduce a plug-and-play pre-process method that normalizes the logits of both the student and teacher across all classes, rather than just the old classes, before distillation. This approach allows the student to focus on both old and new classes, capturing intrinsic inter-class relations from the teacher. By doing so, our method avoids the inherent conflict between KD and CE, maintaining fairness between old and new classes. Additionally, recognizing that overconfident teacher predictions can hinder the transfer of inter-class relations (i.e., dark knowledge), we extend our method to capture intra-class relations among different instances, ensuring fairness within old classes. Our method integrates seamlessly with existing logit-based KD approaches, consistently enhancing their performance across multiple CIL benchmarks without incurring additional training costs. Zijian Gao, Shanhao Han, Xingxing Zhang 0001, Kele Xu, Dulan Zhou, Xinjun Mao, Yong Dou, Huaimin Wang 0001 |
AAAI | 3 |
| 2025 | Knowledge Memorization and Rumination for Pre-trained Model-based Class-Incremental LearningabstractClass-Incremental Learning (CIL) enables models to continuously learn new classes while mitigating catastrophic forgetting. Recently, Pre-Trained Models (PTMs) have greatly enhanced CIL performance, even when fine-tuning is limited to the first task. This advantage is particularly beneficial for CIL methods that freeze the feature extractor after first-task fine-tuning, such as analytic learning-based approaches using a least squares solution-based classification head to acquire knowledge recursively. In this work, we revisit the analytical learning approach combined with PTMs and identify its limitations in adapting to new classes, leading to sub-optimal performance. To address this, we propose the Momentum-based Analytical Learning (MoAL) approach. MoAL achieves robust knowledge memorization via an analytical classification head and improves adaptivity to new classes through momentum-based adapter weight interpolation, leading to forgetting outdated knowledge. Importantly, we introduce a knowledge rumination mechanism that leverages refined adaptivity, allowing the model to revisit and reinforce old knowledge, thereby improving performance on old classes. MoAL facilitates the acquisition of new knowledge and consolidates old knowledge, achieving a win-win outcome between plasticity and stability. Extensive experiments on various incremental settings show MoAL’s state-of-the-art performance1. Zijian Gao, Wangwang Jia, Xingxing Zhang 0001, Dulan Zhou, Kele Xu, Yong Dou, Xinjun Mao, Huaimin Wang 0001 |
CVPR | 3 |
| 2025 | HiDe-PET: Continual Learning via Hierarchical Decomposition of Parameter-Efficient TuningabstractThe deployment of pre-trained models (PTMs) has greatly advanced the field of continual learning (CL), enabling positive knowledge transfer and resilience to catastrophic forgetting. To sustain these advantages for sequentially arriving tasks, a promising direction involves keeping the pre-trained backbone frozen while employing parameter-efficient tuning (PET) techniques to instruct representation learning. Despite the popularity of Prompt-based PET for CL, its empirical design often leads to sub-optimal performance in our evaluation of different PTMs and target tasks. To this end, we propose a unified framework for CL with PTMs and PET that provides both theoretical and empirical advancements. We first perform an in-depth theoretical analysis of the CL objective in a pre-training context, decomposing it into hierarchical components namely within-task prediction, task-identity inference and task-adaptive prediction. We then present Hierarchical Decomposition PET (HiDe-PET), an innovative approach that explicitly optimizes the decomposed objective through incorporating task-specific and task-shared knowledge via mainstream PET techniques along with efficient recovery of pre-trained representations. Leveraging this framework, we delve into the distinct impacts of implementation strategy, PET technique and PET architecture, as well as adaptive knowledge accumulation amidst pronounced distribution changes. Finally, across various CL scenarios, our approach demonstrates remarkably superior performance over a broad spectrum of recent strong baselines. Xingxing Zhang 0001, Hang Su 0006, Jun Zhu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | RobustPrompt: Learning to defend against adversarial attacks with adaptive visual prompts
Chang Liu 0077, Wenzhao Xiang 0001, Yinpeng Dong, Xingxing Zhang 0001, Ranjie Duan, Shibao Zheng, Hang Su 0006 |
Pattern Recognit. Lett. | 4 |
| 2025 | PHI: Bridging Domain Shift in Long-Term Action Quality Assessment via Progressive Hierarchical InstructionabstractLong-term Action Quality Assessment (AQA) aims to evaluate the quantitative performance of actions in long videos. However, existing methods face challenges due to domain shifts between the pre-trained large-scale action recognition backbones and the specific AQA task, thereby hindering their performance. This arises since fine-tuning resource-intensive backbones on small AQA datasets is impractical. We address this by identifying two levels of domain shift: task-level, regarding differences in task objectives, and feature-level, regarding differences in important features. For feature-level shifts, which are more detrimental, we propose Progressive Hierarchical Instruction (PHI) with two strategies. First, Gap Minimization Flow (GMF) leverages flow matching to progressively learn a fast flow path that reduces the domain gap between initial and desired features across shallow to deep layers. Additionally, a temporally-enhanced attention module captures long-range dependencies essential for AQA. Second, List-wise Contrastive Regularization (LCR) facilitates coarse-to-fine alignment by comprehensively comparing batch pairs to learn fine-grained cues while mitigating domain shift. Integrating these modules, PHI offers an effective solution. Experiments demonstrate that PHI achieves state-of-the-art performance on three representative long-term AQA datasets, proving its superiority in addressing the domain shift for long-term AQA. Kanglei Zhou, Hubert P. H. Shum, Frederick W. B. Li, Xingxing Zhang 0001, Xiaohui Liang 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Determinantal Point Processes Guided Crowd-wise Mixture-of-Experts for Recommendation in AlipayabstractFacing the challenges of sparsity and long tail in thousands of Mini-apps recommendation scenarios deployed on Alipay platform, there is a great need for a simple, effective, and easy-to-deploy industrial solution. To address this issue, we follow the strategy of “divide and conquer” and propose a crowd-based recommendation model by using D eterminantal P oint P rocesse s on C rowd-wise M ixture- o f- E xperts (DPPs-CMoE). Specifically, under the guidance of DPPs-based prototypical tags, the user profiling space is sequentially divided into multiple crowds, with each of them taking on a unique latent specificity; Meanwhile, by treating the modeling of crowd specificity as one of multiple tasks, a crowd-wise architecture is adopted to seamlessly unify the multiple expert networks from the overall user space and the gating network from each of independent crowd spaces. The effectiveness of the proposed method has been illustrated in the experimental results on a mini-apps recommendation scenario deployed in Alipay APPs. Youru Li, Zhenfeng Zhu, Shaohu Chen, Kaiming Shen, Xingxing Zhang 0001, Leon Wenliang Zhong, Yao Zhao 0001 |
Trans. Recomm. Syst. | 5 |
| 2024 | MAGR: Manifold-Aligned Graph Regularization for Continual Action Quality Assessment
Kanglei Zhou, Xingxing Zhang 0001, Hubert P. H. Shum, Frederick W. B. Li, Xiaohui Liang 0001 |
ECCV (11) | 3 |
| 2024 | Fourier Controller Networks for Real-Time Decision-Making in Embodied LearningabstractTransformer has shown promise in reinforcement learning to model time-varying features for obtaining generalized low-level robot policies on diverse robotics datasets in embodied learning. However, it still suffers from the issues of low data efficiency and high inference latency. In this paper, we propose to investigate the task from a new perspective of the frequency domain. We first observe that the energy density in the frequency domain of a robot's trajectory is mainly concentrated in the low-frequency part. Then, we present the Fourier Controller Network (FCNet), a new network that uses Short-Time Fourier Transform (STFT) to extract and encode time-varying features through frequency domain interpolation. In order to do real-time decision-making, we further adopt FFT and Sliding DFT methods in the model architecture to achieve parallel training and efficient recurrent inference. Extensive results in both simulated (e.g., D4RL) and real-world environments (e.g., robot locomotion) demonstrate FCNet's substantial efficiency and effectiveness over existing methods such as Transformer, e.g., FCNet outperforms Transformer on multi-environmental robotics datasets of all types of sizes (from 1.9M to 120M). The project page and code can be found https://thkkk.github.io/fcnet. Hengkai Tan, Songming Liu, Chengyang Ying, Xingxing Zhang 0001, Hang Su 0006, Jun Zhu 0001 |
ICML | 5 |
| 2024 | CoFInAl: Enhancing Action Quality Assessment with Coarse-to-Fine Instruction Alignment
Kanglei Zhou, Ruizhi Cai, Xingxing Zhang 0001, Xiaohui Liang 0001 |
IJCAI | 5 |
| 2024 | One-shot In-context Part SegmentationabstractIn this paper, we present the One-shot In-context Part Segmentation (OIParts) framework, designed to tackle the challenges of part segmentation by leveraging visual foundation models (VFMs). Existing training-based one-shot part segmentation methods that utilize VFMs encounter difficulties when faced with scenarios where the one-shot image and test image exhibit significant variance in appearance and perspective, or when the object in the test image is partially visible. We argue that training on the one-shot example often leads to overfitting, thereby compromising the model's generalization capability. Our framework offers a novel approach to part segmentation that is training-free, flexible, and data-efficient, requiring only a single in-context example for precise segmentation with superior generalization ability. By thoroughly exploring the complementary strengths of VFMs, specifically DINOv2 and Stable Diffusion, we introduce an adaptive channel selection approach by minimizing the intra-class distance for better exploiting these two features, thereby enhancing the discriminatory power of the extracted features for the fine-grained parts. We have achieved remarkable segmentation performance across diverse object categories. The OIParts framework not only eliminates the need for extensive labeled data but also demonstrates superior generalization ability. Through comprehensive experimentation on three benchmark datasets, we have demonstrated the superiority of our proposed method over existing part segmentation approaches in one-shot settings. Zhenqi Dai, Ting Liu 0012, Xingxing Zhang 0001, Yunchao Wei, Yanning Zhang 0001 |
ACM Multimedia | 3 |
| 2024 | Stabilizing Zero-Shot Prediction: A Novel Antidote to Forgetting in Continual Vision-Language TasksabstractContinual learning (CL) empowers pre-trained vision-language (VL) models to efficiently adapt to a sequence of downstream tasks. However, these models often encounter challenges in retaining previously acquired skills due to parameter shifts and limited access to historical data. In response, recent efforts focus on devising specific frameworks and various replay strategies, striving for a typical learning-forgetting trade-off. Surprisingly, both our empirical research and theoretical analysis demonstrate that the stability of the model in consecutive zero-shot predictions serves as a reliable indicator of its anti-forgetting capabilities for previously learned tasks.
Motivated by these insights, we develop a novel replay-free CL method named ZAF (Zero-shot Antidote to Forgetting), which preserves acquired knowledge through a zero-shot stability regularization applied to wild data in a plug-and-play manner. To enhance efficiency in adapting to new tasks and seamlessly access historical models, we introduce a parameter-efficient EMA-LoRA neural architecture based on the Exponential Moving Average (EMA). ZAF utilizes new data for low-rank adaptation (LoRA), complemented by a zero-shot antidote on wild data, effectively decoupling learning from forgetting. Our extensive experiments demonstrate ZAF's superior performance and robustness in pre-trained models across various continual VL concept learning tasks, achieving leads of up to 3.70\%, 4.82\%, and 4.38\%, along with at least a 10x acceleration in training speed on three benchmarks, respectively. Additionally, our zero-shot antidote significantly reduces forgetting in existing models by at least 6.37\%. Our code is available at https://github.com/Zi-Jian-Gao/Stabilizing-Zero-Shot-Prediction-ZAF. Zijian Gao, Xingxing Zhang 0001, Kele Xu, Xinjun Mao, Huaimin Wang 0001 |
NeurIPS | 2 |
| 2024 | PEAC: Unsupervised Pre-training for Cross-Embodiment Reinforcement LearningabstractDesigning generalizable agents capable of adapting to diverse embodiments has achieved significant attention in Reinforcement Learning (RL), which is critical for deploying RL agents in various real-world applications. Previous Cross-Embodiment RL approaches have focused on transferring knowledge across embodiments within specific tasks. These methods often result in knowledge tightly coupled with those tasks and fail to adequately capture the distinct characteristics of different embodiments. To address this limitation, we introduce the notion of Cross-Embodiment Unsupervised RL (CEURL), which leverages unsupervised learning to enable agents to acquire embodiment-aware and task-agnostic knowledge through online interactions within reward-free environments. We formulate CEURL as a novel Controlled Embodiment Markov Decision Process (CE-MDP) and systematically analyze CEURL's pre-training objectives under CE-MDP. Based on these analyses, we develop a novel algorithm Pre-trained Embodiment-Aware Control (PEAC) for handling CEURL, incorporating an intrinsic reward function specifically designed for cross-embodiment pre-training. PEAC not only provides an intuitive optimization strategy for cross-embodiment pre-training but also can integrate flexibly with existing unsupervised RL methods, facilitating cross-embodiment exploration and skill discovery. Extensive experiments in both simulated (e.g., DMC and Robosuite) and real-world environments (e.g., legged locomotion) demonstrate that PEAC significantly improves adaptation performance and cross-embodiment generalization, demonstrating its effectiveness in overcoming the unique challenges of CEURL. The project page and code are in https://yingchengyang.github.io/ceurl. Chengyang Ying, Zhongkai Hao, Xinning Zhou, Xuezhou Xu, Hang Su 0006, Xingxing Zhang 0001, Jun Zhu 0001 |
NeurIPS | 6 |
| 2024 | A Comprehensive Survey of Continual Learning: Theory, Method and ApplicationabstractTo cope with real-world dynamics, an intelligent system needs to incrementally acquire, update, accumulate, and exploit knowledge throughout its lifetime. This ability, known as continual learning, provides a foundation for AI systems to develop themselves adaptively. In a general sense, continual learning is explicitly limited by catastrophic forgetting, where learning a new task usually results in a dramatic performance drop of the old tasks. Beyond this, increasingly numerous advances have emerged in recent years that largely extend the understanding and application of continual learning. The growing and widespread interest in this direction demonstrates its realistic significance as well as complexity. In this work, we present a comprehensive survey of continual learning, seeking to bridge the basic settings, theoretical foundations, representative methods, and practical applications. Based on existing theoretical and empirical results, we summarize the general objectives of continual learning as ensuring a proper stability-plasticity trade-off and an adequate intra/inter-task generalizability in the context of resource efficiency. Then we provide a state-of-the-art and elaborated taxonomy, extensively analyzing how representative strategies address continual learning, and how they are adapted to particular challenges in various applications. Through an in-depth discussion of promising directions, we believe that such a holistic perspective can greatly facilitate subsequent exploration in this field and beyond. Xingxing Zhang 0001, Hang Su 0006, Jun Zhu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | To Boost Zero-Shot Generalization for Embodied Reasoning With Vision-Language Pre-TrainingabstractRecently, there exists an increased research interest in embodied artificial intelligence (EAI), which involves an agent learning to perform a specific task when dynamically interacting with the surrounding 3D environment. There into, a new challenge is that many unseen objects may appear due to the increased number of object categories in 3D scenes. It makes developing models with strong zero-shot generalization ability to new objects necessary. Existing work tries to achieve this goal by providing embodied agents with massive high-quality human annotations closely related to the task to be learned, while it is too costly in practice. Inspired by recent advances in pre-trained models in 2D visual tasks, we attempt to boost zero-shot generalization for embodied reasoning with vision-language pre-training that can encode common sense as general prior knowledge. To further improve its performance on a specific task, we rectify the pre-trained representation through masked scene graph modeling (MSGM) in a self-supervised manner, where the task-specific knowledge is learned from iterative message passing. Our method can improve a variety of representative embodied reasoning tasks by a large margin (e.g., over 5.0% w.r.t. answer accuracy on MP3D-EQA dataset that consists of many real-world scenes with a large number of new objects during testing), and achieve the new state-of-the-art performance. Xingxing Zhang 0001, Siyang Zhang, Jun Zhu 0001, Bo Zhang 0010 |
IEEE Trans. Image Process. | 2 |
| 2024 | ATZSL: Defensive Zero-Shot Recognition in the Presence of AdversariesabstractZero-shot learning (ZSL) has received extensive attention recently especially in areas of fine-grained object recognition, retrieval, and image captioning. Due to the complete lack of training samples and high requirement of defense transferability, the ZSL model learned is particularly vulnerable against adversarial attacks. Recent work also showed adversarially robust generalization requires more data. This may significantly affect the robustness of ZSL. However, very few efforts have been devoted towards this direction. In this paper, we take an initial attempt, and propose a generic formulation to provide a systematical solution (namedATZSL) for learning a defensive ZSL model. It is capable of achieving better generalization on various adversarial objects recognition while only losing a negligible performance on clean images for unseen classes, by casting ZSL into a min-max optimization problem. To address it, we design a defensive relation prediction network, which can bridge the seen and unseen class domains via attributes to generalize prediction and defense strategy. Additionally, our framework can be extended to deal with the poisoned scenario of unseen class attributes. An extensive group of experiments are then presented, demonstrating that ATZSL obtains remarkably more favorable trade-off between model transferability and robustness, over currently available alternatives under various settings. Xingxing Zhang 0001, Shupeng Gui, Zhenfeng Zhu, Yao Zhao 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Overcoming Recency Bias of Normalization Statistics in Continual Learning: Balance and AdaptationabstractContinual learning entails learning a sequence of tasks and balancing their knowledge appropriately. With limited access to old training samples, much of the current work in deep neural networks has focused on overcoming catastrophic forgetting of old tasks in gradient-based optimization. However, the normalization layers provide an exception, as they are updated interdependently by the gradient and statistics of currently observed training samples, which require specialized strategies to mitigate recency bias. In this work, we focus on the most popular Batch Normalization (BN) and provide an in-depth theoretical analysis of its sub-optimality in continual learning. Our analysis demonstrates the dilemma between balance and adaptation of BN statistics for incremental tasks, which potentially affects training stability and generalization. Targeting on these particular challenges, we propose Adaptive Balance of BN (AdaB$^2$N), which incorporates appropriately a Bayesian-based strategy to adapt task-wise contributions and a modified momentum to balance BN statistics, corresponding to the training and testing stages. By implementing BN in a continual learning fashion, our approach achieves significant performance gains across a wide range of benchmarks, particularly for the challenging yet realistic online scenarios (e.g., up to 7.68\%, 6.86\% and 4.26\% on Split CIFAR-10, Split CIFAR-100 and Split Mini-ImageNet, respectively). Our code is available at https://github.com/lvyilin/AdaB2N. Yilin Lyu, Xingxing Zhang 0001, Zicheng Sun, Hang Su 0006, Jun Zhu 0001, Liping Jing |
NeurIPS | 3 |
| 2023 | Hierarchical Decomposition of Prompt-Based Continual Learning: Rethinking Obscured Sub-optimalityabstractPrompt-based continual learning is an emerging direction in leveraging pre-trained knowledge for downstream continual learning, and has almost reached the performance pinnacle under supervised pre-training. However, our empirical research reveals that the current strategies fall short of their full potential under the more realistic self-supervised pre-training, which is essential for handling vast quantities of unlabeled data in practice. This is largely due to the difficulty of task-specific knowledge being incorporated into instructed representations via prompt parameters and predicted by uninstructed representations at test time. To overcome the exposed sub-optimality, we conduct a theoretical analysis of the continual learning objective in the context of pre-training, and decompose it into hierarchical components: within-task prediction, task-identity inference, and task-adaptive prediction. Following these empirical and theoretical insights, we propose Hierarchical Decomposition (HiDe-)Prompt, an innovative approach that explicitly optimizes the hierarchical components with an ensemble of task-specific prompts and statistics of both uninstructed and instructed representations, further with the coordination of a contrastive regularization strategy. Our extensive experiments demonstrate the superior performance of HiDe-Prompt and its robustness to pre-training paradigms in continual learning (e.g., up to 15.01% and 9.61% lead on Split CIFAR-100 and Split ImageNet-R, respectively). Xingxing Zhang 0001, Mingyi Huang, Hang Su 0006, Jun Zhu 0001 |
NeurIPS | 3 |
| 2023 | Defensive Few-Shot LearningabstractThis article investigates a new challenging problem called defensive few-shot learning in order to learn a robust few-shot model against adversarial attacks. Simply applying the existing adversarial defense methods to few-shot learning cannot effectively solve this problem. This is because the commonly assumed sample-level distribution consistency between the training and test sets can no longer be met in the few-shot setting. To address this situation, we develop a general defensive few-shot learning (DFSL) framework to answer the following two key questions: (1) how to transfer adversarial defense knowledge from one sample distribution to another? (2) how to narrow the distribution gap between clean and adversarial examples under the few-shot setting? To answer the first question, we propose an episode-based adversarial training mechanism by assuming a task-level distribution consistency to better transfer the adversarial defense knowledge. As for the second question, within each few-shot task, we design two kinds of distribution consistency criteria to narrow the distribution gap between clean and adversarial examples from the feature-wise and prediction-wise perspectives, respectively. Extensive experiments demonstrate that the proposed framework can effectively make the existing few-shot models robust against adversarial attacks. Code is available at https://github.com/WenbinLee/DefensiveFSL.git. Wenbin Li 0006, Lei Wang 0001, Xingxing Zhang 0001, Lei Qi 0001, Jing Huo, Yang Gao 0001, Jiebo Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Auto-Weighted Layer Representation Based View Synthesis Distortion Estimation for 3-D Video CodingabstractRecently, various view synthesis distortion estimation models have been studied to better serve 3-D video coding. However, they can hardly model the relationship quantitatively among different levels of depth changes, texture degeneration, and view synthesis distortion (VSD), which is crucial for rate-distortion optimization and rate allocation. In this paper, an auto-weighted layer representation based view synthesis distortion estimation model is developed. Firstly, sub-VSD (S-VSD) is defined according to the level of depth changes and their associated texture degeneration. After that, a set of theoretical derivations demonstrate that the VSD can be approximately decomposed into the S-VSDs multiplied by their associated weights. To obtain the S-VSDs efficiently, a layer-based representation method is developed, where all the pixels with the same level of depth changes are represented with a layer. It enables the S-VSD calculation at the layer level. Meanwhile, a nonlinear mapping function is learnt to accurately represent the relationship between the VSD and S-VSDs, automatically providing weights for the S-VSDs during VSD estimation. To learn such a function, a dataset of the VSD and its associated S-VSDs are built, termed as VSDSet. Experimental results show that the VSD can be accurately estimated with the weights learnt by the nonlinear mapping function once its associated S-VSDs are available. The proposed method outperforms the relevant state-of-the-art methods in both accuracy and efficiency. The VSDSet and source code of the proposed method will be available athttps://github.com/jianjin008/. Xingxing Zhang 0001, Lili Meng, Weisi Lin, Jie Liang 0001, Huaxiang Zhang 0001, Yao Zhao 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | CoSCL: Cooperation of Small Continual Learners is Stronger Than a Big One
Xingxing Zhang 0001, Qian Li 0040, Jun Zhu 0001 |
ECCV (26) | 2 |
| 2022 | Memory Replay with Data Compression for Continual Learning
Xingxing Zhang 0001, Longhui Yu, Chongxuan Li, Lanqing Hong, Zhenguo Li, Jun Zhu 0001 |
ICLR | 2 |
| 2022 | One-step Low-Rank Representation for ClusteringabstractExisting low-rank representation-based methods adopt a two-step framework, which must employ an extra clustering method to gain labels after representation learning. In this paper, a novel one-step representation-based method, i.e., One-step Low-Rank Representation (OLRR), is proposed to capture multi-subspace structures for clustering. OLRR integrates the low-rank representation model and clustering into a unified framework. Thus it can jointly learn the low-rank subspace structure embedded in the database and gain the clustering results. In particular, by approximating the representation matrix with two same clustering indicator matrices, OLRR can directly show the probability of samples belonging to each cluster. Further, a probability penalty is introduced to ensure that the samples with smaller distances are more inclined to be in the same cluster, thus enhancing the discrimination of the clustering indicator matrix and resulting in a more favorable clustering performance. Moreover, to enhance the robustness against noise, OLRR uses the probability to guide denoising and then performs representation learning and clustering in a recovered clean space. Extensive experiments well demonstrate the robustness and effectiveness of OLRR. Our code is publicly available at: https://github.com/fuzhiqiang1230/OLRR. Zhiqiang Fu, Yao Zhao 0001, Dongxia Chang, Yiming Wang 0007, Jie Wen 0001, Xingxing Zhang 0001, Guodong Guo |
ACM Multimedia | 6 |
| 2022 | The More, The Better? Active Silencing of Non-Positive Transfer for Efficient Multi-Domain Few-Shot ClassificationabstractFew-shot classification refers to recognizing several novel classes given only a few labeled samples. Many recent methods try to gain an adaptation benefit by learning prior knowledge from more base training domains, aka. multi-domain few-shot classification. However, with extensive empirical evidence, we find more is not always better: current models do not necessarily benefit from pre-training on more base classes and domains, since the pre-trained knowledge might be non-positive for a downstream task. In this work, we hypothesize that such redundant pre-training can be avoided without compromising the downstream performance. Inspired by the selective activating/silencing mechanism in the biological memory system, which enables the brain to learn a new concept from a few experiences both quickly and accurately, we propose to actively silence those redundant base classes and domains for efficient multi-domain few-shot classification. Then, a novel data-driven approach named Active Silencing with hierarchical Subset Selection (AS3) is developed to address two problems: 1) finding a subset of base classes that adequately represent novel classes for efficient positive transfer; and 2) finding a subset of base learners (i.e., domains) with confident accurate prediction in a new domain. Both problems are formulated as distance-based sparse subset selection. We extensively evaluate AS3 on the recent META-DATASET benchmark as well as MNIST, CIFAR10, and CIFAR100, where AS3 achieves over 100% acceleration while maintaining or even improving accuracy. Our code and Appendix are available at https://github.com/indussky8/AS3. Xingxing Zhang 0001, Zhizhe Liu, Weikai Yang, Jun Zhu 0001 |
ACM Multimedia | 1 |
| 2022 | Auto-weighted low-rank representation for clustering
Zhiqiang Fu, Yao Zhao 0001, Dongxia Chang, Xingxing Zhang 0001, Yiming Wang 0007 |
Knowl. Based Syst. | 4 |
| 2022 | Just Noticeable Difference for Deep Machine VisionabstractAs an important perceptual characteristic of the Human Visual System (HVS), the Just Noticeable Difference (JND) has been studied for decades with image and video processing (e.g., perceptual visual signal compression). However, there is little exploration on the existence of JND for the Deep Machine Vision (DMV), although the DMV has made great strides in many machine vision tasks. In this paper, we take an initial attempt, and demonstrate that the DMV has the JND, termed as the DMV-JND. We then propose a JND model for the image classification task in the DMV. It has been discovered that the DMV can tolerate distorted images with average PSNR of only 9.56dB (the lower the better), by generating JND via unsupervised learning with the proposed DMV-JND-NET. In particular, a semantic-guided redundancy assessment strategy is designed to restrain the magnitude and spatial distribution of the DMV-JND. Experimental results on image classification demonstrate that we successfully find the JND for deep machine vision. Our DMV-JND facilitates a possible direction for DMV-oriented image and video compression, watermarking, quality assessment, deep neural network security, and so on. Xingxing Zhang 0001, Xin Fu 0009, Huan Zhang 0008, Weisi Lin, Jian Lou 0003, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | MFHI: Taking Modality-Free Human Identification as Zero-Shot LearningabstractHuman identification is an important topic in event detection, person tracking, and public security. There have been numerous methods proposed for human identification, such as face identification, person re-identification, and gait identification. Typically, existing methods predominantly classify a queried image to a specific identity in an image gallery set (I2I). This is seriously limited for the scenario where only a textual description of the query or an attribute gallery set is available in a wide range of video surveillance applications (A2IorI2A). However, very few efforts have been devoted towards modality-free identification, i.e., identifying a query in a gallery set in a scalable way. In this work, we take an initial attempt, and formulate such a novelModality-FreeHumanIdentification (named MFHI) task as a generic zero-shot learning model in a scalable way. Meanwhile, it is capable of bridging the visual and semantic modalities by learning a discriminative prototype of each identity. In addition, the semantics-guided spatial attention is enforced on visual modality to obtain interpretable representations with both high global category-level and local attribute-level discrimination. Finally, we design and conduct an extensive group of experiments on two common challenging identification tasks, including face identification and person re-identification, demonstrating that our method outperforms a wide variety of state-of-the-art methods on modality-free human identification. Zhizhe Liu, Xingxing Zhang 0001, Zhenfeng Zhu, Shuai Zheng 0005, Yao Zhao 0001, Jian Cheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Diagnosing Ensemble Few-Shot ClassifiersabstractThe base learners and labeled samples (shots) in an ensemble few-shot classifier greatly affect the model performance. When the performance is not satisfactory, it is usually difficult to understand the underlying causes and make improvements. To tackle this issue, we propose a visual analysis method, FSLDiagnotor. Given a set of base learners and a collection of samples with a few shots, we consider two problems: 1) finding a subset of base learners that well predict the sample collections; and 2) replacing the low-quality shots with more representative ones to adequately represent the sample collections. We formulate both problems as sparse subset selection and develop two selection algorithms to recommend appropriate learners and shots, respectively. A matrix visualization and a scatterplot are combined to explain the recommended learners and shots in context and facilitate users in adjusting them. Based on the adjustment, the algorithm updates the recommendation results for another round of improvement. Two case studies are conducted to demonstrate that FSLDiagnotor helps build a few-shot classifier efficiently and increases the accuracy by 12% and 21%, respectively. Weikai Yang, Xi Ye 0003, Xingxing Zhang 0001, Lanxi Xiao, Jiazhi Xia, Zhongyuan Wang 0006, Jun Zhu 0001, Hanspeter Pfister, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2021 | Double Low-Rank Representation With Projection Distance Penalty for ClusteringabstractThis paper presents a novel, simple yet robust self-representation method, i.e., Double Low-Rank Representation with Projection Distance penalty (DLRRPD) for clustering. With the learned optimal projected representations, DLRRPD is capable of obtaining an effective similarity graph to capture the multi-subspace structure. Besides the global low-rank constraint, the local geometrical structure is additionally exploited via a projection distance penalty in our DLRRPD, thus facilitating a more favorable graph. Moreover, to improve the robustness of DLRRPD to noises, we introduce a Laplacian rank constraint, which can further encourage the learned graph to be more discriminative for clustering tasks. Meanwhile, Frobenius norm (instead of the popularly used nuclear norm) is employed to enforce the graph to be more block-diagonal with lower complexity. Extensive experiments have been conducted on synthetic, real, and noisy data to show that the proposed method outperforms currently available alternatives by a margin of 1.0%~10.1%. Zhiqiang Fu, Yao Zhao 0001, Dongxia Chang, Xingxing Zhang 0001, Yiming Wang 0007 |
CVPR | 4 |
| 2021 | To See in the Dark: N2DGAN for Background Modeling in Nighttime SceneabstractDue to the deteriorated conditions of illumination lack and uneven lighting, the performance of traditional background modeling methods is greatly limited for the surveillance of nighttime video. To make background modeling under nighttime scene performs as well as in daytime condition, we put forward a promising generation-based background modeling framework for foreground surveillance. With a pre-specified daytime reference image as background frame, the GAN based generation model, called N2DGAN, is trained to transfer each frame of nighttime video to a virtual daytime image with the same scene to the reference image except for the foreground part. Specifically, to balance the preservation of background scene and the foreground object(s) in generating the virtual daytime image, we presented a two-pathway generation model, in which the global and local sub-networks were well combined with spatial and temporal consistency constraints. For the sequence of generated virtual daytime images, a multi-scale Bayes model was further proposed to characterize pertinently the temporal variation of background. We manually labeled ground truth on the collected nightime video datasets for performance evaluation. The impressive results illustrated in both the main paper and supplementary show the effectiveness of our proposed approach. Zhenfeng Zhu, Yingying Meng, Deqiang Kong, Xingxing Zhang 0001, Yandong Guo, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | A Theoretical Revisit to Linear Convergence for Saddle Point ProblemsabstractRecently, convex-concave bilinear Saddle Point Problems (SPP) is widely used in lasso problems, Support Vector Machines, game theory, and so on. Previous researches have proposed many methods to solve SPP, and present their convergence rate theoretically. To achieve linear convergence, analysis in those previouse studies requires strong convexity of φ( z ). But, we find the linear convergence can also be achieved even for a general convex but not strongly convex φ( z ). In the article, by exploiting the strong duality of SPP, we propose a new method to solve SPP, and achieve the linear convergence. We present a new general sufficient condition to achieve linear convergence, but do not require the strong convexity of φ( z ). Furthermore, a more efficient method is also proposed, and its convergence rate is analyzed in theoretical. Our analysis shows that the well conditioned φ( z ) is necessary to improve the efficiency of our method. Finally, we conduct extensive empirical studies to evaluate the convergence performance of our methods. Wendi Wu, En Zhu, Xinwang Liu 0002, Xingxing Zhang 0001, Lailong Luo, Jianping Yin |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2020 | Distribution-Induced Bidirectional Generative Adversarial Network for Graph Representation LearningabstractGraph representation learning aims to encode all nodes of a graph into low-dimensional vectors that will serve as input of many computer vision tasks. However, most existing algorithms ignore the existence of inherent data distribution and even noises. This may significantly increase the phenomenon of over-fitting and deteriorate the testing accuracy. In this paper, we propose a Distribution-induced Bidirectional Generative Adversarial Network (named DBGAN) for graph representation learning. Instead of the widely used Gaussian assumption, the prior distribution of latent representation in our DBGAN is estimated in a structure-aware way, which implicitly bridges the graph and content spaces by prototype learning. Thus discriminative and robust representations are generated for all nodes. Furthermore, to improve their generalization ability while preserving representation ability, the sample-level and distribution-level consistency are well balanced via a bidirectional adversarial learning framework. An extensive group of experiments is then carefully designed and presented, demonstrating that our DBGAN obtains remarkably more favorable trade-off between representation and robustness, and meanwhile is dimension-efficient, over currently available alternatives in various tasks. Shuai Zheng 0005, Zhenfeng Zhu, Xingxing Zhang 0001, Zhizhe Liu, Jian Cheng 0001, Yao Zhao 0001 |
CVPR | 3 |
| 2020 | ProLFA: Representative prototype selection for local feature aggregation
Xingxing Zhang 0001, Zhenfeng Zhu, Yao Zhao 0001 |
Neurocomputing | 1 |
| 2020 | Convolutional prototype learning for zero-shot recognition
Zhizhe Liu, Xingxing Zhang 0001, Zhenfeng Zhu, Shuai Zheng 0005, Yao Zhao 0001, Jian Cheng 0001 |
Image Vis. Comput. | 2 |
| 2020 | Canonical Correlation Analysis With L2, 1-Norm for Multiview Data RepresentationabstractFor many machine learning algorithms, their success heavily depends on data representation. In this paper, we present an l2,1-norm constrained canonical correlation analysis (CCA) model, that is, L2,1-CCA, toward discovering compact and discriminative representation for the data associated with multiple views. To well exploit the complementary and coherent information across multiple views, the l2,1-norm is employed to constrain the canonical loadings and measure the canonical correlation loss term simultaneously. It enables, on the one hand, the canonical loadings to be with the capacity of variable selection for facilitating the interpretability of the learned canonical variables, and on the other hand, the learned canonical common representation keeps highly consistent with the most canonical variables from each view of the data. Meanwhile, the proposed L2,1-CCA can also be provided with the desired insensitivity to noise (outliers) to some degree. To solve the optimization problem, we develop an efficient alternating optimization algorithm and give its convergence analysis both theoretically and experimentally. Considerable experiment results on several real-world datasets have demonstrated that L2,1-CCA can achieve competitive or better performance in comparison with some representative approaches for multiview representation learning. Meixiang Xu, Zhenfeng Zhu, Xingxing Zhang 0001, Yao Zhao 0001, Xuelong Li 0001 |
IEEE Trans. Cybern. | 3 |
| 2020 | Understand Dynamic Regret with Switching Cost for Online Decision MakingabstractAs a metric to measure the performance of an online method, dynamic regret with switching cost has drawn much attention for online decision making problems. Although the sublinear regret has been provided in much previous research, we still have little knowledge about the relation between the dynamic regret and the switching cost . In the article, we investigate the relation for two classic online settings: Online Algorithms (OA) and Online Convex Optimization (OCO). We provide a new theoretical analysis framework that shows an interesting observation; that is, the relation between the switching cost and the dynamic regret is different for settings of OA and OCO. Specifically, the switching cost has significant impact on the dynamic regret in the setting of OA. But it does not have an impact on the dynamic regret in the setting of OCO. Furthermore, we provide a lower bound of regret for the setting of OCO, which is same with the lower bound in the case of no switching cost. It shows that the switching cost does not change the difficulty of online decision making problems in the setting of OCO. Xingxing Zhang 0001, En Zhu, Xinwang Liu 0002, Jianping Yin |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2020 | Hierarchical Prototype Learning for Zero-Shot RecognitionabstractZero-Shot Learning (ZSL) has received extensive attention and successes in recent years especially in areas of fine-grained object recognition, retrieval, and image captioning. Key to ZSL is to transfer knowledge from the seen to the unseen classes via auxiliary semantic prototypes (e.g., word or attribute vectors). However, the popularly learned projection functions in previous works cannot generalize well due to non-visual components included in semantic prototypes. Besides, the incompleteness of provided prototypes and captured images has less been considered by the state-of-the-art approaches in ZSL. In this paper, we propose a hierarchical prototype learning formulation to provide a systematical solution (named HPL) for zero-shot recognition. Specifically, HPL is able to obtain discriminability on both seen and unseen class domains by learning visual prototypes respectively under the transductive setting. To narrow the gap of two domains, we further learn the interpretable super-prototypes in both visual and semantic spaces. Meanwhile, the two spaces are further bridged by maximizing their structural consistency. This not only facilitates the representativeness of visual prototypes, but also alleviates the loss of information of semantic prototypes. An extensive group of experiments are then carefully designed and presented, demonstrating that HPL obtains remarkably more favorable efficiency and effectiveness, over currently available alternatives under various settings. Xingxing Zhang 0001, Shupeng Gui, Zhenfeng Zhu, Yao Zhao 0001, Ji Liu 0002 |
IEEE Trans. Multim. | 1 |
| 2019 | Diffusion induced graph representation learning
Fuzhen Li, Zhenfeng Zhu, Xingxing Zhang 0001, Jian Cheng 0001, Yao Zhao 0001 |
Neurocomputing | 3 |
| 2019 | Seeing All From a Few: ℓ1-Norm-Induced Discriminative Prototype SelectionabstractPrototype selection aims to remove redundancy and irrelevance from large-scale data by selecting an informative subset, which makes it possible to see all data from a few prototypes. However, due to the outliers and uncertain distribution of the data, the selected prototypes are generally less representative and diversified. To alleviate this issue, we develop, in this paper, a ℓ1-norm-induced discriminative prototype selection model (ℓ1-ProSe). Inspired by the good performance of sparse representation, the sparsity property of data is rationally exploited in the formulated model. Meanwhile, to characterize the pairwise similarity in the learned sparse representation space, a more promising ℓ1-norm metric is applied for robust selection instead of the popularly used Euclidean metric in previous works. Considering the convexity of the model to be solved, a composite block coordinate descent solver is presented with rigorous theoretical analysis on its convergence. Furthermore, we extend our model to support online prototype selection by using already obtained prototypes and newly arrived data. Experimental results on synthetic data sets and some applications such as video summarization, motion segmentation, and scene categorization demonstrate that the proposed method is considerably superior to the state-of-the-art methods in the prototype selection. Xingxing Zhang 0001, Zhenfeng Zhu, Yao Zhao 0001, Dongxia Chang, Ji Liu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | Self-Supervised Deep Low-Rank Assignment Model for Prototype SelectionabstractPrototype selection is a promising technique for removing redundancy and irrelevance from large-scale data. Here, we consider it as a task assignment problem, which refers to assigning each element of a source set to one representative, i.e., prototype. However, due to the outliers and uncertain distribution on source, the selected prototypes are generally less representative and interesting. To alleviate this issue, we develop in this paper a Self-supervised Deep Low-rank Assignment model (SDLA). By dynamically integrating a low-rank assignment model with deep representation learning, our model effectively ensures the goodness-of-exemplar and goodness-of-discrimination of selected prototypes. Specifically, on the basis of a denoising autoencoder, dissimilarity metrics on source are continuously self-refined in embedding space with weak supervision from selected prototypes, thus preserving categorical similarity. Conversely, working on this metric space, similar samples tend to select the same prototypes by designing a low-rank assignment model. Experimental results on applications like text clustering and image classification (using prototypes) demonstrate our method is considerably superior to the state-of-the-art methods in prototype selection. Xingxing Zhang 0001, Zhenfeng Zhu, Yao Zhao 0001, Deqiang Kong |
IJCAI | 1 |
| 2018 | Sparsity induced prototype learning via ℓp, 1-norm grouping
Xingxing Zhang 0001, Zhenfeng Zhu, Yao Zhao 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2018 | Learning a General Assignment Model for Video AnalyticsabstractVideo analytics, also known as video content analysis, involves a variety of pivotal tasks, such as video segmentation and recognition. In essence, performing most of these tasks can be viewed as learning an assignment model. Here, assignment model refers to what each element in a target set is assigned to the element in an opposite source set under some kind of constraint. Existing assignment models generally suffer from some limitations. For instance, imperfect results can be obtained when the source set is artificially provided with less representativeness, or part of the target set has expected assignments, thus significantly limiting the application of assignment model. To alleviate these issues, we develop, in this paper, a general assignment model to dynamically learn a source set by minimizing structural dissimilarity in a low-dimensional space. Furthermore, by associating the expected assignment solution with empirical assignment indicator via a consistency boosting strategy, the proposed assignment model is provided with a potential powerful generalization ability to deal flexibly with the unsupervised, semi-supervised, and fully supervised scenarios in assignment model learning. Considering the separability of both objective and constraints to be solved, an alternating direction method of multipliers solver is presented with rigorous theoretical analysis on its convergence. Experimental results on video analysis, including motion segmentation, activities recognition, and scene categorization, demonstrate that the proposed general assignment model is considerable superior to the state-of-the-art methods. Xingxing Zhang 0001, Zhenfeng Zhu, Yao Zhao 0001, Dongxia Chang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |