EDBT 2026 Demo / reviewers in the wild / expert
Min Zhang 0068
dblp:83/5342-68
· DBLP profile ↗
31ranked-venue papers
6as first author
30since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 4 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 14 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Step-GRPO: Internalizing Dynamic Early Exit for Efficient ReasoningabstractLarge reasoning models that use long chainof-thought excel at problem-solving yet waste compute on redundant checks.Curbing this overthinking is hard: training-time length penalties can cripple ability, while inferencetime early-exit adds system overhead.To bridge this gap, we propose Step-GRPO, a novel post-training framework that internalizes dynamic early-exit capabilities directly into the model.Step-GRPO shifts the optimization objective from raw tokens to semantic steps by utilizing linguistic markers to structure reasoning.We introduce a Dynamic Truncated Rollout mechanism that exposes the model to concise high-confidence trajectories during exploration, synergized with a Step-Aware Relative Reward that dynamically penalizes redundancy based on group-level baselines.Extensive experiments across three model sizes on diverse benchmarks demonstrate that Step-GRPO achieves a superior accuracy-efficiency tradeoff.On Qwen3-8B, our method reduces token consumption by 32.0% compared to the vanilla model while avoiding the accuracy degradation observed in traditional length-penalty methods. Benteng Chen, Weida Wang, Shufei Zhang, Mingbao Lin, Min Zhang 0068 |
ACL (1) | 5 |
| 2026 | Tailoring Diagnostic Modeling to Individual Learners: Personalized Distractor Generation via MCTS-Guided Reasoning ReconstructionabstractTao Wu, Jingyuan Chen, Wang Lin, Jian Zhan, Mengze Li, Fangzhou Jin, Min Zhang, Kun Kuang, Fei Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jingyuan Chen 0003, Jian Zhan, Mengze Li 0001, Fangzhou Jin, Min Zhang 0068, Kun Kuang 0001, Fei Wu 0001 |
ACL (1) | 7 |
| 2026 | A survey of slow thinking-based reasoning LLMs using reinforcement learning and test-time scaling law
Qianjun Pan, Wenkai Ji, Yuyang Ding, Junsong Li, Shilian Chen, Jie Zhou 0015, Qin Chen 0001, Min Zhang 0068, Yulan Wu, Liang He 0001 |
Inf. Process. Manag. | 9 |
| 2025 | Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient InferenceabstractIn recent years, applying multi-modal large language models (MLLMs) in various fields has achieved remarkable success. However, as the foundation model for many downstream tasks, MLLMs comprise the well-known Transformer network, which has a less efficient quadratic computation complexity. In this study, we introduce Cobra, a multi-modal large-scale language model built upon a state-space model, which has demonstrated significant potential in efficiently handling long sequences with fast inference and linear scalability concerning sequence length. Specifically, Cobra involves replacing Transformer-based backbone models (e.g., LLaMA or Phi) with pre-trained Mamba language models. We then empirically explore effective strategies for aligning visual and textual modalities and integrating various pre-trained Mamba model variants with visual encoders. Experiments across various multi-modal benchmarks demonstrate that: (i) Cobra performs 3× ∼ 4× faster than the most computationally efficient state-of-the-art methods, e.g., LLaVA-Phi and MobileVLM v2. Additionally, its performance is significantly enhanced thanks to the implementation of linear sequential modeling. (ii) Cobra fine-tunes a small parameter (∼48% of model parameters), leading to a significant improvement in overall performance compared to LLaVA. Han Zhao 0008, Min Zhang 0068, Pengxiang Ding, Siteng Huang |
AAAI | 2 |
| 2025 | FIPO: Free-form Instruction-oriented Prompt Optimization with Preference Dataset and Modular Fine-tuning SchemaabstractWhen carefully optimized by human experts, naive prompts can significantly enhance the task performance of large language models (LLMs). However, such expert-driven prompt optimizations are resource-intensive. To address this, some studies have proposed Automatic Prompt Optimization (APO), which refines naive prompts according to task outputs from in-box testing models, utilizing advanced LLMs (e.g., GPT-4) in an ad-hoc way. Although effective, current approaches face challenges in generalization and privacy risks. To overcome these limitations, we have developed the first large-scale Prompt Optimization Preference (POP) dataset, fine-tuned offline local LLM-based optimizers, and conducted fairly evaluations across various downstream models. Our method, named Free-from Instruction-oriented Prompt Optimization (FIPO), allows precise optimization of the core task instructions in naive prompts in a model-agnostic manner. FIPO uses a modular APO template that dynamically incorporates the naive task instructions, optional instruction responses, and optional ground truth to produce refined prompts. The POP dataset is meticulously constructed using advanced LLMs, undergoing rigorous cross-validation by human experts and analytical models. By leveraging insights from this dataset, along with Tulu2 models and diverse fine-tuning strategies, we validate the efficacy of the FIPO framework across five public benchmarks and six testing models. Our dataset and codes are available at: https://github.com/LuJunru/FIPO_Project. Junru Lu, Siyu An, Min Zhang 0068, Yulan He 0002, Xing Sun 0001 |
COLING | 3 |
| 2025 | VLAS: Vision-Language-Action Model with Speech Instructions for Customized Robot ManipulationabstractVision-language-action models (VLAs) have recently become highly prevalent in robot manipulation due to its end-to-end architecture and impressive performance. However, current VLAs are limited to processing human instructions in textual form, neglecting the more natural speech modality for human interaction. A typical approach of incorporating speech modality into VLA necessitates a separate speech recognition system to transcribe spoken instructions into text. Such a cascading pipeline raises two major concerns for robotic systems. First, the entire model grows in size and complexity, potentially resulting in redundant computations and increased memory consumption. Second, the transcription procedure would lose non-semantic information in the raw speech, such as voiceprint, which is crucial for a robot to successfully understand and complete customized tasks. To this end, we propose VLAS, the fisrt end-to-end policy model that seamlessly integrates speech modality for robot manipulation. We present a three-stage speech instruction tuning strategy leveraging multimodal datasets, including our manually curated SQA and CSI datasets. Furthermore, to facilitate personalized operations, we develop a voice retrieval-augmented generation (RAG) approach to enhance the robot's performance in tasks requiring individual-specific knowledge. Experimental results show that the proposed VLAS, following either textual or speech instructions, can achieve performance comparable to traditional VLAs on the CALVIN benchmark. In addition, we created a benchmark consisting of customization tasks, where our VLAS demonstrates absolute superiority by fully leveraging the auxiliary information in speech. Pengxiang Ding, Min Zhang 0068, Zhefei Gong, Shuanghao Bai, Han Zhao 0008 |
ICLR | 3 |
| 2025 | REMEDY: Recipe Merging Dynamics in Large Vision-Language ModelsabstractModel merging has emerged as a powerful technique for combining task-specific vision models into a unified and multi-functional model. Previous methods represented by task arithmetic, have demonstrated effectiveness and scalability in this domain. When large vision-language models (LVLMs) arise with model size scaling up, this design becomes challenging to fuse different instruction-tuned LVLMs for generalization enhancement. The large scale and multi-modal nature of LVLMs present unique obstacles, including constructing reusable and modular components to accommodate the multi-component architecture of LVLMs and the requirement for dynamic fusion based on multi-modal input tokens. To address these challenges, we propose the \textbf{RE}cipe \textbf{ME}rging \textbf{DY}namics (REMEDY) method, a scalable and flexible paradigm for model merging in LVLMs. We first define reusable modules termed \textit{recipes} including the projector and shallow LLM layers, enhancing visual-language understanding. Then, we introduce a modality-aware allocator dynamically generates weights in a one-shot manner based on input relevance to existing recipes, enabling efficient cross-modal knowledge integration. REMEDY thus offers an adaptive solution for LVLMs to tackle both seen (i.e., multi-task learning) and unseen (i.e., zero-shot generalization) tasks. Experimental results demonstrate that our method consistently improves performance on both seen and unseen tasks, underscoring the effectiveness of REMEDY in diverse multi-modal scenarios. Didi Zhu, Yibing Song, Tao Shen 0002, Ziyu Zhao 0001, Jinluan Yang, Min Zhang 0068, Chao Wu 0001 |
ICLR | 6 |
| 2025 | Strong and Weak Identifiability of Optimization-based Causal Discovery in Non-linear Additive Noise ModelsabstractCausal discovery aims to identify causal relationships from observational data. Recently, optimization-based causal discovery methods have attracted extensive attention in the literature due to their efficiency in handling high-dimensional problems. However, we observe that optimization-based methods often perform well on certain problems but struggle with others. This paper identifies a specific characteristic of causal structural equations that determines the difficulty of identification in causal discovery and, in turn, the performance of optimization-based methods. We conduct an in-depth study of the additive noise model (ANM) and propose to further divide identifiable problems into strongly and weakly identifiable types based on the difficulty of identification. We also provide a sufficient condition to distinguish the two categories. Inspired by these findings, this paper further proposes GENE, a generic method for addressing strongly and weakly identifiable problems in a unified way under the ANM assumption. GENE adopts an order-based search framework that incorporates conditional independence tests into order fitness evaluation, ensuring effectiveness on weakly identifiable problems. In addition, GENE restricts the dimensionality of the effect variables to ensure scale invariance, a property crucial for practical applications. Experiments demonstrate that GENE is uniquely effective in addressing weakly identifiable problems while also remaining competitive with state-of-the-art causal discovery algorithms for strongly identifiable problems. Mingjia Li 0002, Hong Qian, Tian-Zuo Wang, Min Zhang 0068, Aimin Zhou |
ICML | 5 |
| 2025 | ERICT: Enhancing Robustness by Identifying Concept Tokens in Zero-Shot Vision Language ModelsabstractPre-trained vision-language models (VLMs) have revolutionized the field of machine learning, demonstrating exceptional performance across a wide range of tasks. However, their robustness remains vulnerable to the spurious-correlation problem. Existing works often involve fine-tuning the model with labeled data or relying on large language models (LLMs) to generate more complex prompts. Although effective to some extent, these methods introduce new challenges, including additional computational costs and dependence on the quality of prompts without fully utilizing the vision modality. To address these limitations, we propose a novel method named ERICT to Enhance model Robustness by Identifying Concept Tokens. ERICT mitigates spurious correlation directly in the inference stage and comprises two key steps: (1) Identify concept tokens capturing invariant features through auxiliary prompts to generate a token-level mask. (2) Apply the mask to the attention weights of the CLS token in the vision encoder to help the model focus on the relevant image region. Extensive experiments show that ERICT significantly improves the overall performance including that of the worst group, and achieves new state-of-the-art results. Xinpeng Dong, Min Zhang 0068, Didi Zhu, Ye Jun Jian, Keli Zhang, Aimin Zhou, Fei Wu 0001, Kun Kuang 0001 |
ICML | 2 |
| 2025 | Advancing Personalized Learning with Neural Collapse for Long-Tail ChallengeabstractPersonalized learning, especially data-based methods, has garnered widespread attention in recent years, aiming to meet individual student needs. However, many works rely on the implicit assumption that benchmarks are high-quality and well-annotated, which limits their practical applicability. In real-world scenarios, these benchmarks often exhibit long-tail distributions, significantly impacting model performance. To address this challenge, we propose a novel method called Neural-Collapse-Advanced personalized Learning (NCAL), designed to learn features that conform to the same simplex equiangular tight frame (ETF) structure. NCAL introduces Text-modality Collapse (TC) regularization to optimize the distribution of text embeddings within the large language model (LLM) representation space. Notably, NCAL is model-agnostic, making it compatible with various architectures and approaches, thereby ensuring broad applicability. Extensive experiments demonstrate that NCAL effectively enhances existing works, achieving new state-of-the-art performance. Additionally, NCAL mitigates class imbalance, significantly improving the model’s generalization ability. Hanglei Hu, Zhikang Chen, Sen Cui, Fei Wu 0001, Kun Kuang 0001, Min Zhang 0068, Bo Jiang 0016 |
ICML | 7 |
| 2025 | CALM: Consensus-Aware Localized Merging for Multi-Task LearningabstractModel merging aims to integrate the strengths of multiple fine-tuned models into a unified model while preserving task-specific capabilities. Existing methods, represented by task arithmetic, are typically classified into global- and local-aware methods. However, global-aware methods inevitably cause parameter interference, while local-aware methods struggle to maintain the effectiveness of task-specific details in the merged model. To address these limitations, we propose a Consensus Aware Localized Merging (CALM) method which incorporates localized information aligned with global task consensus, ensuring its effectiveness post-merging. CALM consists of three key components: (1) class-balanced entropy minimization
sampling, providing a more flexible and reliable way to leverage unsupervised data; (2) an efficient-aware framework, selecting a small set of tasks for sequential merging with high scalability; (3) a consensus-aware mask optimization, aligning localized binary masks with global task consensus and merging them conflict-free. Experiments demonstrate the superiority and robustness of our CALM, significantly outperforming existing methods and achieving performance close to traditional MTL. Kunda Yan, Min Zhang 0068, Sen Cui, Zikun Qu, Bo Jiang 0016, Changshui Zhang |
ICML | 2 |
| 2025 | A Two-Stage Pretraining-Finetuning Framework for Treatment Effect Estimation with Unmeasured ConfoundingabstractEstimating the conditional average treatment effect (CATE) from observational data plays a crucial role in areas such as e-commerce, healthcare, and economics. Existing studies mainly rely on the strong ignorability assumption that there are no unmeasured confounders, whose presence cannot be tested from observational data and can invalidate any causal conclusion. In contrast, data collected from randomized controlled trials (RCT) do not suffer from confounding, but are usually limited by a small sample size. In this paper, we propose a two-stage pretraining-finetuning (TSPF) framework using both large-scale observational data and small-scale RCT data to estimate the CATE in the presence of unmeasured confounding. In the first stage, a foundational representation of covariates is trained to estimate counterfactual outcomes through large-scale observational data. In the second stage, we propose to train an augmented representation of the covariates, which is concatenated to the foundational representation obtained in the first stage to adjust for the unmeasured confounding. To avoid overfitting caused by the small-scale RCT data in the second stage, we further propose a partial parameter initialization approach, rather than training a separate network. The superiority of our approach is validated on two public datasets with extensive experiments. The code is available at https://github.com/zhouchuanCN/KDD25-TSPF. Chuan Zhou 0013, Yaxuan Li 0002, Chunyuan Zheng 0001, Haiteng Zhang, Min Zhang 0068, Haoxuan Li 0001, Mingming Gong |
KDD (1) | 5 |
| 2024 | Learning to Reweight for Generalizable Graph Neural NetworkabstractGraph Neural Networks (GNNs) show promising results for graph tasks. However, existing GNNs' generalization ability will degrade when there exist distribution shifts between testing and training graph data. The fundamental reason for the severe degeneration is that most GNNs are designed based on the I.I.D hypothesis. In such a setting, GNNs tend to exploit subtle statistical correlations existing in the training set for predictions, even though it is a spurious correlation. In this paper, we study the problem of the generalization ability of GNNs on Out-Of-Distribution (OOD) settings. To solve this problem, we propose the Learning to Reweight for Generalizable Graph Neural Network (L2R-GNN) to enhance the generalization ability for achieving satisfactory performance on unseen testing graphs that have different distributions with training graphs. We propose a novel nonlinear graph decorrelation method, which can substantially improve the out-of-distribution generalization ability and compares favorably to previous methods in restraining the over-reduced sample size. The variables of graph representation are clustered based on the stability of their correlations, and graph decorrelation method learns weights to remove correlations between the variables of different clusters rather than any two variables. Besides, we introduce an effective stochastic algorithm based on bi-level optimization for the L2R-GNN framework, which enables simultaneously learning the optimal weights and GNN parameters, and avoids the over-fitting issue. Experiments show that L2R-GNN greatly outperforms baselines on various graph prediction benchmarks under distribution shifts. Zhengyu Chen 0001, Teng Xiao, Kun Kuang 0001, Zheqi Lv, Min Zhang 0068, Jinluan Yang, Chengqiang Lu, Hongxia Yang, Fei Wu 0001 |
AAAI | 5 |
| 2024 | Prompt-Based Distribution Alignment for Unsupervised Domain AdaptationabstractRecently, despite the unprecedented success of large pre-trained visual-language models (VLMs) on a wide range of downstream tasks, the real-world unsupervised domain adaptation (UDA) problem is still not well explored. Therefore, in this paper, we first experimentally demonstrate that the unsupervised-trained VLMs can significantly reduce the distribution discrepancy between source and target domains, thereby improving the performance of UDA. However, a major challenge for directly deploying such models on downstream UDA tasks is prompt engineering, which requires aligning the domain knowledge of source and target domains, since the performance of UDA is severely influenced by a good domain-invariant representation. We further propose a Prompt-based Distribution Alignment (PDA) method to incorporate the domain knowledge into prompt learning. Specifically, PDA employs a two-branch prompt-tuning paradigm, namely base branch and alignment branch. The base branch focuses on integrating class-related representation into prompts, ensuring discrimination among different classes. To further minimize domain discrepancy, for the alignment branch, we construct feature banks for both the source and target domains and propose image-guided feature tuning (IFT) to make the input attend to feature banks, which effectively integrates self-enhanced and cross-domain features into the model. In this way, these two branches can be mutually promoted to enhance the adaptation of VLMs for UDA. We conduct extensive experiments on three benchmarks to demonstrate that our proposed PDA achieves state-of-the-art performance. The code is available at https://github.com/BaiShuanghao/Prompt-based-Distribution-Alignment. Shuanghao Bai, Min Zhang 0068, Wanqi Zhou, Siteng Huang, Zhirong Luan, Badong Chen |
AAAI | 2 |
| 2024 | Troika: Multi-Path Cross-Modal Traction for Compositional Zero-Shot LearningabstractRecent compositional zero-shot learning (CZSL) methods adapt pre-trained vision-language models (VLMs) by constructing trainable prompts only for composed state-object pairs. Relying on learning the joint representation of seen compositions, these methods ignore the explicit modeling of the state and object, thus limiting the exploitation of pre-trained knowledge and generalization to unseen compositions. With a particular focus on the universality of the solution, in this work, we propose a novel paradigm for CZSL models that establishes three identification branches (i.e., Multi-Path) to jointly model the state, object, and composition. The presented Troika is an outstanding implementation that aligns the branch-specific prompt representations with decomposed visual features. To calibrate the bias between semantically similar multi-modal representations, we further devise a Cross-Modal Traction module into Troika that shifts the prompt representation towards the current visual content. We conduct extensive experiments on three popular benchmarks, where our method significantly outperforms existing methods in both closed-world and open-world settings. The code will be available at https://github.com/bighuang624/Troika. Siteng Huang, Biao Gong, Yutong Feng, Min Zhang 0068, Yiliang Lv |
CVPR | 4 |
| 2024 | QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
Pengxiang Ding, Han Zhao 0008, Wenxuan Song, Min Zhang 0068, Siteng Huang, Ningxi Yang |
ECCV (5) | 5 |
| 2024 | PiTe: Pixel-Temporal Alignment for Large Video-Language Model
Yang Liu 0358, Pengxiang Ding, Siteng Huang, Min Zhang 0068, Han Zhao 0008 |
ECCV (5) | 4 |
| 2024 | Scaling Laws Across Model Architectures: A Comparative Analysis of Dense and MoE Models in Large Language ModelsabstractThe scaling of large language models (LLMs) is a critical research area for the efficiency and effectiveness of model training and deployment.Our work investigates the transferability and discrepancies of scaling laws between Dense Models and Mixture of Experts (MoE) models.Through a combination of theoretical analysis and extensive experiments, including consistent loss scaling, optimal batch size and learning rate scaling, and resource allocation strategies scaling, our findings reveal that the power-law scaling framework also applies to MoE Models, indicating that the fundamental principles governing the scaling behavior of these models are preserved, even though the architecture differs.Additionally, MoE Models demonstrate superior generalization, resulting in lower testing losses with the same training compute budget compared to Dense Models.These findings indicate the scaling consistency and transfer generalization capabilities of MoE Models, providing new insights for optimizing MoE Model training and deployment strategies. Zhengyu Chen 0001, Keqing He 0001, Min Zhang 0068, Jingang Wang |
EMNLP | 5 |
| 2024 | MetaCoCo: A New Few-Shot Classification Benchmark with Spurious CorrelationabstractOut-of-distribution (OOD) problems in few-shot classification (FSC) occur when novel classes sampled from testing distributions differ from base classes drawn from training distributions, which considerably degrades the performance of deep learning models deployed in real-world applications. Recent studies suggest that the OOD problems in FSC mainly including: (a) cross-domain few-shot classification (CD-FSC) and (b) spurious-correlation few-shot classification (SC-FSC). Specifically, CD-FSC occurs when a classifier learns transferring knowledge from base classes drawn from \underline{seen} training distributions but recognizes novel classes sampled from unseen testing distributions. In contrast, SC-FSC arises when a classifier relies on non-causal features (or contexts) that happen to be correlated with the labels (or concepts) in base classes but such relationships no longer hold during the model deployment. Despite CD-FSC has been extensively studied, SC-FSC remains understudied due to lack of the corresponding evaluation benchmarks. To this end, we present Meta Concept Context (MetaCoCo), a benchmark with spurious-correlation shifts collected from real-world scenarios. Moreover, to quantify the extent of spurious-correlation shifts of the presented MetaCoCo, we further propose a metric by using CLIP as a pre-trained vision-language model. Extensive experiments on the proposed benchmark are performed to evaluate the state-of-the-art methods in FSC, cross-domain shifts, and self-supervised learning. The experimental results show that the performance of the existing methods degrades significantly in the presence of spurious-correlation shifts. We open-source all codes of our benchmark and hope that the proposed MetaCoCo can facilitate future research on spurious-correlation shifts problems in FSC. Min Zhang 0068, Haoxuan Li 0001, Fei Wu 0001, Kun Kuang 0001 |
ICLR | 1 |
| 2024 | RotoGBML: Towards Out-of-distribution Generalization for Gradient-based Meta-learningabstractGradient-based meta-learning (GBML) algorithms can quickly adapt to new tasks by transferring the learned meta-knowledge while assuming that all tasks come from the same distribution (in-distribution, ID). However, in the real world, they often grapple with an out-of-distribution (OOD) generalization challenge, where tasks stem from diverse distributions. OOD exacerbates discrepancies in task gradient magnitudes and directions, posing a formidable challenge for GBML in optimizing meta-knowledge by minimizing the sum of task gradients in each minibatch. To address this problem, we propose RotoGBML, a novel approach designed to homogenize OOD task gradients. RotoGBML employs reweighted vectors to dynamically balance diverse magnitudes to a standardized scale and uses rotation matrices to align conflicting directions. To reduce overhead, we homogenize gradients with the features rather than network parameters. Additionally, to circumvent the impact of non-causal features (e.g., backgrounds), we propose an Invariant Self-Information (ISI) module to extract invariant causal features (e.g., the outlines of objects). Finally, task gradients are homogenized based on these invariant causal features. Experiments demonstrate that RotoGBML outperforms state-of-the-art methods across various few-shot benchmarks. Min Zhang 0068, Zifeng Zhuang, Zhitao Wang |
ICME | 1 |
| 2024 | Neural Collapse Anchored Prompt Tuning for Generalizable Vision-Language ModelsabstractLarge-scale vision-language (V-L) models have demonstrated remarkable generalization capabilities for downstream tasks through prompt tuning. However, the mechanisms behind the learned text representations are unknown, limiting further generalization gains, and the limitations are more severe when faced with the prevalent class imbalances seen in web-sourced datasets. Recent advances in the neural collapse (NC) phenomenon of vision-only models suggest that the optimal representation structure is the simplex ETF, which paves the way to study representations in V-L models. In this paper, we make the first attempt to use NC for examining the representations in V-L models via prompt tuning. It is found that NC optimality of text-to-image representations shows a positive correlation with downstream generalizability, which is more severe under class imbalance settings. To improve the representations, we propose Neural-collapse-anchored Prompt Tuning (NPT), a novel method that learns prompts with text and image representations that satisfy the same simplex Equiangular Tight Frame (ETF). NPT incorporates two regularization terms: language-modality collapse and multi-modality isomorphism; and it is compatible with other prompt tuning methods. Extensive experiments show that NPT can consistently help to improve existing prompt tuning techniques across 11 datasets for both balanced and imbalanced settings. Didi Zhu, Zexi Li 0001, Min Zhang 0068, Junkun Yuan, Kun Kuang 0001, Chao Wu 0001 |
KDD | 3 |
| 2024 | ProFD: Prompt-Guided Feature Disentangling for Occluded Person Re-IdentificationabstractTo address the occlusion issues in person Re-Identification (ReID) tasks, many methods have been proposed to extract part features by introducing external spatial information. However, due to missing part appearance information caused by occlusion and noisy spatial information from external model, these purely vision-based approaches fail to correctly learn the features of human body parts from limited training data and struggle in accurately locating body parts, ultimately leading to misaligned part features. To tackle these challenges, we propose a Prompt-guided Feature Disentangling method (ProFD), which leverages the rich pre-trained knowledge in the textual modality facilitate model to generate well-aligned part features. ProFD first designs part-specific prompts and utilizes noisy segmentation mask to preliminarily align visual and textual embedding, enabling the textual prompts to have spatial awareness. Furthermore, to alleviate the noise from external masks, ProFD adopts a hybrid-attention decoder, ensuring spatial and semantic consistency during the decoding process to minimize noise impact. Additionally, to avoid catastrophic forgetting, we employ a self-distillation strategy, retaining pre-trained knowledge of CLIP to mitigate over-fitting. Evaluation results on the Market1501, DukeMTMC-ReID, Occluded-Duke, Occluded-ReID, and P-DukeMTMC datasets demonstrate that ProFD achieves state-of-the-art results. Can Cui 0008, Siteng Huang, Wenxuan Song, Pengxiang Ding, Min Zhang 0068 |
ACM Multimedia | 5 |
| 2024 | Neural Collapse Inspired Feature Alignment for Out-of-Distribution GeneralizationabstractThe spurious correlation between the background features of the image and its label arises due to that the samples labeled with the same class in the training set often co-occurs with a specific background, which will cause the encoder to extract non-semantic features for classification, resulting in poor out-of-distribution generalization performance. Although many studies have been proposed to address this challenge, the semantic and spurious features are still difficult to accurately decouple from the original image and fail to achieve high performance with deep learning models. This paper proposes a novel perspective inspired by neural collapse to solve the spurious correlation problem through the alternate execution of environment partitioning and learning semantic masks. Specifically, we propose to assign an environment to each sample by learning a local model for each environment and using maximum likelihood probability. At the same time, we require that the learned semantic mask neurally collapses to the same simplex equiangular tight frame (ETF) in each environment after being applied to the original input. We conduct extensive experiments on four datasets, and the results demonstrate that our method significantly improves out-of-distribution performance. Zhikang Chen, Min Zhang 0068, Sen Cui, Haoxuan Li 0001, Gang Niu 0001, Mingming Gong, Changshui Zhang, Kun Zhang 0001 |
NeurIPS | 2 |
| 2023 | MAP: Towards Balanced Generalization of IID and OOD through Model-Agnostic AdaptersabstractDeep learning has achieved tremendous success in recent years, but most of these successes are built on an independent and identically distributed (IID) assumption. This somewhat hinders the application of deep learning to the more challenging out-of-distribution (OOD) scenarios. Although many OOD methods have been proposed to address this problem and have obtained good performance on testing data that is of major shifts with training distributions, interestingly, we experimentally find that these methods achieve excellent OOD performance by making a great sacrifice of the IID performance. We call this finding the IID-OOD dilemma. Clearly, in real-world applications, distribution shifts between training and testing data are often uncertain, where shifts could be minor, and even close to the IID scenario, and thus it is truly important to design a deep model with the balanced generalization ability between IID and OOD. To this end, in this paper, we investigate an intriguing problem of balancing IID and OOD generalizations and propose a novel Model Agnostic adaPters (MAP) method, which is more reliable and effective for distribution-shift-agnostic real-world data. Our key technical contribution is to use auxiliary adapter layers to incorporate the inductive bias of IID into OOD methods. To achieve this goal, we apply a bilevel optimization to explicitly model and optimize the coupling relationship between the OOD model and auxiliary adapter layers. We also theoretically give a first-order approximation to save computational time. Experimental results on six datasets successfully demonstrate that MAP can greatly improve the performance of IID while achieving good OOD performance. Min Zhang 0068, Junkun Yuan, Yue He 0001, Zhengyu Chen 0001, Kun Kuang 0001 |
ICCV | 1 |
| 2023 | Multi-Level Correlation Network For Few-Shot Image ClassificationabstractFew-shot image classification(FSIC) aims to recognize novel classes given few labeled images from base classes. Recent works have achieved promising classification performance, especially for metric-learning methods, where a measure at only image feature level is usually used. In this paper, we argue that measure at such a level may not be effective enough to generalize from base to novel classes when using only a few images. Instead, a multi-level descriptor of an image is taken for consideration in this paper. We propose a multi-level correlation network (MLCN) for FSIC to tackle this problem by effectively capturing local information. Concretely, we present the self-correlation module and cross-correlation module to learn the semantic correspondence relation of local information based on learned representations. Moreover, we propose a pattern-correlation module to capture the pattern of fine-grained images and find relevant structural patterns between base classes and novel classes. Extensive experiments and analysis show the effectiveness of our proposed method on four widely-used FSIC benchmarks. Yunkai Dang, Meijun Sun, Min Zhang 0068, Zhengyu Chen 0001, Zheng Wang 0008 |
ICME | 3 |
| 2023 | Quantitatively Measuring and Contrastively Exploring Heterogeneity for Domain GeneralizationabstractDomain generalization (DG) is a prevalent problem in real-world applications, which aims to train well-generalized models for unseen target domains by utilizing several source domains. Since domain labels, i.e., which domain each data point is sampled from, naturally exist, most DG algorithms treat them as a kind of supervision information to improve the generalization performance. However, the original domain labels may not be the optimal supervision signal due to the lack of domain heterogeneity, i.e., the diversity among domains. For example, a sample in one domain may be closer to another domain, its original label thus can be the noise to disturb the generalization learning. Although some methods try to solve it by re-dividing domains and applying the newly generated dividing pattern, the pattern they choose may not be the most heterogeneous due to the lack of the metric for heterogeneity. In this paper, we point out that domain heterogeneity mainly lies in variant features under the invariant learning framework. With contrastive learning, we propose a learning potential-guided metric for domain heterogeneity by promoting learning variant features. Then we notice the differences between seeking variance-based heterogeneity and training invariance-based generalizable model. We thus propose a novel method called H eterogeneity-based Two-stage Contrastive Learning (HTCL) for the DG task. In the first stage, we generate the most heterogeneous dividing pattern with our contrastive metric. In the second stage, we employ an invariance-aimed contrastive learning by re-building pairs with the stable relation hinted by domains and classes, which better utilizes generated domain labels for generalization learning. Extensive experiments show HTCL better digs heterogeneity and yields great generalization performance. Yunze Tong, Junkun Yuan, Min Zhang 0068, Didi Zhu, Keli Zhang, Fei Wu 0001, Kun Kuang 0001 |
KDD | 3 |
| 2023 | TAGM: Task-Aware Graph Model for Few-shot Node ClassificationabstractGraph representation learning has attracted tremendous attention due to its remarkable performance in variety of real-world applications. However, because data labeling is always time and resource intensive, current supervised graph representation learning models for particular tasks frequently suffer from label sparsity issues. In light of this, graph few-shot learning has been proposed to tackle the performance degradation in face of limited annotated data challenge. While recent advances in graph few shot learning achieve promising performance, they typically force to use a generic feature embedding across various tasks. Ideally, we want to construct feature embeddings that are tuned for the given task because of the differences in distribution between tasks. In this work, we propose a novel Task-Aware Graph Model (TAGM) to learn task-aware node embedding. Specifically, we provide a new graph cell design that includes a graph convolution layer for aggregating and updating graph information as well as a two-layer linear transformation for node feature transformation. On this basis, we encode task information to learn the binary weight mask set and gradient mask set, where the weight mask set selects different network parameters for different tasks and the gradient mask set can dynamically update the selected network parameters in a different manner during the optimization process. Our model is more sensitive to task identity and performs better for a task graph input. Our extensive experiments on three graph-structured datasets demonstrate that our proposed method generally outperforms the state-of-the-art baselines in few-shot learning. Feng Zhao 0014, Min Zhang 0068 |
ICMR | 2 |
| 2022 | Tree Structure-Aware Few-Shot Image Classification via Hierarchical Aggregation
Min Zhang 0068, Siteng Huang |
ECCV (20) | 1 |
| 2022 | Domain Generalized Few-Shot Image Classification via Meta Regularization NetworkabstractIn few-shot image classification scenarios, meta-learning methods aim to learn transferable feature representations extracted from seen domains (base classes) in the meta-training phase and quickly adapt to unseen domains (novel classes) in the meta-testing phase. However, when seen and unseen domains have a large discrepancy, existing approaches do not perform well due to the incapability of generalizing to unseen domains. In this paper, we investigate the challenging domain generalized few-shot image classification problem. We design an Meta Regularization Network (MRN) to learn a domain-invariant discriminative feature space, where a learning to learn update strategy is used to simulate domain shifts caused by seen and unseen domains. The simulation trains the model to learn to reorganize the feature knowledge acquired from seen domains to represent unseen domains. Extensive experiments and analysis show that our proposed MRN can significantly improve the generalization ability of various meta-learning methods to achieve state-of-the-art performance in domain generalized few-shot learning. Min Zhang 0068, Siteng Huang |
ICASSP | 1 |
| 2021 | Attributes-Guided and Pure-Visual Attention Alignment for Few-Shot RecognitionabstractThe purpose of few-shot recognition is to recognize novel categories with a limited number of labeled examples in each class. To encourage learning from a supplementary view, recent approaches have introduced auxiliary semantic modalities into effective metric-learning frameworks that aim to learn a feature similarity between training samples (support set) and test samples (query set). However, these approaches only augment the representations of samples with available semantics while ignoring the query set, which loses the potential for the improvement and may lead to a shift between the modalities combination and the pure-visual representation. In this paper, we devise an attributes-guided attention module (AGAM) to utilize human-annotated attributes and learn more discriminative features. This plug-and-play module enables visual contents and corresponding attributes to collectively focus on important channels and regions for the support set. And the feature selection is also achieved for query set with only visual information while the attributes are not available. Therefore, representations from both sets are improved in a fine-grained manner. Moreover, an attention alignment mechanism is proposed to distill knowledge from the guidance of attributes to the pure-visual branch for samples without attributes. Extensive experiments and analysis show that our proposed module can significantly improve simple metric-based approaches to achieve state-of-the-art performance on different datasets and settings. Siteng Huang, Min Zhang 0068, Yachen Kang |
AAAI | 2 |
| 2020 | Knowledge Distillation for Model-Agnostic Meta-LearningabstractRecently, model-agnostic meta-learning (MAML) and its variants have drawn much attention in few-shot learning. In this paper, we investigate how to improve the performance of a portable MAML network so that it can be used in handheld devices, such as small robots, mobile phones, and laptops. We propose a novel approach named portable model-agnostic meta-learning (P-MAML), where valuable knowledge is distilled from a teacher MAML network to a portable student MAML. Moreover, data augmentation and ResNet architecture are employed in the teacher MAML network so as to avoid overfitting and enhance efficiency. To the best of our knowledge, this is the first work to consider a portable meta-learning model through knowledge distillation (KD) to learn a good initialization. Extensive experimental results on three real datasets show that our P-MAML algorithm greatly enhances the accuracy through KD from the teacher network. As shown, P-MAML with KD improves the performance of one-shot learning as high as 10% in comparison to that without KD. Min Zhang 0068, Sibo Gai |
ECAI | 1 |