EDBT 2026 Demo / reviewers in the wild / expert
Peng Ye 0006
dblp:53/930-6
· DBLP profile ↗
64ranked-venue papers
6as first author
64since 2021 · last 2026
0000-0002-8486-7562ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 5 first-author · 46 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 2 first-author · 26 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mitigating Low-Quality Reasoning in MLLMs: Self-Driven Refined Multimodal CoT with Selective Thinking and Step-wise Visual EnhancementabstractCurrent Multimodal Chain-of-Thought (MCoT) methods suffer from low-quality multimodal reasoning, characterized by overthinking on simple queries and inefficient utilization of visual information, resulting in vast inefficient and ineffective computations. In this paper, we discover that Multimodal Large Language Models (MLLMs) possess inherent capabilities to distinguish between simple and difficult queries and enhance task-related visual information, which remain underutilized by existing approaches. Based on this insight, we propose Self-Driven Refined Multimodal CoT (SDR-MCoT), a training-free framework that mitigates these issues through two self-driven modules. First, our selective thinking module employs entropy-based confidence estimation to determine whether queries require detailed reasoning, preventing overthinking on simple questions. Second, our step-wise visual enhancement module strengthens attention to relevant visual regions at each reasoning step without inserting additional tokens, achieving fine-grained visual grounding and enhancement with minimal overhead. Moreover, SDR-MCoT can be seamlessly integrated into various MLLMs, offering a practical solution for improving multimodal reasoning. Comprehensive experiments across eight benchmarks from diverse domains (multimodal reasoning, visual understanding, hallucination, and mathematical reasoning) demonstrate that SDR-MCoT consistently outperforms existing MCoT methods on four different base models with reduced overhead. For instance, on Qwen2-VL-7B, our method improves average accuracy by over 6% while reducing token consumption by approximately 60% compared to zero-shot CoT. Chongjun Tu, Peng Ye 0006, Dongzhan Zhou, Tao Chen 0003, Wanli Ouyang |
AAAI | 2 |
| 2026 | The Avengers: A Routing Recipe for Collective Intelligence in Language ModelsabstractProprietary models are increasingly dominating the race for ever-larger language models. Can open-source, smaller models remain competitive across a broad range of tasks? In this paper, we present the Avengers---a lightweight framework that leverages the collective intelligence of these smaller models. The Avengers builds upon four lightweight operations: (i) embedding: encode queries using a text embedding model; (ii) clustering: group queries based on their semantic similarity; (iii) scoring: scores each model's performance within each cluster; and (iv) voting: improve outputs via repeated sampling and voting. At inference time, each query is embedded and assigned to its nearest cluster. The top-performing model(s) within that cluster are selected to generate the response with repeated sampling. Remarkably, with 10 open-source models (~7B parameters each), the Avengers surpasses GPT-4o, 4.1, and 4.5 in average performance across 15 diverse datasets spanning mathematics, coding, logical reasoning, general knowledge, and affective tasks. In particular, it surpasses GPT-4.1 on mathematics tasks by 18.21% and on code tasks by 7.46%. Furthermore, the Avengers delivers superior out-of-distribution generalization, and remains robust across various embedding models, clustering algorithms, ensemble strategies, data efficiency, and values of its sole parameter---the number of clusters. Hao Li 0069, Linyao Chen, Qiaosheng Zhang 0002, Peng Ye 0006, Shi Feng 0001, Xinrun Wang, Xu Jia 0012, Lei Bai 0001, Shuyue Hu |
AAAI | 6 |
| 2026 | A Scalable Multi-LLM Collaboration System with Retrieval-based Selection and Exploration-Exploitation-Driven EnhancementabstractShengji Tang, Jianjian Cao, Weihao Lin, Jiale Hong, Bo Zhang, Shuyue Hu, Lei Bai, Tao Chen, Wanli Ouyang, Peng Ye. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shengji Tang, Jianjian Cao, Weihao Lin 0002, Jiale Hong, Bo Zhang 0069, Shuyue Hu, Lei Bai 0001, Tao Chen 0003, Wanli Ouyang, Peng Ye 0006 |
ACL (1) | 10 |
| 2026 | Nature-Inspired Population-Based Evolution of Large Language ModelsabstractYiqun Zhang, Peng Ye, Xiaocui Yang, Shi Feng, Shufei Zhang, Lei Bai, Wanli Ouyang, Shuyue Hu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Peng Ye 0006, Xiaocui Yang, Shi Feng 0001, Shufei Zhang, Lei Bai 0001, Wanli Ouyang, Shuyue Hu |
ACL (1) | 2 |
| 2026 | Δ-DiT: Accelerating Diffusion Transformers without Training via Denoising Property Alignment
Pengtao Chen, Mingzhu Shen, Peng Ye 0006, Jianjian Cao, Chongjun Tu, Christos-Savvas Bouganis, Tao Chen 0003 |
Int. J. Comput. Vis. | 3 |
| 2026 | Attention Reallocation: Towards Zero-cost and Controllable Hallucination Mitigation of MLLMs
Chongjun Tu, Peng Ye 0006, Dongzhan Zhou, Lei Bai 0001, Gang Yu 0002, Tao Chen 0003, Wanli Ouyang |
Int. J. Comput. Vis. | 2 |
| 2026 | Adapter-X: A general parameter-efficient fine-tuning framework for 2D and 3D vision
Peng Ye 0006, Lin Zhang 0055, Bizhe Bai, Tao Chen 0003 |
Neurocomputing | 2 |
| 2026 | A knowledge-driven self-supervised learning method for enhancing EEG-based emotion recognition
Hanqi Wang, Peng Ye 0006, Kun Yang 0010, Jichuan Xiong, Tao Chen 0003 |
Neural Networks | 3 |
| 2026 | MADTP++: Bridge the Gap Between Token and Weight Pruning for Accelerating VLTsabstractVision-Language Transformers (VLTs) have achieved remarkable success, yet their high computational costs remain challenging due to numerous input tokens and large model parameters. Existing VLT compression methods primarily rely on single-modality-based token pruning or coarse-grained weight pruning techniques. However, these methods face significant obstacles, such as ignoring the critical alignment of different modalities and lacking layer-wise dynamic token pruning flexibility, exhibiting inevitable performance degradation due to coarsegrained weight pruning, and struggling with the simultaneous compression of both input tokens and model parameters. To address those limitations, we propose MADTP++, a novel approach that integrates custom-made token and weight pruning processes into a unified framework, achieving superior compression in both parameter counts and computational costs. Specifically, for the token pruning process, we introduce the Multi-modality Alignment Guidance (MAG) module and the Dynamic Token Pruning (DTP) module to align semantic features across different modalities and guide the dynamic elimination of redundant tokens based on different input instances. For the weight pruning process, we propose a Hardware-aware Weight Pruning (HWP) module that leverages the Sparse Tensor Cores across diverse hardware setups to enable fine-grained parameter pruning within VLTs. To further unify token and weight pruning, we also propose a Cooperative Optimization Training Strategy that automatically allocates GFLOPs and parameter reductions per branch before pruning and employs Knowledge Distillation Constraints to facilitate joint optimization of both pruning dimensions. Extensive experiments conducted on various VLT models and datasets demonstrate that MADTP++ can significantly reduce model parameters and computational costs while maintaining competitive performance. Jianjian Cao, Chong Yu 0001, Peng Ye 0006, Tao Chen 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | StructChart: On the Schema, Metric, and Augmentation for Visual Chart UnderstandingabstractCharts are common in literature across various scientific fields, conveying rich information easily accessible to readers. Current chart-related tasks focus on either chart perception that extracts information from the visual charts, or chart reasoning given the extracted data, e.g. in a tabular form. In this paper, we introduce StructChart, a novel framework that leverages Structured Triplet Representations (STR) to achieve a unified and label-efficient approach to chart perception and reasoning tasks, which is generally applicable to different downstream tasks, beyond the question-answering task as specifically studied in peer works. Specifically, StructChart first reformulates the chart data from the tubular form (linearized CSV) to STR, which can friendlily reduce the task gap between chart perception and reasoning. We then propose a Structuring Chart-oriented Representation Metric (SCRM) to quantitatively evaluate the chart perception task performance. To augment the training, we further explore the potential of Large Language Models (LLMs) to enhance the diversity in both chart visual style and statistical information. Extensive experiments on various chart-related tasks demonstrate the effectiveness and potential of a unified chart perception-reasoning paradigm to push the frontier of chart understanding. Renqiu Xia, Haoyang Peng, Hancheng Ye, Mingsheng Li, Xiangchao Yan, Peng Ye 0006, Botian Shi, Yu Qiao 0001, Junchi Yan, Bo Zhang 0069 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | $\beta $-DARTS++: Bi-Level Regularization for Proxy-Robust Differentiable Architecture SearchabstractNeural Architecture Search (NAS) has attracted increasing attention in recent years because of its capability to design neural networks automatically. Among them, differential NAS approaches such as DARTS, have gained popularity for search efficiency. However, they still suffer from three main issues, that are, the weak stability due to the performance collapse, the poor generalization ability of the searched architectures, and the inferior robustness to different kinds of proxies (i.e., computationally reduced search configurations). To solve the search stability and searched architecture's generalization problems, a simple-but-effective regularization method, termed as Beta-Decay, is proposed to regularize the DARTS-based NAS searching process (referred as $\beta$β-DARTS). Specifically, Beta-Decay regularization can impose constraints to keep the value and variance of activated architecture parameters from being too large, thereby ensuring fair competition among architecture parameters and making the supernet less sensitive to the impact of input on the operation set. In-depth theoretical analyses on how it works and why it works are provided, and comprehensive experiments on a variety of search spaces and datasets validate that Beta-Decay regularization can help to stabilize the searching process and make the searched network more transferable across different datasets. To address the proxy robustness problem, we first benchmark differentiable NAS methods under a wide range of proxy data, proxy channels, proxy layers, and proxy epochs, since the robustness of NAS under different kinds of proxies has not been explored before. We then conclude some interesting findings and find that $\beta$β-DARTS always achieves the best result among all compared NAS methods under almost all proxy settings. We further introduce the novel flooding regularization to the weight optimization of $\beta$β-DARTS (termed as Bi-level regularization), and experimentally and theoretically verify its effectiveness for improving the proxy robustness of differentiable NAS. Peng Ye 0006, Tong He 0001, Baopu Li, Tao Chen 0003, Lei Bai 0001, Wanli Ouyang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Temporally Coherent Dynamic Surface Reconstruction With Planar Gaussian SplattingabstractDynamic surface reconstruction requires not only photorealistic rendering but also temporally stable geometry. In this paper, we present DynaSurfGS, a dynamic surface reconstruction framework built on planar Gaussian splatting with a new temporal path consistency constraint. The key idea is to treat temporal coherence as a path-invariant property of deformation. For a given canonical primitive, its future state should be consistent, regardless of whether it is obtained by direct prediction or temporal propagation through an intermediate state. This trajectory-level formulation regularizes deformation in motion space rather than only supervising per-frame geometry, thereby suppressing temporal instability at its source. Built upon planar dynamic Gaussians and complemented by local geometric supervision, our framework produces both high-quality rendering and geometrically stable dynamic surfaces. Experiments demonstrate that DynaSurfGS achieves strong rendering performance while producing more coherent dynamic surfaces over time. Weicai Ye, Peng Ye 0006, Tong He 0001, Tao Chen 0003 |
IEEE Signal Process. Lett. | 3 |
| 2026 | Mamba-Based Progressive-Recovery Framework for Multimodal Low Light Image Enhancement
Mohammad Mahdizadeh, Jianjian Cao, Peng Ye 0006, Tao Chen 0003 |
IEEE Trans. Multim. | 3 |
| 2025 | All-in-One: Transferring Vision Foundation Models into Stereo MatchingabstractAs a fundamental vision task, stereo matching has made remarkable progress. While recent iterative optimization-based methods have achieved promising performance, their feature extraction capabilities still have room for improvement. Inspired by the ability of vision foundation models (VFMs) to extract general representations, in this work, we propose AIO-Stereo which can flexibly select and transfer knowledge from multiple heterogeneous VFMs to a single stereo matching model. To better reconcile features between heterogeneous VFMs and the stereo matching model and fully exploit prior knowledge from VFMs, we proposed a dual-level feature utilization mechanism that aligns heterogeneous features and transfers multi-level knowledge. Based on the mechanism, a dual-level selective knowledge transfer module is designed to selectively transfer knowledge and integrate the advantages of multiple VFMs. Experimental results show that AIO-Stereo achieves start-of-the-art performance on multiple datasets and ranks 1st on the Middlebury dataset and outperforms all the published work on the ETH3D benchmark. Jiakang Yuan, Peng Ye 0006, Tao Chen 0003, Hao Jiang 0013, Meiya Chen |
AAAI | 4 |
| 2025 | DeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts ModelsabstractUpcycled Mixture-of-Experts (MoE) models have shown great potential in various tasks by converting the original Feed-Forward Network (FFN) layers in pre-trained dense models into MoE layers. However, these models still suffer from significant parameter inefficiency due to the introduction of multiple experts. In this work, we propose a novel DeRS (Decompose, Replace, and Synthesis) paradigm to overcome this shortcoming, which is motivated by our observations about the unique redundancy mechanisms of upcycled MoE experts. Specifically, DeRS decomposes the experts into one expert-shared base weight and multiple expert-specific delta weights, and subsequently represents these delta weights in lightweight forms. Our proposed DeRS paradigm can be applied to enhance parameter efficiency in two different scenarios, including: 1) DeRS Compression for inference stage, using sparsification or quantization to compress vanilla upcycled MoE models; and 2) DeRS Up-cycling for training stage, employing lightweight sparse or low-rank matrixes to efficiently upcycle dense models into MoE models. Extensive experiments across three different tasks show that the proposed methods can achieve extreme parameter efficiency while maintaining the performance for both training and compression of upcycled MoE models. Yongqi Huang, Peng Ye 0006, Chenyu Huang 0001, Jianjian Cao, Lin Zhang 0055, Baopu Li, Gang Yu 0002, Tao Chen 0003 |
CVPR | 2 |
| 2025 | Less is More: Efficient Model Merging with Binary Task SwitchabstractAs an effective approach to equip models with multitask capabilities without additional training, model merging has garnered significant attention. However, existing merging methods face challenges of redundant parameter conflicts and the excessive storage burden of fine-tuned parameters. In this work, through controlled experiments, we reveal that for fine-tuned task vectors, only those parameters with magnitudes above a certain threshold contribute positively to the task, exhibiting a pulse-like characteristic. We then attempt leveraging this pulse-like characteristic to binarize the task vectors and reduce storage overhead. Further controlled experiments show that the binarized task vectors incur almost no decrease in fine-tuning and merging performance, and even exhibit stronger performance improvements as the proportion of redundant parameters increases. Based on these insights, we propose Task Switch (T-Switch), which decomposes task vectors into three components: 1) an activation switch instantiated by a binarized mask vector, 2) a polarity switch instantiated by a binarized sign vector, and 3) a scaling knob instantiated by a scalar coefficient. By storing task vectors in a binarized form, T-Switch alleviates parameter conflicts while ensuring efficient task parameter storage. Furthermore, to enable automated switch combination in T-Switch, we further introduce Auto-Switch, which enables training-free switch combination via retrieval from a small query set. Experiments indicate that our methods achieve significant performance improvements over existing baselines, requiring only 1-3% of the storage space of full-precision parameters. Biqing Qi, Zhen Wang 0004, Junqi Gao, Dong Li 0016, Peng Ye 0006, Bowen Zhou 0002 |
CVPR | 6 |
| 2025 | Critic-V: VLM Critics Help Catch VLM Errors in Multimodal ReasoningabstractVision-language models (VLMs) have shown remarkable advancements in multimodal reasoning tasks. However, they still often generate inaccurate or irrelevant responses due to issues like hallucinated image understandings or unrefined reasoning paths. To address these challenges, we introduce Critic-V, a novel framework inspired by the Actor-Critic paradigm to boost the reasoning capability of VLMs. This framework decouples the reasoning process and critic process by integrating two independent components: the Reasoner, which generates reasoning paths based on visual and textual inputs, and the Critic, which provides constructive critique to refine these paths. In this approach, the Reasoner generates reasoning responses according to text prompts, which can evolve iteratively as a policy based on feedback from the Critic. This interaction process was theoretically driven by a reinforcement learning framework where the Critic offers natural language critiques instead of scalar rewards, enabling more nuanced feedback to boost the Reasoner’s capability on complex reasoning tasks. The Critic model is trained using Direct Preference Optimization (DPO), leveraging a preference dataset of critiques ranked by Rule-based Reward (RBR) to enhance its critic capabilities. Evaluation results show that the Critic-V framework significantly outperforms existing methods, including GPT-4V, on 5 out of 8 benchmarks, especially regarding reasoning accuracy and efficiency. Combining a dynamic text-based policy for the Reasoner and constructive feedback from the preference-optimized Critic enables a more reliable and context-sensitive multimodal reasoning process. Our approach provides a promising solution to enhance the reliability of VLMs, improving their performance in real-world reasoning-heavy multimodal applications such as autonomous driving and embodied intelligence. Our data and code are released at https://github.com/kyrieLei/Critic-V. Di Zhang 0026, Jingdi Lei, Junxian Li 0001, Xunzhi Wang, Zonglin Yang 0001, Jiatong Li 0003, Weida Wang, Suorong Yang, Peng Ye 0006, Wanli Ouyang, Dongzhan Zhou |
CVPR | 11 |
| 2025 | Consistency-aware Self-Training for Iterative-based Stereo MatchingabstractIterative-based methods have become mainstream in stereo matching due to their high performance. However, these methods heavily rely on labeled data and face challenges with unlabeled real-world data. To this end, we propose a consistency-aware self-training framework for iterative-based stereo matching for the first time, leveraging real-world unlabeled data in a teacher-student manner. We first observe that regions with larger errors tend to exhibit more pronounced oscillation characteristics during model prediction. Based on this, we introduce a novel consistency-aware soft filtering module to evaluate the reliability of teacher-predicted pseudo-labels, which consists of a multi-resolution prediction consistency filter and an iterative prediction consistency filter to assess the prediction fluctuations of multiple resolutions and iterative optimization respectively. Further, we introduce a consistency-aware soft-weighted loss to adjust the weight of pseudo-labels accordingly, relieving the error accumulation and performance degradation problem due to incorrect pseudo-labels. Extensive experiments demonstrate that our method can improve the performance of various iterative-based stereo matching approaches in various scenarios. In particular, our method can achieve further enhancements over the current SOTA methods on several benchmark datasets. Peng Ye 0006, Jiakang Yuan, Rao Qiang, Yangchenxu Liu, Wu Cailin, Feng Xu 0001, Tao Chen 0003 |
CVPR | 2 |
| 2025 | Beyond GPT-5: Making LLMs Cheaper and Better via Performance-Efficiency Optimized RoutingabstractBalancing performance and efficiency is a central challenge in large language model (LLM) advancement. GPT-5 addresses this with test-time routing, dynamically assigning queries to either an efficient or a high-capacity model during inference. In this work, we present Avengers-Pro, a test-time routing framework that ensembles LLMs of varying capacities and efficiencies, providing a unified solution for all performance-efficiency tradeoffs. The Avengers-Pro embeds and clusters incoming queries, then routes each to the most suitable model based on a performance-efficiency score. Across 6 challenging benchmarks and 8 leading models—including GPT-5-medium, Gemini-2.5-pro, and Claude-opus-4.1—Avengers-Pro achieves state-of-the-art results: by varying a performance-efficiency trade-off parameter, it can surpass the strongest single model (GPT-5-medium) by +7% in average accuracy. Moreover, it can match the average accuracy of the strongest single model at 27% lower cost, and reach ∼ 90% of that performance at 63% lower cost. Last but not least, it achieves a Pareto frontier, consistently yielding the highest accuracy for any given cost, and the lowest cost for any given accuracy, among all single models. Code is available at https://github.com/ZhangYiqun018/AvengersPro. Hao Li 0069, Jianhao Chen 0001, Hangfan Zhang, Peng Ye 0006, Lei Bai 0001, Shuyue Hu |
DAI | 5 |
| 2025 | Revisiting Convolution Architecture in the Realm of DNA Foundation ModelsabstractIn recent years, A variety of methods based on Transformer and state space model (SSM) architectures have been proposed, advancing foundational DNA language models.
However, there is a lack of comparison between these recent approaches and the classical architecture—convolutional networks (CNNs)—on foundation model benchmarks.
This raises the question: are CNNs truly being surpassed by these recent approaches based on transformer and SSM architectures? In this paper, we develop a simple but well-designed CNN-based method, termed ConvNova. ConvNova identifies and proposes three effective designs: 1) dilated convolutions, 2) gated convolutions, and 3) a dual-branch framework for gating mechanisms.
Through extensive empirical experiments, we demonstrate that ConvNova significantly outperforms recent methods on more than half of the tasks across several foundation model benchmarks. For example, in histone-related tasks, ConvNova exceeds the second-best method by an average of 5.8\%, while generally utilizing fewer parameters and enabling faster computation. In addition, the experiments observed findings that may be related to biological characteristics. This indicates that CNNs are still a strong competitor compared to Transformers and SSMs. We anticipate that this work will spark renewed interest in CNN-based methods for DNA foundation models. Yu Bo, Weian Mao, Yanjun Shao, Weiqiang Bai, Peng Ye 0006, Xinzhu Ma, Hao Chen 0041, Chunhua Shen |
ICLR | 5 |
| 2025 | HiSplat: Hierarchical 3D Gaussian Splatting for Generalizable Sparse-View ReconstructionabstractReconstructing 3D scenes from multiple viewpoints is a fundamental task in stereo vision. Recently, advances in generalizable 3D Gaussian Splatting have enabled high-quality novel view synthesis for unseen scenes from sparse input views by feed-forward predicting per-pixel Gaussian parameters without extra optimization. However, existing methods typically generate single-scale 3D Gaussians, which lack representation of both large-scale structure and texture details, resulting in mislocation and artefacts. In this paper, we propose a novel framework, HiSplat, which introduces a hierarchical manner in generalizable 3D Gaussian Splatting to construct hierarchical 3D Gaussians via a coarse-to-fine strategy. Specifically, HiSplat generates large coarse-grained Gaussians to capture large-scale structures, followed by fine-grained Gaussians to enhance delicate texture details. To promote inter-scale interactions, we propose an Error Aware Module for Gaussian compensation and a Modulating Fusion Module for Gaussian repair. Our method achieves joint optimization of hierarchical representations, allowing for novel view synthesis using only two-view reference images. Comprehensive experiments on various datasets demonstrate that HiSplat significantly enhances reconstruction quality and cross-dataset generalization compared to prior single-scale methods. The corresponding ablation study and analysis of different-scale 3D Gaussians reveal the mechanism behind the effectiveness. Code is at https://github.com/Open3DVLab/HiSplat. Shengji Tang, Weicai Ye, Peng Ye 0006, Weihao Lin 0002, Tao Chen 0003, Wanli Ouyang |
ICLR | 3 |
| 2025 | A CLIP-Powered Framework for Robust and Generalizable Data SelectionabstractLarge-scale datasets have been pivotal to the advancements of deep learning models in recent years, but training on such large datasets inevitably incurs substantial storage and computational overhead.
Meanwhile, real-world datasets often contain redundant and noisy data, imposing a negative impact on training efficiency and model performance.
Data selection has shown promise in identifying the most representative samples from the entire dataset, which aims to minimize the performance gap with reduced training costs.
Existing works typically rely on single-modality information to assign importance scores for individual samples, which may lead to inaccurate assessments, especially when dealing with noisy or corrupted samples.
To address this limitation, we propose a novel CLIP-powered data selection framework that leverages multimodal information for more robust and generalizable sample selection.
Specifically, our framework consists of three key modules—dataset adaptation, sample scoring, and selection optimization—that together harness extensive pre-trained multimodal knowledge to comprehensively assess sample influence and optimize the selection results through multi-objective optimization.
Extensive experiments demonstrate that our approach consistently outperforms existing state-of-the-art baselines on various benchmark datasets. Notably, our method effectively removes noisy or damaged samples from the dataset, enabling it to achieve even higher performance with less data. This indicates that it is not only a way to accelerate training but can also improve overall data quality.
The implementation is available at https://github.com/Jackbrocp/clip-powered-data-selection. Suorong Yang, Peng Ye 0006, Wanli Ouyang, Dongzhan Zhou, Furao Shen |
ICLR | 2 |
| 2025 | When Dynamic Data Selection Meets Data Augmentation: Achieving Enhanced Training AccelerationabstractDynamic data selection aims to accelerate training with lossless performances. However, reducing training data inherently limits data diversity, potentially hindering generalization. While data augmentation is widely used to enhance diversity, it is typically not optimized in conjunction with selection. As a result, directly combining these techniques fails to fully exploit their synergies. To tackle the challenge, we propose a novel online data training framework that, for the first time, unifies dynamic data selection and augmentation, achieving both training efficiency and enhanced performance. Our method estimates each sample’s joint distribution of local density and multimodal semantic consistency, allowing for the targeted selection of augmentation-suitable samples while suppressing the inclusion of noisy or ambiguous data. This enables a more significant reduction in dataset size without sacrificing model generalization. Experimental results demonstrate that our method outperforms existing state-of-the-art approaches on various benchmark datasets and architectures, e.g., reducing 50% training costs on ImageNet-1k with lossless performance. Furthermore, our approach enhances noise resistance and improves model robustness, reinforcing its practical utility in real-world scenarios. Suorong Yang, Peng Ye 0006, Furao Shen, Dongzhan Zhou |
ICML | 2 |
| 2025 | MLLM-ISU: The First-Ever Comprehensive Benchmark for Multimodal Large Language Models based Intrusion Scene UnderstandingabstractVision-based intrusion detection has multiple applications in practical scenarios, e.g., autonomous driving, intelligent monitoring, and security. Previous works mainly focus on improving the intrusion detection performance, without a comprehensive and in-depth understanding of the intrusion scene. To fill this gap, we explore a novel task called Multimodal Large Language Models based Intrusion Scene Understanding (MLLM-ISU) and report a comprehensive benchmark for the task. Specifically, we first design an effective and automatic visual question-answer generation strategy, constructing a new MLLM-ISU dataset, with 3000 VQA evaluation Pairs, 8925 training Pairs, and six relevant subtasks. Then, we perform a comprehensive assessment on various state-of-the-art proprietary and open-source MLLMs, e.g., DeepSeek-VL2, GPT-4o, Qwen2.5-VL, etc, and find that current MLLMs have weak abilities for this task. Further, in order to improve the intrusion understanding capabilities of current MLLMs, we propose a Post-Training Framework with three sequential training stages, i.e., Intrusion-aware Visual Instruction Pre-training, Intrusion Chain of Thought tuning, and Intrusion-centric VQA tuning, and sufficient experiments and comparisons are conducted to verify the effectiveness of the proposed three-stages training framework. Available datasets and codes: https://github.com/1012537710/MLLM-ISU. Fujun Han, Peng Ye 0006 |
NeurIPS | 2 |
| 2025 | GoRA: Gradient-driven Adaptive Low Rank AdaptationabstractLow-Rank Adaptation (LoRA) is a crucial method for efficiently fine-tuning large language models (LLMs), with its effectiveness influenced by two key factors: rank selection and weight initialization. While numerous LoRA variants have been proposed to improve performance by addressing one of these aspects, they often compromise usability or computational efficiency. In this paper, we analyze and identify the core limitations of existing approaches and propose a novel framework—**GoRA** (**G**radient-driven Adaptive L**o**w **R**ank **A**daptation)—that simultaneously adapts both the rank and initialization strategy within a unified framework. GoRA leverages gradient information during training to dynamically assign optimal ranks and initialize low-rank adapter weights in an adaptive manner. To our knowledge, GoRA is the first method that not only addresses the limitations of prior approaches—which often focus on either rank selection or initialization in isolation—but also unifies both aspects within a single framework, enabling more effective and efficient adaptation. Extensive experiments across various architectures and modalities show that GoRA consistently outperforms existing LoRA-based methods while preserving the efficiency of vanilla LoRA. For example, when fine-tuning Llama3.1-8B-Base for mathematical reasoning, GoRA achieves a 5.13-point improvement over standard LoRA and even outperforms full fine-tuning by 2.05 points under high-rank settings. Code is available at: https://github.com/hhnqqq/MyTransformers. Peng Ye 0006, Yuchen Ren 0001, Luyang Zhou, Shucun Ju |
NeurIPS | 2 |
| 2025 | PaceLLM: Brain-Inspired Large Language Models for Long-Context UnderstandingabstractWhile Large Language Models (LLMs) demonstrate strong performance across domains, their long-context capabilities are limited by transient neural activations causing information decay and unstructured feed-forward network (FFN) weights leading to semantic fragmentation. Inspired by the brain’s working memory and cortical modularity, we propose PaceLLM, featuring two innovations: (1) a Persistent Activity (PA) Mechanism that mimics prefrontal cortex (PFC) neurons’ persistent firing by introducing an activation-level memory bank to dynamically retrieve, reuse, and update critical FFN states, addressing contextual decay; and (2) Cortical Expert (CE) Clustering that emulates task-adaptive neural specialization to reorganize FFN weights into semantic modules, establishing cross-token dependencies and mitigating fragmentation. Extensive evaluations show that PaceLLM achieves 6% improvement on LongBench’s Multi-document QA and 12.5–17.5% performance gains on $\infty$-Bench tasks, while extending measurable context length to 200K tokens in Needle-In-A-Haystack (NIAH) tests. This work pioneers brain-inspired LLM optimization and is complementary to other works. Besides, it can be generalized to any model and enhance their long-context performance and interpretability without structural overhauls. Kangcong Li, Peng Ye 0006, Chongjun Tu, Lin Zhang 0055, Chunfeng Song, Qihao Zheng, Tao Chen 0003 |
NeurIPS | 2 |
| 2025 | scMRDR: A scalable and flexible framework for unpaired single-cell multi-omics data integrationabstractAdvances in single-cell sequencing have enabled high-resolution profiling of diverse molecular modalities, while integrating unpaired multi-omics single-cell data remains challenging. Existing approaches either rely on pair information or prior correspondences, or require computing a global pairwise coupling matrix, limiting their scalability and flexibility. In this paper, we introduce a scalable and flexible generative framework called single-cell Multi-omics Regularized Disentangled Representations (scMRDR) for unpaired multi-omics integration. Specifically, we disentangle each cell’s latent representations into modality-shared and modality-specific components using a well-designed $\beta$-VAE architecture, which are augmented with isometric regularization to preserve intra-omics biological heterogeneity, adversarial objective to encourage cross-modal alignment, and masked reconstruction loss strategy to address the issue of missing features across modalities. Our method achieves excellent performance on benchmark datasets in terms of batch correction, modality alignment, and biological signal preservation. Crucially, it scales effectively to large-scale datasets and supports integration of more than two omics, offering a powerful and flexible solution for large-scale multi-omics data integration and downstream biological discovery. Jianle Sun, Chaoqi Liang, Peng Zheng 0004, Lei Bai 0001, Wanli Ouyang, Hongliang Yan, Peng Ye 0006 |
NeurIPS | 8 |
| 2025 | FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion UnderstandingabstractMultimodal Large Language Models (MLLMs) have shown impressive video content understanding capabilities but struggle with fine-grained motion comprehension. To comprehensively assess the motion understanding ability of existing MLLMs, we introduce FAVOR-Bench, which comprises 1,776 videos from both ego-centric and third-person perspectives and enables assessment through both close-ended and open-ended tasks. For close-ended evaluation, we carefully design 8,184 multiple-choice question-answer pairs spanning six distinct sub-tasks. For open-ended evaluation, we employ the GPT-assisted evaluation and develop a novel cost-efficient LLM-free assessment method, where the latter can enhance benchmarking interpretability and accessibility. Comprehensive experiments with21 state-of-the-art MLLMs reveal significant limitations in their ability to comprehend and describe detailed temporal dynamics in video motions. To alleviate this limitation, we further build FAVOR-Train, a dataset of 17,152 videos with fine-grained motion annotations. Finetuning Qwen2.5-VL on FAVOR-Train yields consistent improvements on motion-related tasks across TVBench, MotionBenchand our FAVOR-Bench. Our assessment results demonstrate that the proposed FAVOR-Bench and FAVOR-Train provide valuable tools for the community to develop more powerful video understanding models. Chongjun Tu, Lin Zhang 0055, Pengtao Chen, Peng Ye 0006, Xianfang Zeng, Gang Yu 0002, Tao Chen 0003 |
NeurIPS | 4 |
| 2025 | Breaking the Compression Ceiling: Data-Free Pipeline for Ultra-Efficient Delta CompressionabstractWith the rise of the fine-tuned–pretrained paradigm, storing numerous fine-tuned models for multi-tasking creates significant storage overhead.
Delta compression alleviates this by storing only the pretrained model and the highly compressed delta weights (the differences between fine-tuned and pretrained model weights).
However, existing methods fail to maintain both high compression and performance, and often rely on data.
To address these challenges, we propose UltraDelta, the first data-free delta compression pipeline that achieves both ultra-high compression and strong performance.
UltraDelta is designed to minimize redundancy, maximize information, and stabilize performance across inter-layer, intra-layer, and global dimensions, using three key components:
(1) Variance-Based Mixed Sparsity Allocation assigns sparsity based on variance, giving lower sparsity to high-variance layers to preserve inter-layer information.
(2) Distribution-Aware Compression applies uniform quantization and then groups parameters by value, followed by group-wise pruning, to better preserve intra-layer distribution.
(3) Trace-Norm-Guided Rescaling uses the trace norm of delta weights to estimate a global rescaling factor, improving model stability under higher compression.
Extensive experiments across
(a) large language models (fine-tuned on LLaMA-2 7B and 13B) with up to 50$\times$ compression,
(b) general NLP models (RoBERTa-base, T5-base) with up to 224$\times$ compression,
(c) vision models (ViT-B/32, ViT-L/14) with up to 132$\times$ compression, and
(d) multi-modal models (BEiT-3) with 18$\times$ compression,
demonstrate that UltraDelta consistently outperforms existing methods, especially under ultra-high compression.
Code is available at https://github.com/xiaohuiwang000/UltraDelta. Peng Ye 0006, Chenyu Huang 0001, Shenghe Zheng, Bo Zhang 0069, Lei Bai 0001, Wanli Ouyang, Tao Chen 0003 |
NeurIPS | 2 |
| 2025 | Scaling Physical Reasoning with the PHYSICS DatasetabstractLarge Language Models (LLMs) have achieved remarkable progress on advanced reasoning tasks such as mathematics and coding competitions. Meanwhile, physics, despite being both reasoning-intensive and essential to real-world understanding, received limited academic and industrial attention. This paper introduces PHYSICS, a dataset containing 16,568 high-quality physics problems spanning subjects and difficulty levels, to facilitate this issue. Specifically, PHYSICS is curated with exercises from over 100 textbooks through a carefully designed pipeline for quality control. It covers five major physics domains: Mechanics, Electromagnetism, Thermodynamics, Optics, and Modern Physics. It also spans a wide range of difficulty levels, from high school to graduate-level physics courses. To utilize the data for improving and evaluating the model's physical reasoning capabilities, we split the dataset into training and test sets, and provide reasoning paths generated by powerful reasoning models for the training data to facilitate model training. In addition, for the evaluation part, we find that existing evaluation frameworks exhibit biases in aspects such as units, simplification, and precision in physics domain. To balance efficiency and accuracy, we introduce a Rule+Model evaluation framework tailored to physics problems. Our evaluations on current state-of-the-art open-source and proprietary models highlight the limitations of current models in handling physics-related tasks. We hope that our dataset and evaluation methodology will jointly advance the development of LLMs in the field of physics. The code and data can be found at: https://github.com/Zhengsh123/PHYSICS. Shenghe Zheng, Qianjia Cheng, Junchi Yao, Mengsong Wu, Ning Ding 0002, Yu Cheng 0001, Shuyue Hu, Lei Bai 0001, Dongzhan Zhou, Ganqu Cui, Peng Ye 0006 |
NeurIPS | 12 |
| 2025 | ClipSAM: CLIP and SAM collaboration for zero-shot anomaly segmentation
Shengze Li, Jianjian Cao, Peng Ye 0006, Yuhan Ding, Chongjun Tu, Tao Chen 0003 |
Neurocomputing | 3 |
| 2025 | Enhancing visibility in hazy conditions: A multimodal multispectral image dehazing approach
Mohammad Mahdizadeh, Peng Ye 0006, Shaoqing Zhao |
J. Vis. Commun. Image Represent. | 2 |
| 2025 | Stimulative Training++: Go Beyond the Performance Limits of Residual NetworksabstractResidual networks have shown great success and become indispensable in recent deep neural network models. In this work, we aim to re-investigate the training process of residual networks from a novel perspective of loafing, and further propose a new training scheme as well as three improved strategies for boosting residual networks beyond their performance limits. Previous research has suggested that residual networks can be considered as ensembles of shallow networks, which implies that the final performance of a residual network is influenced by a group of subnetworks. Furthermore, we identify a previously overlooked problem, where subnetworks within a residual network are prone to exert less effort when working as part of a group compared to working alone. We define this problem as network loafing. Since network loafing may inevitably cause the sub-par performance of the residual network, we propose a novel training scheme called stimulative training, which randomly samples a residual subnetwork and calculates the KL divergence loss between the sampled subnetwork and the given residual network for extra supervision. In order to unleash the potential of stimulative training, we further propose three simple-yet-effective strategies, including a novel KL- loss that only aligns the network logits direction, random smaller inputs for subnetworks, and inter-stage sampling rules. Comprehensive experiments and analysis verify the effectiveness of stimulative training as well as its three improved strategies. For example, the proposed method can boost the performance of ResNet50 on ImageNet to 80.5% Top1 accuracy without using any extra data, model, trick, or changing the structure. With only uniform augment, the performance can be further improved to 81.0% Top1 accuracy, better than the best training recipes provided by Timm library and PyTorch official version. We also verify its superiority on various typical models, datasets, and tasks and give some theoretical analysis. As such, we advocate utilizing the proposed method as a general and next-generation technology to train residual networks. Peng Ye 0006, Tong He 0001, Shengji Tang, Baopu Li, Tao Chen 0003, Lei Bai 0001, Wanli Ouyang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | BridgeNet: Comprehensive and Effective Feature Interactions via Bridge Feature for Multi-Task Dense PredictionsabstractMulti-task dense prediction aims at handling multiple pixel-wise prediction tasks within a unified network simultaneously for visual scene understanding. However, cross-task feature interactions of current methods are still suffering from incomplete levels of representations, less discriminative semantics in feature participants, and inefficient pair-wise task interaction processes. To tackle these under-explored issues, we propose a novel BridgeNet framework, which extracts comprehensive and discriminative intermediate Bridge Features, and conducts interactions based on them. Specifically, a Task Pattern Propagation (TPP) module is first applied to ensure highly semantic task-specific feature participants are prepared for subsequent interactions, and a Bridge Feature Extractor (BFE) is specially designed to selectively integrate both high-level and low-level representations to generate the comprehensive bridge features. Then, instead of conducting heavy pair-wise cross-task interactions, a Task-Feature Refiner (TFR) is developed to efficiently take guidance from bridge features and form final task predictions. To the best of our knowledge, this is the first work considering the completeness and quality of feature participants in cross-task interactions. Extensive experiments are conducted on NYUD-v2, Cityscapes and PASCAL Context benchmarks, and the superior performance shows the proposed architecture is effective and powerful in promoting different dense prediction tasks simultaneously. Jingdong Zhang 0003, Jiayuan Fan 0001, Peng Ye 0006, Bo Zhang 0069, Hancheng Ye, Baopu Li, Yancheng Cai, Tao Chen 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | MF-ID: A Benchmark and Approach for Multi-Category Fine-Grained Intrusion DetectionabstractWith the development of computer vision, the task of vision-based intrusion detection has been widely applied to various important fields such as intelligent monitoring, autonomous driving, and security. Previous vision-based intrusion detection tasks aim at detecting whetherpedestriansinvade a restricted Area-of-Interest (AoI) from a static or dynamic view. However, for real application scenarios, we need to detect whether various types of intrusion objects, not just pedestrians, intrude into the AoI of dynamic view, and the fine-grained categories of intrusion objects need to be accurately given, such a dynamic-viewmulti-category fine-grainedintrusion detection (namely MF-ID) task is important but not yet explored. In this paper, we propose a new benchmark and approach to address this task. Firstly, due to the current lack of relevant benchmark, we develop a new publicly available dataset Cityintrusion-Multicategory, conduct statistical analysis on this dataset, and design three evaluation metrics. Secondly, we propose an end-to-end framework MF-YOLOV5, with five improvements: (1) We modify YOLOV5 to make it more suitable for our task, with a lower branch for object detection and an upper branch for segmenting AoI. (2) A new multi-category fine-grained loss (MFLoss) is designed to improve the fine-grained classification capability. (3) We improve the YOLOV5 C3 modules by enhancing the ability of cross-channel interaction. (4) A detection layer for the tiny objects is integrated into the network to improve its ability of detecting tiny objects. (5) A lightweight module with bottleneck transformer is introduced to reduce the network parameters. Finally, comprehensive experiments and comparisons demonstrate the validity of the proposed approach, and MF-YOLOV5 can reach the level of current SOTA, with 97.46% Miou, 55.29% [email protected] and 42.97% MF-IDAcc. The relevant datasets and codes are available at https://ieee-dataport.org/documents/mf-id-1.Note to Practitioners—The motivation of this paper is to address the pitfall of the dynamic-view intrusion detection system. The existing methods mainly focus on pedestrian intrusion detection in dynamic view, ignoring the more practical and valuable task of multi-category fine-grained intrusion detection (MF-ID). Based on this, we propose a new benchmark and an advanced approach to meet the requirement of MF-ID task in dynamic view. Extensive experiments show that the proposed approach can not only reach promising performance but also maintain a high real-time intrusion detection speed. The proposed approach can be deployed in realistic scenes, e.g., autonomous driving, intelligent monitoring, security, and intelligent transportation management. In future research, we will explore a more comprehensive benchmark and more efficient approach for intrusion detection. Fujun Han, Peng Ye 0006, Shukai Duan 0001, Lidan Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Sparse-to-Dense Training: A Novel Training Scheme to Enhance Vision TransformersabstractAs Vision Transformers (ViTs) become increasingly popular in various vision tasks, one may question:if a new training scheme for ViTs exists that can improve performance without increasing training and inference computation cost?In this paper, we affirmatively answer this question with a novel Sparse-to-Dense (S2D) training scheme. Specifically, we decouple the training and inference phases of ViTs. During training, we replace some Feed-Forward Network (FFN) layers of ViTs with computationally efficient RUP-Mixture-of-FFN (RUP-MoF) layers, each comprising multiple FFN experts and allocating tokens to experts via Random Uniform Partition (RUP). Furthermore, an additional Experts Weights Averaging (EWA) update is performed specifically on these RUP-MoF layers after each gradient update. After training, we convert each RUP-MoF layer back to a single FFN layer by averaging the experts, transforming the training-time sparse model back to the original dense ViT model for inference. We further provide theoretical analysis to illustrate why and how it works. Comprehensive experiments across various 2D and 3D vision tasks, ViT architectures and datasets validate the effectiveness and generalization ability of the proposed S2D training scheme. Besides, we show that, S2D training scheme can also be applied to improve the performance of Transformer-based language models, and EWA update technique can also significantly improve the effectiveness of classic Mixture-of-Experts on various 2D vision small-scale datasets and 3D vision tasks. Yongqi Huang, Peng Ye 0006, Chongjun Tu, Tao Chen 0003, Tong He 0001, Wanli Ouyang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Efficient Architecture Search via Bi-Level Data PruningabstractImproving the efficiency of Neural Architecture Search (NAS) is a challenging but significant task that has received much attention. Previous studies mainly adopt the Differentiable Architecture Search (DARTS) and improve its search strategies or modules to enhance search efficiency. Recently, some methods have started considering data reduction for speedup, but they are not tightly coupled with the architecture search process and cannot capture the training dynamics of DARTS well, resulting in sub-optimal performances. To this end, this work pioneers an exploration into the critical role of dataset characteristics in the bi-level optimization of DARTS, and then proposes a novel Bi-level Data Pruning (BDP) paradigm that targets the weights and architecture levels of DARTS to enhance efficiency from a data perspective. Specifically, we introduce a progressive bi-level data pruning strategy that utilizes supernet prediction dynamics as the metric to gradually prune unsuitable samples for DARTS during the search. An effective automatic class balance constraint is also integrated into BDP, to suppress potential class imbalances resulting from data-efficient algorithms. Comprehensive evaluations on the NAS-Bench-201 search space, DARTS search space, and MobileNet-like search space validate that BDP reduces search costs by over 50% while achieving superior performance when applied to the baseline DARTS. Besides, we demonstrate that BDP can harmoniously integrate with advanced DARTS variants, like P-DARTS, PC-DARTS, EG-NAS, and$\beta $-DARTS, offering an approximately$2\times $speedup with minimal performance compromise. Chongjun Tu, Peng Ye 0006, Weihao Lin 0002, Hancheng Ye, Chong Yu 0001, Tao Chen 0003, Baopu Li, Wanli Ouyang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Dynamic Model Merging With Mixture of WeightsabstractThe pretrain-finetune paradigm brings about the release of numerous model weights. Under this background, model merging is becoming increasingly popular, as it enables a model to handle multiple tasks by fusing model weights from these tasks, without the need for labeled data, additional training, or high training costs. Though with great potential, model merging suffers from severe performance degradation due to the interference among model weights. And existing model merging methods (i.e., static merging) commonly provide a single set of merging coefficients for all the input samples and do not distinguish layers based on the severity of weight interference, which may not be the optimal solution. In this paper, we propose MoW-Merging, a dynamic model merging method based on Mixture of Weights. First, we apply a gating network to adaptively generate merging coefficients depending on the input samples, realizing sample-wisely dynamic merging and automated classifier selection. The gating network is lightweight and is trained with only a small number of unlabeled data. Further, we utilize a weight similarity metric to judge the severity of weight interference of each layer and apply suitable merging methods to different layers. The proposed MoW-Merging shows plug-and-play capabilities and can be seamlessly combined with various model merging methods to greatly boost their performance. The effectiveness of MoW-Merging is validated by comprehensive experiments on various classical and newly-established benchmarks under multiple settings. The code is available athttps://github.com/harveyhuang18/Mixture_of_Weights. Peng Ye 0006, Chenyu Huang 0001, Mingzhu Shen, Tao Chen 0003, Yongqi Huang, Wanli Ouyang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | DualMamba: A Lightweight Spectral-Spatial Mamba-Convolution Network for Hyperspectral Image ClassificationabstractThe effectiveness and efficiency of modeling complex spectral–spatial relations are crucial for hyperspectral image (HSI) classification. Most existing methods based on convolution neural networks (CNNs) and transformers still suffer from heavy computational burdens and have room for improvement in capturing the global–local spectral–spatial feature representation. To this end, we propose a novel lightweight parallel design called a lightweight dual-stream Mamba-convolution network (DualMamba) for HSI classification. Specifically, a parallel lightweight Mamba and CNN block are developed to extract global and local spectral–spatial features. First, the cross-attention spectral–spatial Mamba module (CAS2MM) is proposed to leverage the global modeling of Mamba at linear complexity. In this module, dynamic positional embedding (DPE) is designed to enhance the spatial location information of visual sequences. The lightweight spectral–spatial Mamba blocks comprise an efficient scanning strategy and a lightweight Mamba design to efficiently extract global spectral–spatial features. And the cross-attention spectral–spatial fusion (CAS2F) is designed to learn cross correlation and fuse spectral–spatial features. Second, the lightweight spectral–spatial residual convolution module is proposed with lightweight spectral and spatial branches to extract local spectral–spatial features through residual learning. Finally, the adaptive global–local fusion is proposed to dynamically combine global Mamba features and local convolution features for a global–local spectral–spatial representation. Compared with state-of-the-art HSI classification methods, experimental results demonstrate that DualMamba achieves significant classification accuracy on three public HSI datasets and a superior reduction in model parameters and floating-point operations (FLOPs). Jiamu Sheng, Peng Ye 0006, Jiayuan Fan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Boosting Residual Networks with Group KnowledgeabstractRecent research understands the residual networks from a new perspective of the implicit ensemble model. From this view, previous methods such as stochastic depth and stimulative training have further improved the performance of the residual network by sampling and training of its subnets. However, they both use the same supervision for all subnets of different capacities and neglect the valuable knowledge generated by subnets during training. In this manuscript, we mitigate the significant knowledge distillation gap caused by using the same kind of supervision and advocate leveraging the subnets to provide diverse knowledge. Based on this motivation, we propose a group knowledge based training framework for boosting the performance of residual networks. Specifically, we implicitly divide all subnets into hierarchical groups by subnet-in-subnet sampling, aggregate the knowledge of different subnets in each group during training, and exploit upper-level group knowledge to supervise lower-level subnet group. Meanwhile, we also develop a subnet sampling strategy that naturally samples larger subnets, which are found to be more helpful than smaller subnets in boosting performance for hierarchical groups. Compared with typical subnet training and other methods, our method achieves the best efficiency and performance trade-offs on multiple datasets and network structures. The code is at https://github.com/tsj-001/AAAI24-GKT. Shengji Tang, Peng Ye 0006, Baopu Li, Weihao Lin 0002, Tao Chen 0003, Tong He 0001, Chong Yu 0001, Wanli Ouyang |
AAAI | 2 |
| 2024 | MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language TransformerabstractVision-Language Transformers (VLTs) have shown great success recently, but are meanwhile accompanied by heavy computation costs, where a major reason can be attributed to the large number of visual and language tokens. Existing token pruning research for compressing VLTs mainly follows a single-modality-based scheme yet ignores the critical role of aligning different modalities for guiding the token pruning process, causing the important tokens for one modality to be falsely pruned in another modality branch. Meanwhile, existing VLT pruning works also lack the flexibility to dynamically compress each layer based on different input samples. To this end, we propose a novel framework named Multimodal Alignment-Guided Dynamic Token Pruning (MADTP) for accelerating various VLTs. Specifically, we first introduce a well-designed Multi-modality Alignment Guidance (MAG) module that can align features of the same semantic concept from different modalities, to ensure the pruned tokens are less important for all modalities. We further design a novel Dynamic Token Pruning (DTP) module, which can adaptively adjust the token compression ratio in each layer based on different input instances. Extensive experiments on various benchmarks demonstrate that MADTP significantly reduces the computational complexity of kinds of multimodal models while preserving competitive performance. Notably, when applied to the BLIP model in the NLVR2 dataset, MADTP can reduce the GFLOPs by 80% with less than 4% performance degradation. The code is available at https://github.com/double125IMADTP. Jianjian Cao, Peng Ye 0006, Shengze Li, Chong Yu 0001, Yansong Tang, Jiwen Lu, Tao Chen 0003 |
CVPR | 2 |
| 2024 | Once for Both: Single Stage of Importance and Sparsity Search for Vision Transformer CompressionabstractRecent Vision Transformer Compression (VTC) works mainly follow a two-stage scheme, where the importance score of each model unit is first evaluated or preset in each submodule, followed by the sparsity score evaluation ac-cording to the target sparsity constraint. Such a separate evaluation process induces the gap between importance and sparsity score distributions, thus causing high search costs for VTC. In this work, for the first time, we investigate how to integrate the evaluations of importance and sparsity scores into a single stage, searching the optimal subnets in an effi-cient manner. Specifically, we present OFB, a cost-efficient approach that simultaneously evaluates both importance and sparsity scores, termed Once for Both (OFB), for VTC. First, a bi-mask scheme is developed by entangling the importance score and the differentiable sparsity score to jointly deter-mine the pruning potential (prunability) of each unit. Such a bi-mask search strategy is further used together with a proposed adaptive one-hot loss to realize the progressive-and-efficient search for the most important subnet. Finally, Progressive Masked Image Modeling (PMIM) is proposed to regularize the feature space to be more representative during the search process, which may be degraded by the dimension reduction. Extensive experiments demonstrate that OFB can achieve superior compression performance over state-of-the-art searching-based and pruning-based methods under various Vision Transformer architectures, meanwhile pro-moting search efficiency significantly, e.g., costing one GPU search day for the compression of DeiT-S on ImageNet-1K. Hancheng Ye, Chong Yu 0001, Peng Ye 0006, Renqiu Xia, Yansong Tang, Jiwen Lu, Tao Chen 0003, Bo Zhang 0069 |
CVPR | 3 |
| 2024 | Enhanced Sparsification via Stimulative Training
Shengji Tang, Weihao Lin 0002, Hancheng Ye, Peng Ye 0006, Chong Yu 0001, Baopu Li, Tao Chen 0003 |
ECCV (51) | 4 |
| 2024 | CasCast: Skillful High-resolution Precipitation Nowcasting via Cascaded ModellingabstractPrecipitation nowcasting based on radar data plays a crucial role in extreme weather prediction and has broad implications for disaster management. Despite progresses have been made based on deep learning, two key challenges of precipitation nowcasting are not well-solved: (i) the modeling of complex precipitation system evolutions with different scales, and (ii) accurate forecasts for extreme precipitation. In this work, we propose CasCast, a cascaded framework composed of a deterministic and a probabilistic part to decouple the predictions for mesoscale precipitation distributions and small-scale patterns. Then, we explore training the cascaded framework at the high resolution and conducting the probabilistic modeling in a low dimensional latent space with a frame-wise-guided diffusion transformer for enhancing the optimization of extreme events while reducing computational costs. Extensive experiments on three benchmark radar precipitation datasets show that CasCast achieves competitive performance. Especially, CasCast significantly surpasses the baseline (up to +91.8%) for regional extreme-precipitation nowcasting. Junchao Gong, Lei Bai 0001, Peng Ye 0006, Wanghan Xu, Xiaokang Yang 0001, Wanli Ouyang |
ICML | 3 |
| 2024 | Multi-dimensional Search with Strip Convolution and R-Squared Loss for Lane DetectionabstractExisting deep learning methods for lane detection mainly rely on commonly used backbones such as ResNet, equipped with various hand-craft modules for feature extraction and modeling. Such manually designed backbones and modules heavily rely on human expert experience, and it is difficult to guarantee their extracted features always fit well with various lanes. Besides backbones, the imbalanced number of lanes with different curvature in datasets and the ways modeling lane lines incur a strong curvature bias. Therefore, inspired by the recent popularity of AutoML techniques such as Neural Architecture Search (NAS), we propose an automatic lane detection architecture design framework, namely StripLaneNet-NAS, to solve the above problem. To enable the searched model structure well capture various kinds of lane features, we propose a multi-dimensional search space equipped with specially designed lane-specific Strip Convolution Modules (SCM), and correspondingly propose an adaptive solution space regularization loss, to accelerate and optimize the multi-dimensional search process. A novel R-squared loss is further proposed in the optimization objective to alleviate the curvature bias problem as mentioned above. Experiments are conducted on two benchmarks, i.e., TuSimple and CULane, and results show that our method outperforms state-of-the-art lane detection baselines, achieving the fastest inference speed with a maximum of 74.0% parameter reduction over the baselines. Peng Ye 0006, Tao Chen 0003, Shengji Tang |
IJCNN | 2 |
| 2024 | Ada-iD: Active Domain Adaptation for Intrusion DetectionabstractVision-based intrusion detection has many applications in life environments, e.g., security, intelligent monitoring, and autonomous driving. Previous works improve the performance of intrusion detection under unknown environments by introducing unsupervised domain adaptation (UDA) methods. However, these works do not fully fulfill the practical requirements due to the performance gap between UDA and fully supervised methods. To address the problem, we develop a new and vital active domain adaptation intrusion detection task, namely Ada-iD. Our aim is to query and annotate the most informative samples of the target domain at the lowest possible cost, striving for a balance between achieving high performance and keeping low annotation expenses. Specifically, we propose a multi-task joint active domain adaptation intrusion detection framework, namely ADAID-YOLO. It consists of a lower branch for detection and an upper branch for segmentation. Further, three effective strategies are designed to better achieve the Ada-iD task: 1) An efficient Dynamic Diffusion Pseudo-Labeling method (DDPL) is introduced to get Pseudo ground truth to help identify areas of uncertainty in segmentation. 2) An Enhanced Region Impurity and Prediction Uncertainty sampling strategy (Enhanced-RIPU) is proposed to better capture the uncertainty of the segmentation region. 3) A Multi-Element Joint sampling strategy (MEJ) is designed to calculate the uncertainty of the detection comprehensively. Finally, comprehensive experiments and comparisons are conducted on multiple dominant intrusion detection datasets. The results show that our method can outperform other classic and promising active domain adaptation methods and reach current SOTA performance, even surpassing the performance of UDA and full supervision on Normal-Foggy with only 0.1% and 10% data annotation, respectively. Available code: https://github.com/1012537710/Ada-iD. Fujun Han, Peng Ye 0006, Shukai Duan 0001, Lidan Wang 0001 |
ACM Multimedia | 2 |
| 2024 | S2HPruner: Soft-to-Hard Distillation Bridges the Discretization Gap in PruningabstractRecently, differentiable mask pruning methods optimize the continuous relaxation architecture (soft network) as the proxy of the pruned discrete network (hard network) for superior sub-architecture search. However, due to the agnostic impact of the discretization process, the hard network struggles with the equivalent representational capacity as the soft network, namely discretization gap, which severely spoils the pruning performance. In this paper, we first investigate the discretization gap and propose a novel structural differentiable mask pruning framework named S2HPruner to bridge the discretization gap in a one-stage manner. In the training procedure, SH2Pruner forwards both the soft network and its corresponding hard network, then distills the hard network under the supervision of the soft network. To optimize the mask and prevent performance degradation, we propose a decoupled bidirectional knowledge distillation. It blocks the weight updating from the hard to the soft network while maintaining the gradient corresponding to the mask. Compared with existing pruning arts, S2HPruner achieves surpassing pruning performance without fine-tuning on comprehensive benchmarks, including CIFAR-100, Tiny ImageNet, and ImageNet with a variety of network architectures. Besides, investigation and analysis experiments explain the effectiveness of S2HPruner. Codes will be released soon. Weihao Lin 0002, Shengji Tang, Chong Yu 0001, Peng Ye 0006, Tao Chen 0003 |
NeurIPS | 4 |
| 2024 | FNP: Fourier Neural Processes for Arbitrary-Resolution Data AssimilationabstractData assimilation is a vital component in modern global medium-range weather forecasting systems to obtain the best estimation of the atmospheric state by combining the short-term forecast and observations. Recently, AI-based data assimilation approaches have attracted increasing attention for their significant advantages over traditional techniques in terms of computational consumption. However, existing AI-based data assimilation methods can only handle observations with a specific resolution, lacking the compatibility and generalization ability to assimilate observations with other resolutions. Considering that complex real-world observations often have different resolutions, we propose the Fourier Neural Processes (FNP) for arbitrary-resolution data assimilation in this paper. Leveraging the efficiency of the designed modules and flexible structure of neural processes, FNP achieves state-of-the-art results in assimilating observations with varying resolutions, and also exhibits increasing advantages over the counterparts as the resolution and the amount of observations increase. Moreover, our FNP trained on a fixed resolution can directly handle the assimilation of observations with out-of-distribution resolutions and the observational information reconstruction task without additional fine-tuning, demonstrating its excellent generalization ability across data resolutions as well as across tasks. Code is available at https://github.com/OpenEarthLab/FNP. Kun Chen 0004, Peng Ye 0006, Hao Chen 0045, Tao Han 0002, Wanli Ouyang, Tao Chen 0003, Lei Bai 0001 |
NeurIPS | 2 |
| 2024 | EMR-Merging: Tuning-Free High-Performance Model MergingabstractThe success of pretrain-finetune paradigm brings about the release of numerous model weights. In this case, merging models finetuned on different tasks to enable a single model with multi-task capabilities is gaining increasing attention for its practicability. Existing model merging methods usually suffer from (1) significant performance degradation or (2) requiring tuning by additional data or training. In this paper, we rethink and analyze the existing model merging paradigm. We discover that using a single model's weights can hardly simulate all the models' performance. To tackle this issue, we propose Elect, Mask & Rescale-Merging (EMR-Merging). We first (a) elect a unified model from all the model weights and then (b) generate extremely lightweight task-specific modulators, including masks and rescalers, to align the direction and magnitude between the unified model and each specific model, respectively. EMR-Merging is tuning-free, thus requiring no data availability or any additional training while showing impressive performance. We find that EMR-Merging shows outstanding performance compared to existing merging methods under different classical and newly-established settings, including merging different numbers of vision models (up to 30), NLP models, PEFT models, and multi-modal models. Chenyu Huang 0001, Peng Ye 0006, Tao Chen 0003, Tong He 0001, Xiangyu Yue 0001, Wanli Ouyang |
NeurIPS | 2 |
| 2024 | Model Decides How to Tokenize: Adaptive DNA Sequence Tokenization with MxDNAabstractFoundation models have made significant strides in understanding the genomic language of DNA sequences. However, previous models typically adopt the tokenization methods designed for natural language, which are unsuitable for DNA sequences due to their unique characteristics. In addition, the optimal approach to tokenize DNA remains largely under-explored, and may not be intuitively understood by humans even if discovered. To address these challenges, we introduce MxDNA, a novel framework where the model autonomously learns an effective DNA tokenization strategy through gradient decent. MxDNA employs a sparse Mixture of Convolution Experts coupled with a deformable convolution to model the tokenization process, with the discontinuous, overlapping, and ambiguous nature of meaningful genomic segments explicitly considered. On Nucleotide Transformer Benchmarks and Genomic Benchmarks, MxDNA demonstrates superior performance to existing methods with less pretraining data and time, highlighting its effectiveness. Finally, we show that MxDNA learns unique tokenization strategy distinct to those of previous methods and captures genomic functionalities at a token level during self-supervised pretraining. Our MxDNA aims to provide a new perspective on DNA tokenization, potentially offering broad applications in various domains and yielding profound insights. Code is available at https://github.com/qiaoqiaoLF/MxDNA. Lifeng Qiao, Peng Ye 0006, Yuchen Ren 0001, Weiqiang Bai, Chaoqi Liang, Xinzhu Ma, Nanqing Dong, Wanli Ouyang |
NeurIPS | 2 |
| 2024 | BEACON: Benchmark for Comprehensive RNA Tasks and Language ModelsabstractRNA plays a pivotal role in translating genetic instructions into functional outcomes, underscoring its importance in biological processes and disease mechanisms. Despite the emergence of numerous deep learning approaches for RNA, particularly universal RNA language models, there remains a significant lack of standardized benchmarks to assess the effectiveness of these methods. In this study, we introduce the first comprehensive RNA benchmark BEACON BEnchmArk for COmprehensive RNA Task and Language Models).First, BEACON comprises 13 distinct tasks derived from extensive previous work covering structural analysis, functional studies, and engineering applications, enabling a comprehensive assessment of the performance of methods on various RNA understanding tasks. Second, we examine a range of models, including traditional approaches like CNNs, as well as advanced RNA foundation models based on language models, offering valuable insights into the task-specific performances of these models. Third, we investigate the vital RNA language model components from the tokenizer and positional encoding aspects. Notably, our findings emphasize the superiority of single nucleotide tokenization and the effectiveness of Attention with Linear Biases (ALiBi) over traditional positional encoding methods. Based on these insights, a simple yet strong baseline called BEACON-B is proposed, which can achieve outstanding performance with limited data and computational resources. The datasets and source code of our benchmark are available at https://github.com/terry-r123/RNABenchmark. Yuchen Ren 0001, Lifeng Qiao, Hongtai Jing, Peng Ye 0006, Xinzhu Ma, Hongliang Yan, Wanli Ouyang, Xihui Liu |
NeurIPS | 7 |
| 2024 | Latency-Aware Neural Architecture Performance Predictor With Query-to-Tier TechniqueabstractNeural Architecture Search (NAS) is a powerful tool for automating effective image and video processing DNN designing. The ranking of the accuracy has been advocated to design an efficient performance predictor for NAS. The previous contrastive method solves the ranking problem by comparing pairs of architectures and predicting their relative performance. However, it only focuses on the rankings between the two involved architectures and neglects the overall quality distributions of the search space, which may suffer generalization issues. On the contrary, we propose to let the performance predictor concentrate on the global quality level of specific architecture, and learn the tier embeddings of the whole search space automatically with learnable queries. The proposed method, dubbed as Neural Architecture Ranker with Query-to-Tier technique (NARQ2T), explores the quality tiers of the search space globally and classifies each individual to the tier they belong to. Thus, the predictor gains knowledge of the performance distributions of the search space which helps to generalize its ranking ability to the datasets more easily. Thanks to the encoder-decoder design, our method is able to predict the latency of the searched model without deteriorating the performance prediction. Meanwhile, the global quality distribution facilitates the search phase by directly sampling candidates according to the statistics of quality tiers, which is free of training a search algorithm, e.g., Reinforcement Learning or Evolutionary Algorithm, thus it simplifies the NAS pipeline and saves the computational overheads. The proposed NARQ2T achieves state-of-the-art performance on two widely used datasets for NAS research. Moreover, extensive experiments have validated the efficacy of the designed method. Bicheng Guo, Lilin Xu, Tao Chen 0003, Peng Ye 0006, Shibo He, Haoyu Liu 0002, Jiming Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | DeNKD: Decoupled Non-Target Knowledge Distillation for Complementing Transformer-Based Unsupervised Domain AdaptationabstractThere is a growing need to explore the potential of transformers in Unsupervised Domain Adaptation (UDA) due to their increasing success in various vision tasks. However, the application of transformers in UDA has yet to be thoroughly investigated and requires further research. In this study, our primary focus is to design a novel pipeline specifically tailored for transformer-based UDA, to address a crucial challenge: the overemphasis on the transfer of target-oriented information, mainly caused by the self-attention blocks in transformers and the cross-domain adversarial learning scheme. First, we show that non-target information, including semantic contextual information such as background features and non-target classes, must be addressed in the domain adaptation process. Recognizing the importance of incorporating non-target knowledge, we propose a decoupled non-target knowledge distillation method called DeNKD. DeNKD decouples non-target information across domains at both feature and logit levels. This decoupling is achieved through a bi-directional knowledge distillation approach that facilitates the interaction and exchange of non-target knowledge to facilitate an effective transformer-based cross-domain knowledge transfer. We perform extensive evaluations on several well-established UDA benchmark datasets. The results consistently show that DeNKD outperforms other methods, achieving the best performance across the board. For example, on the Office-Home dataset, DeNKD achieves an accuracy of 85.54%, while on the VisDA-2017 dataset, it achieves an accuracy of 89.95%. These results highlight the effectiveness of DeNKD in transformer-based UDA and its potential for improving cross-domain adaptation performance. Peng Ye 0006, Baopu Li, Tao Chen 0003, Jiayuan Fan 0001, Wanli Ouyang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | U²ConvFormer: Marrying and Evolving Nested U-Net and Scale-Aware Transformer for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification plays an important role in the human exploration of the Earth. Recent research of deep learning-based HSI classification has been fast-growing, but still suffers from three obstacles: First, existing deep learning-based HSI works lack of extraction and utilization of multigrained multiscale information and multiscale local-to-global information. Second, most previous works have too fixed-sized receptive fields in their convolutional network parts to handle HSI classification problems, and pay no attention to the existence of asymmetries in the spectral-spatial dimension of the HSI data. Third, most networks for HSI classification are hand-craft. To this end, we propose a novel architecture in this article, which is the first to combine the advantages of nested U-Net and scale-aware Transformer, named U2ConvFormer. Specifically, the nested U-Net structure can fully extract and aggregate multiscale spectral-spatial features at both inter- and inner stage granularity. The scale-aware Transformer takes multiscale local spectral-spatial features from the encoder of nested U-Net and produces multiscale global spectral-spatial features for its decoder. After that, we design a novel plug-and-play searchable operation called asymmetric spectral-spatial convolution (A2SConv), where asymmetric spectral-spatial feature pooling and multiscale feature extraction can be concurrently searched. Furthermore, we develop a customized search strategy to automatically design U2ConvFormer, which uses advanced neural architecture search (NAS) methods to enable the customization of suitable models for different hyperspectral datasets. Experimental results on three benchmark datasets, including Indian Pines, Pavia University and Houston University 2018, validate the superiority of our proposed U2ConvFormer, which achieves new state-of-the-art performance across different benchmark datasets. Lin Zhan, Peng Ye 0006, Jiayuan Fan 0001, Tao Chen 0003 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Exploring Multi-Timestep Multi-Stage Diffusion Features for Hyperspectral Image ClassificationabstractThe effectiveness of spectral-spatial feature learning is crucial for the hyperspectral image (HSI) classification task. Diffusion models, as a new class of groundbreaking generative models, have the ability to learn both contextual semantics and textual details from the distinct timestep dimension, enabling the modeling of complex spectral-spatial relations in HSIs. However, existing diffusion-based HSI classification methods only utilize manually selected single-timestep single-stage features, limiting the full exploration and exploitation of rich contextual semantics and textual information hidden in the diffusion model. To address this issue, we propose a novel diffusion-based feature learning framework that explores Multi-Timestep Multi-Stage Diffusion features for HSI classification for the first time, called MTMSD. Specifically, the diffusion model is first pretrained with unlabeled HSI patches to mine the connotation of unlabeled data, and then is used to extract the multi-timestep multi-stage diffusion features. To effectively and efficiently leverage multi-timestep multi-stage features, two strategies are further developed. One strategy is class & timestep-oriented multi-stage feature purification module with the inter-class and inter-timestep prior for reducing the redundancy of multi-stage features and alleviating memory constraints. The other one is selective timestep feature fusion module with the guidance of global features to adaptively select different timestep features for integrating texture and semantics. Both strategies facilitate the generality and adaptability of the MTMSD framework for diverse patterns of different HSI data. Extensive experiments are conducted on four public HSI datasets, and the results demonstrate that our method outperforms state-of-the-art methods for HSI classification, especially on the challenging Houston 2018 dataset. The codes are available at https://github.com/zjyaccount/MTMSD. Jiamu Sheng, Peng Ye 0006, Jiayuan Fan 0001, Tong He 0001, Bin Wang 0008, Tao Chen 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | JNDMix: Jnd-Based Data Augmentation for No-Reference Image Quality AssessmentabstractDespite substantial progress in no-reference image quality assessment (NR-IQA), previous training models often suffer from over-fitting due to the limited scale of used datasets, resulting in model performance bottlenecks. To tackle this challenge, we explore the potential of leveraging data augmentation to improve data efficiency and enhance model robustness. However, most existing data augmentation methods incur a serious issue, namely that it alters the image quality and leads to training images mismatching with their original labels. Additionally, although only a few data augmentation methods are available for NR-IQA task, their ability to enrich dataset diversity is still insufficient. To address these issues, we propose a effective and general data augmentation based on just noticeable difference (JND) noise mixing for NR-IQA task, named JNDMix. In detail, we randomly inject the JND noise, imperceptible to the human visual system (HVS), into the training image without any adjustment to its label. Extensive experiments demonstrate that JNDMix significantly improves the performance and data efficiency of various state-of-the-art NR-IQA models and the commonly used baseline models, as well as the generalization ability. More importantly, JNDMix facilitates MANIQA to achieve the state-of-the-art performance on LIVEC and KonIQ-10k. Jiamu Sheng, Jiayuan Fan 0001, Peng Ye 0006, Jianjian Cao |
ICASSP | 3 |
| 2023 | A2S-NAS: Asymmetric Spectral-Spatial Neural Architecture Search for Hyperspectral Image ClassificationabstractExisting deep learning-based hyperspectral image (HSI) classification works still suffer from the limitation of the fixed-sized receptive field, leading to difficulties in distinctive spectral-spatial features for ground objects with various sizes and arbitrary shapes. Meanwhile, plenty of previous works ignore asymmetric spectral-spatial dimensions in HSI. To address the above issues, we propose a multi-stage search architecture in order to overcome asymmetric spectral-spatial dimensions and capture significant features. First, the asymmetric pooling on the spectral-spatial dimension maximally retains the essential features of HSI. Then, the 3D convolution with a selectable range of receptive fields overcomes the constraints of fixed-sized convolution kernels. Finally, we extend these two searchable operations to different layers of each stage to build the final architecture. Extensive experiments are conducted on two challenging HSI benchmarks including Indian Pines and Houston University, and results demonstrate the effectiveness of the proposed method with superior performance compared with the related works. Lin Zhan, Jiayuan Fan 0001, Peng Ye 0006, Jianjian Cao |
ICASSP | 3 |
| 2023 | RFD-ECNet: Extreme Underwater Image Compression with Reference to Feature DictionaryabstractThriving underwater applications demand efficient extreme compression technology to realize the transmission of underwater images (UWIs) in very narrow underwater bandwidth. However, existing image compression methods achieve inferior performance on UWIs because they do not consider the characteristics of UWIs: (1) Multifarious underwater styles of color shift and distance-dependent clarity, caused by the unique underwater physical imaging; (2) Massive redundancy between different UWIs, caused by the fact that different UWIs contain several common ocean objects, which have plenty of similarities in structures and semantics. To remove redundancy among UWIs, we first construct an exhaustive underwater multi-scale feature dictionary to provide coarse-to-fine reference features for UWI compression. Subsequently, an extreme UWI compression network with reference to the feature dictionary (RFD-ECNet)1is creatively proposed, which utilizes feature match and reference feature variant to significantly remove redundancy among UWIs. To align the multifarious underwater styles and improve the accuracy of feature match, an underwater style normalized block (USNB) is proposed, which utilizes underwater physical priors extracted from the underwater physical imaging model to normalize the underwater styles of dictionary features toward the input. Moreover, a reference feature variant module (RFVM) is designed to adaptively morph the reference features, improving the similarity between the reference and input features. Experimental results on four UWI datasets show that our RFD-ECNet is the first work that achieves a significant BDrate saving of 31% over the most advanced VVC. Liquan Shen, Peng Ye 0006, Guorui Feng, Zheyin Wang |
ICCV | 3 |
| 2023 | Automatic Loss Function Search for Adversarial Unsupervised Domain AdaptationabstractUnsupervised domain adaption (UDA) aims to reduce the domain gap between labeled source and unlabeled target domains. Many prior works exploit adversarial learning that leverages pre-designed discriminators to drive the network for aligning distributions between domains. However, most of them do not consider the degeneration of the domain discriminators caused by the gradually dominating gradients of aligned target samples during training, and they still suffer from the cross-domain semantic mismatch problem in the learned feature space. Hence, this paper attempts to understand and solve both issues from the lens of optimization loss and propose an automatic loss function search for adversarial domain adaptation (ALSDA). First, we extend the common adversarial loss by adding an adjustable hyper-parameter that can re-weight the gradients assigned to target samples, so that the domain discriminator can impose consecutive and influential driving forces for domain alignment. Meanwhile, we upgrade the traditional orthogonality loss with class-wisely adjustable hyper-parameters that can strengthen the cross-domain feature separation. Since manually determining the optimal loss functions requires expensive expert efforts, we leverage the popular AutoML to automatically search for the optimal loss functions from a pre-defined novel and unique search space for UDA. Further, to enable the loss function search when the target domain is unlabeled, we introduce a simple-but-effective entropy-guided search strategy with the aid of REINFORCE learning. Extensive experiments on various typical baselines and benchmark datasets such as Office-Home, Office-31, and Birds-31 have been conducted, and the results validate the generalization and superiority of the proposed ALSDA. Peng Ye 0006, Hancheng Ye, Baopu Li, Jinyang Guo 0002, Tao Chen 0003, Wanli Ouyang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Hyperspectral Image Classification Using Spectral-Spatial Token Enhanced Transformer With Hash-Based Positional EmbeddingabstractHyperspectral image (HSI) classification aims to distinguish the category of a land coverage object for each pixel. In an effective way, the transformer architecture has been successfully introduced for the HSI classification task with promising performance. However, existing transformer-based HSI classification methods still suffer from the inability to fully explore both spectral information and spatial information in HSIs. To this end, we propose a Spectral-Spatial Token Enhanced Transformer (SSTE-Former) method with the hash-based positional embedding, which is the first to exploit multiscale spectral-spatial information for transformer-based HSI classification in-depth. Specifically, SSTE-Former accepts multiscale HSI cubes centered on the target pixel, that are preprocessed by PCA. Then, a designed multiscale CNN architecture is utilized to extract short-range spectral-spatial features and generate token embeddings. In parallel, a novel hash-based spatially enhanced positional embedding tailored for HSI cubes is developed to model the correlations within and across multiscale token embeddings. Finally, multiscale token embeddings and hash-based positional embeddings are concatenated and flattened into the transformer encoder for long-range spectral-spatial feature fusion. We conduct extensive experiments on four benchmark HSI datasets and achieve superior performance compared with the state-of-the-art HSI classification methods. Jiayuan Fan 0001, Peng Ye 0006, Mingzhen Zhu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | β-DARTS: Beta-Decay Regularization for Differentiable Architecture SearchabstractNeural Architecture Search (NAS) has attracted increasingly more attention in recent years because of its capability to design deep neural network automatically. Among them, differential NAS approaches such as DARTS, have gained popularity for the search efficiency. However, they suffer from two main issues, the weak robustness to the performance collapse and the poor generalization ability of the searched architectures. To solve these two problems, a simple-but-efficient regularization method, termed as Beta-Decay, is proposed to regularize the DARTS-based NAS searching process. Specifically, Beta-Decay regularization can impose constraints to keep the value and variance of activated architecture parameters from too large. Furthermore, we provide in-depth theoretical analysis on how it works and why it works. Experimental results on NAS-Bench-201 show that our proposed method can help to stabilize the searching process and makes the searched network more transferable across different datasets. In addition, our search scheme shows an outstanding property of being less dependent on training time and data. Comprehensive experiments on a variety of search spaces and datasets validate the effectiveness of the proposed method. The code is available at https://github.com/Sunshine-Ye/Beta-DARTS. Peng Ye 0006, Baopu Li, Yikang Li 0002, Tao Chen 0003, Jiayuan Fan 0001, Wanli Ouyang |
CVPR | 1 |
| 2022 | Generalized Global Ranking-Aware Neural Architecture Ranker for Efficient Image Classifier SearchabstractNeural Architecture Search (NAS) is a powerful tool for automating effective image processing DNN designing. The ranking has been advocated to design an efficient performance predictor for NAS. The previous contrastive method solves the ranking problem by comparing pairs of architectures and predicting their relative performance. However, it only focuses on the rankings between two involved architectures and neglects the overall quality distributions of the search space, which may suffer generalization issues. A predictor, namely Neural Architecture Ranker (NAR) which concentrates on the global quality tier of specific architecture, is proposed to tackle such problems caused by the local perspective. The NAR explores the quality tiers of the search space globally and classifies each individual to the tier they belong to according to its global ranking. Thus, the predictor gains the knowledge of the performance distributions of the search space which helps to generalize its ranking ability to the datasets more easily. Meanwhile, the global quality distribution facilitates the search phase by directly sampling candidates according to the statistics of quality tiers, which is free of training a search algorithm, e.g., Reinforcement Learning (RL) or Evolutionary Algorithm (EA), thus it simplifies the NAS pipeline and saves the computational overheads. The proposed NAR achieves better performance than the state-of-the-art methods on two widely used datasets for NAS research. On the vast search space of NAS-Bench-101, the NAR easily finds the architecture with top 0.01 performance only by sampling. It also generalizes well to different image datasets of NAS-Bench-201, i.e., CIFAR-10, CIFAR-100, and ImageNet-16-120 by identifying the optimal architectures for each of them. Bicheng Guo, Tao Chen 0003, Shibo He, Haoyu Liu 0002, Lilin Xu, Peng Ye 0006, Jiming Chen 0001 |
ACM Multimedia | 6 |
| 2022 | Stimulative Training of Residual Networks: A Social Psychology Perspective of LoafingabstractResidual networks have shown great success and become indispensable in today’s deep models. In this work, we aim to re-investigate the training process of residual networks from a novel social psychology perspective of loafing, and further propose a new training strategy to strengthen the performance of residual networks. As residual networks can be viewed as ensembles of relatively shallow networks (i.e., unraveled view) in prior works, we also start from such view and consider that the final performance of a residual network is co-determined by a group of sub-networks. Inspired by the social loafing problem of social psychology, we find that residual networks invariably suffer from similar problem, where sub-networks in a residual network are prone to exert less effort when working as part of the group compared to working alone. We define this previously overlooked problem as network loafing. As social loafing will ultimately cause the low individual productivity and the reduced overall performance, network loafing will also hinder the performance of a given residual network and its sub-networks. Referring to the solutions of social psychology, we propose stimulative training, which randomly samples a residual sub-network and calculates the KL-divergence loss between the sampled sub-network and the given residual network, to act as extra supervision for sub-networks and make the overall goal consistent. Comprehensive empirical results and theoretical analyses verify that stimulative training can well handle the loafing problem, and improve the performance of a residual network by improving the performance of its sub-networks. The code is available at https://github.com/Sunshine-Ye/NIPS22-ST. Peng Ye 0006, Shengji Tang, Baopu Li, Tao Chen 0003, Wanli Ouyang |
NeurIPS | 1 |
| 2022 | Efficient Joint-Dimensional Search with Solution Space Regularization for Real-Time Semantic Segmentation
Peng Ye 0006, Baopu Li, Tao Chen 0003, Jiayuan Fan 0001, Chen Lin 0003, Chongyan Zuo, Qinghua Chi, Wanli Ouyang |
Int. J. Comput. Vis. | 1 |