EDBT 2026 Demo / reviewers in the wild / expert
Suorong Yang
dblp:318/9341
· DBLP profile ↗
15ranked-venue papers
9as first author
15since 2021 · last 2026
0000-0001-8788-6382ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 8 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dual prototypes for adaptive pre-trained model in class-incremental learning
Suorong Yang, Baile Xu, Furao Shen, Jian Zhao 0013 |
Neural Networks | 2 |
| 2026 | IPF-RDA: An Information-Preserving Framework for Robust Data AugmentationabstractData augmentation is widely utilized as an effective technique to enhance the generalization performance of deep models. However, data augmentation may inevitably introduce distribution shifts and noises, which significantly constrain the potential and deteriorate the performance of deep networks. To this end, we propose a novel information-preserving framework, namely IPF-RDA, to enhance the robustness of data augmentations in this paper. IPF-RDA combines the proposal of (i) a new class-discriminative information estimation algorithm that identifies the points most vulnerable to data augmentation operations and corresponding importance scores; And (ii) a new information-preserving scheme that preserves the critical information in the augmented samples and ensures the diversity of augmented data adaptively. We divide data augmentation methods into three categories according to the operation types and integrate these approaches into our framework accordingly. After being integrated into our framework, the robustness of data augmentation methods can be enhanced and their full potential can be unleashed. Extensive experiments demonstrate that although being simple, IPF-RDA consistently improves the performance of numerous commonly used state-of-the-art data augmentation methods with popular deep models on a variety of datasets, including CIFAR-10, CIFAR-100, Tiny-ImageNet, CUHK03, Market1501, Oxford Flower, and MNIST, where its performance and scalability are stressed. Suorong Yang, Hongchao Yang, Suhan Guo, Furao Shen, Jian Zhao 0013 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Critic-V: VLM Critics Help Catch VLM Errors in Multimodal ReasoningabstractVision-language models (VLMs) have shown remarkable advancements in multimodal reasoning tasks. However, they still often generate inaccurate or irrelevant responses due to issues like hallucinated image understandings or unrefined reasoning paths. To address these challenges, we introduce Critic-V, a novel framework inspired by the Actor-Critic paradigm to boost the reasoning capability of VLMs. This framework decouples the reasoning process and critic process by integrating two independent components: the Reasoner, which generates reasoning paths based on visual and textual inputs, and the Critic, which provides constructive critique to refine these paths. In this approach, the Reasoner generates reasoning responses according to text prompts, which can evolve iteratively as a policy based on feedback from the Critic. This interaction process was theoretically driven by a reinforcement learning framework where the Critic offers natural language critiques instead of scalar rewards, enabling more nuanced feedback to boost the Reasoner’s capability on complex reasoning tasks. The Critic model is trained using Direct Preference Optimization (DPO), leveraging a preference dataset of critiques ranked by Rule-based Reward (RBR) to enhance its critic capabilities. Evaluation results show that the Critic-V framework significantly outperforms existing methods, including GPT-4V, on 5 out of 8 benchmarks, especially regarding reasoning accuracy and efficiency. Combining a dynamic text-based policy for the Reasoner and constructive feedback from the preference-optimized Critic enables a more reliable and context-sensitive multimodal reasoning process. Our approach provides a promising solution to enhance the reliability of VLMs, improving their performance in real-world reasoning-heavy multimodal applications such as autonomous driving and embodied intelligence. Our data and code are released at https://github.com/kyrieLei/Critic-V. Di Zhang 0026, Jingdi Lei, Junxian Li 0001, Xunzhi Wang, Zonglin Yang 0001, Jiatong Li 0003, Weida Wang, Suorong Yang, Peng Ye 0006, Wanli Ouyang, Dongzhan Zhou |
CVPR | 9 |
| 2025 | Reinforcement Learning-Guided Data Selection Via Redundancy Assessment
Suorong Yang, Peijia Li, Furao Shen, Jian Zhao 0013 |
ICCV | 1 |
| 2025 | A CLIP-Powered Framework for Robust and Generalizable Data SelectionabstractLarge-scale datasets have been pivotal to the advancements of deep learning models in recent years, but training on such large datasets inevitably incurs substantial storage and computational overhead.
Meanwhile, real-world datasets often contain redundant and noisy data, imposing a negative impact on training efficiency and model performance.
Data selection has shown promise in identifying the most representative samples from the entire dataset, which aims to minimize the performance gap with reduced training costs.
Existing works typically rely on single-modality information to assign importance scores for individual samples, which may lead to inaccurate assessments, especially when dealing with noisy or corrupted samples.
To address this limitation, we propose a novel CLIP-powered data selection framework that leverages multimodal information for more robust and generalizable sample selection.
Specifically, our framework consists of three key modules—dataset adaptation, sample scoring, and selection optimization—that together harness extensive pre-trained multimodal knowledge to comprehensively assess sample influence and optimize the selection results through multi-objective optimization.
Extensive experiments demonstrate that our approach consistently outperforms existing state-of-the-art baselines on various benchmark datasets. Notably, our method effectively removes noisy or damaged samples from the dataset, enabling it to achieve even higher performance with less data. This indicates that it is not only a way to accelerate training but can also improve overall data quality.
The implementation is available at https://github.com/Jackbrocp/clip-powered-data-selection. Suorong Yang, Peng Ye 0006, Wanli Ouyang, Dongzhan Zhou, Furao Shen |
ICLR | 1 |
| 2025 | When Dynamic Data Selection Meets Data Augmentation: Achieving Enhanced Training AccelerationabstractDynamic data selection aims to accelerate training with lossless performances. However, reducing training data inherently limits data diversity, potentially hindering generalization. While data augmentation is widely used to enhance diversity, it is typically not optimized in conjunction with selection. As a result, directly combining these techniques fails to fully exploit their synergies. To tackle the challenge, we propose a novel online data training framework that, for the first time, unifies dynamic data selection and augmentation, achieving both training efficiency and enhanced performance. Our method estimates each sample’s joint distribution of local density and multimodal semantic consistency, allowing for the targeted selection of augmentation-suitable samples while suppressing the inclusion of noisy or ambiguous data. This enables a more significant reduction in dataset size without sacrificing model generalization. Experimental results demonstrate that our method outperforms existing state-of-the-art approaches on various benchmark datasets and architectures, e.g., reducing 50% training costs on ImageNet-1k with lossless performance. Furthermore, our approach enhances noise resistance and improves model robustness, reinforcing its practical utility in real-world scenarios. Suorong Yang, Peng Ye 0006, Furao Shen, Dongzhan Zhou |
ICML | 1 |
| 2025 | Neural-Driven Image EditingabstractTraditional image editing typically relies on manual prompting, making it labor-intensive and inaccessible to individuals with limited motor control or language abilities. Leveraging recent advances in brain-computer interfaces (BCIs) and generative models, we propose LoongX, a hands-free image editing approach driven by multimodal neurophysiological signals.
LoongX utilizes state-of-the-art diffusion models trained on a comprehensive dataset of 23,928 image editing pairs, each paired with synchronized electroencephalography (EEG), functional near-infrared spectroscopy (fNIRS), photoplethysmography (PPG), and head motion signals that capture user intent.
To effectively address the heterogeneity of these signals, LoongX integrates two key modules. The cross-scale state space (CS3) module encodes informative modality-specific features. The dynamic gated fusion (DGF) module further aggregates these features into a unified latent space, which is then aligned with edit semantics via fine-tuning on a diffusion transformer (DiT).
Additionally, we pre-train the encoders using contrastive learning to align cognitive states with semantic intentions from embedded natural language.
Extensive experiments demonstrate that LoongX achieves performance comparable to text-driven methods (CLIP-I: 0.6605 vs. 0.6558; DINO: 0.4812 vs. 0.4637) and outperforms them when neural signals are combined with speech (CLIP-T: 0.2588 vs. 0.2549). These results highlight the promise of neural-driven generative models in enabling accessible, intuitive image editing and open new directions for cognitive-driven creative technologies. The code and dataset are released on the project website: https://loongx1.github.io. Xiaopeng Peng 0001, Wangbo Zhao, Zilong Ye, Suorong Yang, Jiadong Pan, Yuanxiang Chen, Kai Wang 0036, Xiaojun Chang, Gang Pan 0001, Shurong Dong, Kaipeng Zhang, Yang You 0001 |
NeurIPS | 7 |
| 2025 | GMNI: Achieve good data augmentation in unsupervised graph contrastive learning
Xin Xiong 0012, Suorong Yang, Furao Shen, Jian Zhao 0013 |
Neural Networks | 3 |
| 2025 | CS-QCFS: Bridging the performance gap in ultra-low latency spiking neural networks
Hongchao Yang, Suorong Yang, Hui Dou 0001, Furao Shen, Jian Zhao 0013 |
Neural Networks | 2 |
| 2025 | Supervised contrastive learning with prototype distillation for data incremental learning
Suorong Yang, Peijia Li, Baile Xu, Furao Shen, Jian Zhao 0013 |
Neural Networks | 1 |
| 2025 | AdaAugment: A Tuning-Free and Adaptive Approach to Enhance Data AugmentationabstractData augmentation (DA) is widely employed to improve the generalization performance of deep models. However, most existing DA methods employ augmentation operations with fixed or random magnitudes throughout the training process. While this fosters data diversity, it can also inevitably introduce uncontrolled variability in augmented data, which could potentially cause misalignment with the evolving training status of the target models. Both theoretical and empirical findings suggest that this misalignment increases the risks of both underfitting and overfitting. To address these limitations, we propose AdaAugment, an innovative and tuning-free adaptive augmentation method that leverages reinforcement learning to dynamically and adaptively adjust augmentation magnitudes for individual training samples based on real-time feedback from the target network. Specifically, AdaAugment features a dual-model architecture consisting of a policy network and a target network, which are jointly optimized to adapt augmentation magnitudes in accordance with the model's training progress effectively. The policy network optimizes the variability within the augmented data, while the target network utilizes the adaptively augmented samples for training. These two networks are jointly optimized and mutually reinforce each other. Extensive experiments across benchmark datasets and deep architectures demonstrate that AdaAugment consistently outperforms other state-of-the-art DA methods in effectiveness while maintaining remarkable efficiency. Code is available at https://github.com/Jackbrocp/AdaAugment. Suorong Yang, Peijia Li, Xin Xiong 0012, Furao Shen, Jian Zhao 0013 |
IEEE Trans. Image Process. | 1 |
| 2024 | EntAugment: Entropy-Driven Adaptive Data Augmentation Framework for Image Classification
Suorong Yang, Furao Shen, Jian Zhao 0013 |
ECCV (66) | 1 |
| 2024 | Investigating the effectiveness of data augmentation from similarity and diversity: An empirical study
Suorong Yang, Suhan Guo, Jian Zhao 0013, Furao Shen |
Pattern Recognit. | 1 |
| 2023 | Sensitivity pruner: Filter-Level compression algorithm for deep neural networks
Suhan Guo, Bilan Lai, Suorong Yang, Jian Zhao 0013, Furao Shen |
Pattern Recognit. | 3 |
| 2023 | AdvMask: A sparse adversarial attack-based data augmentation method for image classification
Suorong Yang, Jinqiao Li, Jian Zhao 0013, Furao Shen |
Pattern Recognit. | 1 |