EDBT 2026 Demo / reviewers in the wild / expert
Yi Xin 0003
dblp:33/1127-3
· DBLP profile ↗
20ranked-venue papers
5as first author
20since 2021 · last 2026
0000-0001-9526-1323ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 4 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 11 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TIDE: Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image GenerationabstractDiffusion Transformers (DiTs) are a powerful yet underexplored class of generative models compared to U-Net-based diffusion architectures. We propose TIDE—Temporal-aware sparse autoencoders for Interpretable Diffusion transformErs—a framework designed to extract sparse, interpretable activation features across timesteps in DiTs. TIDE effectively captures temporally-varying representations and reveals that DiTs naturally learn hierarchical semantics (e.g., 3D structure, object class, and fine-grained concepts) during large-scale pretraining. Experiments show that TIDE enhances interpretability and controllability while maintaining reasonable generation quality, enabling applications such as safe image editing and style transfer. Victor Shea-Jay Huang, Le Zhuo, Yi Xin 0003, Zhaokai Wang, Fu-Yun Wang, Yuchi Wang, Renrui Zhang, Peng Gao 0007, Hongsheng Li 0001 |
AAAI | 3 |
| 2026 | Benchmarking Multimodal Knowledge Conflict for Large Multimodal ModelsabstractLarge Multimodal Models (LMMs) face notable challenges when encountering multimodal knowledge conflicts, particularly under retrieval-augmented generation (RAG) frameworks, where the contextual information from external sources may contradict the model’s internal parametric knowledge, leading to unreliable outputs. However, existing benchmarks fail to reflect such realistic conflict scenarios. Most focus solely on intra-memory conflicts, while context-memory and inter-context conflicts remain largely unaddressed. Furthermore, commonly used factual knowledge-based evaluations are often overlooked, and existing datasets lack a thorough investigation into conflict detection capabilities.To bridge this gap, we propose MMKC-Bench, a benchmark designed to evaluate factual knowledge conflicts in both context-memory and inter-context scenarios. MMKC-Bench encompasses four types of multimodal knowledge conflicts and includes 1,881 knowledge instances and 3,997 images across 32 broad types, collected through automated pipelines with human verification. We evaluate four representative series of LMMs on both model behavior analysis and conflict detection tasks. Our findings show that while current LMMs are capable of recognizing knowledge conflicts, they tend to favor internal parametric knowledge over external evidence. We hope MMKC-Bench will foster further research in multimodal knowledge conflict and enhance the development of multimodal RAG systems. Yuntao Du 0001, Kailin Jiang, Yuyang Liang, Qihan Ren, Yi Xin 0003, Fenze Feng, Mingcai Chen, Hengyang Lu, Haozhe Wang 0002, Xiaoye Qu, Qian Li 0043, Dongrui Liu |
AAAI | 6 |
| 2026 | Lumina-mGPT: Flexible Photorealistic Autoregressive Text-to-Image Generation
Yi Xin 0003, Shitian Zhao, Le Zhuo, Weifeng Lin, Xinyue Li 0001, Guangtao Zhai, Xiaohong Liu 0001, Hongsheng Li 0001, Yu Qiao 0001, Peng Gao 0007 |
Int. J. Comput. Vis. | 2 |
| 2026 | Parameter-Efficient Fine-Tuning for Pre-Trained Vision Models: A Survey and Benchmark
Yi Xin 0003, Jianjiang Yang, Yuntao Du 0001, Haoxing Chen, Kangrui Cen, Yangfan He, Yuewen Cao, Junjun He, Xiaokang Yang 0001, Guangtao Zhai, Ming-Hsuan Yang 0001, Xiaohong Liu 0001 |
Int. J. Comput. Vis. | 1 |
| 2026 | M2IST: Multi-Modal Interactive Side-Tuning for Efficient Referring Expression ComprehensionabstractReferring expression comprehension (REC) is a vision-language task to locate a target object in an image based on a language expression. Fully fine-tuning general-purpose pre-trained vision-language foundation models for REC yields impressive performance but becomes increasingly costly. Parameter-efficient transfer learning (PETL) methods have shown strong performance with fewer tunable parameters. However, directly applying PETL to REC faces two challenges: (1) insufficient multi-modal interaction between pre-trained vision-language foundation models, and (2) high GPU memory usage due to gradients passing through the heavy vision-language foundation models. To this end, we present M2IST: Multi-Modal Interactive Side-Tuning with M3ISAs: Mixture of Multi-Modal Interactive Side-Adapters. During fine-tuning, we fix the pre-trained uni-modal encoders and update M3ISAs to enable efficient vision-language alignment for REC. Empirical results reveal that M2IST achieves better performance-efficiency trade-off than full fine-tuning and other PETL methods, requiring only 2.11% tunable parameters, 39.61% GPU memory, and 63.46% training time while maintaining competitive performance. Our code is released at https://github.com/xuyang-liu16/M2IST. Xuyang Liu 0002, Ting Liu 0018, Siteng Huang, Yi Xin 0003, Yue Hu 0016, Long Qin 0004, Yuanyuan Wu 0001, Honggang Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Robust Logit Adjustment for Learning with Long-Tailed Noisy DataabstractLearning with noisy labels (LNL) methods have enabled the deployment of machine learning systems with imperfectly labeled data. However, these methods often struggle to identify noise in the presence of long-tailed (LT) class distributions, where the memorization effect becomes class-dependent. Conversely, LT methods are suboptimal under label noise, as it hinders access to accurate label frequency statistics. This study aims to address the long-tailed noisy data by bridging the methodological gap between LNL and LT approaches. We propose a direct solution, termed Robust Logit Adjustment, which estimates ground-truth labels through label refurbishment, thereby mitigating the impact of label noise. Simultaneously, our method incorporates the distribution of training-time corrected target labels into the LT method logit adjustment, providing class-rebalanced supervision. Extensive experiments on both synthetic and real-world long-tailed noisy datasets demonstrate the superior performance of our method. Mingcai Chen, Yuntao Du 0001, Baoming Zhang, Yi Xin 0003, Chong-Jun Wang |
AAAI | 6 |
| 2025 | Knowledge Is Powerful: Art Knowledge-Driven Framework for Painting Style Classification Integrating Multimodal KnowledgeabstractPaintings possess profound cultural and historical backgrounds. Unlike real-life images, they convey complex semantics beyond simple visual features. This diversity and complexity make painting style classification highly challenging, and many popular visual models struggle with it. To address this issue, we propose an art knowledge-driven framework(AKDF) to improve models’ comprehension of art knowledge. AKDF utilizes multimodal models and prompts to extract style-related textual descriptions from images. Text and image features will be fused in enhanced bilinear pooling module, thus integrating art knowledge into the output features. Additionally, we design a contrastive learning auxiliary task based on label embeddings to introduce further art knowledge about style labels. Besides, AKDF incorporates texture features and adds another genre classification auxiliary task to provide more painting information. We constructs two datasets based on WikiArt due to its comprehensive and challenging nature. The extensive experiment results demonstrate the superiority of the AKDF, which effectively develops the performance of various models by over three percentage points across two datasets. Haoyang Chen 0001, Chong-Jun Wang, Yi Xin 0003, Lei Zhang 0086 |
ICASSP | 3 |
| 2025 | TR-PTS: Task-Relevant Parameter and Token Selection for Efficient Tuning
Yi Xin 0003, Mingyang Yi, Guangyang Wu, Guangtao Zhai, Xiaohong Liu 0001 |
ICCV | 3 |
| 2025 | Lumina-Image 2.0: a Unified and Efficient Image Generative FrameworkabstractWe introduce Lumina-Image 2.0, an advanced text-to-image generation framework that achieves significant progress compared to previous work, Lumina-Next. Lumina-Image 2.0 is built upon two key principles: (1) Unification - it adopts a unified architecture (Unified Next-DiT) that treats text and image tokens as a joint sequence, enabling natural cross-modal interactions and allowing seamless task expansion. Besides, since high-quality captioners can provide semantically well-aligned text-image training pairs, we introduce a unified captioning system, Unified Captioner (UniCap), specifically designed for T2I generation tasks. UniCap excels at generating comprehensive and accurate captions, accelerating convergence and enhancing prompt adherence. (2) Efficiency - to improve the efficiency of our proposed model, we develop multi-stage progressive training strategies and introduce inference acceleration techniques without compromising image quality. Extensive evaluations on academic benchmarks and public text-to-image arenas show that Lumina-Image 2.0 delivers strong performances even with only 2.6B parameters, highlighting its scalability and design efficiency. We have released our training details, code, and models at https://github.com/Alpha-VLLM/Lumina-Image-2.0. Le Zhuo, Yi Xin 0003, Ruoyi Du, Zhen Li 0026, Yiting Lu, Xinyue Li 0001, Will Beddow, Erwann Millon, Victor Perez 0005, Wenhai Wang, Yu Qiao 0001, Bo Zhang 0069, Xiaohong Liu 0001, Hongsheng Li 0001, Chang Xu 0002, Peng Gao 0007 |
ICCV | 3 |
| 2025 | From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection TuningabstractRecent text-to-image diffusion models achieve impressive visual quality through extensive scaling of training data and model parameters, yet they often struggle with complex scenes and fine-grained details. Inspired by the self-reflection capabilities emergent in large language models, we propose ReflectionFlow, an inference-time framework enabling diffusion models to iteratively reflect upon and refine their outputs. ReflectionFlow introduces three complementary inference-time scaling axes: (1) noise-level scaling to optimize latent initialization; (2) prompt-level scaling for precise semantic guidance; and most notably, (3) reflection-level scaling, which explicitly provides actionable reflections to iteratively assess and correct previous generations. To facilitate reflection-level scaling, we construct GenRef, a large-scale dataset comprising 1 million triplets, each containing a reflection, a flawed image, and an enhanced image. Leveraging this dataset, we efficiently perform reflection tuning on state-of-the-art diffusion transformer, FLUX.1-dev, by jointly modeling multimodal inputs within a unified framework. Experimental results show that ReflectionFlow significantly outperforms naive noise-level scaling methods, offering a scalable and compute-efficient solution toward higher-quality image synthesis on challenging tasks. Le Zhuo, Liangbing Zhao, Sayak Paul, Yue Liao, Renrui Zhang, Yi Xin 0003, Peng Gao 0007, Mohamed Elhoseiny 0001, Hongsheng Li 0001 |
ICCV | 6 |
| 2025 | D2O: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models
Zhongwei Wan, Xinjian Wu, Yu Zhang 0133, Yi Xin 0003, Chaofan Tao, Zhihong Zhu 0001, Xin Wang 0120, Longyue Wang, Mi Zhang 0002 |
ICLR | 4 |
| 2025 | Towards Emotional Insights in Art: A Knowledge-Driven Framework Integrating Artistic Emotion Knowledge into Vision Language ModelsabstractUnderstanding emotions in art paintings and generating comments about emotions is a highly challenging task due to their rich semantics and complex expressions. However, existing related methods tend to neglect the critical role of artistic emotion features in artworks at the visual level, while general vision language models(VLMs) lack a comprehensive grasp of art domain knowledge. To address this, we propose a knowledge-driven framework based on artistic emotion knowledge, which helps VLMs better comprehend emotions in artworks. The framework primarily comprises the artistic-emotion visual tower and the ANP cross-attention module. Specifically, the artistic-emotion visual tower introduces additional emotion tokens before visual features are fed into transformer layers. It integrate low-level art features into emotion tokens and infuses high-level art features into emotion tokens in the penultimate transformer layer. Moreover, the attention mechanism of visual tower is optimized by emotion bias, which enhances the model’s sensitivity to art features related to emotions. The ANP cross-attention module will extract affective entities directly from paintings as adjective-noun pairs and apply cross-attention with visual tower outputs. We have conducted extensive experiments on Artemis 1.0 and Artemis 2.0 datasets. The results demonstrate that our framework effectively improves the performance of VLMs in the task of understanding artistic emotions on eight metrics, outperforming existing methods. Haoyang Chen 0001, Hengyang Lu, Yi Xin 0003, Dian Kong, Chong-Jun Wang |
IJCNN | 3 |
| 2025 | PgM: Partitioner Guided Modal Learning FrameworkabstractMultimodal learning benefits from multiple modal information, and each learned modal representations can be divided into uni-modal that can be learned from uni-modal training and paired-modal features that can be learned from cross-modal interaction. Building on this perspective, we propose a partitioner-guided modal learning framework, PgM, which consists of the modal partitioner, uni-modal learner, paired-modal learner, and uni-paired modal decoder. Modal partitioner segments the learned modal representation into uni-modal and paired-modal features. Modal learner incorporates two dedicated components for uni-modal and paired-modal learning. Uni-paired modal decoder reconstructs modal representation based on uni-modal and paired-modal features. PgM offers three key benefits: 1) thorough learning of uni-modal and paired-modal features, 2) flexible distribution adjustment for uni-modal and paired-modal representations to suit diverse downstream tasks, and 3) different learning rates across modalities and partitions. Extensive experiments demonstrate the effectiveness of PgM across four multimodal tasks and further highlight its transferability to existing models. Additionally, we visualize the distribution of uni-modal and paired-modal features across modalities and tasks, offering insights into their respective contributions. Guimin Hu, Yi Xin 0003, Lijie Hu, Zhihong Zhu 0001, Hasti Seifi |
ACM Multimedia | 2 |
| 2025 | SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement LearningabstractMultimodal large language models (MLLMs) have shown promising capabilities in reasoning tasks, yet still struggle significantly with complex problems requiring explicit self-reflection and self-correction, especially compared to their unimodal text-based counterparts. Existing reflection methods are simplistic and struggle to generate meaningful, instructive feedback, as the reasoning ability and knowledge limits of pre-trained models are largely fixed during initial training. To overcome these challenges, we propose \textit{multimodal \textbf{S}elf-\textbf{R}eflection enhanced reasoning with Group Relative \textbf{P}olicy \textbf{O}ptimization} \textbf{SRPO}, a two-stage reflection-aware reinforcement learning (RL) framework explicitly designed to enhance multimodal LLM reasoning. In the first stage, we construct a high-quality, reflection-focused dataset under the guidance of an advanced MLLM, which generates reflections based on initial responses to help the policy model to learn both reasoning and self-reflection. In the second stage, we introduce a novel reward mechanism within the GRPO framework that encourages concise and cognitively meaningful reflection while avoiding redundancy. Extensive experiments across multiple multimodal reasoning benchmarks—including MathVista, MathVision, Mathverse, and MMMU-Pro—using Qwen-2.5-VL-7B and Qwen-2.5-VL-32B demonstrate that SRPO significantly outperforms state-of-the-art models, achieving notable improvements in both reasoning accuracy and reflection quality. Zhongwei Wan, Zhihao Dou, Che Liu 0002, Yu Zhang 0133, Dongfei Cui, Qinjian Zhao, Hui Shen 0008, Yi Xin 0003, Chaofan Tao, Yangfan He, Mi Zhang 0002, Shen Yan 0008 |
NeurIPS | 9 |
| 2025 | Multi-source fully test-time adaptation
Yuntao Du 0001, Yi Xin 0003, Mingcai Chen, Mujie Zhang, Chong-Jun Wang |
Neural Networks | 3 |
| 2024 | VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene UnderstandingabstractLarge-scale pre-trained models have achieved remarkable success in various computer vision tasks. A standard approach to leverage these models is to fine-tune all model parameters for downstream tasks, which poses challenges in terms of computational and storage costs. Recently, inspired by Natural Language Processing (NLP), parameter-efficient transfer learning has been successfully applied to vision tasks. However, most existing techniques primarily focus on single-task adaptation, and despite limited research on multi-task adaptation, these methods often exhibit suboptimal training/inference efficiency. In this paper, we first propose an once-for-all Vision Multi-Task Adapter (VMT-Adapter), which strikes approximately O(1) training and inference efficiency w.r.t task number. Concretely, VMT-Adapter shares the knowledge from multiple tasks to enhance cross-task interaction while preserves task-specific knowledge via independent knowledge extraction modules. Notably, since task-specific modules require few parameters, VMT-Adapter can handle an arbitrary number of tasks with a negligible increase of trainable parameters. We also propose VMT-Adapter-Lite, which further reduces the trainable parameters by learning shared parameters between down- and up-projections. Extensive experiments on four dense scene understanding tasks demonstrate the superiority of VMT-Adapter(-Lite), achieving a 3.96% (1.34%) relative improvement compared to single-task full fine-tuning, while utilizing merely ~1% (0.36%) trainable parameters of the pre-trained model. Yi Xin 0003, Junlong Du, Zhiwen Lin |
AAAI | 1 |
| 2024 | MmAP: Multi-Modal Alignment Prompt for Cross-Domain Multi-Task LearningabstractMulti-Task Learning (MTL) is designed to train multiple correlated tasks simultaneously, thereby enhancing the performance of individual tasks. Typically, a multi-task network structure consists of a shared backbone and task-specific decoders. However, the complexity of the decoders increases with the number of tasks. To tackle this challenge, we integrate the decoder-free vision-language model CLIP, which exhibits robust zero-shot generalization capability. Recently, parameter-efficient transfer learning methods have been extensively explored with CLIP for adapting to downstream tasks, where prompt tuning showcases strong potential. Nevertheless, these methods solely fine-tune a single modality (text or visual), disrupting the modality structure of CLIP. In this paper, we first propose Multi-modal Alignment Prompt (MmAP) for CLIP, which aligns text and visual modalities during fine-tuning process. Building upon MmAP, we develop an innovative multi-task prompt learning framework. On the one hand, to maximize the complementarity of tasks with high similarity, we utilize a gradient-driven task grouping method that partitions tasks into several disjoint groups and assign a group-shared MmAP to each group. On the other hand, to preserve the unique characteristics of each task, we assign an task-specific MmAP to each task. Comprehensive experiments on two large multi-task learning datasets demonstrate that our method achieves significant performance improvements compared to full fine-tuning while only utilizing approximately ~ 0.09% of trainable parameters. Yi Xin 0003, Junlong Du, Shouhong Ding |
AAAI | 1 |
| 2024 | V-PETL Bench: A Unified Visual Parameter-Efficient Transfer Learning BenchmarkabstractParameter-efficient transfer learning (PETL) methods show promise in adapting a pre-trained model to various downstream tasks while training only a few parameters. In the computer vision (CV) domain, numerous PETL algorithms have been proposed, but their direct employment or comparison remains inconvenient. To address this challenge, we construct a Unified Visual PETL Benchmark (V-PETL Bench) for the CV domain by selecting 30 diverse, challenging, and comprehensive datasets from image recognition, video action recognition, and dense prediction tasks. On these datasets, we systematically evaluate 25 dominant PETL algorithms and open-source a modular and extensible codebase for fair evaluation of these algorithms. V-PETL Bench runs on NVIDIA A800 GPUs and requires approximately 310 GPU days. We release all the benchmark, making it more efficient and friendly to researchers. Additionally, V-PETL Bench will be continuously updated for new PETL algorithms and CV tasks. Yi Xin 0003, Xuyang Liu 0002, Yuntao Du 0001, Haodi Zhou, Christina E. Lee, Junlong Du, Haozhe Wang 0002, Mingcai Chen, Ting Liu 0018, Guimin Hu, Zhongwei Wan, Rongchao Zhang, Aoxue Li, Mingyang Yi, Xiaohong Liu 0001 |
NeurIPS | 1 |
| 2024 | Generation, augmentation, and alignment: a pseudo-source domain based method for source-free domain adaptation
Yuntao Du 0001, Haiyang Yang, Mingcai Chen, Hongtao Luo, Juan Jiang, Yi Xin 0003, Chong-Jun Wang |
Mach. Learn. | 6 |
| 2023 | Self-Training with Label-Feature-Consistency for Domain Adaptation
Yi Xin 0003, Pengsheng Jin, Yuntao Du 0001, Chong-Jun Wang |
DASFAA (4) | 1 |