EDBT 2026 Demo / reviewers in the wild / expert
Jiacheng Ruan
dblp:274/2808
· DBLP profile ↗
21ranked-venue papers
10as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 8 first-author · 15 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language ModelsabstractRecently, multimodal large language models (MLLMs) have achieved significant advancements across various domains, and corresponding evaluation benchmarks have been continuously refined and improved. In this process, benchmarks in the scientific domain have played an important role in assessing the reasoning capabilities of MLLMs. However, existing benchmarks still face three key challenges: 1) Insufficient evaluation of models' reasoning abilities in multilingual scenarios; 2) Inadequate assessment of MLLMs' comprehensive modality coverage; 3) Lack of fine-grained annotation of scientific knowledge points. To address these gaps, we propose MME-SCI, a comprehensive and challenging benchmark. We carefully collected 1,019 high-quality question-answer pairs, which involve 3 distinct evaluation modes. These pairs cover four subjects, namely mathematics, physics, chemistry, and biology, and support five languages: Chinese, English, French, Spanish, and Japanese. We conducted extensive experiments on 16 open-source models and 4 closed-source models, and the results demonstrate that MME-SCI is widely challenging for existing MLLMs. For instance, under the Image-only evaluation mode, o4-mini achieved accuracy of only 52.11%, 24.73%, 36.57%, and 29.80% in mathematics, physics, chemistry, and biology, respectively, indicating a significantly higher difficulty level compared to existing benchmarks. More importantly, using MME-SCI's multilingual and fine-grained knowledge attributes, we analyzed existing models' performance in depth and identified their weaknesses in specific domains. For example, in questions related to "Magnetic Field", o4-mini correctly answered only 5 out of 33 questions, thereby fine-grainedly exposing the model's vulnerabilities. These findings highlight the urgent need to enhance the scientific reasoning capabilities of MLLMs. Jiacheng Ruan, Ting Liu 0016, Yuzhuo Fu, Yangyang Kang |
AAAI | 1 |
| 2026 | Unnoticed Yet Effective: A Hybrid Physical Camouflage Framework Against DNNs and Human PerceptionabstractWhile adversarial attacks can effectively deceive deep neural networks, their real-world applicability is often limited by complex and conspicuous patterns that reveal their attack intent to human observers. To overcome this limitation, we propose UYE, a novel camouflage framework designed to simultaneously mislead DNNs and evade human perception. UYE incorporates two key components: an attention refiner leveraging a pre-trained vision encoder to optimize adversarial patterns for robust attacks across diverse environments, and a perception evaluator trained on a preference dataset curated using tailored prompts from human-aligned large multimodal models to ensure natural and unobtrusive camouflage generation. Extensive experiments demonstrate that UYE outperforms state-of-the-art methods in achieving an optimal balance between human stealth and model deception while maintaining effectiveness in real-world scenarios. Mingye Xie, Jiacheng Ruan, Ting Liu 0016, Yuzhuo Fu |
AAAI | 2 |
| 2026 | DiverSeed: Integrating Active Learning for Target Domain Data Generation in Instruction Tuning
Jingsheng Gao, Mengnan Qi, Suncheng Xiang, Ke Ji, Jiacheng Ruan, Ting Liu 0016, Yuzhuo Fu |
Mach. Learn. | 6 |
| 2025 | MM-CamObj: A Comprehensive Multimodal Dataset for Camouflaged Object ScenariosabstractLarge visual-language models (LVLMs) have achieved great success in multiple applications. However, they still encounter challenges in complex scenes, especially those involving camouflaged objects. This is primarily due to the lack of samples related to camouflaged scenes in the training dataset. To mitigate this issue, we construct the MM-CamObj dataset for the first time, comprising two subsets: CamObj-Align and CamObj-Instruct. Specifically, CamObj-Align contains 11,363 image-text pairs, and it is designed for VL alignment and injecting rich knowledge of camouflaged scenes into LVLMs. CamObj-Instruct is collected for fine-tuning the LVLMs with improved instruction-following capabilities, and it includes 11,363 images and 68,849 conversations with diverse instructions. Based on the MM-CamObj dataset, we propose the CamObj-Llava, an LVLM specifically designed for addressing tasks in camouflaged scenes. To facilitate our model's effective acquisition of knowledge about camouflaged objects and scenes, we introduce a curriculum learning strategy with six distinct modes. Additionally, we construct the CamObj-Bench to evaluate the existing LVLMs' capabilities of understanding, recognition, localization and count in camouflage scenes. This benchmark includes 600 images and 7 tasks, with a total of 9,449 questions. Extensive experiments are conducted on the CamObj-Bench with CamObj-Llava, 8 existing open-source and 3 closed-source LVLMs. Surprisingly, the results indicate that our model achieves a 25.84% improvement in 4 out of 7 tasks compared to GPT-4o. Jiacheng Ruan, Wenzhen Yuan 0002, Zehao Lin, Ning Liao, Feiyu Xiong, Ting Liu 0016, Yuzhuo Fu |
AAAI | 1 |
| 2025 | TTE: Two Tokens Are Enough to Improve Parameter-Efficient TuningabstractExisting fine-tuning paradigms are predominantly characterized by Full Parameter Tuning (FPT) and Parameter-Efficient Tuning (PET). FPT fine-tunes all parameters of a pre-trained model on downstream tasks, whereas PET freezes the pre-trained model and employs only a minimal number of learnable parameters for fine-tuning. However, both approaches face issues of overfitting, especially in scenarios where downstream samples are limited. This issue has been thoroughly explored in FPT, but less so in PET. To this end, this paper investigates overfitting in PET, representing a pioneering study in the field. Specifically, across 19 image classification datasets, we employ three classic PET methods (e.g., VPT, Adapter/Adaptformer, and LoRA) and explore various regularization techniques to mitigate overfitting. Regrettably, the results suggest that existing regularization techniques are incompatible with the PET process and may even lead to performance degradation. Consequently, we introduce a new framework named TTE (Two Tokens are Enough), which effectively alleviates overfitting in PET through a novel constraint function based on the learnable tokens. Experiments conducted on 24 datasets across image and few-shot classification tasks demonstrate that our fine-tuning framework not only mitigates overfitting but also significantly enhances PET's performance. Notably, our TTE framework surpasses the highest-performing FPT framework (DR-Tune), utilizing significantly fewer parameters (0.15M vs. 85.84M) and achieving an improvement of 1%. Jiacheng Ruan, Mingye Xie, Jingsheng Gao, Suncheng Xiang, Ting Liu 0016, Yuzhuo Fu |
AAAI | 1 |
| 2025 | GPA: Enhancing Generalizable Physical Adversarial Attacks Across Multiple Vision TasksabstractAdversarial attacks pose a significant challenge in deep learning, as carefully crafted perturbations can severely degrade even the most advanced models. In real-world scenarios, where the target models are often unknown, previous works often focus on creating adversarial patterns for specific known models, with the goal of generalizing these patterns to other models. However, such attacks rely heavily on prior model information, leading to poor generalization. To overcome this, we propose a novel method called GPA. Our solution includes an attention extraction module based on a pre-trained vision encoder, which captures precise and generalizable features of model attention on objects. We also introduce attack loss functions that divert attention away from target objects. Compared to state-of-the-art methods, our approach achieves superior attack performance across various downstream vision tasks, including object detection, instance segmentation, and depth estimation. Moreover, the adversarial patterns generated by GPA maintain their effectiveness in real-world scenarios. Mingye Xie, Suncheng Xiang, Jiacheng Ruan, Zefang Yu, Ting Liu 0016, Yuzhuo Fu |
ICASSP | 3 |
| 2025 | VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward Models
Jiacheng Ruan, Wenzhen Yuan 0002, Xiqi Gao 0001, Daoxin Zhang, Yao Hu 0002, Ting Liu 0016, Yuzhuo Fu |
ICCV | 1 |
| 2025 | NLOSdiffuser: Generalized Steady-State Non-Line-of-sight Imaging toward Indoor ScenariosabstractNon-Line-of-sight (NLOS) imaging broadens the field of view to observe objects beyond direct sightlines, attracting considerable research interest. However, previous studies often overlook real-world environmental impacts, focusing on idealized scenarios, which limits their practical applicability. We introduce a steady-state NLOS imaging method tailored for indoor environments, showcasing robust generalization across varied settings despite training on static scenes. Our method employs a synthetic indoor dataset mimicking real-world conditions for training and diverse scenarios for evaluation. A two-stage network initially conducts pixel-level coarse reconstruction to convert diffuse reflection images into intermediate feature maps, followed by semantic-level reconstruction using a pre-trained diffusion model, yielding highly recognizable images. Comprehensive experiments demonstrate superior reconstruction accuracy, image quality, and adaptability compared to conventional methods. Verified in real-world scenarios, our method effectively images various concealed objects, marking significant progress in NLOS imaging. Jiacheng Ruan, Zongyun Zhang, Ting Liu 0016, Yuzhuo Fu |
ICME | 3 |
| 2025 | EIAD: Explainable Industrial Anomaly Detection Via Multi-Modal Large Language ModelsabstractIndustrial Anomaly Detection (IAD) is critical to ensure product quality during manufacturing. Although existing zero-shot defect segmentation and detection methods have shown effectiveness, they cannot provide detailed descriptions of the defects. Furthermore, the application of large multi-modal models in IAD remains in its infancy, facing challenges in balancing question-answering (QA) performance and mask-based grounding capabilities, often owing to overfitting during the fine-tuning process. To address these challenges, we propose a novel approach that introduces a dedicated multi-modal defect localization module to decouple the dialog functionality from the core feature extraction. This decoupling is achieved through independent optimization objectives and tailored learning strategies. Additionally, we contribute to the first multi-modal industrial anomaly detection training dataset, named Defect Detection Question Answering (DDQA), encompassing a wide range of defect types and industrial scenarios. Unlike conventional datasets that rely on GPT-generated data, DDQA ensures authenticity and reliability and offers a robust foundation for model training. Experimental results demonstrate that our proposed method, Explainable Industrial Anomaly Detection Assistant (EIAD), achieves outstanding performance in defect detection and localization tasks. It not only significantly enhances accuracy but also improves interpretability. These advancements highlight the potential of EIAD for practical applications in industrial settings. Zongyun Zhang, Jiacheng Ruan, Ting Liu 0016, Yuzhuo Fu |
ICME | 2 |
| 2025 | MPI-CD: Multi-Path Information Contrastive Decoding for Mitigating Hallucinations in Large Vision-Language ModelsabstractIn recent years, despite substantial advancements in large vision-language models (LVLMs), they still encounter the issue of ''hallucinations''-where generated results appears reasonable but often deviates from the visual input or actual facts. In contrast, the human cognitive system, when processing visual input, initially relies on visual perception to distinguish between the salient region and non-salient region, integrating relevant information. Subsequently, it recalls pertinent memory details, ultimately generating a comprehensive cognitive outcome. Inspired by this process, we propose a novel, training-free decoding approach, dubbed as Multi-Path Information Contrastive Decoding (MPI-CD). Specifically, to simulate the human information integration process, we design a three-branch structure called the Tri-Branch Integrator (TBI), which contrasts the original, salient region, and non-salient region images to effectively improve the reliability of the LVLMs' output. Furthermore, to mimic the human memory recall mechanism, we further investigate the importance of hidden layer features and propose the Memory Recall Module (MRM). This module adaptively extracts meaningful memory information from the hidden layers and incorporates it into the decoding process, thereby effectively alleviating the hallucination issue. We conduct extensive experiments on three widely used benchmarks (e.g. POPE, AMBER, and MME) using two classic LVLMs. The experimental results demonstrate that our MPI-CD significantly mitigates hallucinations in LVLMs without requiring additional training. Jiacheng Ruan, Zongyun Zhang, Jingsheng Gao, Wenzhen Yuan 0002, Ting Liu 0016, Yuzhuo Fu |
ACM Multimedia | 1 |
| 2025 | Dynamic Data Mixing Maximizes Instruction Tuning for Mixture-of-ExpertsabstractTong Zhu, Daize Dong, Xiaoye Qu, Jiacheng Ruan, Wenliang Chen, Yu Cheng. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Tong Zhu 0002, Daize Dong, Xiaoye Qu, Jiacheng Ruan, Wenliang Chen, Yu Cheng 0001 |
NAACL (Long Papers) | 4 |
| 2025 | Learning multi-axis representation in frequency domain for medical image segmentation
Jiacheng Ruan, Jingsheng Gao, Mingye Xie, Suncheng Xiang |
Mach. Learn. | 1 |
| 2025 | Learning Visual-Semantic Embedding for Generalizable Person Re-Identification: A Unified PerspectiveabstractGeneralizable person Re-Identification (Re-ID) is a very hot research topic in machine learning and computer vision, which plays a significant role in realistic scenarios due to its various applications in public security and video surveillance. However, previous methods mainly focus on the visual representation learning, while neglect to explore the potential of semantic features during training, which easily leads to poor generalization capability when adapted to the new domain. In this article, we present a unified perspective called MMET for more robust visual-semantic embedding learning on generalizable Re-ID. To further enhance the robust feature learning in the context of transformer, a dynamic masking mechanism called Masked Multimodal Modeling (MMM) strategy is introduced to mask both the image patches and the text tokens, which can jointly work on multimodal or unimodal data and significantly boost the performance of generalizable person Re-ID. Extensive experiments on benchmark datasets demonstrate the competitive performance of our method over previous approaches. We hope this method could advance the research towards visual-semantic representation learning. Our source code is also publicly available at https://github.com/JeremyXSC/MMET . Suncheng Xiang, Jingsheng Gao, Mingye Xie, Mengyuan Guan, Jiacheng Ruan, Yuzhuo Fu |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | LAMM: Label Alignment for Multi-Modal Prompt LearningabstractWith the success of pre-trained visual-language (VL) models such as CLIP in visual representation tasks, transferring pre-trained models to downstream tasks has become a crucial paradigm. Recently, the prompt tuning paradigm, which draws inspiration from natural language processing (NLP), has made significant progress in VL field. However, preceding methods mainly focus on constructing prompt templates for text and visual inputs, neglecting the gap in class label representations between the VL models and downstream tasks. To address this challenge, we introduce an innovative label alignment method named \textbf{LAMM}, which can dynamically adjust the category embeddings of downstream datasets through end-to-end training. Moreover, to achieve a more appropriate label distribution, we propose a hierarchical loss, encompassing the alignment of the parameter space, feature space, and logits space. We conduct experiments on 11 downstream vision datasets and demonstrate that our method significantly improves the performance of existing multi-modal prompt learning models in few-shot scenarios, exhibiting an average accuracy improvement of 2.31(\%) compared to the state-of-the-art methods on 16 shots. Moreover, our methodology exhibits the preeminence in continual learning compared to other prompt tuning methods. Importantly, our method is synergistic with existing prompt tuning methods and can boost the performance on top of them. Our code and dataset will be publicly available at https://github.com/gaojingsheng/LAMM. Jingsheng Gao, Jiacheng Ruan, Suncheng Xiang, Zefang Yu, Ke Ji, Mingye Xie, Ting Liu 0016, Yuzhuo Fu |
AAAI | 2 |
| 2024 | LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-TrainingabstractMixture-of-Experts (MoE) has gained increasing popularity as a promising framework for scaling up large language models (LLMs).However, training MoE from scratch in a largescale setting still suffers from data-hungry and instability problems.Motivated by this limit, we investigate building MoE models from existing dense large language models.Specifically, based on the well-known LLaMA-2 7B model, we obtain an MoE model by: (1) Expert Construction, which partitions the parameters of original Feed-Forward Networks (FFNs) into multiple experts; (2) Continual pretraining, which further trains the transformed MoE model and additional gate networks.In this paper, we comprehensively explore different methods for expert construction and various data sampling strategies for continual pretraining.After these stages, our LLaMA-MoE models could maintain language abilities and route the input tokens to specific experts with part of the parameters activated.Empirically, by training 200B tokens, LLaMA-MoE-3.5Bmodels significantly outperform dense models that contain similar activation parameters. Tong Zhu 0002, Xiaoye Qu, Daize Dong, Jiacheng Ruan, Jingqi Tong, Conghui He, Yu Cheng 0001 |
EMNLP | 4 |
| 2024 | VT-ReID: Learning Discriminative Visual-Text Representation for Polyp Re-IdentificationabstractColonoscopic Polyp Re-Identification (ReID) aims to match a specific polyp in a large gallery with different cameras and views, which plays a key role in the prevention and treatment of colorectal cancer in the computer-aided diagnosis. However, traditional methods mainly focus on the visual representation learning, while neglecting to explore the potential of semantic features during training, which may easily lead to poor generalization capability when adapting the pre-trained model to the new scenarios. To relieve this dilemma, we propose a simple but effective training method named VT-ReID, which can remarkably enrich the representation of polyp videos with the interchange of high-level semantic information. Moreover, a dynamic mechanism named DCM is introduced to leverage contrastive learning to promote better separation between different categories. Empirical results show that our method significantly outperforms current state-of-the art methods with a clear margin. Suncheng Xiang, Cang Liu, Jiacheng Ruan, Shilun Cai, Sijia Du, Dahong Qian |
ICASSP | 3 |
| 2024 | Oceanship: A Large-Scale Dataset for Underwater Audio Target Recognition
Suncheng Xiang, Jingsheng Gao, Jiacheng Ruan, Yanping Hu, Ting Liu 0016, Yuzhuo Fu |
ICIC (4) | 5 |
| 2024 | iDAT: inverse Distillation Adapter-TuningabstractAdapter-Tuning (AT) method involves freezing a pre-trained model and introducing trainable adapter modules to acquire downstream knowledge, thereby calibrating the model for better adaptation to downstream tasks. This paper proposes a distillation framework for the AT method instead of crafting a carefully designed adapter module, which aims to improve fine-tuning performance. For the first time, we explore the possibility of combining the AT method with knowledge distillation. Via statistical analysis, we observe significant differences in the knowledge acquisition between adapter modules of different models. Leveraging these differences, we propose a simple yet effective framework called inverse Distillation Adapter-Tuning (iDAT). Specifically, we designate the smaller model as the teacher and the larger model as the student. The two are jointly trained, and online knowledge distillation is applied to inject knowledge of different perspective to student model, and significantly enhance the fine-tuning performance on downstream tasks. Extensive experiments on the VTAB-1K benchmark with 19 image classification tasks demonstrate the effectiveness of iDAT. The results show that using existing AT method within our iDAT framework can further yield a 2.66% performance gain, with only an additional 0.07M trainable parameters. Our approach compares favorably with state-of-the-arts without bells and whistles. Our code is available at https://github.com/JCruan519/iDAT. Jiacheng Ruan, Jingsheng Gao, Mingye Xie, Daize Dong, Suncheng Xiang, Ting Liu 0016, Yuzhuo Fu |
ICME | 1 |
| 2024 | GIST: Improving Parameter Efficient Fine-Tuning via Knowledge InteractionabstractRecently, the Parameter Efficient Fine-Tuning (PEFT) method, which adjusts or introduces fewer trainable parameters to calibrate pre-trained models on downstream tasks, has been a hot research topic. However, existing PEFT methods within the traditional fine-tuning framework have two main shortcomings: 1) They overlook the explicit association between trainable parameters and downstream knowledge. 2) They neglect the interaction between the intrinsic task-agnostic knowledge of pre-trained models and the task-specific knowledge of downstream tasks. These oversights lead to insufficient utilization of knowledge and suboptimal performance. To address these issues, we propose a novel fine-tuning framework, named GIST, that can be seamlessly integrated into the current PEFT methods in a plug-and-play manner. Specifically, our framework first introduces a trainable token, called the Gist token, when applying PEFT methods on downstream tasks. This token serves as an aggregator of the task-specific knowledge learned by the PEFT methods and builds an explicit association with downstream tasks. Furthermore, to facilitate explicit interaction between task-agnostic and task-specific knowledge, we introduce the concept of knowledge interaction via a Bidirectional Kullback-Leibler Divergence objective. As a result, PEFT methods within our framework can enable the pre-trained model to understand downstream tasks more comprehensively by fully leveraging both types of knowledge. Extensive experiments on the 35 datasets demonstrate the universality and scalability of our framework. Notably, the PEFT method within our GIST framework achieves up to a 2.25% increase on the VTAB-1K benchmark with an addition of just 0.8K parameters (0.009 of ViT-B/16). The code is available at https://github.com/JCruan519/GIST. Jiacheng Ruan, Jingsheng Gao, Mingye Xie, Suncheng Xiang, Zefang Yu, Ting Liu 0016, Yuzhuo Fu, Xiaoye Qu |
ACM Multimedia | 1 |
| 2023 | EGE-UNet: An Efficient Group Enhanced UNet for Skin Lesion Segmentation
Jiacheng Ruan, Mingye Xie, Jingsheng Gao, Ting Liu 0016, Yuzhuo Fu |
MICCAI (4) | 1 |
| 2022 | MALUNet: A Multi-Attention and Light-weight UNet for Skin Lesion SegmentationabstractRecently, some pioneering works have preferred applying more complex modules to improve segmentation performances. However, it is not friendly for actual clinical environments due to limited computing resources. To address this challenge, we propose a light-weight model to achieve competitive performances for skin lesion segmentation at the lowest cost of parameters and computational complexity so far. Briefly, we propose four modules: (1) DGA consists of dilated convolution and gated attention mechanisms to extract global and local feature information; (2) IEA, which is based on external attention to characterize the overall datasets and enhance the connection between samples; (3) CAB is composed of 1D convolution and fully connected layers to perform a global and local fusion of multi-stage features to generate attention maps at channel axis; (4) SAB, which operates on multi-stage features by a shared 2D convolution to generate attention maps at spatial axis. We combine four modules with our U-shape architecture and obtain a light-weight medical image segmentation model dubbed as MALUNet. Compared with UNet, our model improves the mIoU and DSC metrics by 2.39% and 1.49%, respectively, with a 44x and 166x reduction in the number of parameters and computational complexity. In addition, we conduct comparison experiments on two skin lesion segmentation datasets (ISIC2017 and ISIC2018). Experimental results show that our model achieves state-of-the-art in balancing the number of parameters, computational complexity and segmentation performances. Code is available at https://github.com/JCruan519/MALUNet. Jiacheng Ruan, Suncheng Xiang, Mingye Xie, Ting Liu 0016, Yuzhuo Fu |
BIBM | 1 |