Dong Chen 0017

dblp:44/3371-17 · DBLP profile ↗
← Back
13ranked-venue papers
8as first author
13since 2021 · last 2026
0000-0002-4859-1757ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Improving large models with small models: Lower costs and better performance
Dong Chen 0017, Shuo Zhang 0014, Yueting Zhuang, Siliang Tang, Qidong Liu 0001, Xin Yang 0011, Mingliang Xu 0001
Neural Networks1
2026 Collaborative Small and Large Models for Crowd Simulation With Incomplete Trajectory Data
abstract
Crowd simulation plays a crucial role in various domains, including entertainment, urban planning, and safety assessment. Data-driven methods offer significant advantages in simulating natural and diverse crowd behaviors, enabling highly realistic simulations. However, existing methods often face challenges due to incomplete trajectory data and limited generalization to unfamiliar scenarios. To address these limitations, we propose a novel crowd simulation framework based on the collaboration of a small model and a large model. Inspired by the dual-process decision-making mechanism in cognitive psychology, this framework enables efficient handling of familiar scenarios while leveraging the reasoning capabilities of large models in complex or unfamiliar environments. The small model, responsible for generating fast and reactive behaviors, is trained on real-world incomplete trajectory data to learn movement patterns. The large model, which performs simulation correction to refine failed behaviors, leverages past successful and failed experiences to enhance behavior generation in complex scenarios. Experimental results demonstrate that our framework significantly improves simulation accuracy in the presence of missing trajectory segments and enhances cross-scene generalization.
Dong Chen 0017, Shuo He 0002, Yingcai Wu, Mingliang Xu 0001
IEEE Trans. Vis. Comput. Graph.4
2025 Wavelet and Prototype Augmented Query-based Transformer for Pixel-level Surface Defect Detection
abstract
As an important part of intelligent manufacturing, pixel-level surface defect detection (SDD) aims to locate defect areas through mask prediction. Previous methods adopt the image-independent static convolution to indiscriminately classify per-pixel features for mask prediction, which leads to suboptimal results for some challenging scenes such as weak defects and cluttered backgrounds. In this paper, inspired by query-based methods, we propose a Wavelet and Prototype Augmented Query-based Transformer (WP-Former) for surface defect detection. Specifically, a set of dynamic queries for mask prediction is updated through the dual-domain transformer decoder. Firstly, a Wavelet-enhanced Cross-Attention (WCA) is proposed, which aggregates meaningful high-and low-frequency information of image features in the wavelet domain to refine queries. WCA enhances the representation of high-frequency components by capturing multi-scale relationships between different frequency components, enabling queries to focus more on defect details. Secondly, a Prototype-guided Cross-Attention (PCA) is proposed to refine queries through meta-prototypes in the spatial domain. The prototypes aggregate semantically meaningful tokens from image features, facilitating queries to aggregate crucial defect information under the cluttered backgrounds. Extensive experiments on three defect detection datasets (i.e., ESDIs-SOD, CrackSeg9k, and ZJU-Leaper) demonstrate that the proposed method achieves state-of-the-art performance in defect detection. The code will be available at https://github.com/yfhdm/WPFormer.
Xiaoheng Jiang, Yang Lu 0016, Jiale Cao, Dong Chen 0017, Mingliang Xu 0001
CVPR5
2025 Logic Distillation: Learning from Code Function by Function for Decision-making Tasks
abstract
Large language models (LLMs) have garnered increasing attention owing to their powerful comprehension and generation capabilities. Generally, larger LLMs (L-LLMs) that require paid interfaces exhibit significantly superior performance compared to smaller LLMs (S-LLMs) that can be deployed on a variety of devices. Knowledge distillation (KD) aims to empower S-LLMs with the capabilities of L-LLMs, while S-LLMs merely mimic the outputs of L-LLMs, failing to get the powerful decision-making capability for new situations. Consequently, S-LLMs are helpless when it comes to continuous decision-making tasks that require logical reasoning. To tackle the identified challenges, we propose a novel framework called Logic Distillation (LD). Initially, LD employs L-LLMs to instantiate complex instructions into discrete functions and illustrates their usage to establish a function base. Subsequently, LD fine-tunes S-LLMs based on the function base to learn the logic employed by L-LLMs in decision-making. During testing, S-LLMs will yield decision-making outcomes, function by function, based on current states. Experiments demonstrate that with the assistance of LD, S-LLMs can achieve outstanding results in continuous decision-making tasks, comparable to, or even surpassing, those of L-LLMs. The code and data for the proposed method are provided for research purposes https://github.com/Anfeather/Logic-Distillation.
Dong Chen 0017, Yueting Zhuang, Siliang Tang, Qidong Liu 0001, Mingliang Xu 0001
IJCAI1
2025 LLAUS: A High-Quality Instruction-Tuned Large Vision Language Assistant for UltraSound
abstract
In recent years, multimodal large models in the medical field have garnered widespread attention. However, this focus has primarily been on CT and MRI imaging, inadvertently neglecting the needs of economically underdeveloped regions and specific populations, such as pregnant women. These groups are often unable to utilize CT and MRI due to their prohibitive costs and potential harm to the body. Meanwhile, ultrasound, an economically viable and very low side effects medical imaging technique, has been largely overlooked by researchers. This study introduces a high-quality instruction-tuned Large vision Language Assistant for UltraSound (LLAUS), designed to answer questions about medical ultrasound images, aiming to assist clinicians in impoverished areas to improve the provision of healthcare services. To address the challenge of missing high-quality ultrasound data, we propose the Adaptive Caption Enhancement(ACE) and Adaptive Caption Optimization (ACO) strategies and have developed a high-quality instruction-following dataset. Subsequently, we fine-tune a Large Vision-Language Model (LVLM) using a novel Zoom-In method. By training on high-quality instruction-following datas, LLAUS demonstrates exceptional multimodal ultrasound communication capabilities, assisting in querying ultrasound images based on open-ended instructions. On tasks related to question-answering and caption generation for ultrasound images, LLAUS exhibits strong performance.
Junhao Guo, XueFeng Shan, Guoming Wang, Dong Chen 0017, Rongxing Lu, Siliang Tang
ICMR4
2025 MIRAGE25: ACM MM25 Multimodal Interleaved Reasoning and Generation Challenge
abstract
We introduce the MIRAGE Challenge, a comprehensive benchmark for multimodal interleaved reasoning and generation, to ACM MM 2025. The challenge aims to evaluate models' abilities to both understand and generate content from complex, multimodal contexts consisting of interlinked images and text. The challenge is accompanied by the MIRAGE Dataset, comprising 263.7K high-quality instruction-response pairs across 35 tasks in two tracks: reasoning and generation. These pairs span 20 diverse scenarios, from surveillance to artistic creation, ensuring broad coverage. The challenge includes seven major categories: Multi-Image Reasoning, Document and Knowledge-Based Understanding, Interactive Multi-Modal Communication, Multi-Image Discrimination, Sequential Visual Generation, Material-based Image Coloring, and Visual Reference Customization. Hosting the MIRAGE Challenge at MM 2025 will drive significant progress in unified multimodal learning and inspire broad involvement in developing more versatile AI systems capable of both understanding and generating multimodal content. Challenge details and participation information are available at https://mm25mirage.github.io/mirage/.
Dong Chen 0017, Zhengqing Hu, Xiaojun Chang
ACM Multimedia1
2025 MM-CARP: Multimodal Model with Cross-Modal Retrieval-Augmented and Visual Region Perception
Junhao Guo, Chenhan Fu, Guoming Wang, Rongxing Lu, Dong Chen 0017, Siliang Tang
MMM (2)5
2025 Improving Vision Anomaly Detection With the Guidance of Language Modality
abstract
Recent years have seen a surge of interest in anomaly detection. However, existing unsupervised anomaly detectors, particularly those for the vision modality, face significant challenges due to redundant information and sparse latent space. In contrast, anomaly detectors demonstrate superior performance in the language modality due to the unimodal nature of the data. This paper tackles the aforementioned challenges for vision modality from a multimodal point of view. Specifically, we propose Cross-modal Guidance (CMG), comprising of Cross-modal Entropy Reduction (CMER) and Cross-modal Linear Embedding (CMLE), to address the issues of redundant information and sparse latent space, respectively. CMER involves masking portions of the raw image and computing the matching score with the corresponding text. Essentially, CMER eliminates irrelevant pixels to direct the detector's focus towards critical content. To learn a more compact latent space for the vision anomaly detection, CMLE learns a correlation structure matrix from the language modality. Then, the acquired matrix compels the distribution of images to resemble that of texts in the latent space. Extensive experiments demonstrate the effectiveness of the proposed methods. Particularly, compared to the baseline that only utilizes images, the performance of CMG has been improved by 16.81%. Ablation experiments further confirm the synergy among the proposed CMER and CMLE, as each component depends on the other to achieve optimal performance.
Dong Chen 0017, Kaihang Pan, Guangyu Dai, Guoming Wang, Yueting Zhuang, Siliang Tang, Mingliang Xu 0001
IEEE Trans. Multim.1
2025 FADngs: Federated Learning for Anomaly Detection
abstract
With the increasing demand for data privacy, federated learning (FL) has gained popularity for various applications. Most existing FL works focus on the classification task, overlooking those scenarios where anomaly detection may also require privacy-preserving. Traditional anomaly detection algorithms cannot be directly applied to the FL setting due to false and missing detection issues. Moreover, with common aggregation methods used in FL (e.g., averaging model parameters), the global model cannot keep the capacities of local models in discriminating anomalies deviating from local distributions, which further degrades the performance. For the aforementioned challenges, we propose Federated Anomaly Detection with Noisy Global Density Estimation, and Self-supervised Ensemble Distillation (FADngs). Specifically, FADngs aligns the knowledge of data distributions from each client by sharing processed density functions. Besides, FADngs trains local models in an improved contrastive learning way that learns more discriminative representations specific for anomaly detection based on the shared density functions. Furthermore, FADngs aggregates capacities by ensemble distillation, which distills the knowledge learned from different distributions to the global model. Our experiments demonstrate that the proposed method significantly outperforms state-of-the-art federated anomaly detection methods. We also empirically show that the shared density function is privacy-preserving. The code for the proposed method is provided for research purposes https://github.com/kanade00/Federated_Anomaly_detection.
Boyu Dong, Dong Chen 0017, Yu Wu 0011, Siliang Tang, Yueting Zhuang
IEEE Trans. Neural Networks Learn. Syst.2
2024 Data Shunt: Collaboration of Small and Large Models for Lower Costs and Better Performance
abstract
Pretrained large models, particularly large language models, have garnered increasing attention, as they have demonstrated remarkable abilities through contextual learning. Pretrained large models are increasingly recognized as fundamental tools for solving various tasks. However, the substantial computational demands of large models have dissuaded most product teams and individuals from running them. In such scenarios, to leverage the exceptional performance of large models, one must solely depend on costly APIs, further burdening product teams and individuals. On the other hand, despite the overall inferior performance of small models compared to large models, there are certain distributions where small models can achieve comparable or even superior results. For instance, during training, small models may become trapped in a local optimum that is unique to certain distributions, leading to superior performance. Hence, we propose Data Shunt (DS), a general paradigm for collaboration of small and large models. DS not only substantially reduces the cost associated with deploying large models but also effectively enhances overall performance. Specifically, DS determines the shunting direction by evaluating the confidence level of small models. When the confidence level falls below a specific threshold, the input data is forwarded to large models. To further leverage the advantages of the small and large models, we introduce Prompt Pruning (PP) and 2-Stage Confidence Distillation (2CD), which facilitate mutual collaboration, leading to better results and less cost. The remarkable performance across diverse modalities and tasks demonstrates the superiority of the proposed DS over large models. For instance, ChatGPT achieves an accuracy of 94.43% on Amazon Product sentiment analysis, and DS achieves an accuracy of 95.64%, while the cost has been reduced to only 31.18%. The code for the proposed method are provided for research purposes https://github.com/Anfeather/Data-Shunt.
Dong Chen 0017, Yueting Zhuang, Shuo Zhang 0014, Jinfeng Liu 0007, Su Dong 0002, Siliang Tang
AAAI1
2023 FedAA: Using Non-sensitive Modalities to Improve Federated Learning while Preserving Image Privacy
abstract
Federated learning aims to train a better global model without sharing the sensitive training samples (usually images) of local clients. Since the sample distributions in local clients tend to be different from each other (i.e., non-IID), one of the major challenges for federated learning is to alleviate model degradation when aggregating local models. The degradation can be attributed to the weight divergence that quantifies the difference of local models from different training processes. Furthermore, non-IID also results in feature space heterogeneity during local training, making neurons of local models in the same location have different functions and further exacerbating weight divergence. In this paper, we demonstrate that the problem can be solved by sharing information from the non-sensitive modality (e.g., metadata, non-sensitive descriptions, etc.) while keeping the sensitive information of images protected. In particular, we propose Federated Learning with Adversarial Example and Adversarial Identifier (FedAA) that trains adversarial examples based on the shared non-sensitive modality to fine-tune local models before global aggregation. The training of local models is enhanced by client identifiers that discriminate the source of inputs to force different local models to get similar outputs and be more homogeneous during the local training. Experiments show that FedAA significantly outperforms recent non-IID federated learning algorithms while preserving image privac, by sharing information from non-sensitive modalities.
Dong Chen 0017, Siliang Tang, Zijin Shen, Guoming Wang, Jun Xiao 0001, Yueting Zhuang, Carl Yang 0001
ACM Multimedia1
2023 Cross-Modal Data Augmentation for Tasks of Different Modalities
abstract
Data augmentation has become one of the keys to alleviating the over-fitting of models on training data and improving the generalization capabilities on testing data. Most existing data augmentation methods only focus on one modality, which is incapable when facing multiple data modalities. Some prior works try to interpolate with random coefficients in the latent space to generate new samples, which can generically work for any data modality. However, these works ignore the extra information conveyed by multimodality data. In fact, the extra information in one modality can provide semantic directions to generate more meaningful samples in another modality. This paper proposes Cross-modal Data Augmentation (CMDA), a simple yet effective data augmentation method to alleviate the over-fitting issue and improve the generalization performance. We evaluate CMDA on unsupervised and supervised tasks of different modalities, on which CMDA consistently and significantly outperforms baselines. For instance, CMDA improves the unsupervised anomaly detection baseline in vision modality from the AUROC$76.46\%, 73.07\%$and 64.36% to$83.25\%, 76.22\%$and 70.57% on three different datasets, respectively. Besides, extensive experiments demonstrate that CMDA is applicable to various neural network architectures. Furthermore, prior methods that interpolate in the latent space need to work with downstream tasks to construct the latent space. In contrast, CMDA can work with or without downstream tasks, which makes the applicability of CMDA more extensive. The source code is publicly available for non-commercial or research use athttps://github.com/Anfeather/CMDA
Dong Chen 0017, Yueting Zhuang, Zijin Shen, Carl Yang 0001, Guoming Wang, Siliang Tang, Yi Yang 0001
IEEE Trans. Multim.1
2022 Robust Meta-learning with Sampling Noise and Label Noise via Eigen-Reptile
abstract
Recent years have seen a surge of interest in meta-learning techniques for tackling the few-shot learning (FSL) problem. However, the meta-learner is prone to overfitting since there are only a few available samples, which can be identified as sampling noise on a clean dataset. Besides, when handling the data with noisy labels, the meta-learner could be extremely sensitive to label noise on a corrupted dataset. To address these two challenges, we present Eigen-Reptile (ER) that updates the meta-parameters with the main direction of historical task-specific parameters. Specifically, the main direction is computed in a fast way, where the scale of the calculated matrix is related to the number of gradient steps for the specific task instead of the number of parameters. Furthermore, to obtain a more accurate main direction for Eigen-Reptile in the presence of many noisy labels, we further propose Introspective Self-paced Learning (ISPL). We have theoretically and experimentally demonstrated the soundness and effectiveness of the proposed Eigen-Reptile and ISPL. Particularly, our experiments on different tasks show that the proposed method is able to outperform or achieve highly competitive performance compared with other gradient-based methods with or without noisy labels. The code and data for the proposed method are provided for research purposes https://github.com/Anfeather/Eigen-Reptile.
Dong Chen 0017, Lingfei Wu 0001, Siliang Tang, Xiao Yun, Bo Long, Yueting Zhuang
ICML1