Jianjian Cao

dblp:261/3612 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-1473-3956ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Efficient and distributed learning · 53% Generative modeling · 16% Language models and text generation · 13%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 18 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
2.942026
MADTP++: Bridge the Gap Between Token and Weight Pruning for Accelerating VLTs · IEEE Trans. Pattern Anal. Mach. Intell. 2026
DeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts Models · CVPR 2025
MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer · CVPR 2024
Machine learning › Efficient and distributed learning › model compression
token pruning
1.822026
MADTP++: Bridge the Gap Between Token and Weight Pruning for Accelerating VLTs · IEEE Trans. Pattern Anal. Mach. Intell. 2026
MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer · CVPR 2024
Machine learning › Generative modeling
diffusion model
1.012026
Δ-DiT: Accelerating Diffusion Transformers without Training via Denoising Property Alignment · Int. J. Comput. Vis. 2026
Machine learning › Generative modeling › diffusion model › diffusion model acceleration
diffusion transformer acceleration
1.012026
Δ-DiT: Accelerating Diffusion Transformers without Training via Denoising Property Alignment · Int. J. Comput. Vis. 2026
Machine learning › Reinforcement learning › exploration
exploration-exploitation tradeoff
1.012026
A Scalable Multi-LLM Collaboration System with Retrieval-based Selection and Exploration-Exploitation-Driven Enhancement · ACL (1) 2026
Natural language and speech › Language models and text generation › LLM agents
LLM collaboration
1.012026
A Scalable Multi-LLM Collaboration System with Retrieval-based Selection and Exploration-Exploitation-Driven Enhancement · ACL (1) 2026
Machine learning › Efficient and distributed learning › model compression › pruning
weight pruning
1.012026
MADTP++: Bridge the Gap Between Token and Weight Pruning for Accelerating VLTs · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Image and video processing
image enhancement
1.012026
Mamba-Based Progressive-Recovery Framework for Multimodal Low Light Image Enhancement · IEEE Trans. Multim. 2026
Image and video processing
image restoration
1.012026
Mamba-Based Progressive-Recovery Framework for Multimodal Low Light Image Enhancement · IEEE Trans. Multim. 2026
Image and video processing › image enhancement
low-light image enhancement
1.012026
Mamba-Based Progressive-Recovery Framework for Multimodal Low Light Image Enhancement · IEEE Trans. Multim. 2026
Machine learning › Efficient and distributed learning › model compression › parameter compression
mixture-of-experts compression
0.912025
DeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts Models · CVPR 2025
Natural language and speech › Language models and text generation
instruction tuning
0.712023
LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark · NeurIPS 2023
Computer vision › Vision and language › vision-language model
multimodal large language model
0.712023
LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark · NeurIPS 2023
Machine learning › Efficient and distributed learning › inference acceleration
training-free acceleration
0.312026
Δ-DiT: Accelerating Diffusion Transformers without Training via Denoising Property Alignment · Int. J. Comput. Vis. 2026
GPUs and heterogeneous computing › GPU computing › tensor cores
sparse tensor core
0.312026
MADTP++: Bridge the Gap Between Token and Weight Pruning for Accelerating VLTs · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Machine learning › Deep learning architectures and training
mixture of experts
0.312025
DeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts Models · CVPR 2025
Computer vision › Vision and language › vision-language model › vision-language model architecture
vision-language transformer
0.212024
MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer · CVPR 2024
Computer vision › Vision and language
multimodal benchmark
0.212023
LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

knowledge distillation · 2.0cooperative optimization · 2.0retrieval-based selection · 1.0progressive recovery · 1.0mamba · 1.0exploration-exploitation · 1.0denoising property alignment · 1.0sparsification · 0.9quantization · 0.9low-rank approximation · 0.9decomposition · 0.9dynamic token pruning · 0.8
YearPublicationVenuePosition
2026 A Scalable Multi-LLM Collaboration System with Retrieval-based Selection and Exploration-Exploitation-Driven Enhancement
abstract
Shengji Tang, Jianjian Cao, Weihao Lin, Jiale Hong, Bo Zhang, Shuyue Hu, Lei Bai, Tao Chen, Wanli Ouyang, Peng Ye. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Shengji Tang, Jianjian Cao, Weihao Lin 0002, Jiale Hong, Bo Zhang 0069, Shuyue Hu, Lei Bai 0001, Tao Chen 0003, Wanli Ouyang, Peng Ye 0006
ACL (1)2
2026 Δ-DiT: Accelerating Diffusion Transformers without Training via Denoising Property Alignment
Pengtao Chen, Mingzhu Shen, Peng Ye 0006, Jianjian Cao, Chongjun Tu, Christos-Savvas Bouganis, Tao Chen 0003
Int. J. Comput. Vis.4
2026 MADTP++: Bridge the Gap Between Token and Weight Pruning for Accelerating VLTs
abstract
Vision-Language Transformers (VLTs) have achieved remarkable success, yet their high computational costs remain challenging due to numerous input tokens and large model parameters. Existing VLT compression methods primarily rely on single-modality-based token pruning or coarse-grained weight pruning techniques. However, these methods face significant obstacles, such as ignoring the critical alignment of different modalities and lacking layer-wise dynamic token pruning flexibility, exhibiting inevitable performance degradation due to coarsegrained weight pruning, and struggling with the simultaneous compression of both input tokens and model parameters. To address those limitations, we propose MADTP++, a novel approach that integrates custom-made token and weight pruning processes into a unified framework, achieving superior compression in both parameter counts and computational costs. Specifically, for the token pruning process, we introduce the Multi-modality Alignment Guidance (MAG) module and the Dynamic Token Pruning (DTP) module to align semantic features across different modalities and guide the dynamic elimination of redundant tokens based on different input instances. For the weight pruning process, we propose a Hardware-aware Weight Pruning (HWP) module that leverages the Sparse Tensor Cores across diverse hardware setups to enable fine-grained parameter pruning within VLTs. To further unify token and weight pruning, we also propose a Cooperative Optimization Training Strategy that automatically allocates GFLOPs and parameter reductions per branch before pruning and employs Knowledge Distillation Constraints to facilitate joint optimization of both pruning dimensions. Extensive experiments conducted on various VLT models and datasets demonstrate that MADTP++ can significantly reduce model parameters and computational costs while maintaining competitive performance.
Jianjian Cao, Chong Yu 0001, Peng Ye 0006, Tao Chen 0003
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 Mamba-Based Progressive-Recovery Framework for Multimodal Low Light Image Enhancement
Mohammad Mahdizadeh, Jianjian Cao, Peng Ye 0006, Tao Chen 0003
IEEE Trans. Multim.2
2025 DeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts Models
abstract
Upcycled Mixture-of-Experts (MoE) models have shown great potential in various tasks by converting the original Feed-Forward Network (FFN) layers in pre-trained dense models into MoE layers. However, these models still suffer from significant parameter inefficiency due to the introduction of multiple experts. In this work, we propose a novel DeRS (Decompose, Replace, and Synthesis) paradigm to overcome this shortcoming, which is motivated by our observations about the unique redundancy mechanisms of upcycled MoE experts. Specifically, DeRS decomposes the experts into one expert-shared base weight and multiple expert-specific delta weights, and subsequently represents these delta weights in lightweight forms. Our proposed DeRS paradigm can be applied to enhance parameter efficiency in two different scenarios, including: 1) DeRS Compression for inference stage, using sparsification or quantization to compress vanilla upcycled MoE models; and 2) DeRS Up-cycling for training stage, employing lightweight sparse or low-rank matrixes to efficiently upcycle dense models into MoE models. Extensive experiments across three different tasks show that the proposed methods can achieve extreme parameter efficiency while maintaining the performance for both training and compression of upcycled MoE models.
Yongqi Huang, Peng Ye 0006, Chenyu Huang 0001, Jianjian Cao, Lin Zhang 0055, Baopu Li, Gang Yu 0002, Tao Chen 0003
CVPR4
2025 ClipSAM: CLIP and SAM collaboration for zero-shot anomaly segmentation
Shengze Li, Jianjian Cao, Peng Ye 0006, Yuhan Ding, Chongjun Tu, Tao Chen 0003
Neurocomputing2
2025 Bilateral Cross-Modality Graph Matching Attention for Feature Fusion in Visual Question Answering
abstract
Answering semantically complicated questions according to an image is challenging in a visual question answering (VQA) task. Although the image can be well represented by deep learning, the question is always simply embedded and cannot well indicate its meaning. Besides, the visual and textual features have a gap for different modalities, it is difficult to align and utilize the cross-modality information. In this article, we focus on these two problems and propose a graph matching attention (GMA) network. First, it not only builds graph for the image but also constructs graph for the question in terms of both syntactic and embedding information. Next, we explore the intramodality relationships by a dual-stage graph encoder and then present a bilateral cross-modality GMA to infer the relationships between the image and the question. The updated cross-modality features are then sent into the answer prediction module for final answer prediction. Experiments demonstrate that our network achieves the state-of-the-art performance on the GQA dataset and the VQA 2.0 dataset. The ablation studies verify the effectiveness of each module in our GMA network.
Jianjian Cao, Xiameng Qin, Sanyuan Zhao, Jianbing Shen
IEEE Trans. Neural Networks Learn. Syst.1
2024 MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer
abstract
Vision-Language Transformers (VLTs) have shown great success recently, but are meanwhile accompanied by heavy computation costs, where a major reason can be attributed to the large number of visual and language tokens. Existing token pruning research for compressing VLTs mainly follows a single-modality-based scheme yet ignores the critical role of aligning different modalities for guiding the token pruning process, causing the important tokens for one modality to be falsely pruned in another modality branch. Meanwhile, existing VLT pruning works also lack the flexibility to dynamically compress each layer based on different input samples. To this end, we propose a novel framework named Multimodal Alignment-Guided Dynamic Token Pruning (MADTP) for accelerating various VLTs. Specifically, we first introduce a well-designed Multi-modality Alignment Guidance (MAG) module that can align features of the same semantic concept from different modalities, to ensure the pruned tokens are less important for all modalities. We further design a novel Dynamic Token Pruning (DTP) module, which can adaptively adjust the token compression ratio in each layer based on different input instances. Extensive experiments on various benchmarks demonstrate that MADTP significantly reduces the computational complexity of kinds of multimodal models while preserving competitive performance. Notably, when applied to the BLIP model in the NLVR2 dataset, MADTP can reduce the GFLOPs by 80% with less than 4% performance degradation. The code is available at https://github.com/double125IMADTP.
Jianjian Cao, Peng Ye 0006, Shengze Li, Chong Yu 0001, Yansong Tang, Jiwen Lu, Tao Chen 0003
CVPR1
2023 JNDMix: Jnd-Based Data Augmentation for No-Reference Image Quality Assessment
abstract
Despite substantial progress in no-reference image quality assessment (NR-IQA), previous training models often suffer from over-fitting due to the limited scale of used datasets, resulting in model performance bottlenecks. To tackle this challenge, we explore the potential of leveraging data augmentation to improve data efficiency and enhance model robustness. However, most existing data augmentation methods incur a serious issue, namely that it alters the image quality and leads to training images mismatching with their original labels. Additionally, although only a few data augmentation methods are available for NR-IQA task, their ability to enrich dataset diversity is still insufficient. To address these issues, we propose a effective and general data augmentation based on just noticeable difference (JND) noise mixing for NR-IQA task, named JNDMix. In detail, we randomly inject the JND noise, imperceptible to the human visual system (HVS), into the training image without any adjustment to its label. Extensive experiments demonstrate that JNDMix significantly improves the performance and data efficiency of various state-of-the-art NR-IQA models and the commonly used baseline models, as well as the generalization ability. More importantly, JNDMix facilitates MANIQA to achieve the state-of-the-art performance on LIVEC and KonIQ-10k.
Jiamu Sheng, Jiayuan Fan 0001, Peng Ye 0006, Jianjian Cao
ICASSP4
2023 A2S-NAS: Asymmetric Spectral-Spatial Neural Architecture Search for Hyperspectral Image Classification
abstract
Existing deep learning-based hyperspectral image (HSI) classification works still suffer from the limitation of the fixed-sized receptive field, leading to difficulties in distinctive spectral-spatial features for ground objects with various sizes and arbitrary shapes. Meanwhile, plenty of previous works ignore asymmetric spectral-spatial dimensions in HSI. To address the above issues, we propose a multi-stage search architecture in order to overcome asymmetric spectral-spatial dimensions and capture significant features. First, the asymmetric pooling on the spectral-spatial dimension maximally retains the essential features of HSI. Then, the 3D convolution with a selectable range of receptive fields overcomes the constraints of fixed-sized convolution kernels. Finally, we extend these two searchable operations to different layers of each stage to build the final architecture. Extensive experiments are conducted on two challenging HSI benchmarks including Indian Pines and Houston University, and results demonstrate the effectiveness of the proposed method with superior performance compared with the related works.
Lin Zhan, Jiayuan Fan 0001, Peng Ye 0006, Jianjian Cao
ICASSP4
2023 LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark
abstract
Large language models have emerged as a promising approach towards achieving general-purpose AI agents. The thriving open-source LLM community has greatly accelerated the development of agents that support human-machine dialogue interaction through natural language processing. However, human interaction with the world extends beyond only text as a modality, and other modalities such as vision are also crucial. Recent works on multi-modal large language models, such as GPT-4V and Bard, have demonstrated their effectiveness in handling visual modalities. However, the transparency of these works is limited and insufficient to support academic research. To the best of our knowledge, we present one of the very first open-source endeavors in the field, LAMM, encompassing a Language-Assisted Multi-Modal instruction tuning dataset, framework, and benchmark. Our aim is to establish LAMM as a growing ecosystem for training and evaluating MLLMs, with a specific focus on facilitating AI agents capable of bridging the gap between ideas and execution, thereby enabling seamless human-AI interaction. Our main contribution is three-fold: 1) We present a comprehensive dataset and benchmark, which cover a wide range of vision tasks for 2D and 3D vision. Extensive experiments validate the effectiveness of our dataset and benchmark. 2) We outline the detailed methodology of constructing multi-modal instruction tuning datasets and benchmarks for MLLMs, enabling rapid scaling and extension of MLLM research to diverse domains, tasks, and modalities. 3) We provide a primary but potential MLLM training framework optimized for modality extension. We also provide baseline models, comprehensive experimental observations, and analysis to accelerate future research. Our baseline model is trained within 24 A100 GPU hours, framework supports training with V100 and RTX3090 is available thanks to the open-source society. Codes and data are now available at https://openlamm.github.io.
Zhenfei Yin, Jianjian Cao, Zhelun Shi, Dingning Liu, Mukai Li, Xiaoshui Huang, Zhiyong Wang 0001, Lu Sheng, Lei Bai 0001, Wanli Ouyang
NeurIPS3