VLDB 2026 Research / reviewers in the wild / expert
Jianjian Cao
dblp:261/3612
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-1473-3956ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Efficient and distributed learning · 53% Generative modeling · 16% Language models and text generation · 13% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% |
Topics — the 18 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
model compression |
2.9 | 4 | 2026 | MADTP++: Bridge the Gap Between Token and Weight Pruning for Accelerating VLTs · IEEE Trans. Pattern Anal. Mach. Intell. 2026 DeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts Models · CVPR 2025 MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer · CVPR 2024 |
Machine learning › Efficient and distributed learning › model compression
token pruning |
1.8 | 2 | 2026 | MADTP++: Bridge the Gap Between Token and Weight Pruning for Accelerating VLTs · IEEE Trans. Pattern Anal. Mach. Intell. 2026 MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer · CVPR 2024 |
Machine learning › Generative modeling
diffusion model |
1.0 | 1 | 2026 | Δ-DiT: Accelerating Diffusion Transformers without Training via Denoising Property Alignment · Int. J. Comput. Vis. 2026 |
Machine learning › Generative modeling › diffusion model › diffusion model acceleration
diffusion transformer acceleration |
1.0 | 1 | 2026 | Δ-DiT: Accelerating Diffusion Transformers without Training via Denoising Property Alignment · Int. J. Comput. Vis. 2026 |
Machine learning › Reinforcement learning › exploration
exploration-exploitation tradeoff |
1.0 | 1 | 2026 | A Scalable Multi-LLM Collaboration System with Retrieval-based Selection and Exploration-Exploitation-Driven Enhancement · ACL (1) 2026 |
Natural language and speech › Language models and text generation › LLM agents
LLM collaboration |
1.0 | 1 | 2026 | A Scalable Multi-LLM Collaboration System with Retrieval-based Selection and Exploration-Exploitation-Driven Enhancement · ACL (1) 2026 |
Machine learning › Efficient and distributed learning › model compression › pruning
weight pruning |
1.0 | 1 | 2026 | MADTP++: Bridge the Gap Between Token and Weight Pruning for Accelerating VLTs · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Image and video processing
image enhancement |
1.0 | 1 | 2026 | Mamba-Based Progressive-Recovery Framework for Multimodal Low Light Image Enhancement · IEEE Trans. Multim. 2026 |
Image and video processing
image restoration |
1.0 | 1 | 2026 | Mamba-Based Progressive-Recovery Framework for Multimodal Low Light Image Enhancement · IEEE Trans. Multim. 2026 |
Image and video processing › image enhancement
low-light image enhancement |
1.0 | 1 | 2026 | Mamba-Based Progressive-Recovery Framework for Multimodal Low Light Image Enhancement · IEEE Trans. Multim. 2026 |
Machine learning › Efficient and distributed learning › model compression › parameter compression
mixture-of-experts compression |
0.9 | 1 | 2025 | DeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts Models · CVPR 2025 |
Natural language and speech › Language models and text generation
instruction tuning |
0.7 | 1 | 2023 | LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark · NeurIPS 2023 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.7 | 1 | 2023 | LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark · NeurIPS 2023 |
Machine learning › Efficient and distributed learning › inference acceleration
training-free acceleration |
0.3 | 1 | 2026 | Δ-DiT: Accelerating Diffusion Transformers without Training via Denoising Property Alignment · Int. J. Comput. Vis. 2026 |
GPUs and heterogeneous computing › GPU computing › tensor cores
sparse tensor core |
0.3 | 1 | 2026 | MADTP++: Bridge the Gap Between Token and Weight Pruning for Accelerating VLTs · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Machine learning › Deep learning architectures and training
mixture of experts |
0.3 | 1 | 2025 | DeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts Models · CVPR 2025 |
Computer vision › Vision and language › vision-language model › vision-language model architecture
vision-language transformer |
0.2 | 1 | 2024 | MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer · CVPR 2024 |
Computer vision › Vision and language
multimodal benchmark |
0.2 | 1 | 2023 | LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
knowledge distillation · 2.0cooperative optimization · 2.0retrieval-based selection · 1.0progressive recovery · 1.0mamba · 1.0exploration-exploitation · 1.0denoising property alignment · 1.0sparsification · 0.9quantization · 0.9low-rank approximation · 0.9decomposition · 0.9dynamic token pruning · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Scalable Multi-LLM Collaboration System with Retrieval-based Selection and Exploration-Exploitation-Driven EnhancementabstractShengji Tang, Jianjian Cao, Weihao Lin, Jiale Hong, Bo Zhang, Shuyue Hu, Lei Bai, Tao Chen, Wanli Ouyang, Peng Ye. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shengji Tang, Jianjian Cao, Weihao Lin 0002, Jiale Hong, Bo Zhang 0069, Shuyue Hu, Lei Bai 0001, Tao Chen 0003, Wanli Ouyang, Peng Ye 0006 |
ACL (1) | 2 |
| 2026 | Δ-DiT: Accelerating Diffusion Transformers without Training via Denoising Property Alignment
Pengtao Chen, Mingzhu Shen, Peng Ye 0006, Jianjian Cao, Chongjun Tu, Christos-Savvas Bouganis, Tao Chen 0003 |
Int. J. Comput. Vis. | 4 |
| 2026 | MADTP++: Bridge the Gap Between Token and Weight Pruning for Accelerating VLTsabstractVision-Language Transformers (VLTs) have achieved remarkable success, yet their high computational costs remain challenging due to numerous input tokens and large model parameters. Existing VLT compression methods primarily rely on single-modality-based token pruning or coarse-grained weight pruning techniques. However, these methods face significant obstacles, such as ignoring the critical alignment of different modalities and lacking layer-wise dynamic token pruning flexibility, exhibiting inevitable performance degradation due to coarsegrained weight pruning, and struggling with the simultaneous compression of both input tokens and model parameters. To address those limitations, we propose MADTP++, a novel approach that integrates custom-made token and weight pruning processes into a unified framework, achieving superior compression in both parameter counts and computational costs. Specifically, for the token pruning process, we introduce the Multi-modality Alignment Guidance (MAG) module and the Dynamic Token Pruning (DTP) module to align semantic features across different modalities and guide the dynamic elimination of redundant tokens based on different input instances. For the weight pruning process, we propose a Hardware-aware Weight Pruning (HWP) module that leverages the Sparse Tensor Cores across diverse hardware setups to enable fine-grained parameter pruning within VLTs. To further unify token and weight pruning, we also propose a Cooperative Optimization Training Strategy that automatically allocates GFLOPs and parameter reductions per branch before pruning and employs Knowledge Distillation Constraints to facilitate joint optimization of both pruning dimensions. Extensive experiments conducted on various VLT models and datasets demonstrate that MADTP++ can significantly reduce model parameters and computational costs while maintaining competitive performance. Jianjian Cao, Chong Yu 0001, Peng Ye 0006, Tao Chen 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Mamba-Based Progressive-Recovery Framework for Multimodal Low Light Image Enhancement
Mohammad Mahdizadeh, Jianjian Cao, Peng Ye 0006, Tao Chen 0003 |
IEEE Trans. Multim. | 2 |
| 2025 | DeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts ModelsabstractUpcycled Mixture-of-Experts (MoE) models have shown great potential in various tasks by converting the original Feed-Forward Network (FFN) layers in pre-trained dense models into MoE layers. However, these models still suffer from significant parameter inefficiency due to the introduction of multiple experts. In this work, we propose a novel DeRS (Decompose, Replace, and Synthesis) paradigm to overcome this shortcoming, which is motivated by our observations about the unique redundancy mechanisms of upcycled MoE experts. Specifically, DeRS decomposes the experts into one expert-shared base weight and multiple expert-specific delta weights, and subsequently represents these delta weights in lightweight forms. Our proposed DeRS paradigm can be applied to enhance parameter efficiency in two different scenarios, including: 1) DeRS Compression for inference stage, using sparsification or quantization to compress vanilla upcycled MoE models; and 2) DeRS Up-cycling for training stage, employing lightweight sparse or low-rank matrixes to efficiently upcycle dense models into MoE models. Extensive experiments across three different tasks show that the proposed methods can achieve extreme parameter efficiency while maintaining the performance for both training and compression of upcycled MoE models. Yongqi Huang, Peng Ye 0006, Chenyu Huang 0001, Jianjian Cao, Lin Zhang 0055, Baopu Li, Gang Yu 0002, Tao Chen 0003 |
CVPR | 4 |
| 2025 | ClipSAM: CLIP and SAM collaboration for zero-shot anomaly segmentation
Shengze Li, Jianjian Cao, Peng Ye 0006, Yuhan Ding, Chongjun Tu, Tao Chen 0003 |
Neurocomputing | 2 |
| 2025 | Bilateral Cross-Modality Graph Matching Attention for Feature Fusion in Visual Question AnsweringabstractAnswering semantically complicated questions according to an image is challenging in a visual question answering (VQA) task. Although the image can be well represented by deep learning, the question is always simply embedded and cannot well indicate its meaning. Besides, the visual and textual features have a gap for different modalities, it is difficult to align and utilize the cross-modality information. In this article, we focus on these two problems and propose a graph matching attention (GMA) network. First, it not only builds graph for the image but also constructs graph for the question in terms of both syntactic and embedding information. Next, we explore the intramodality relationships by a dual-stage graph encoder and then present a bilateral cross-modality GMA to infer the relationships between the image and the question. The updated cross-modality features are then sent into the answer prediction module for final answer prediction. Experiments demonstrate that our network achieves the state-of-the-art performance on the GQA dataset and the VQA 2.0 dataset. The ablation studies verify the effectiveness of each module in our GMA network. Jianjian Cao, Xiameng Qin, Sanyuan Zhao, Jianbing Shen |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language TransformerabstractVision-Language Transformers (VLTs) have shown great success recently, but are meanwhile accompanied by heavy computation costs, where a major reason can be attributed to the large number of visual and language tokens. Existing token pruning research for compressing VLTs mainly follows a single-modality-based scheme yet ignores the critical role of aligning different modalities for guiding the token pruning process, causing the important tokens for one modality to be falsely pruned in another modality branch. Meanwhile, existing VLT pruning works also lack the flexibility to dynamically compress each layer based on different input samples. To this end, we propose a novel framework named Multimodal Alignment-Guided Dynamic Token Pruning (MADTP) for accelerating various VLTs. Specifically, we first introduce a well-designed Multi-modality Alignment Guidance (MAG) module that can align features of the same semantic concept from different modalities, to ensure the pruned tokens are less important for all modalities. We further design a novel Dynamic Token Pruning (DTP) module, which can adaptively adjust the token compression ratio in each layer based on different input instances. Extensive experiments on various benchmarks demonstrate that MADTP significantly reduces the computational complexity of kinds of multimodal models while preserving competitive performance. Notably, when applied to the BLIP model in the NLVR2 dataset, MADTP can reduce the GFLOPs by 80% with less than 4% performance degradation. The code is available at https://github.com/double125IMADTP. Jianjian Cao, Peng Ye 0006, Shengze Li, Chong Yu 0001, Yansong Tang, Jiwen Lu, Tao Chen 0003 |
CVPR | 1 |
| 2023 | JNDMix: Jnd-Based Data Augmentation for No-Reference Image Quality AssessmentabstractDespite substantial progress in no-reference image quality assessment (NR-IQA), previous training models often suffer from over-fitting due to the limited scale of used datasets, resulting in model performance bottlenecks. To tackle this challenge, we explore the potential of leveraging data augmentation to improve data efficiency and enhance model robustness. However, most existing data augmentation methods incur a serious issue, namely that it alters the image quality and leads to training images mismatching with their original labels. Additionally, although only a few data augmentation methods are available for NR-IQA task, their ability to enrich dataset diversity is still insufficient. To address these issues, we propose a effective and general data augmentation based on just noticeable difference (JND) noise mixing for NR-IQA task, named JNDMix. In detail, we randomly inject the JND noise, imperceptible to the human visual system (HVS), into the training image without any adjustment to its label. Extensive experiments demonstrate that JNDMix significantly improves the performance and data efficiency of various state-of-the-art NR-IQA models and the commonly used baseline models, as well as the generalization ability. More importantly, JNDMix facilitates MANIQA to achieve the state-of-the-art performance on LIVEC and KonIQ-10k. Jiamu Sheng, Jiayuan Fan 0001, Peng Ye 0006, Jianjian Cao |
ICASSP | 4 |
| 2023 | A2S-NAS: Asymmetric Spectral-Spatial Neural Architecture Search for Hyperspectral Image ClassificationabstractExisting deep learning-based hyperspectral image (HSI) classification works still suffer from the limitation of the fixed-sized receptive field, leading to difficulties in distinctive spectral-spatial features for ground objects with various sizes and arbitrary shapes. Meanwhile, plenty of previous works ignore asymmetric spectral-spatial dimensions in HSI. To address the above issues, we propose a multi-stage search architecture in order to overcome asymmetric spectral-spatial dimensions and capture significant features. First, the asymmetric pooling on the spectral-spatial dimension maximally retains the essential features of HSI. Then, the 3D convolution with a selectable range of receptive fields overcomes the constraints of fixed-sized convolution kernels. Finally, we extend these two searchable operations to different layers of each stage to build the final architecture. Extensive experiments are conducted on two challenging HSI benchmarks including Indian Pines and Houston University, and results demonstrate the effectiveness of the proposed method with superior performance compared with the related works. Lin Zhan, Jiayuan Fan 0001, Peng Ye 0006, Jianjian Cao |
ICASSP | 4 |
| 2023 | LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and BenchmarkabstractLarge language models have emerged as a promising approach towards achieving general-purpose AI agents. The thriving open-source LLM community has greatly accelerated the development of agents that support human-machine dialogue interaction through natural language processing. However, human interaction with the world extends beyond only text as a modality, and other modalities such as vision are also crucial. Recent works on multi-modal large language models, such as GPT-4V and Bard, have demonstrated their effectiveness in handling visual modalities. However, the transparency of these works is limited and insufficient to support academic research. To the best of our knowledge, we present one of the very first open-source endeavors in the field, LAMM, encompassing a Language-Assisted Multi-Modal instruction tuning dataset, framework, and benchmark. Our aim is to establish LAMM as a growing ecosystem for training and evaluating MLLMs, with a specific focus on facilitating AI agents capable of bridging the gap between ideas and execution, thereby enabling seamless human-AI interaction. Our main contribution is three-fold: 1) We present a comprehensive dataset and benchmark, which cover a wide range of vision tasks for 2D and 3D vision. Extensive experiments validate the effectiveness of our dataset and benchmark. 2) We outline the detailed methodology of constructing multi-modal instruction tuning datasets and benchmarks for MLLMs, enabling rapid scaling and extension of MLLM research to diverse domains, tasks, and modalities. 3) We provide a primary but potential MLLM training framework optimized for modality extension. We also provide baseline models, comprehensive experimental observations, and analysis to accelerate future research. Our baseline model is trained within 24 A100 GPU hours, framework supports training with V100 and RTX3090 is available thanks to the open-source society. Codes and data are now available at https://openlamm.github.io. Zhenfei Yin, Jianjian Cao, Zhelun Shi, Dingning Liu, Mukai Li, Xiaoshui Huang, Zhiyong Wang 0001, Lu Sheng, Lei Bai 0001, Wanli Ouyang |
NeurIPS | 3 |