VLDB 2026 Research / reviewers in the wild / expert
Jiajun Cao
dblp:139/0685
· DBLP profile ↗
11ranked-venue papers
6as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 4 · 3 first-authorSecurity and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Efficient and distributed learning · 37% Vision and language · 23% Autonomous driving · 14% | |
| Software engineering, system software, and programming languages
2 papers |
Software maintenance and evolution · 58% Empirical software engineering · 42% | |
| Network and information security
2 papers |
Systems and software security · 100% |
Topics — the 26 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
model compression |
2.0 | 3 | 2026 | Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs · ICCV 2025 MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders · CVPR 2025 FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning · AAAI 2026 |
Machine learning › Efficient and distributed learning › model compression
token pruning |
1.2 | 2 | 2026 | Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs · ICCV 2025 FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning · AAAI 2026 |
Computer vision › Vision and language
cross-modal alignment |
1.0 | 1 | 2026 | Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation · AAAI 2026 |
Robotics › Autonomous driving
end-to-end driving |
1.0 | 1 | 2026 | FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning · AAAI 2026 |
Natural language and speech › Information extraction and text analysis
keyphrase generation |
1.0 | 1 | 2026 | Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation · AAAI 2026 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
1.0 | 1 | 2026 | Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation · AAAI 2026 |
Robotics › Autonomous driving
perception |
1.0 | 1 | 2026 | FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning · AAAI 2026 |
Robotics › Robot manipulation › embodied foundation models
vision-language-action model |
1.0 | 1 | 2026 | FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning · AAAI 2026 |
Machine learning › Efficient and distributed learning › model compression › token pruning
visual token pruning |
1.0 | 1 | 2026 | FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning · AAAI 2026 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.9 | 1 | 2025 | MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders · CVPR 2025 |
Computer vision › Vision and language
vision-language model |
0.9 | 1 | 2025 | MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders · CVPR 2025 |
Computer vision › Segmentation and scene understanding
medical image segmentation |
0.8 | 1 | 2024 | I-MedSAM: Implicit Medical Image Segmentation with Segment Anything · ECCV (10) 2024 |
Empirical software engineering
mining software repositories |
0.6 | 1 | 2022 | Understanding the Practice of Security Patch Management across Multiple Branches in OSS Projects · WWW 2022 |
Empirical software engineering
open source software |
0.6 | 1 | 2022 | Understanding the Practice of Security Patch Management across Multiple Branches in OSS Projects · WWW 2022 |
Software maintenance and evolution › software updates
security patch management |
0.6 | 1 | 2022 | Understanding the Practice of Security Patch Management across Multiple Branches in OSS Projects · WWW 2022 |
Systems and software security
vulnerability discovery |
0.5 | 1 | 2021 | Locating the Security Patches for Disclosed OSS Vulnerabilities with Vulnerability-Commit Correlation Ranking · CCS 2021 |
Software maintenance and evolution › software updates
security patch identification |
0.5 | 1 | 2021 | Locating the Security Patches for Disclosed OSS Vulnerabilities with Vulnerability-Commit Correlation Ranking · CCS 2021 |
Software maintenance and evolution
vulnerability management |
0.5 | 1 | 2021 | Locating the Security Patches for Disclosed OSS Vulnerabilities with Vulnerability-Commit Correlation Ranking · CCS 2021 |
Machine learning › Trustworthy machine learning › dataset bias
modality bias |
0.3 | 1 | 2026 | Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation · AAAI 2026 |
Machine learning › Trustworthy machine learning
robustness |
0.3 | 1 | 2026 | Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation · AAAI 2026 |
Computer vision › Vision and language › vision-language model
vision-language model inference |
0.3 | 1 | 2025 | Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs · ICCV 2025 |
Computer vision › Segmentation and scene understanding
prompt-based segmentation |
0.2 | 1 | 2024 | I-MedSAM: Implicit Medical Image Segmentation with Segment Anything · ECCV (10) 2024 |
Distributed systems › fault tolerance
checkpointing |
0.2 | 1 | 2014 | Transparent checkpoint-restart over infiniband · HPDC 2014 |
Systems and software security
vulnerability management |
0.2 | 1 | 2022 | Understanding the Practice of Security Patch Management across Multiple Branches in OSS Projects · WWW 2022 |
Parallel and multicore computing
MPI |
0.1 | 1 | 2014 | Transparent checkpoint-restart over infiniband · HPDC 2014 |
Parallel and multicore computing › parallel computing › parallel programming languages
unified parallel c |
0.1 | 1 | 2014 | Transparent checkpoint-restart over infiniband · HPDC 2014 |
Methods — techniques the papers use, named apart from their topics
empirical study · 1.1vulnerability-commit correlation ranking · 1.0progressive modality masking · 1.0masked autoencoder reconstruction · 1.0gradient-based filtering · 1.0adversarial foreground-background reconstruction · 1.0visual cue exploitation · 0.9mixture of experts · 0.9low-rank adaptation · 0.9attention-based distillation · 0.9attention analysis · 0.9implicit neural representation · 0.8system-initiated checkpointing · 0.2kernel module avoidance · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token PruningabstractVision-Language-Action (VLA) models have demonstrated significant potential in complex scene understanding and action reasoning, leading to their increasing adoption in end-to-end autonomous driving systems. However, the long visual tokens of VLA models greatly increase computational costs. Current visual token pruning methods in Vision-Language Models (VLM) rely on either visual token similarity or visual-text attention, but both have shown poor performance in autonomous driving scenarios. Given that human drivers concentrate on relevant foreground areas while driving, we assert that retaining visual tokens containing this foreground information is essential for effective decision-making. Inspired by this, we propose FastDriveVLA, a novel reconstruction-based vision token pruning framework designed specifically for autonomous driving. FastDriveVLA includes a plug-and-play visual token pruner called ReconPruner, which prioritizes foreground information through MAE-style pixel reconstruction. A novel adversarial foreground-background reconstruction strategy is designed to train ReconPruner for the visual encoder of VLA models. Once trained, ReconPruner can be seamlessly applied to different VLA models with the same visual encoder without retraining. To train ReconPruner, we also introduce a large-scale dataset called nuScenes-FG, consisting of 241K image-mask pairs with annotated foreground regions. Our approach achieves state-of-the-art results on the nuScenes open-loop planning benchmark across different pruning ratios. Jiajun Cao, Qizhe Zhang, Peidong Jia, Xiaoan Zhang, Lizhuo, Xiaobao Wei, Sixiang Chen, Liyun Li, Ming Lu 0002, Shanghang Zhang |
AAAI | 1 |
| 2026 | Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase GenerationabstractMultimodal keyphrase generation (MKP) aims to extract a concise set of keyphrases that capture the essential meaning of paired image–text inputs, enabling structured understanding, indexing, and retrieval of multimedia data across the web and social platforms. Success in this task demands effectively bridging the semantic gap between heterogeneous modalities. While multimodal large language models (MLLMs) achieve superior cross-modal understanding by leveraging massive pretraining on image-text corpora, we observe that they often struggle with modality bias and fine-grained intra-modal feature extraction. This oversight leads to a lack of robustness in real-world scenarios where multimedia data is noisy, along with incomplete or misaligned modalities. To address this problem, we propose AimKP, a novel framework that explicitly reinforces intra-modal semantic learning in MLLMs while preserving cross-modal alignment. AimKP incorporates two core innovations: (i) Progressive Modality Masking, which forces fine-grained feature extraction from corrupted inputs by progressively masking modality information during training; (ii) Gradient-based Filtering, that identifies and discards noisy samples, preventing them from corrupting the model’s core cross-modal learning. Extensive experiments validate AimKP’s effectiveness in multimodal keyphrase generation and its robustness across different scenarios. Jiajun Cao, Qinggang Zhang, Yunbo Tang, Zhishang Xiang, Jinsong Su |
AAAI | 1 |
| 2025 | MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual EncodersabstractVisual encoders are fundamental components in vision-language models (VLMs), each showcasing unique strengths derived from various pre-trained visual foundation models. To leverage the various capabilities of these encoders, recent studies incorporate multiple encoders within a single VLM, leading to a considerable increase in computational cost. In this paper, we present Mixture-of-Visual-Encoder Knowledge Distillation (MoVEKD), a novel framework that distills the unique proficiencies of multiple vision encoders into a single, efficient encoder model. Specifically, to mitigate conflicts and retain the unique characteristics of each teacher encoder, we employ low-rank adaptation (LoRA) and mixture-of-experts (MoEs) to selectively activate specialized knowledge based on input features, enhancing both adaptability and efficiency. To regularize the KD process and enhance performance, we propose an attention-based distillation strategy that adaptively weighs the different encoders and emphasizes valuable visual tokens, reducing the burden of replicating comprehensive but distinct features from multiple teachers. Comprehensive experiments on popular VLMs, such as LLaVA and LLaVA-NeXT, validate the effectiveness of our method. Our code is available at: https://github.com/hey-cjj/MoVE-KD. Jiajun Cao, Yuan Zhang 0020, Tao Huang 0020, Ming Lu 0002, Qizhe Zhang, Ruichuan An, Ningning Ma, Shanghang Zhang |
CVPR | 1 |
| 2025 | Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
Qizhe Zhang, Aosong Cheng, Ming Lu 0002, Renrui Zhang, Zhiyong Zhuo, Jiajun Cao, Shaobo Guo, Qi She, Shanghang Zhang |
ICCV | 6 |
| 2024 | I-MedSAM: Implicit Medical Image Segmentation with Segment Anything
Xiaobao Wei, Jiajun Cao, Yizhu Jin, Ming Lu 0002, Shanghang Zhang |
ECCV (10) | 2 |
| 2022 | Understanding the Practice of Security Patch Management across Multiple Branches in OSS ProjectsabstractSince the users of open source software (OSS) projects may not use the latest version all the time, OSS development teams often support code maintenance for old versions through maintaining multiple stable branches. Typically, the developers create a stable branch for each old stable version, deploy security patches on the branch, and release fixed versions at regular intervals. As such, old-version applications in production environments are protected from the disclosed vulnerabilities in a long time. However, the rapidly growing number of OSS vulnerabilities has greatly strained this patch deployment model, and a critical need has arisen for the security community to understand the practice of security patch management across stable branches. In this work, we conduct a large-scale empirical study of stable branches in OSS projects and the security patches deployed on them via investigating 608 stable branches belonging to 26 popular OSS projects as well as more than 2,000 security fixes for 806 CVEs deployed on stable branches. Yuan Zhang 0009, Jiajun Cao, Kun Sun 0001, Mi Zhang 0001, Min Yang 0002 |
WWW | 3 |
| 2021 | Locating the Security Patches for Disclosed OSS Vulnerabilities with Vulnerability-Commit Correlation RankingabstractSecurity patches play an important role in defending against the security threats brought by the increasing OSS vulnerabilities. However, the collection of security patches still remains a challenging problem. Existing works mainly adopt a matching-based design that uses auxiliary information in CVE/NVD to reduce the search scope of patch commits. However, our preliminary study shows that these approaches can only cover a small part of disclosed OSS vulnerabilities (about 12%-53%) even with manual assistance. Yuan Zhang 0009, Chenyuan Mi, Jiajun Cao, Kun Sun 0001, Min Yang 0002 |
CCS | 4 |
| 2019 | Job migration in HPC clusters by means of checkpoint/restart
Manuel Aurelio Rodriguez Pascual, Jiajun Cao, José A. Moríñigo, Gene Cooperman, Rafael Mayo 0001 |
J. Supercomput. | 2 |
| 2016 | System-Level Scalable Checkpoint-Restart for Petascale ComputingabstractFault tolerance for the upcoming exascale generation has long been an area of active research. One of the components of a fault tolerance strategy is checkpointing. Petascale-level checkpointing is demonstrated through a new mechanism for virtualization of the InfiniBand UD (unreliable datagram) mode, and for updating the remote address on each UD-based send, due to lack of a fixed peer. Note that InfiniBand UD is required to support modern MPI implementations. An extrapolation from the current results to future SSD-based storage systems provides evidence that the current approach will remain practical in the exascale generation. This transparent checkpointing approach is evaluated using a framework of the DMTCP checkpointing package. Results are shown for HPCG (linear algebra), NAMD (molecular dynamics), and the NAS NPB benchmarks. In tests up to 32,752 MPI processes on 32,752 CPU cores, checkpointing of a computation with a 38 TB memory footprint in 11 minutes is demonstrated. Runtime overhead is reduced to less than 1%. The approach is also evaluated across three widely used MPI implementations. Jiajun Cao, Kapil Arya, Rohan Garg 0001, L. Shawn Matott, Dhabaleswar K. Panda 0001, Hari Subramoni, Jérôme Vienne, Gene Cooperman |
ICPADS | 1 |
| 2015 | Checkpointing as a Service in Heterogeneous Cloud EnvironmentsabstractA non-invasive, cloud-agnostic approach is demonstrated for extending existing cloud platforms to include checkpoint-restart capability. Most cloud platforms currently rely on each application to provide its own fault tolerance. A uniform mechanism within the cloud itself serves two purposes: (a) direct support for long-running jobs, which would otherwise require a custom fault-tolerant mechanism for each application, and (b) the administrative capability to manage an over-subscribed cloud by temporarily swapping out jobs when higher priority jobs arrive. An advantage of this uniform approach is that it also supports parallel and distributed computations, over both TCP and InfiniBand, thus allowing traditional HPC applications to take advantage of an existing cloud infrastructure. Additionally, an integrated health-monitoring mechanism detects when long-running jobs either fail or incur exceptionally low performance, perhaps due to resource starvation, and proactively suspends the job. The cloud-agnostic feature is demonstrated by applying the implementation to two very different cloud platforms: Snooze and Open Stack. The use of a cloud-agnostic architecture also enables, for the first time, migration of applications from one cloud platform to another. Jiajun Cao, Matthieu Simonin, Gene Cooperman, Christine Morin |
CCGRID | 1 |
| 2014 | Transparent checkpoint-restart over infinibandabstractTransparently saving the state of the InfiniBand network as part of distributed checkpointing has been a long-standing challenge for researchers. The lack of a solution has forced typical MPI implementations to include custom checkpoint-restart services that "tear down" the network, checkpoint each node in isolation, and then re-connect the network again. This work presents the first example of transparent, system-initiated checkpoint-restart that directly supports InfiniBand. The new approach simplifies current practice by avoiding the need for a privileged kernel module. The generality of this approach is demonstrated by applying it both to MPI and to Berkeley UPC (Unified Parallel C), in its native mode (without MPI). Scalability is shown by checkpointing 2,048 MPI processes across 128 nodes (with 16 cores per node). The run-time overhead varies between 0.8% and 1.7%. While checkpoint times dominate, the network-only portion of the implementation is shown to require less than 100 milliseconds (not including the time to locally write application memory to stable storage). Jiajun Cao, Gregory Kerr, Kapil Arya, Gene Cooperman |
HPDC | 1 |