Jiajun Cao

dblp:139/0685 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 4 · 3 first-authorSecurity and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Efficient and distributed learning · 37% Vision and language · 23% Autonomous driving · 14%
Software engineering, system software, and programming languages
2 papers
Software maintenance and evolution · 58% Empirical software engineering · 42%
Network and information security
2 papers
Systems and software security · 100%

Topics — the 26 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
2.032026
Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs · ICCV 2025
MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders · CVPR 2025
FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning · AAAI 2026
Machine learning › Efficient and distributed learning › model compression
token pruning
1.222026
Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs · ICCV 2025
FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning · AAAI 2026
Computer vision › Vision and language
cross-modal alignment
1.012026
Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation · AAAI 2026
Robotics › Autonomous driving
end-to-end driving
1.012026
FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning · AAAI 2026
Natural language and speech › Information extraction and text analysis
keyphrase generation
1.012026
Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation · AAAI 2026
Computer vision › Vision and language › vision-language model
multimodal large language model
1.012026
Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation · AAAI 2026
Robotics › Autonomous driving
perception
1.012026
FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning · AAAI 2026
Robotics › Robot manipulation › embodied foundation models
vision-language-action model
1.012026
FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning · AAAI 2026
Machine learning › Efficient and distributed learning › model compression › token pruning
visual token pruning
1.012026
FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning · AAAI 2026
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.912025
MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders · CVPR 2025
Computer vision › Vision and language
vision-language model
0.912025
MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders · CVPR 2025
Computer vision › Segmentation and scene understanding
medical image segmentation
0.812024
I-MedSAM: Implicit Medical Image Segmentation with Segment Anything · ECCV (10) 2024
Empirical software engineering
mining software repositories
0.612022
Understanding the Practice of Security Patch Management across Multiple Branches in OSS Projects · WWW 2022
Empirical software engineering
open source software
0.612022
Understanding the Practice of Security Patch Management across Multiple Branches in OSS Projects · WWW 2022
Software maintenance and evolution › software updates
security patch management
0.612022
Understanding the Practice of Security Patch Management across Multiple Branches in OSS Projects · WWW 2022
Systems and software security
vulnerability discovery
0.512021
Locating the Security Patches for Disclosed OSS Vulnerabilities with Vulnerability-Commit Correlation Ranking · CCS 2021
Software maintenance and evolution › software updates
security patch identification
0.512021
Locating the Security Patches for Disclosed OSS Vulnerabilities with Vulnerability-Commit Correlation Ranking · CCS 2021
Software maintenance and evolution
vulnerability management
0.512021
Locating the Security Patches for Disclosed OSS Vulnerabilities with Vulnerability-Commit Correlation Ranking · CCS 2021
Machine learning › Trustworthy machine learning › dataset bias
modality bias
0.312026
Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation · AAAI 2026
Machine learning › Trustworthy machine learning
robustness
0.312026
Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation · AAAI 2026
Computer vision › Vision and language › vision-language model
vision-language model inference
0.312025
Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs · ICCV 2025
Computer vision › Segmentation and scene understanding
prompt-based segmentation
0.212024
I-MedSAM: Implicit Medical Image Segmentation with Segment Anything · ECCV (10) 2024
Distributed systems › fault tolerance
checkpointing
0.212014
Transparent checkpoint-restart over infiniband · HPDC 2014
Systems and software security
vulnerability management
0.212022
Understanding the Practice of Security Patch Management across Multiple Branches in OSS Projects · WWW 2022
Parallel and multicore computing
MPI
0.112014
Transparent checkpoint-restart over infiniband · HPDC 2014
Parallel and multicore computing › parallel computing › parallel programming languages
unified parallel c
0.112014
Transparent checkpoint-restart over infiniband · HPDC 2014

Methods — techniques the papers use, named apart from their topics

empirical study · 1.1vulnerability-commit correlation ranking · 1.0progressive modality masking · 1.0masked autoencoder reconstruction · 1.0gradient-based filtering · 1.0adversarial foreground-background reconstruction · 1.0visual cue exploitation · 0.9mixture of experts · 0.9low-rank adaptation · 0.9attention-based distillation · 0.9attention analysis · 0.9implicit neural representation · 0.8system-initiated checkpointing · 0.2kernel module avoidance · 0.2
YearPublicationVenuePosition
2026 FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning
abstract
Vision-Language-Action (VLA) models have demonstrated significant potential in complex scene understanding and action reasoning, leading to their increasing adoption in end-to-end autonomous driving systems. However, the long visual tokens of VLA models greatly increase computational costs. Current visual token pruning methods in Vision-Language Models (VLM) rely on either visual token similarity or visual-text attention, but both have shown poor performance in autonomous driving scenarios. Given that human drivers concentrate on relevant foreground areas while driving, we assert that retaining visual tokens containing this foreground information is essential for effective decision-making. Inspired by this, we propose FastDriveVLA, a novel reconstruction-based vision token pruning framework designed specifically for autonomous driving. FastDriveVLA includes a plug-and-play visual token pruner called ReconPruner, which prioritizes foreground information through MAE-style pixel reconstruction. A novel adversarial foreground-background reconstruction strategy is designed to train ReconPruner for the visual encoder of VLA models. Once trained, ReconPruner can be seamlessly applied to different VLA models with the same visual encoder without retraining. To train ReconPruner, we also introduce a large-scale dataset called nuScenes-FG, consisting of 241K image-mask pairs with annotated foreground regions. Our approach achieves state-of-the-art results on the nuScenes open-loop planning benchmark across different pruning ratios.
Jiajun Cao, Qizhe Zhang, Peidong Jia, Xiaoan Zhang, Lizhuo, Xiaobao Wei, Sixiang Chen, Liyun Li, Ming Lu 0002, Shanghang Zhang
AAAI1
2026 Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation
abstract
Multimodal keyphrase generation (MKP) aims to extract a concise set of keyphrases that capture the essential meaning of paired image–text inputs, enabling structured understanding, indexing, and retrieval of multimedia data across the web and social platforms. Success in this task demands effectively bridging the semantic gap between heterogeneous modalities. While multimodal large language models (MLLMs) achieve superior cross-modal understanding by leveraging massive pretraining on image-text corpora, we observe that they often struggle with modality bias and fine-grained intra-modal feature extraction. This oversight leads to a lack of robustness in real-world scenarios where multimedia data is noisy, along with incomplete or misaligned modalities. To address this problem, we propose AimKP, a novel framework that explicitly reinforces intra-modal semantic learning in MLLMs while preserving cross-modal alignment. AimKP incorporates two core innovations: (i) Progressive Modality Masking, which forces fine-grained feature extraction from corrupted inputs by progressively masking modality information during training; (ii) Gradient-based Filtering, that identifies and discards noisy samples, preventing them from corrupting the model’s core cross-modal learning. Extensive experiments validate AimKP’s effectiveness in multimodal keyphrase generation and its robustness across different scenarios.
Jiajun Cao, Qinggang Zhang, Yunbo Tang, Zhishang Xiang, Jinsong Su
AAAI1
2025 MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders
abstract
Visual encoders are fundamental components in vision-language models (VLMs), each showcasing unique strengths derived from various pre-trained visual foundation models. To leverage the various capabilities of these encoders, recent studies incorporate multiple encoders within a single VLM, leading to a considerable increase in computational cost. In this paper, we present Mixture-of-Visual-Encoder Knowledge Distillation (MoVEKD), a novel framework that distills the unique proficiencies of multiple vision encoders into a single, efficient encoder model. Specifically, to mitigate conflicts and retain the unique characteristics of each teacher encoder, we employ low-rank adaptation (LoRA) and mixture-of-experts (MoEs) to selectively activate specialized knowledge based on input features, enhancing both adaptability and efficiency. To regularize the KD process and enhance performance, we propose an attention-based distillation strategy that adaptively weighs the different encoders and emphasizes valuable visual tokens, reducing the burden of replicating comprehensive but distinct features from multiple teachers. Comprehensive experiments on popular VLMs, such as LLaVA and LLaVA-NeXT, validate the effectiveness of our method. Our code is available at: https://github.com/hey-cjj/MoVE-KD.
Jiajun Cao, Yuan Zhang 0020, Tao Huang 0020, Ming Lu 0002, Qizhe Zhang, Ruichuan An, Ningning Ma, Shanghang Zhang
CVPR1
2025 Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
Qizhe Zhang, Aosong Cheng, Ming Lu 0002, Renrui Zhang, Zhiyong Zhuo, Jiajun Cao, Shaobo Guo, Qi She, Shanghang Zhang
ICCV6
2024 I-MedSAM: Implicit Medical Image Segmentation with Segment Anything
Xiaobao Wei, Jiajun Cao, Yizhu Jin, Ming Lu 0002, Shanghang Zhang
ECCV (10)2
2022 Understanding the Practice of Security Patch Management across Multiple Branches in OSS Projects
abstract
Since the users of open source software (OSS) projects may not use the latest version all the time, OSS development teams often support code maintenance for old versions through maintaining multiple stable branches. Typically, the developers create a stable branch for each old stable version, deploy security patches on the branch, and release fixed versions at regular intervals. As such, old-version applications in production environments are protected from the disclosed vulnerabilities in a long time. However, the rapidly growing number of OSS vulnerabilities has greatly strained this patch deployment model, and a critical need has arisen for the security community to understand the practice of security patch management across stable branches. In this work, we conduct a large-scale empirical study of stable branches in OSS projects and the security patches deployed on them via investigating 608 stable branches belonging to 26 popular OSS projects as well as more than 2,000 security fixes for 806 CVEs deployed on stable branches.
Yuan Zhang 0009, Jiajun Cao, Kun Sun 0001, Mi Zhang 0001, Min Yang 0002
WWW3
2021 Locating the Security Patches for Disclosed OSS Vulnerabilities with Vulnerability-Commit Correlation Ranking
abstract
Security patches play an important role in defending against the security threats brought by the increasing OSS vulnerabilities. However, the collection of security patches still remains a challenging problem. Existing works mainly adopt a matching-based design that uses auxiliary information in CVE/NVD to reduce the search scope of patch commits. However, our preliminary study shows that these approaches can only cover a small part of disclosed OSS vulnerabilities (about 12%-53%) even with manual assistance.
Yuan Zhang 0009, Chenyuan Mi, Jiajun Cao, Kun Sun 0001, Min Yang 0002
CCS4
2019 Job migration in HPC clusters by means of checkpoint/restart
Manuel Aurelio Rodriguez Pascual, Jiajun Cao, José A. Moríñigo, Gene Cooperman, Rafael Mayo 0001
J. Supercomput.2
2016 System-Level Scalable Checkpoint-Restart for Petascale Computing
abstract
Fault tolerance for the upcoming exascale generation has long been an area of active research. One of the components of a fault tolerance strategy is checkpointing. Petascale-level checkpointing is demonstrated through a new mechanism for virtualization of the InfiniBand UD (unreliable datagram) mode, and for updating the remote address on each UD-based send, due to lack of a fixed peer. Note that InfiniBand UD is required to support modern MPI implementations. An extrapolation from the current results to future SSD-based storage systems provides evidence that the current approach will remain practical in the exascale generation. This transparent checkpointing approach is evaluated using a framework of the DMTCP checkpointing package. Results are shown for HPCG (linear algebra), NAMD (molecular dynamics), and the NAS NPB benchmarks. In tests up to 32,752 MPI processes on 32,752 CPU cores, checkpointing of a computation with a 38 TB memory footprint in 11 minutes is demonstrated. Runtime overhead is reduced to less than 1%. The approach is also evaluated across three widely used MPI implementations.
Jiajun Cao, Kapil Arya, Rohan Garg 0001, L. Shawn Matott, Dhabaleswar K. Panda 0001, Hari Subramoni, Jérôme Vienne, Gene Cooperman
ICPADS1
2015 Checkpointing as a Service in Heterogeneous Cloud Environments
abstract
A non-invasive, cloud-agnostic approach is demonstrated for extending existing cloud platforms to include checkpoint-restart capability. Most cloud platforms currently rely on each application to provide its own fault tolerance. A uniform mechanism within the cloud itself serves two purposes: (a) direct support for long-running jobs, which would otherwise require a custom fault-tolerant mechanism for each application, and (b) the administrative capability to manage an over-subscribed cloud by temporarily swapping out jobs when higher priority jobs arrive. An advantage of this uniform approach is that it also supports parallel and distributed computations, over both TCP and InfiniBand, thus allowing traditional HPC applications to take advantage of an existing cloud infrastructure. Additionally, an integrated health-monitoring mechanism detects when long-running jobs either fail or incur exceptionally low performance, perhaps due to resource starvation, and proactively suspends the job. The cloud-agnostic feature is demonstrated by applying the implementation to two very different cloud platforms: Snooze and Open Stack. The use of a cloud-agnostic architecture also enables, for the first time, migration of applications from one cloud platform to another.
Jiajun Cao, Matthieu Simonin, Gene Cooperman, Christine Morin
CCGRID1
2014 Transparent checkpoint-restart over infiniband
abstract
Transparently saving the state of the InfiniBand network as part of distributed checkpointing has been a long-standing challenge for researchers. The lack of a solution has forced typical MPI implementations to include custom checkpoint-restart services that "tear down" the network, checkpoint each node in isolation, and then re-connect the network again. This work presents the first example of transparent, system-initiated checkpoint-restart that directly supports InfiniBand. The new approach simplifies current practice by avoiding the need for a privileged kernel module. The generality of this approach is demonstrated by applying it both to MPI and to Berkeley UPC (Unified Parallel C), in its native mode (without MPI). Scalability is shown by checkpointing 2,048 MPI processes across 128 nodes (with 16 cores per node). The run-time overhead varies between 0.8% and 1.7%. While checkpoint times dominate, the network-only portion of the implementation is shown to require less than 100 milliseconds (not including the time to locally write application memory to stable storage).
Jiajun Cao, Gregory Kerr, Kapil Arya, Gene Cooperman
HPDC1