VLDB 2026 Research / reviewers in the wild / expert
Jiajun Fan
dblp:250/5193
· DBLP profile ↗
13ranked-venue papers
5as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Systems, architecture and hardware · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Interaction makes better segmentation: An interaction-based framework for temporal action segmentation
Minjie Xu, Jiajun Fan, Chenyu Xiao, Shenglan Liu 0001, Lin Feng 0001 |
Knowl. Based Syst. | 3 |
| 2026 | PRANCE: Joint Token-Optimization and Structural Channel-Pruning for Adaptive ViT InferenceabstractThe troublesome model size and quadratic computational complexity associated with token quantity pose significant deployment challenges for Vision Transformers (ViTs) in practical applications. Despite recent advancements in model pruning and token reduction techniques speed up the inference speed of ViTs, these approaches either adopt a fixed sparsity ratio or overlook the meaningful interplay between architectural optimization and token selection. Consequently, this static and single-dimension compression often leads to pronounced accuracy degradation under aggressive compression rates, as they fail to fully explore redundancies across these two orthogonal dimensions. Therefore, we introduce PRANCE, a framework which can jointly optimize activated channels and tokens on a per-sample basis, aiming to accelerate ViTs' inference process from a unified data and architectural perspective. However, the joint framework poses challenges to both architectural and decision-making aspects. First, while ViTs inherently support variable-token inference, they do not facilitate dynamic computations for variable channels. To overcome this limitation, we propose a meta-network using weight-sharing techniques to support arbitrary channels of the Multi-Head Self-Attention (MHSA) and Multi-Layer Perceptron (MLP) layers, serving as a foundational model for architectural decision-making. Second, simultaneously optimizing the model structure and input data constitutes a combinatorial optimization problem with an extremely large decision space, reaching up to around $10^{14}$1014, making supervised learning infeasible. To this end, we design a lightweight selector employing Proximal Policy Optimization algorithm (PPO) for efficient decision-making. Furthermore, we introduce a novel "Result-to-Go" training mechanism that models ViTs' inference process as a Markov decision process, significantly reducing action space and mitigating delayed-reward issues during training. Additionally, our framework simultaneously supports different kinds of token optimization methods such as pruning, merging, and sequential pruning-merging strategies. Extensive experiments demonstrate the effectiveness of PRANCE in reducing FLOPs by approximately 50%, retaining only about 10% of tokens while achieving lossless Top-1 accuracy. Ye Li 0016, Jiajun Fan, Zenghao Chai, Xinzhu Ma, Zhi Wang 0001, Wenwu Zhu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein RegularizationabstractRecent advancements in reinforcement learning (RL) have achieved great success in fine-tuning diffusion-based generative models. However, fine-tuning continuous flow-based generative models to align with arbitrary user-defined reward functions remains challenging, particularly due to issues such as policy collapse from overoptimization and the prohibitively high computational cost of likelihoods in continuous-time flows. In this paper, we propose an easy-to-use and theoretically sound RL fine-tuning method, which we term Online Reward-Weighted Conditional Flow Matching with Wasserstein-2 Regularization (ORW-CFM-W2). Our method integrates RL into the flow matching framework to fine-tune generative models with arbitrary reward functions, without relying on gradients of rewards or filtered datasets. By introducing an online reward-weighting mechanism, our approach guides the model to prioritize high-reward regions in the data manifold. To prevent policy collapse and maintain diversity, we incorporate Wasserstein-2 (W2) distance regularization into our method and derive a tractable upper bound for it in flow matching, effectively balancing exploration and exploitation of policy optimization. We provide theoretical analyses to demonstrate the convergence properties and induced data distributions of our method, establishing connections with traditional RL algorithms featuring Kullback-Leibler (KL) regularization and offering a more comprehensive understanding of the underlying mechanisms and learning behavior of our approach. Extensive experiments on tasks including target image generation, image compression, and text-image alignment demonstrate the effectiveness of our method, where our method achieves optimal policy convergence while allowing controllable trade-offs between reward maximization and diversity preservation. Jiajun Fan, Shuaike Shen, Chaoran Cheng, Chumeng Liang |
ICLR | 1 |
| 2025 | Adaptive Divergence Regularized Policy Optimization for Fine-tuning Generative ModelsabstractBalancing exploration and exploitation during reinforcement learning fine-tuning of generative models presents a critical challenge, as existing approaches rely on fixed divergence regularization that creates an inherent dilemma: strong regularization preserves model capabilities but limits reward optimization, while weak regularization enables greater alignment but risks instability or reward hacking. We introduce Adaptive Divergence Regularized Policy Optimization (ADRPO), which automatically adjusts regularization strength based on advantage estimates—reducing regularization for high-value samples while applying stronger regularization to poor samples, enabling policies to navigate between exploration and aggressive exploitation according to data quality. Our implementation with Wasserstein-2 regularization for flow matching generative models achieves remarkable results on text-to-image generation, achieving better semantic alignment and diversity than offline methods like DPO and online methods with fixed regularization like ORW-CFM-W2. ADRPO enables a 2B parameter SD3 model to surpass much larger models with 4.8B and 12B parameters in attribute binding, semantic consistency, artistic style transfer, and compositional control while maintaining generation diversity. ADRPO generalizes to KL-regularized fine-tuning of both text-only LLMs and multi-modal reasoning models, enhancing existing online RL methods like GRPO while requiring no additional networks or complex architectural changes. In LLM fine-tuning, ADRPO demonstrates an emergent ability to escape local optima through active exploration, while in multi-modal audio reasoning, it outperforms GRPO through superior step-by-step reasoning, enabling a 7B model to outperform substantially larger commercial models including Gemini 2.5 Pro and GPT-4o Audio, offering an effective plug-and-play solution to the exploration-exploitation challenge across diverse generative architectures and modalities. Jiajun Fan, Chaoran Cheng |
NeurIPS | 1 |
| 2025 | Variational Supervised Contrastive LearningabstractContrastive learning has proven to be highly efficient and adaptable in shaping representation spaces across diverse modalities by pulling similar samples together and pushing dissimilar ones apart. However, two key limitations persist: (1) Without explicit regulation of the embedding distribution, semantically related instances can inadvertently be pushed apart unless complementary signals guide pair selection, and (2) excessive reliance on large in-batch negatives and tailored augmentations hinders generalization. To address these limitations, we propose Variational Supervised Contrastive Learning (VarCon), which reformulates supervised contrastive learning as variational inference over latent class variables and maximizes a posterior-weighted evidence lower bound (ELBO) that replaces exhaustive pair-wise comparisons for efficient class-aware matching and grants fine-grained control over intra-class dispersion in the embedding space. Trained exclusively on image data, our experiments on CIFAR-10, CIFAR-100, ImageNet-100, and ImageNet-1K show that VarCon (1) achieves state-of-the-art performance for contrastive learning frameworks, reaching 79.36% Top-1 accuracy on ImageNet-1K and 78.29% on CIFAR-100 with a ResNet-50 encoder while converging in just 200 epochs; (2) yields substantially clearer decision boundaries and semantic organization in the embedding space, as evidenced by KNN classification, hierarchical clustering results, and transfer-learning assessments; and (3) demonstrates superior performance in few-shot learning than supervised baseline and superior robustness across various augmentation strategies. Jiajun Fan, Heng Ji 0001 |
NeurIPS | 2 |
| 2025 | Bridging the Point to Boundary Gap for Point-Supervised Temporal Action Localization with Single-Stage Inference
Junshi Yang, Shenglan Liu 0001, Xuhan Sheng, Yiheng Zhou, Lin Feng 0001, Jiajun Fan |
PRCV (7) | 7 |
| 2025 | An End-to-End Deep Learning QoS Prediction Model Based on Temporal Context and Feature Fusion
Peiyun Zhang, Jiajun Fan, Haibin Zhu 0001, Qinglin Zhao |
IEEE Trans. Serv. Comput. | 2 |
| 2024 | Dynamic Neural Dowker Network: Approximating Persistent Homology in Dynamic Directed GraphsabstractPersistent homology, a fundamental technique within Topological Data Analysis (TDA), captures structural and shape characteristics of graphs, yet encounters computational difficulties when applied to dynamic directed graphs. This paper introduces the Dynamic Neural Dowker Network (DNDN), a novel framework specifically designed to approximate the results of dynamic Dowker filtration, aiming to capture the high-order topological features of dynamic directed graphs. Our approach creatively uses line graph transformations to produce both source and sink line graphs, highlighting the shared neighbor structures that Dowker complexes focus on. The DNDN incorporates a Source-Sink Line Graph Neural Network (SSLGNN) layer to effectively capture the neighborhood relationships among dynamic edges. Additionally, we introduce an innovative duality edge fusion mechanism, ensuring that the results for both the sink and source line graphs adhere to the duality principle intrinsic to Dowker complexes. Our approach is validated through comprehensive experiments on real-world datasets, demonstrating DNDN's capability not only to effectively approximate dynamic Dowker filtration results but also to perform exceptionally in dynamic graph classification tasks. Hao Li 0080, Hao Jiang 0010, Jiajun Fan, Dongsheng Ye, Liang Du 0006 |
KDD | 3 |
| 2024 | Low-rank persistent probability representation for higher-order role discovery
Dongsheng Ye, Hao Jiang 0010, Jiajun Fan, Qiang Wang 0027 |
Expert Syst. Appl. | 3 |
| 2023 | Learnable Behavior Control: Breaking Atari Human World Records via Sample-Efficient Behavior Selection
Jiajun Fan, Yuzheng Zhuang, Yuecheng Liu, Jianye Hao, Bin Wang 0034, Jiangcheng Zhu, Hao Wang 0049, Shutao Xia |
ICLR | 1 |
| 2023 | Optimal Transport for Treatment Effect EstimationabstractEstimating individual treatment effects from observational data is challenging due to treatment selection bias. Prevalent methods mainly mitigate this issue by aligning different treatment groups in the latent space, the core of which is the calculation of distribution discrepancy. However, two issues that are often overlooked can render these methods invalid:
(1) mini-batch sampling effects (MSE), where the calculated discrepancy is erroneous in non-ideal mini-batches with outcome imbalance and outliers;
(2) unobserved confounder effects (UCE), where the unobserved confounders are not considered in the discrepancy calculation.
Both of these issues invalidate the calculated discrepancy, mislead the training of estimators, and thus impede the handling of treatment selection bias.
To tackle these issues, we propose Entire Space CounterFactual Regression (ESCFR), which is a new take on optimal transport technology in the context of causality.
Specifically, based on the canonical optimal transport framework, we propose a relaxed mass-preserving regularizer to address the MSE issue and design a proximal factual outcome regularizer to handle the UCE issue.
Extensive experiments demonstrate that ESCFR estimates distribution discrepancy accurately, handles the treatment selection bias effectively, and outperforms prevalent competitors significantly. Hao Wang 0049, Jiajun Fan, Zhichao Chen 0001, Haoxuan Li 0001, Weiming Liu 0005, Tianqiao Liu, Quanyu Dai, Yichao Wang 0002, Zhenhua Dong, Ruiming Tang |
NeurIPS | 2 |
| 2022 | Generalized Data Distribution IterationabstractTo obtain higher sample efficiency and superior final performance simultaneously has been one of the major challenges for deep reinforcement learning (DRL). Previous work could handle one of these challenges but typically failed to address them concurrently. In this paper, we try to tackle these two challenges simultaneously. To achieve this, we firstly decouple these challenges into two classic RL problems: data richness and exploration-exploitation trade-off. Then, we cast these two problems into the training data distribution optimization problem, namely to obtain desired training data within limited interactions, and address them concurrently via i) explicit modeling and control of the capacity and diversity of behavior policy and ii) more fine-grained and adaptive control of selective/sampling distribution of the behavior policy using a monotonic data distribution optimization. Finally, we integrate this process into Generalized Policy Iteration (GPI) and obtain a more general framework called Generalized Data Distribution Iteration (GDI). We use the GDI framework to introduce operator-based versions of well-known RL methods from DQN to Agent57. Theoretical guarantee of the superiority of GDI compared with GPI is concluded. We also demonstrate our state-of-the-art (SOTA) performance on Arcade Learning Environment (ALE), wherein our algorithm has achieved 9620.98% mean human normalized score (HNS), 1146.39% median HNS, and surpassed 22 human world records using only 200M training frames. Our performance is comparable to Agent57’s while we consume 500 times less data. We argue that there is still a long way to go before obtaining real superhuman agents in ALE. Jiajun Fan, Changnan Xiao |
ICML | 1 |
| 2020 | Multi-View Facial Expression Recognition based o n Multitask Learning and Generative Adversarial NetworkabstractFacial expressions contain rich emotional information, which is an important method in human communication. At present, most of the researches on facial expression recognition is conducted on the frontal faces. However, in real life, the captu re device may capture facial expression data from various poses. Different from the existing techniques, in this paper, We propose a multitask deep learning method that uses the links between var ious poses and expressions to improve the accuracy of expression recognition. We use the adversarial network to supplement the e xpression information of the face for the head poses with a sever e lack of expression information. There are some advantages: Firstly, we use vgg16 to train the expression images of each deflecti on angle separately and find that the expression recognition acc uracy rate for small offset angle (-30 °, -15 °, 15°, 30 °) is larger t han 0 ° angle when the face angle is greater than 45 °, the accur acy of recognition decreases sharply. So we use multitask learnin g to jointly train these small offsets angle images and frontal ima ges, the multitask learning can learn the emotion-preserving repr esentations at various poses to predict the expression class label f rom the input face, and bring again to its recognition accuracy r ate. Secondly, for the poses of a severe lack of expression inform ation, we use TP-GAN to convert a large deflection pose image i nto a frontal face and supplement its expression information. Th e experimental results show that our proposed algorithm has a g ood recognition effect on facial expressions for all poses. Compa red with the most advanced expression recognition methods, this paper has also achieved the state-of-the-art recognition results. Jiajun Fan, Shipu Wang, Po Yang 0001, Yun Yang 0003 |
INDIN | 1 |