VLDB 2026 Research / reviewers in the wild / expert
Linjiajie Fang
dblp:372/6474
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Reinforcement learning · 41% Generative modeling · 20% Transfer learning and domain adaptation · 19% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
1.6 | 2 | 2025 | Diffusion Actor-Critic: Formulating Constrained Policy Iteration as Diffusion Noise Regression for Offline Reinforcement Learning · ICLR 2025 D2R2: Diffusion-based Representation with Random Distance Matching for Tabular Few-shot Learning · NeurIPS 2024 |
Machine learning › Reinforcement learning
offline reinforcement learning |
1.6 | 2 | 2025 | Diffusion Actor-Critic: Formulating Constrained Policy Iteration as Diffusion Noise Regression for Offline Reinforcement Learning · ICLR 2025 Q-Distribution guided Q-learning for offline reinforcement learning: Uncertainty penalized Q-value via consistency model · NeurIPS 2024 |
Robotics › Robot manipulation
diffusion policy |
0.9 | 1 | 2025 | Diffusion Actor-Critic: Formulating Constrained Policy Iteration as Diffusion Noise Regression for Offline Reinforcement Learning · ICLR 2025 |
Machine learning › Reinforcement learning › regularization for reinforcement learning
policy regularization |
0.9 | 1 | 2025 | Diffusion Actor-Critic: Formulating Constrained Policy Iteration as Diffusion Noise Regression for Offline Reinforcement Learning · ICLR 2025 |
Machine learning › Trustworthy machine learning › robustness
distribution shift |
0.8 | 1 | 2024 | Q-Distribution guided Q-learning for offline reinforcement learning: Uncertainty penalized Q-value via consistency model · NeurIPS 2024 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
0.8 | 1 | 2024 | D2R2: Diffusion-based Representation with Random Distance Matching for Tabular Few-shot Learning · NeurIPS 2024 |
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning |
0.8 | 1 | 2024 | Q-Distribution guided Q-learning for offline reinforcement learning: Uncertainty penalized Q-value via consistency model · NeurIPS 2024 |
Machine learning › Transfer learning and domain adaptation › few-shot learning
tabular few-shot learning |
0.8 | 1 | 2024 | D2R2: Diffusion-based Representation with Random Distance Matching for Tabular Few-shot Learning · NeurIPS 2024 |
Data mining › predictive modeling › classification
tabular data classification |
0.2 | 1 | 2024 | D2R2: Diffusion-based Representation with Random Distance Matching for Tabular Few-shot Learning · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
diffusion model · 2.4random distance matching · 1.5prototype learning · 1.5actor-critic · 0.9KL constraint · 0.9uncertainty estimation · 0.8pessimistic q-value · 0.8consistency model · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Diffusion Actor-Critic: Formulating Constrained Policy Iteration as Diffusion Noise Regression for Offline Reinforcement LearningabstractIn offline reinforcement learning, it is necessary to manage out-of-distribution actions to prevent overestimation of value functions. One class of methods, the policy-regularized method, addresses this problem by constraining the target policy to stay close to the behavior policy. Although several approaches suggest representing the behavior policy as an expressive diffusion model to boost performance, it remains unclear how to regularize the target policy given a diffusion-modeled behavior sampler. In this paper, we propose Diffusion Actor-Critic (DAC) that formulates the Kullback-Leibler (KL) constraint policy iteration as a diffusion noise regression problem, enabling direct representation of target policies as diffusion models. Our approach follows the actor-critic learning paradigm in which we alternatively train a diffusion-modeled target policy and a critic network. The actor training loss includes a soft Q-guidance term from the Q-gradient. The soft Q-guidance is based on the theoretical solution of the KL constraint policy iteration, which prevents the learned policy from taking out-of-distribution actions. We demonstrate that such diffusion-based policy constraint, along with the coupling of the lower confidence bound of the Q-ensemble as value targets, not only preserves the multi-modality of target policies, but also contributes to stable convergence and strong performance in DAC. Our approach is evaluated on D4RL benchmarks and outperforms the state-of-the-art in nearly all environments. Linjiajie Fang, Ruoxue Liu, Wenjia Wang 0005, Bing-Yi Jing |
ICLR | 1 |
| 2024 | D2R2: Diffusion-based Representation with Random Distance Matching for Tabular Few-shot LearningabstractTabular data is widely utilized in a wide range of real-world applications. The challenge of few-shot learning with tabular data stands as a crucial problem in both industry and academia, due to the high cost or even impossibility of annotating additional samples. However, the inherent heterogeneity of tabular features, combined with the scarcity of labeled data, presents a significant challenge in tabular few-shot classification. In this paper, we propose a novel approach named Diffusion-based Representation with Random Distance matching (D2R2) for tabular few-shot learning. D2R2 leverages the powerful expression ability of diffusion models to extract essential semantic knowledge crucial for denoising process. This semantic knowledge proves beneficial in few-shot downstream tasks. During the training process of our designed diffusion model, we introduce a random distance matching to preserve distance information in the embeddings, thereby improving effectiveness for classification. During the classification stage, we introduce an instance-wise iterative prototype scheme to improve performance by accommodating the multimodality of embeddings and increasing clustering robustness. Our experiments reveal the significant efficacy of D2R2 across various tabular few-shot learning benchmarks, demonstrating its state-of-the-art performance in this field. Ruoxue Liu, Linjiajie Fang, Wenjia Wang 0005, Bing-Yi Jing |
NeurIPS | 2 |
| 2024 | Q-Distribution guided Q-learning for offline reinforcement learning: Uncertainty penalized Q-value via consistency modelabstract``Distribution shift'' is the primary obstacle to the success of offline reinforcement learning. As a learning policy may take actions beyond the knowledge of the behavior policy (referred to as Out-of-Distribution (OOD) actions), the Q-values of these OOD actions can be easily overestimated. Consequently, the learning policy becomes biasedly optimized using the incorrect recovered Q-value function. One commonly used idea to avoid the overestimation of Q-value is to make a pessimistic adjustment. Our key idea is to penalize the Q-values of OOD actions that correspond to high uncertainty. In this work, we propose Q-Distribution guided Q-learning (QDQ) which pessimistic Q-value on OOD regions based on uncertainty estimation. The uncertainty measure is based on the conditional Q-value distribution, which is learned via a high-fidelity and efficient consistency model. On the other hand, to avoid the overly conservative problem, we introduce an uncertainty-aware optimization objective to update the Q-value function. The proposed QDQ demonstrates solid theoretical guarantees for the accuracy of Q-value distribution learning and uncertainty measurement, as well as the performance of the learning policy. QDQ consistently exhibits strong performance in the D4RL benchmark and shows significant improvements for many tasks. Our code can be found at <code link>. Linjiajie Fang, Wenjia Wang 0005, Bing-Yi Jing |
NeurIPS | 2 |