VLDB 2026 Research / reviewers in the wild / expert
Shuaike Shen
dblp:359/3334
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Generative modeling · 47% Transfer learning and domain adaptation · 19% Representation and self-supervised learning · 15% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Bioinformatics and computational biology · 81% Computational science and engineering · 19% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
protein design |
1.5 | 2 | 2024 | Floating Anchor Diffusion Model for Multi-motif Scaffolding · ICML 2024 De novo Protein Design Using Geometric Vector Field Networks · ICLR 2024 |
Machine learning › Generative modeling
diffusion model |
1.0 | 2 | 2025 | Floating Anchor Diffusion Model for Multi-motif Scaffolding · ICML 2024 Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization · ICLR 2025 |
Machine learning › Generative modeling
flow matching |
0.9 | 1 | 2025 | Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization · ICLR 2025 |
Machine learning › Reinforcement learning
policy optimization |
0.9 | 1 | 2025 | Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization · ICLR 2025 |
Machine learning › Transfer learning and domain adaptation › fine-tuning
regularization for fine-tuning |
0.9 | 1 | 2025 | Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization · ICLR 2025 |
Machine learning › Generative modeling › diffusion model › diffusion model adaptation
reward fine-tuning |
0.9 | 1 | 2025 | Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization · ICLR 2025 |
Computational science and engineering
materials science |
0.9 | 1 | 2025 | A Denoising Pre-training Framework for Accelerating Novel Material Discovery · AAAI 2025 |
Bioinformatics and computational biology › protein design
de novo protein design |
0.8 | 1 | 2024 | De novo Protein Design Using Geometric Vector Field Networks · ICLR 2024 |
Bioinformatics and computational biology › protein design
motif scaffolding |
0.8 | 1 | 2024 | Floating Anchor Diffusion Model for Multi-motif Scaffolding · ICML 2024 |
Bioinformatics and computational biology › structural bioinformatics
protein structure |
0.8 | 1 | 2024 | De novo Protein Design Using Geometric Vector Field Networks · ICLR 2024 |
Machine learning › Transfer learning and domain adaptation › pre-training and adaptation
pre-training and fine-tuning |
0.3 | 1 | 2025 | A Denoising Pre-training Framework for Accelerating Novel Material Discovery · AAAI 2025 |
Computer vision › 3D vision
geometric deep learning |
0.2 | 1 | 2024 | De novo Protein Design Using Geometric Vector Field Networks · ICLR 2024 |
Methods — techniques the papers use, named apart from their topics
masked atom modeling · 1.7fine-tuning · 1.7denoising reconstruction · 1.7protein diffusion · 1.5floating anchor · 1.5diffusion model · 1.5attention aggregation · 1.5ESM · 1.5wasserstein-2 regularization · 0.9online reward weighting · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Denoising Pre-training Framework for Accelerating Novel Material DiscoveryabstractCrystal materials play an important role in the development of society. The discovery of new materials is critical to achieving sustainable development goals (SDGs), such as climate change mitigation, affordable and clean energy, and fostering innovation in industry and infrastructure. Recent advances in deep learning for crystal property prediction have accelerated material discovery, but these methods typically rely on labeled data, which is often limited and varies across different properties. This limitation hinders the full utilization of the vast amount of unlabeled data in materials science. To overcome this challenge, we introduce an unsupervised Denoising Pre-training Framework (DPF) tailored for crystal structures. DPF trains a model to reconstruct the original crystal structure by recovering the masked atom types, perturbed atom positions, and perturbed crystal lattices. Through pre-training, models learn the intrinsic features of crystal structures and capture the key features influencing crystal properties. We pre-train models on a dataset of 380,743 unlabeled crystal structures and fine-tune them on downstream property prediction tasks. Extensive experiments demonstrate the effectiveness of our framework, showing its potential to significantly advance material science and contribute to the development of society by accelerating the discovery of materials crucial for sustainable technologies. Shuaike Shen, Ke Liu 0012, Muzhi Zhu, Hao Chen 0041 |
AAAI | 1 |
| 2025 | Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein RegularizationabstractRecent advancements in reinforcement learning (RL) have achieved great success in fine-tuning diffusion-based generative models. However, fine-tuning continuous flow-based generative models to align with arbitrary user-defined reward functions remains challenging, particularly due to issues such as policy collapse from overoptimization and the prohibitively high computational cost of likelihoods in continuous-time flows. In this paper, we propose an easy-to-use and theoretically sound RL fine-tuning method, which we term Online Reward-Weighted Conditional Flow Matching with Wasserstein-2 Regularization (ORW-CFM-W2). Our method integrates RL into the flow matching framework to fine-tune generative models with arbitrary reward functions, without relying on gradients of rewards or filtered datasets. By introducing an online reward-weighting mechanism, our approach guides the model to prioritize high-reward regions in the data manifold. To prevent policy collapse and maintain diversity, we incorporate Wasserstein-2 (W2) distance regularization into our method and derive a tractable upper bound for it in flow matching, effectively balancing exploration and exploitation of policy optimization. We provide theoretical analyses to demonstrate the convergence properties and induced data distributions of our method, establishing connections with traditional RL algorithms featuring Kullback-Leibler (KL) regularization and offering a more comprehensive understanding of the underlying mechanisms and learning behavior of our approach. Extensive experiments on tasks including target image generation, image compression, and text-image alignment demonstrate the effectiveness of our method, where our method achieves optimal policy convergence while allowing controllable trade-offs between reward maximization and diversity preservation. Jiajun Fan, Shuaike Shen, Chaoran Cheng, Chumeng Liang |
ICLR | 2 |
| 2024 | De novo Protein Design Using Geometric Vector Field NetworksabstractAdvances like protein diffusion have marked revolutionary progress in $\textit{de novo}$ protein design, a central topic in life science. These methods typically depend on protein structure encoders to model residue backbone frames, where atoms do not exist. Most prior encoders rely on atom-wise features, such as angles and distances between atoms, which are not available in this context. Only a few basic encoders, like IPA, have been proposed for this scenario, exposing the frame modeling as a bottleneck. In this work, we introduce the Vector Field Network (VFN), that enables network layers to perform learnable vector computations between coordinates of frame-anchored virtual atoms, thus achieving a higher capability for modeling frames. The vector computation operates in a manner similar to a linear layer, with each input channel receiving 3D virtual atom coordinates instead of scalar values. The multiple feature vectors output by the vector computation are then used to update the residue representations and virtual atom coordinates via attention aggregation. Remarkably, VFN also excels in modeling both frames and atoms, as the real atoms can be treated as the virtual atoms for modeling, positioning VFN as a potential $\textit{universal encoder}$. In protein diffusion (frame modeling), VFN exhibits a impressive performance advantage over IPA, excelling in terms of both designability ($\textbf{67.04}$\% vs. 53.58\%) and diversity ($\textbf{66.54}$\% vs. 51.98\%). In inverse folding(frame and atom modeling), VFN outperforms the previous SoTA model, PiFold ($\textbf{54.7}$\% vs. 51.66\%), on sequence recovery rate; we also propose a method of equipping VFN with the ESM model, which significantly surpasses the previous ESM-based SoTA ($\textbf{62.67}$\% vs. 55.65\%), LM-Design, by a substantial margin. Code is available at https://github.com/aim-uofa/VFN Weian Mao, Muzhi Zhu, Shuaike Shen, Lin Wu 0001, Hao Chen 0041, Chunhua Shen |
ICLR | 4 |
| 2024 | Floating Anchor Diffusion Model for Multi-motif ScaffoldingabstractMotif scaffolding seeks to design scaffold structures for constructing proteins with functions derived from the desired motif, which is crucial for the design of vaccines and enzymes. Previous works approach the problem by inpainting or conditional generation. Both of them can only scaffold motifs with fixed positions, and the conditional generation cannot guarantee the presence of motifs. However, prior knowledge of the relative motif positions in a protein is not readily available, and constructing a protein with multiple functions in one protein is more general and significant because of the synergies between functions. We propose a Floating Anchor Diffusion (FADiff) model. FADiff allows motifs to float rigidly and independently in the process of diffusion, which guarantees the presence of motifs and automates the motif position design. Our experiments demonstrate the efficacy of FADiff with high success rates and designable novel scaffolds. To the best of our knowledge, FADiff is the first work to tackle the challenge of scaffolding multiple motifs without relying on the expertise of relative motif positions in the protein. Code is available at https://github.com/aim-uofa/FADiff. Ke Liu 0012, Weian Mao, Shuaike Shen, Xiaoran Jiao, Hao Cheng 0012, Chunhua Shen |
ICML | 3 |