Wanjia Zhao

dblp:352/8879 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Generative modeling · 24% Language models and text generation · 13% Reinforcement learning · 8%
Theoretical computer science
1 paper
Automated reasoning and model checking · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 22 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › molecular generation
3d molecule generation
0.912025
GeoAda: Efficiently Finetune Geometric Diffusion Models with Equivariant Adapters · NeurIPS 2025
Machine learning › Generative modeling
diffusion model
0.912025
GeoAda: Efficiently Finetune Geometric Diffusion Models with Equivariant Adapters · NeurIPS 2025
Machine learning › Generative modeling › diffusion model › geometric diffusion model
equivariant diffusion model
0.912025
GeoAda: Efficiently Finetune Geometric Diffusion Models with Equivariant Adapters · NeurIPS 2025
Robotics › Motion planning and robot control › robot control › nonlinear control
geometric control
0.912025
GeoAda: Efficiently Finetune Geometric Diffusion Models with Equivariant Adapters · NeurIPS 2025
Machine learning › Graph learning › graph neural network › continuous graph neural network
graph neural ordinary differential equations
0.912025
Rethink GraphODE Generalization within Coupled Dynamical System · ICML 2025
Natural language and speech › Language models and text generation
large language model
0.912025
DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search · ICLR 2025
Natural language and speech › Language models and text generation
large language model reasoning
0.912025
SiriuS: Self-improving Multi-agent Systems via Bootstrapped Reasoning · NeurIPS 2025
Machine learning › Trustworthy machine learning
out-of-distribution generalization
0.912025
Rethink GraphODE Generalization within Coupled Dynamical System · ICML 2025
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.912025
GeoAda: Efficiently Finetune Geometric Diffusion Models with Equivariant Adapters · NeurIPS 2025
Machine learning › Reinforcement learning
policy learning
0.912025
DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search · ICLR 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › automated reasoning
proof search
0.912025
DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search · ICLR 2025
Automated reasoning and model checking › theorem proving
proof generation
0.912025
DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search · ICLR 2025
Automated reasoning and model checking
theorem proving
0.912025
DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search · ICLR 2025
Bioinformatics and computational biology › systems biology
dynamical system modeling
0.812024
Physics-Informed Regularization for Domain-Agnostic Dynamical System Modeling · NeurIPS 2024
Machine learning › Generative modeling
normalizing flow
0.712023
Positive Distribution Pollution: Rethinking Positive Unlabeled Learning from a Unified Perspective · AAAI 2023
Machine learning › Learning paradigms › weakly supervised learning
positive-unlabeled learning
0.712023
Positive Distribution Pollution: Rethinking Positive Unlabeled Learning from a Unified Perspective · AAAI 2023
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.712023
Positive Distribution Pollution: Rethinking Positive Unlabeled Learning from a Unified Perspective · AAAI 2023
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
adapter tuning
0.312025
GeoAda: Efficiently Finetune Geometric Diffusion Models with Equivariant Adapters · NeurIPS 2025
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
0.312025
Rethink GraphODE Generalization within Coupled Dynamical System · ICML 2025
Machine learning › Reinforcement learning › multi-agent reinforcement learning
self-play
0.312025
SiriuS: Self-improving Multi-agent Systems via Bootstrapped Reasoning · NeurIPS 2025
Machine learning › Deep learning architectures and training › neural differential equations
neural ordinary differential equations
0.212024
Physics-Informed Regularization for Domain-Agnostic Dynamical System Modeling · NeurIPS 2024
Machine learning › Learning theory › statistical estimation
risk estimation
0.212023
Positive Distribution Pollution: Rethinking Positive Unlabeled Learning from a Unified Perspective · AAAI 2023

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.7monte carlo tree search · 1.7large language model · 1.7variational inference · 1.5zero-initialized convolution · 0.9structural causal model · 0.9orthogonal subspace projection · 0.9coupling operator · 0.9bootstrapped reasoning · 0.9SE(3)-equivariant adapter · 0.9regularization · 0.8neural ODE · 0.8
YearPublicationVenuePosition
2025 DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search
abstract
Lean is an advanced proof assistant designed to facilitate formal theorem proving by providing a variety of interactive feedback. In this paper, we explore methodologies to leverage proof assistant feedback to augment the capabilities of large language models in constructing formal proofs. First, we deploy online reinforcement learning using Lean verification outcomes as the reward signal to improve the proof completion policy. This straightforward approach shows great promise in enhancing the model's alignment with the formal verification system. In addition, we propose RMaxTS, a variant of Monte-Carlo tree search that employs an intrinsic-reward-driven exploration strategy to generate diverse proof paths. The tree structure is organized to represent the transitions of intermediate tactic states, extracted from the compilation messages given by Lean's tactic mode. The intrinsic reward is constructed to incentivize the discovery of novel tactic states, which helps to to mitigate the sparse-reward problem inherent in proof search. These techniques lead to a more efficient planning scheme for formal proof generation, achieving new state-of-the-art results on both miniF2F and ProofNet benchmarks.
Huajian Xin, Z. Z. Ren, Junxiao Song, Zhihong Shao, Wanjia Zhao, Qiushi Du, Qihao Zhu, Dejian Yang, Zhibin Gou, Z. F. Wu, Fuli Luo, Chong Ruan
ICLR5
2025 Rethink GraphODE Generalization within Coupled Dynamical System
abstract
Coupled dynamical systems govern essential phenomena across physics, biology, and engineering, where components interact through complex dependencies. While Graph Ordinary Differential Equations (GraphODE) offer a powerful framework to model these systems, their generalization capabilities degrade severely under limited observational training data due to two fundamental flaws: (i) the entanglement of static attributes and dynamic states in the initialization process, and (ii) the reliance on context-specific coupling patterns during training, which hinders performance in unseen scenarios. In this paper, we propose a Generalizable GraphODE with disentanglement and regularization (GREAT) to address these challenges. Through systematic analysis via the Structural Causal Model, we identify backdoor paths that undermine generalization and design two key modules to mitigate their effects. The Dynamic-Static Equilibrium Decoupler (DyStaED) disentangles static and dynamic states via orthogonal subspace projections, ensuring robust initialization. Furthermore, the Causal Mediation for Coupled Dynamics (CMCD) employs variational inference to estimate latent causal factors, reducing spurious correlations and enhancing universal coupling dynamics. Extensive experiments across diverse dynamical systems demonstrate that ours outperforms state-of-the-art methods within both in-distribution and out-of-distribution.
Guancheng Wan, Zijie Huang 0002, Wanjia Zhao, Xiao Luo 0001, Yizhou Sun, Wei Wang 0010
ICML3
2025 Don't Forget the Enjoin: FocalLoRA for Instruction Hierarchical Alignment in Large Language Models
abstract
Recent studies reveal that large language models (LLMs) often struggle to resolve conflicting instructions embedded within hierarchical prompts, resulting in decreased compliance with system-level directives and compromising the reliability of safety-critical applications. While earlier approaches attempt to improve instruction hierarchy awareness through prompt engineering or embedding-level modifications, they typically lack structural modeling and either offer limited gains or require extensive fine-tuning. In this work, we introduce $\textbf{FocalLoRA}$, a parameter-efficient and structure-aware framework that strengthens hierarchical instruction adherence by selectively optimizing structurally critical attention heads, referred to as $\textit{focal heads}$, which exhibit heightened sensitivity to instruction conflicts. Experiments across multiple models and a dedicated benchmark demonstrate that FocalLoRA markedly enhances system instruction compliance with minimal tuning cost. For instance, on Llama-8B, fine-tuning only 0.0188\% of parameters yields a 35.52\% $\uparrow$ in system instruction compliance.
Zitong Shi, Frank Wan, Haixin Wang 0003, Ruoyan Li, Zijie Huang 0002, Wanjia Zhao, Yijia Xiao, Xiao Luo 0001, Carl Yang 0001, Yizhou Sun, Wei Wang 0010
NeurIPS6
2025 GeoAda: Efficiently Finetune Geometric Diffusion Models with Equivariant Adapters
abstract
Geometric diffusion models have shown remarkable success in molecular dynamics and structure generation. However, efficiently fine-tuning them for downstream tasks with varying geometric controls remains underexplored. In this work, we propose an SE(3)-equivariant adapter framework (GeoAda) that enables flexible and parameter-efficient fine-tuning for controlled generative tasks without modifying the original model architecture. GeoAda introduces a structured adapter design: control signals are first encoded through coupling operators, then processed by a trainable copy of selected base model layers, and finally projected back via decoupling operators followed by an equivariant zero-initialized convolution. By fine-tuning only these lightweight adapter modules, GeoAda preserves the model’s geometric consistency while mitigating overfitting and catastrophic forgetting. We theoretically prove that the proposed adapters maintain SE(3)-equivariance, ensuring that the geometric inductive biases of the pretrained diffusion model remain intact during adaptation. We demonstrate the wide applicability of \method across diverse geometric control types, including frame control, global control, subgraph control, and a broad range of application domains such as particle dynamics, molecular dynamics, human motion prediction, and molecule generation. Empirical results show that GeoAda achieves state-of-the-art fine-tuning performance while preserving original task accuracy, whereas other baselines experience significant performance degradation due to overfitting and catastrophic forgetting.
Wanjia Zhao, Jiaqi Han 0001, Siyi Gu, Mingjian Jiang, James Zou 0001, Stefano Ermon
NeurIPS1
2025 SiriuS: Self-improving Multi-agent Systems via Bootstrapped Reasoning
abstract
Multi-agent AI systems powered by large language models (LLMs) are increasingly applied to solve complex tasks. However, these systems often rely on fragile, manually designed prompts and heuristics, making optimization difficult. A key challenge in optimizing multi-agent systems is acquiring suitable training data for specialized agents. We introduce SiriuS, a self-improving, reasoning-driven optimization framework for multi-agent systems. Central to our approach is the construction of an experience library: a repository of high-quality reasoning trajectories. The library is built by retaining reasoning steps that lead to successful outcomes, providing a robust training set for optimizing multi-agent system. Additionally, we introduce a library augmentation procedure that refines unsuccessful trajectories, further enriching the library. SiriuS boosts performance by 2.86% to 21.88% on reasoning and biomedical QA and enhances agent negotiation in competitive settings. Our results show that SiriuS enhances multi-agent performance while generating reusable data for self-correction and self-play enhancement in the future.
Wanjia Zhao, Mert Yüksekgönül, Shirley Wu, James Zou 0001
NeurIPS1
2024 Physics-Informed Regularization for Domain-Agnostic Dynamical System Modeling
abstract
Learning complex physical dynamics purely from data is challenging due to the intrinsic properties of systems to be satisfied. Incorporating physics-informed priors, such as in Hamiltonian Neural Networks (HNNs), achieves high-precision modeling for energy-conservative systems. However, real-world systems often deviate from strict energy conservation and follow different physical priors. To address this, we present a framework that achieves high-precision modeling for a wide range of dynamical systems from the numerical aspect, by enforcing Time-Reversal Symmetry (TRS) via a novel regularization term. It helps preserve energies for conservative systems while serving as a strong inductive bias for non-conservative, reversible systems. While TRS is a domain-specific physical prior, we present the first theoretical proof that TRS loss can universally improve modeling accuracy by minimizing higher-order Taylor terms in ODE integration, which is numerically beneficial to various systems regardless of their properties, even for irreversible systems. By integrating the TRS loss within neural ordinary differential equation models, the proposed model TREAT demonstrates superior performance on diverse physical systems. It achieves a significant 11.5% MSE improvement in a challenging chaotic triple-pendulum scenario, underscoring TREAT’s broad applicability and effectiveness.
Zijie Huang 0002, Wanjia Zhao, Jingdong Gao, Ziniu Hu, Xiao Luo 0001, Yadi Cao, Yuanzhou Chen, Yizhou Sun, Wei Wang 0010
NeurIPS2
2023 Positive Distribution Pollution: Rethinking Positive Unlabeled Learning from a Unified Perspective
abstract
Positive Unlabeled (PU) learning, which has a wide range of applications, is becoming increasingly prevalent. However, it suffers from problems such as data imbalance, selection bias, and prior agnostic in real scenarios. Existing studies focus on addressing part of these problems, which fail to provide a unified perspective to understand these problems. In this paper, we first rethink these problems by analyzing a typical PU scenario and come up with an insightful point of view that all these problems are inherently connected to one problem, i.e., positive distribution pollution, which refers to the inaccuracy in estimating positive data distribution under very little labeled data. Then, inspired by this insight, we devise a variational model named CoVPU, which addresses all three problems in a unified perspective by targeting the positive distribution pollution problem. CoVPU not only accurately separates the positive data from the unlabeled data based on discrete normalizing flows, but also effectively approximates the positive distribution based on our derived unbiased rebalanced risk estimator and supervises the approximation based on a novel prior-free variational loss. Rigorous theoretical analysis proves the convergence of CoVPU to an optimal Bayesian classifier. Extensive experiments demonstrate the superiority of CoVPU over the state-of-the-art PU learning methods under these problems.
Qianqiao Liang, Mengying Zhu, Yan Wang 0002, Xiuyuan Wang 0002, Wanjia Zhao, Mengyuan Yang 0002
AAAI5