VLDB 2026 Research / reviewers in the wild / expert
Zhongjian Wang
dblp:149/5050
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Generative modeling · 33% 3D vision · 29% Optimization for machine learning · 19% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 100% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Optimization for machine learning
convergence analysis |
0.9 | 1 | 2025 | Global Well-posedness and Convergence Analysis of Score-based Generative Models via Sharp Lipschitz Estimates · ICLR 2025 |
Machine learning › Generative modeling › diffusion model
score-based generative model |
0.9 | 1 | 2025 | Global Well-posedness and Convergence Analysis of Score-based Generative Models via Sharp Lipschitz Estimates · ICLR 2025 |
Machine learning › Learning theory
well-posedness |
0.9 | 1 | 2025 | Global Well-posedness and Convergence Analysis of Score-based Generative Models via Sharp Lipschitz Estimates · ICLR 2025 |
Computer vision › 3D vision › neural radiance field
deformable neural radiance field |
0.7 | 1 | 2023 | One-Shot High-Fidelity Talking-Head Synthesis with Deformable Neural Radiance Field · CVPR 2023 |
Computer vision › 3D vision
neural radiance field |
0.7 | 1 | 2023 | One-Shot High-Fidelity Talking-Head Synthesis with Deformable Neural Radiance Field · CVPR 2023 |
Machine learning › Generative modeling › face synthesis
talking face generation |
0.7 | 1 | 2023 | One-Shot High-Fidelity Talking-Head Synthesis with Deformable Neural Radiance Field · CVPR 2023 |
Visual content generation and editing › image generation
face image generation |
0.7 | 1 | 2023 | One-Shot High-Fidelity Talking-Head Synthesis with Deformable Neural Radiance Field · CVPR 2023 |
Visual content generation and editing
talking head generation |
0.7 | 1 | 2023 | One-Shot High-Fidelity Talking-Head Synthesis with Deformable Neural Radiance Field · CVPR 2023 |
Methods — techniques the papers use, named apart from their topics
multi-scale volume features · 1.3deformable neural radiance field · 1.3stochastic differential equation · 0.9lipschitz estimate · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Position-Based Taxonomy of In-Generation Watermarking for Latent Diffusion Models
Zhongjian Wang, Keke Gai, Jing Yu 0007 |
KSEM (2) | 1 |
| 2026 | Research on precision recommendation algorithm based on the integration of deep learning and self-attention mechanism
Wenxin Zhao, Zhongjian Wang, Zhenbin Liu |
J. Supercomput. | 3 |
| 2025 | Global Well-posedness and Convergence Analysis of Score-based Generative Models via Sharp Lipschitz EstimatesabstractWe establish global well-posedness and convergence of the score-based generative models (SGM) under minimal general assumptions of initial data for score estimation. For the smooth case, we start from a Lipschitz bound of the score function with optimal time length. The optimality is validated by an example whose Lipschitz constant of scores is bounded at initial but blows up in finite time. This necessitates the separation of time scales in conventional bounds for non-log-concave distributions. In contrast, our follow up analysis only relies on a local Lipschitz condition and is valid globally in time. This leads to the convergence of numerical scheme without time separation. For the non-smooth case, we show that the optimal Lipschitz bound is $O(1/t)$ in the point-wise sense for distributions supported on a compact, smooth and low-dimensional manifold with boundary. Connor Mooney, Zhongjian Wang, Jack Xin |
ICLR | 2 |
| 2025 | OmniTalker: One-shot Real-time Text-Driven Talking Audio-Video Generation With Multimodal Style MimickingabstractAlthough significant progress has been made in audio-driven talking head generation, text-driven methods remain underexplored. In this work, we present OmniTalker, a unified framework that jointly generates synchronized talking audio-video content from input text while emulating the target identity's speaking and facial movement styles, including speech characteristics, head motion, and facial dynamics. Our framework adopts a dual-branch diffusion transformer (DiT) architecture, with one branch dedicated to audio generation and the other to video synthesis.
At the shallow layers, cross-modal fusion modules are introduced to integrate information between the two modalities. In deeper layers, each modality is processed independently, with the generated audio decoded by a vocoder and the video rendered using a GAN-based high-quality visual renderer. Leveraging DiT’s in-context learning capability through a masked-infilling strategy, our model can simultaneously capture both audio and visual styles without requiring explicit style extraction modules. Thanks to the efficiency of the DiT backbone and the optimized visual renderer, OmniTalker achieves real-time inference at 25 FPS.
To the best of our knowledge, OmniTalker is the first one-shot framework capable of jointly modeling speech and facial styles in real time. Extensive experiments demonstrate its superiority over existing methods in terms of generation quality, particularly in preserving style consistency and ensuring precise audio-video synchronization, all while maintaining efficient inference. Zhongjian Wang, Peng Zhang 0080, Jinwei Qi, Sheng Xu 0007, Bang Zhang |
NeurIPS | 1 |
| 2023 | One-Shot High-Fidelity Talking-Head Synthesis with Deformable Neural Radiance FieldabstractTalking head generation aims to generate faces that maintain the identity information of the source image and imitate the motion of the driving image. Most pioneering methods rely primarily on 2D representations and thus will inevitably suffer from face distortion when large head rotations are encountered. Recent works instead employ explicit 3D structural representations or implicit neural rendering to improve performance under large pose changes. Nevertheless, the fidelity of identity and expression is not so desirable, especially for novel-view synthesis. In this paper, we propose HiDe-NeRF, which achieves high-fidelity and free-view talking-head synthesis. Drawing on the recently proposed Deformable Neural Radiance Fields, HiDe-NeRF represents the 3D dynamic scene into a canonical appearance field and an implicit deformation field, where the former comprises the canonical source face and the latter models the driving pose and expression. In particular, we improve fidelity from two aspects: (i) to enhance identity expressiveness, we design a generalized appearance module that leverages multi-scale volume features to preserve face shape and details; (ii) to improve expression preciseness, we propose a lightweight deformation module that explicitly decouples the pose and expression to enable precise expression modeling. Extensive experiments demonstrate that our proposed approach can generate better results than previous works. Project page: https://www.waytron.net/hidenerf/ Weichuang Li, Longhao Zhang, Dong Wang 0028, Bin Zhao 0001, Zhigang Wang 0002, Mulin Chen, Bang Zhang, Zhongjian Wang, Liefeng Bo, Xuelong Li 0001 |
CVPR | 8 |