VLDB 2026 Research / reviewers in the wild / expert
Yang Zhao 0003
dblp:50/2082-3
· DBLP profile ↗
17ranked-venue papers
4as first author
12since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Generative modeling · 50% Representation and self-supervised learning · 14% 3D vision · 14% | |
| Computer graphics and multimedia
2 papers |
Image and video processing · 50% Visual content generation and editing · 50% |
Topics — the 22 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
video generation |
2.6 | 3 | 2025 | Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation · NeurIPS 2025 How Far Is Video Generation from World Model: A Physical Law Perspective · ICML 2025 Videoauteur: Towards Long Narrative Video Generation · ICCV 2025 |
Machine learning › Generative modeling
diffusion model |
1.7 | 2 | 2025 | Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation · NeurIPS 2025 SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video Restoration · CVPR 2025 |
Machine learning › Generative modeling › video generation
autoregressive video generation |
0.9 | 1 | 2025 | Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model
diffusion transformer |
0.9 | 1 | 2025 | SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video Restoration · CVPR 2025 |
Machine learning › Generative modeling › diffusion model › video diffusion model
latent video diffusion model |
0.9 | 1 | 2025 | Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation · NeurIPS 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations · NeurIPS 2025 |
Computer vision › 3D vision
physical law discovery |
0.9 | 1 | 2025 | How Far Is Video Generation from World Model: A Physical Law Perspective · ICML 2025 |
Computer vision › Vision and language
video captioning |
0.9 | 1 | 2025 | Videoauteur: Towards Long Narrative Video Generation · ICCV 2025 |
Machine learning › Reinforcement learning › model-based reinforcement learning
world model |
0.9 | 1 | 2025 | How Far Is Video Generation from World Model: A Physical Law Perspective · ICML 2025 |
Visual content generation and editing
image generation |
0.9 | 1 | 2025 | Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations · NeurIPS 2025 |
Image and video processing
video restoration |
0.9 | 1 | 2025 | SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video Restoration · CVPR 2025 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › reconstruction-based representation learning
feature reconstruction |
0.8 | 1 | 2024 | Image Understanding Makes for A Good Tokenizer for Image Generation · NeurIPS 2024 |
Machine learning › Generative modeling
image generation |
0.8 | 1 | 2024 | Image Understanding Makes for A Good Tokenizer for Image Generation · NeurIPS 2024 |
Machine learning › Generative modeling
image tokenization |
0.8 | 1 | 2024 | Image Understanding Makes for A Good Tokenizer for Image Generation · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning › representation learning › visual representation learning
visual tokenizer |
0.8 | 1 | 2024 | Image Understanding Makes for A Good Tokenizer for Image Generation · NeurIPS 2024 |
Computer vision › 3D vision
3d face reconstruction |
0.5 | 1 | 2021 | 3D Face Reconstruction From A Single Image Assisted by 2D Face Images in the Wild · IEEE Trans. Multim. 2021 |
Computer vision › 3D vision › 3d face reconstruction
dense face alignment |
0.5 | 1 | 2021 | 3D Face Reconstruction From A Single Image Assisted by 2D Face Images in the Wild · IEEE Trans. Multim. 2021 |
Computer vision › Face, body and person analysis
face alignment |
0.5 | 1 | 2021 | 3D Face Reconstruction From A Single Image Assisted by 2D Face Images in the Wild · IEEE Trans. Multim. 2021 |
Computer vision › 3D vision › 3d face reconstruction
single-image 3d face reconstruction |
0.5 | 1 | 2021 | 3D Face Reconstruction From A Single Image Assisted by 2D Face Images in the Wild · IEEE Trans. Multim. 2021 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.3 | 1 | 2025 | Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation · NeurIPS 2025 |
Natural language and speech › Language models and text generation › language modeling
scaling behavior |
0.3 | 1 | 2025 | How Far Is Video Generation from World Model: A Physical Law Perspective · ICML 2025 |
Machine learning › Trustworthy machine learning › interpretability
visual explanation |
0.2 | 1 | 2024 | Image Understanding Makes for A Good Tokenizer for Image Generation · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
diffusion model · 2.6progressive training · 1.7causal video autoencoder · 1.7autoregressive model · 1.7vision-language model · 0.9simulation testbed · 0.9shifted-window attention · 0.9shifted window attention · 0.9fine-tuning · 0.9embedding alignment · 0.9adversarial post-training · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video RestorationabstractVideo restoration poses non-trivial challenges in maintaining fidelity while recovering temporally consistent details from unknown degradations in the wild. Despite recent advances in diffusion-based restoration, these methods often face limitations in generation capability and sampling efficiency. In this work, we present SeedVR, a diffusion transformer designed to handle real-world video restoration with arbitrary length and resolution. The core design of SeedVR lies in the shifted window attention that facilitates effective restoration on long video sequences. SeedVR further supports variable-sized windows near the boundary of both spatial and temporal dimensions, overcoming the resolution constraints of traditional window attention. Equipped with contemporary practices, including causal video autoencoder, mixed image and video training, and progressive training, SeedVR achieves highly-competitive performance on both synthetic and real-world benchmarks, as well as AI-generated videos. Extensive experiments demonstrate SeedVR’s superiority over existing methods for generic video restoration. Jianyi Wang, Zhijie Lin 0001, Meng Wei 0007, Yang Zhao 0003, Ceyuan Yang, Chen Change Loy |
CVPR | 4 |
| 2025 | Videoauteur: Towards Long Narrative Video GenerationabstractRecent video generation models have shown promising results in producing high-quality video clips lasting several seconds. However, these models face challenges in generating long sequences that convey clear and informative events, limiting their ability to support coherent narrations. In this paper, we present a large-scale cooking video dataset designed to advance long-form narrative generation in the cooking domain. We validate the quality of our proposed dataset in terms of visual fidelity and textual caption accuracy using state-of-the-art Vision-Language Models (VLMs) and video generation models, respectively. We further introduce a Long Narrative Video Director to enhance both visual and semantic coherence in generated videos and emphasize the role of aligning visual embeddings to achieve improved overall video quality. Our method demonstrates substantial improvements in generating visually detailed and semantically aligned keyframes, supported by finetuning techniques that integrate text and image embeddings within the video generation process. Project page: https://videoauteur.github.io/ Junfei Xiao, Liangke Gui, Yang Zhao 0003, Shanchuan Lin, Jiepeng Cen, Zhibei Ma, Alan L. Yuille |
ICCV | 5 |
| 2025 | How Far Is Video Generation from World Model: A Physical Law PerspectiveabstractScaling video generation models is believed to be promising in building world models that adhere to fundamental physical laws. However, whether these models can discover physical laws purely from vision can be questioned. A world model learning the true law should give predictions robust to nuances and correctly extrapolate on unseen scenarios. In this work, we evaluate across three key scenarios: in-distribution, out-of-distribution, and combinatorial generalization. We developed a 2D simulation testbed for object movement and collisions to generate videos deterministically governed by one or more classical mechanics laws. We focus on the scaling behavior of training diffusion-based video generation models to predict object movements based on initial frames. Our scaling experiments show perfect generalization within the distribution, measurable scaling behavior for combinatorial generalization, but failure in out-of-distribution scenarios. Further experiments reveal two key insights about the generalization mechanisms of these models: (1) the models fail to abstract general physical rules and instead exhibit "case-based" generalization behavior, i.e., mimicking the closest training example; (2) when generalizing to new cases, models are observed to prioritize different factors when referencing training data: color $>$ size $>$ velocity $>$ shape. Our study suggests that scaling alone is insufficient for video generation models to uncover fundamental physical laws. Bingyi Kang, Rui Lu 0001, Zhijie Lin 0001, Yang Zhao 0003, Gao Huang 0001, Jiashi Feng |
ICML | 5 |
| 2025 | Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned RepresentationsabstractThis paper presents a multimodal framework that attempts to unify visual understanding and generation within a shared discrete semantic representation. At its core is the Text-Aligned Tokenizer (TA-Tok), which converts images into discrete tokens using a text-aligned codebook projected from a large language model's (LLM) vocabulary. By integrating vision and text into a unified space with an expanded vocabulary, our multimodal LLM, **Tar**, enables cross-modal input and output through a shared interface, without the need for modality-specific designs. Additionally, we propose scale-adaptive encoding and decoding to balance efficiency and visual detail, along with a
generative de-tokenizer to produce high-fidelity visual outputs. To address diverse decoding needs, we utilize two complementary de-tokenizers: a fast autoregressive model and a diffusion-based model. To enhance modality fusion, we investigate advanced pre-training tasks, demonstrating improvements in both visual understanding and generation. Experiments across benchmarks show that **Tar** matches or surpasses existing multimodal LLM methods, achieving faster convergence and greater training efficiency. All code, models, and data will be made publicly available. Jiaming Han, Yang Zhao 0003, Qi Zhao 0001, Xiangyu Yue 0001 |
NeurIPS | 3 |
| 2025 | Autoregressive Adversarial Post-Training for Real-Time Interactive Video GenerationabstractExisting large-scale video generation models are computationally intensive, preventing adoption in real-time and interactive applications. In this work, we propose autoregressive adversarial post-training (AAPT) to turn a pre-trained latent video diffusion model into
a real-time, interactive, streaming video generator. Our model autoregressively generates a latent frame at a time using a single neural function evaluation (1NFE). The model can stream the result to the user in real time and receive interactive responses as control to generate the next latent frame. Unlike existing approaches, our method explores adversarial training as an effective paradigm for autoregressive generation. This allows us to design a more efficient architecture for one-step generation and to train the model in a student-forcing way to mitigate error accumulation. The adversarial approach also enables us to train the model for long-duration generation fully utilizing the KV cache. As a result, our 8B model achieves real-time, 24fps, nonstop, streaming video generation at 736x416 resolution on a single H100, or 1280x720 on 8xH100 up to a minute long (1440 frames). Shanchuan Lin, Ceyuan Yang, Jianwen Jiang, Yuxi Ren, Xin Xia 0005, Yang Zhao 0003, Xuefeng Xiao 0001 |
NeurIPS | 7 |
| 2024 | Image Understanding Makes for A Good Tokenizer for Image GenerationabstractModern image generation (IG) models have been shown to capture rich semantics valuable for image understanding (IU) tasks. However, the potential of IU models to improve IG performance remains uncharted. We address this issue using a token-based IG framework, which relies on effective tokenizers to project images into token sequences. Currently, **pixel reconstruction** (e.g., VQGAN) dominates the training objective for image tokenizers. In contrast, our approach adopts the **feature reconstruction** objective, where tokenizers are trained by distilling knowledge from pretrained IU encoders. Comprehensive comparisons indicate that tokenizers with strong IU capabilities achieve superior IG performance across a variety of metrics, datasets, tasks, and proposal networks. Notably, VQ-KD CLIP achieves $4.10$ FID on ImageNet-1k (IN-1k). Visualization suggests that the superiority of VQ-KD can be partly attributed to the rich semantics within the VQ-KD codebook. We further introduce a straightforward pipeline to directly transform IU encoders into tokenizers, demonstrating exceptional effectiveness for IG tasks. These discoveries may energize further exploration into image tokenizer research and inspire the community to reassess the relationship between IU and IG. The code is released at https://github.com/magic-research/vector_quantization. Luting Wang 0001, Yang Zhao 0003, Jiashi Feng, Si Liu 0001, Bingyi Kang |
NeurIPS | 2 |
| 2022 | WRMatch: Improving FixMatch With Weighted Nuclear-Norm Regularization for Few-Shot Remote Sensing Scene ClassificationabstractSemisupervised learning (SSL), such as FixMatch, has been successfully applied to remote sensing scene classification to relieve the burden of data annotation. However, in some extreme settings, only very few samples available, e.g., one to ten labels per remote sensing scene, can be used. When meeting this “few-shot” scenario, the deep model may be overfitting and prone to generate confusing predictions due to the lack of labels and strong augmentation-based perturbations. Thus, the prediction’s diversity may collapse, and the discriminability exceeds the reasonable interval. How to improve the performance of few-shot learning is underexplored for the remote sensing scene classification in previous studies. In this article, we present a novel framework for the task by utilizing the improved FixMatch and the weighted nuclear-norm regularization (WNNR). Specifically, we regularize the prediction matrix by exploiting the nuclear-norm, which is an approximation of the matrix rank and a relaxed boundary for the Shannon entropy. We further provide two weighting schemes to improve the nuclear-norm-based regularization. First, the random-weighting scheme for nuclear-norm (RWNNR) is proposed based on the Dirichlet distribution to improve the model’s generalization. Second, we present the self-weighting scheme (SWNNR) to weight the singular values according to singular values themselves and adjust the relaxed degree for the boundary between the nuclear-norm and the Shannon entropy. Maximizing the weighted nuclear-norm can improve the prediction diversity and optimize the prediction discriminability simultaneously. Combining the advantages of SSL and the aforementioned improvements, we can reliably classify the remote sensing scene image with very limited annotated datasets. To empirically demonstrate the proposed method’s effectiveness, we comprehensively evaluate the method on three publicly available benchmark datasets. The results show that the proposed method outperforms the baseline methods by a large margin and achieves superior performance on all three datasets. Our method can be an effective alternative to metalearning in few-shot scene classification, with the advantage lying in the competitive performance and the absence of metatraining stage associated with a large number of labels. Yunsheng Xiong, Kele Xu, Yong Dou, Yang Zhao 0003, Zikai Gao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | MOOD 2020: A Public Benchmark for Out-of-Distribution Detection and Localization on Medical ImagesabstractDetecting Out-of-Distribution (OoD) data is one of the greatest challenges in safe and robust deployment of machine learning algorithms in medicine. When the algorithms encounter cases that deviate from the distribution of the training data, they often produce incorrect and over-confident predictions. OoD detection algorithms aim to catch erroneous predictions in advance by analysing the data distribution and detecting potential instances of failure. Moreover, flagging OoD cases may support human readers in identifying incidental findings. Due to the increased interest in OoD algorithms, benchmarks for different domains have recently been established. In the medical imaging domain, for which reliable predictions are often essential, an open benchmark has been missing. We introduce the Medical-Out-Of-Distribution-Analysis-Challenge (MOOD) as an open, fair, and unbiased benchmark for OoD methods in the medical imaging domain. The analysis of the submitted algorithms shows that performance has a strong positive correlation with the perceived difficulty, and that all algorithms show a high variance for different anomalies, making it yet hard to recommend them for clinical practice. We also see a strong correlation between challenge ranking and performance on a simple toy test set, indicating that this might be a valuable addition as a proxy dataset during anomaly detection algorithm development. David Zimmerer, Peter M. Full, Fabian Isensee, Paul F. Jaeger, Tim Adler, Jens Petersen, Gregor Köhler, Tobias Roß, Annika Reinke, Antanas Kascenas, Bjørn Sand Jensen, Alison O'Neil, Jeremy Tan, Benjamin Hou, James Batten, Huaqi Qiu, Bernhard Kainz, Nina Shvetsova, Irina Fedulova, Dmitry V. Dylov, Baolun Yu, Jianyang Zhai, Jingtao Hu, Runxuan Si, Sihang Zhou 0001, Siqi Wang 0001, Xuerun Chen, Yang Zhao 0003, Sergio Naval Marimont, Giacomo Tarroni, Victor Saase, Lena Maier-Hein, Klaus H. Maier-Hein |
IEEE Trans. Medical Imaging | 29 |
| 2021 | RFC-HyPGCN: A Runtime Sparse Feature Compress Accelerator for Skeleton-Based GCNs Action Recognition Model with Hybrid PruningabstractSkeleton-based Graph Convolutional Networks (GCNs) models for action recognition have achieved excellent prediction accuracy in the field. However, limited by large model and computation complexity, GCNs for action recognition like 2s-AGCN have insufficient power-efficiency and throughput on GPU. Thus, the demand of model reduction and hardware acceleration for low-power GCNs action recognition application becomes continuously higher.To address challenges above, this paper proposes a runtime sparse feature compress accelerator with hybrid pruning method: RFC-HyPGCN. First, this method skips both graph and spatial convolution workloads by reorganizing the multiplication order. Following spatial convolutions channel-pruning dataflow, a coarse-grained pruning method on temporal filters is designed, together with sampling-like fine-grained pruning on time dimension. Later, we come up with an architecture where all convolutional layers are mapped on chip to pursue high throughput. To further reduce storage resource utilization, online sparse feature compress format is put forward. Features are divided and encoded into several banks according to presented format, then bank storage is split into depth-variable mini-banks. Furthermore, this work applies quantization, input-skipping and intra-PE dynamic data scheduling to accelerate the model. In experiments, proposed pruning method is conducted on 2s-AGCN, acquiring 3.0x-8.4x model compression ratio and 73.20% graph-skipping efficiency with balancing weight pruning. Implemented on Xilinx XCKU-115 FPGA, the proposed architecture has the peak performance of 1142 GOP/s and achieves up to 9.19x and 3.91x speedup over high-end GPU NVIDIA 2080Ti and NVIDIA V100, respectively. Compared with latest accelerator for action recognition GCNs models, our design reaches 22.9x speedup and 28.93% improvement on DSP efficiency. Dong Wen 0004, Jingfei Jiang, Jinwei Xu, Yang Zhao 0003, Yong Dou |
ASAP | 6 |
| 2021 | Local and Non-local Context Graph Convolutional Networks for Skeleton-Based Action Recognition
Zikai Gao, Yang Zhao 0003, Yong Dou |
ICANN (3) | 2 |
| 2021 | COVID Edge-Net: Automated COVID-19 Lung Lesion Edge Detection in Chest CT Images
Yang Zhao 0003, Yong Dou, Dong Wen 0004, Zikai Gao |
ECML/PKDD (4) | 2 |
| 2021 | 3D Face Reconstruction From A Single Image Assisted by 2D Face Images in the Wildabstract3D face reconstruction from a single image is an important task in many multimedia applications. Recent works typically learn a CNN-based 3D face model that regresses coefficients of a 3D Morphable Model (3DMM) from 2D images to perform 3D face reconstruction. However, the shortage of training data with 3D annotations considerably limits performance of these methods. To alleviate this issue, we propose a novel 2D-Assisted Learning (2DAL) method that can effectively use “in the wild” 2D face images with noisy landmark information to substantially improve 3D face model learning. Specifically, taking the sparse 2D facial landmark heatmaps as additional information, 2DAL introduces four novel self-supervision schemes that view the 2D landmark and 3D landmark prediction as a self-mapping process, including the landmark self-prediction consistency for 2D and 3D faces respectively, cycle-consistency over the 2D landmark prediction and self-critic over the predicted 3DMM coefficients based on landmark prediction. Using these four self-supervision schemes, 2DAL significantly relieves the demands for the the conventional paired 2D-to-3D annotations and gives much higher-quality 3D face models without requiring any additional 3D annotations. Experiments on AFLW2000-3D, AFLW-LFPA and Florence benchmarks show that our method outperforms state-of-the-arts for both 3D face reconstruction and dense face alignment by a large margin. Xiaoguang Tu, Jian Zhao 0006, Mei Xie, Zihang Jiang, Akshaya Balamurugan, Yao Luo, Yang Zhao 0003, Lingxiao He, Zheng Ma 0005, Jiashi Feng |
IEEE Trans. Multim. | 7 |
| 2020 | Temporally Refined Graph U-Nets for Human Shape and Pose Estimation From Monocular VideosabstractThis work addresses a challenging problem of estimating the full 3D human shape and pose from monocular videos. Since real-world 3D mesh-labeled datasets are limited, most current methods in 3D human shape reconstruction only focus on single RGB images, losing all the temporal information. In contrast, we propose temporally refined Graph U-Nets, including an image-level module and a video-level module, to solve this problem. The image-level module is Graph U-Nets for human shape and pose estimation from images, where the Graph Convolutional Neural Network (Graph CNN) helps the information communication of neighboring vertices, and the U-Nets architecture enlarges the receptive field of each vertex and fuses high-level and low-level features. The video-level module is a small Residual Temporal Graph CNN (Residual TG-CNN), which learns temporal dynamics from both structural and temporal neighbors. The temporal dynamics of each vertex are continuous in the temporal dimension and highly relevant to the structural neighbors, so it is helpful to diminish the ambiguity of the body in single images by fusing temporal dynamics. Our algorithm makes full use of labels from image-level datasets and refines the image-level results through video-level module. Evaluated on Human3.6 M and 3DPW datasets, our model produces accurate 3D human meshes and achieves superior 3D human pose estimation accuracy when compared with state-of-the-art methods. Yang Zhao 0003, Yong Dou, Jiashi Feng |
IEEE Signal Process. Lett. | 1 |
| 2017 | Multiple kernel clustering with corrupted kernels
Teng Li 0010, Yong Dou, Xinwang Liu 0002, Yang Zhao 0003 |
Neurocomputing | 4 |
| 2016 | Improved Survey Propagation on Graphics Processing Units
Yang Zhao 0003, Jingfei Jiang, Pengbo Wu |
GPC | 1 |
| 2016 | ELM based multiple kernel k-means with diversity-induced regularizationabstractMultiple-kernel k-means (MKKM) clustering has demonstrated good clustering performance by combining pre-specified kernels. In this paper, we argue that deep relationships within data and the complementary information among them can improve the performance of MKKM. To illustrate this idea, we propose a diversity-induced MKKM algorithm with extreme learning machine (ELM)-based feature extracting method. First, ELM, which has randomly chosen weights of hidden and output nodes, is applied to thoroughly extract features from data by generating different numbers of hidden nodes and using different functions. Second, an MKKM algorithm with diversity-induced regularization is utilized to explore the complementary information among kernels constructed from features. The problem could be solved efficiently by alternating optimization. Experimental results demonstrate that the proposed method outperforms state-of-the-art kernel methods. Yang Zhao 0003, Yong Dou, Xinwang Liu 0002, Teng Li 0010 |
IJCNN | 1 |
| 2016 | A novel multi-view clustering method via low-rank and matrix-induced regularization
Yang Zhao 0003, Yong Dou, Xinwang Liu 0002, Teng Li 0010 |
Neurocomputing | 1 |