EDBT 2026 Demo / reviewers in the wild / expert
Zekun Hao
dblp:202/2193
· DBLP profile ↗
10ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0002-7234-6962ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Generative modeling · 70% 3D vision · 12% Deep learning architectures and training · 7% | |
| Computer graphics and multimedia
7 papers |
Visual content generation and editing · 43% Rendering · 24% Geometric modeling and processing · 19% | |
| Network and information security
1 paper |
Digital forensics and information hiding · 100% |
Topics — the 26 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling › autoregressive model
autoregressive mesh generation |
0.9 | 1 | 2025 | EdgeRunner: Auto-regressive Auto-encoder for Artistic Mesh Generation · ICLR 2025 |
Machine learning › Generative modeling
autoregressive model |
0.9 | 1 | 2025 | EdgeRunner: Auto-regressive Auto-encoder for Artistic Mesh Generation · ICLR 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | EdgeRunner: Auto-regressive Auto-encoder for Artistic Mesh Generation · ICLR 2025 |
Machine learning › Generative modeling › diffusion model
latent diffusion model |
0.9 | 1 | 2025 | EdgeRunner: Auto-regressive Auto-encoder for Artistic Mesh Generation · ICLR 2025 |
Visual content generation and editing
3d content creation |
0.9 | 1 | 2025 | EdgeRunner: Auto-regressive Auto-encoder for Artistic Mesh Generation · ICLR 2025 |
Computational photography and imaging › active illumination
coded illumination |
0.9 | 1 | 2025 | Noise-Coded Illumination for Forensic and Photometric Video Analysis · ACM Trans. Graph. 2025 |
Geometric modeling and processing
mesh generation |
0.9 | 1 | 2025 | EdgeRunner: Auto-regressive Auto-encoder for Artistic Mesh Generation · ICLR 2025 |
Digital forensics and information hiding › digital forensics › multimedia forensics
video forensics |
0.9 | 1 | 2025 | Noise-Coded Illumination for Forensic and Photometric Video Analysis · ACM Trans. Graph. 2025 |
Machine learning › Deep learning architectures and training › feedforward neural network › MLP-based architecture
coordinate network |
0.6 | 1 | 2022 | Implicit Neural Representations with Levels-of-Experts · NeurIPS 2022 |
Computer vision › 3D vision
implicit neural representation |
0.6 | 1 | 2022 | Implicit Neural Representations with Levels-of-Experts · NeurIPS 2022 |
Rendering
neural radiance fields |
0.6 | 1 | 2022 | Implicit Neural Representations with Levels-of-Experts · NeurIPS 2022 |
Rendering
novel view synthesis |
0.6 | 1 | 2022 | Implicit Neural Representations with Levels-of-Experts · NeurIPS 2022 |
Visual content generation and editing
3d content generation |
0.5 | 1 | 2021 | GANcraft: Unsupervised 3D Neural Rendering of Minecraft Worlds · ICCV 2021 |
Rendering
neural rendering |
0.5 | 1 | 2021 | GANcraft: Unsupervised 3D Neural Rendering of Minecraft Worlds · ICCV 2021 |
Visual content generation and editing › 3d content editing
3d object manipulation |
0.4 | 1 | 2020 | DualSDF: Semantic Shape Manipulation Using a Two-Level Representation · CVPR 2020 |
Visual content generation and editing
3d shape generation |
0.4 | 1 | 2020 | Learning Gradient Fields for Shape Generation · ECCV (3) 2020 |
Geometric modeling and processing
shape representation |
0.4 | 1 | 2020 | DualSDF: Semantic Shape Manipulation Using a Two-Level Representation · CVPR 2020 |
Machine learning › Generative modeling › normalizing flow
continuous normalizing flow |
0.4 | 1 | 2019 | PointFlow: 3D Point Cloud Generation With Continuous Normalizing Flows · ICCV 2019 |
Machine learning › Generative modeling
normalizing flow |
0.4 | 1 | 2019 | PointFlow: 3D Point Cloud Generation With Continuous Normalizing Flows · ICCV 2019 |
Computer vision › 3D vision › 3d generation
point cloud generation |
0.4 | 1 | 2019 | PointFlow: 3D Point Cloud Generation With Continuous Normalizing Flows · ICCV 2019 |
Visual content generation and editing › video generation
controllable video generation |
0.3 | 1 | 2018 | Controllable Video Generation With Sparse Trajectories · CVPR 2018 |
Visual content generation and editing
video generation |
0.3 | 1 | 2018 | Controllable Video Generation With Sparse Trajectories · CVPR 2018 |
Computer vision › Face, body and person analysis
face detection |
0.3 | 1 | 2017 | Scale-Aware Face Detection · CVPR 2017 |
Computer vision › Image recognition and object detection › object detection
scale-aware detection |
0.3 | 1 | 2017 | Scale-Aware Face Detection · CVPR 2017 |
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training |
0.1 | 1 | 2021 | GANcraft: Unsupervised 3D Neural Rendering of Minecraft Worlds · ICCV 2021 |
Machine learning › Generative modeling
generative adversarial network |
0.1 | 1 | 2021 | GANcraft: Unsupervised 3D Neural Rendering of Minecraft Worlds · ICCV 2021 |
Methods — techniques the papers use, named apart from their topics
tokenization · 1.7illumination coding · 1.7autoencoder · 1.7mixture of experts · 1.1MLP · 1.1volumetric function · 1.0adversarial training · 1.0pseudo-ground truth · 0.5pseudo ground truth · 0.5variational autoencoder · 0.4score-based generative model · 0.4latent space learning · 0.4variational inference · 0.4convolutional neural network · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EdgeRunner: Auto-regressive Auto-encoder for Artistic Mesh GenerationabstractCurrent auto-regressive mesh generation methods suffer from issues such as incompleteness, insufficient detail, and poor generalization.
In this paper, we propose an Auto-regressive Auto-encoder (ArAE) model capable of generating high-quality 3D meshes with up to 4,000 faces at a spatial resolution of $512^3$.
We introduce a novel mesh tokenization algorithm that efficiently compresses triangular meshes into 1D token sequences, significantly enhancing training efficiency.
Furthermore, our model compresses variable-length triangular meshes into a fixed-length latent space, enabling training latent diffusion models for better generalization.
Extensive experiments demonstrate the superior quality, diversity, and generalization capabilities of our model in both point cloud and image-conditioned mesh generation tasks. Jiaxiang Tang, Zhaoshuo Li, Zekun Hao, Ming-Yu Liu 0001, Qinsheng Zhang |
ICLR | 3 |
| 2025 | Efficient Part-level 3D Object Generation via Dual Volume PackingabstractRecent progress in 3D object generation has greatly improved both the quality and efficiency.
However, most existing methods generate a single mesh with all parts fused together, which limits the ability to edit or manipulate individual parts.
A key challenge is that different objects may have a varying number of parts.
To address this, we propose a new end-to-end framework for part-level 3D object generation.
Given a single input image, our method generates high-quality 3D objects with an arbitrary number of complete and semantically meaningful parts.
We introduce a dual volume packing strategy that organizes all parts into two complementary volumes, allowing for the creation of complete and interleaved parts that assemble into the final object.
Experiments show that our model achieves better quality, diversity, and generalization than previous image-based part-level generation methods.
Our project page is at \url{https://research.nvidia.com/labs/dir/partpacker/}. Jiaxiang Tang, Ruijie Lu, Max Li, Zekun Hao, Xuan Li 0015, Fangyin Wei, Shuran Song, Ming-Yu Liu 0001, Tsung-Yi Lin |
NeurIPS | 4 |
| 2025 | Noise-Coded Illumination for Forensic and Photometric Video AnalysisabstractThe proliferation of advanced tools for manipulating video has led to an arms race, pitting those who wish to sow disinformation against those who want to detect and expose it. Unfortunately, time favors the ill-intentioned in this race, with fake videos growing increasingly difficult to distinguish from real ones. At the root of this trend is a fundamental advantage held by those manipulating media: equal access to a distribution of what we consider authentic (i.e., “natural”) video. In this paper, we show how coding very subtle, noise-like modulations into the illumination of a scene can help combat this advantage by creating an information asymmetry that favors verification. Our approach effectively adds a temporal watermark to any video recorded under coded illumination. However, rather than encoding a specific message, this watermark encodes an image of the unmanipulated scene as it would appear lit only by the coded illumination. We show that even when an adversary knows that our technique is being used, creating a plausible coded fake video amounts to solving a second, more difficult version of the original adversarial content creation problem at an information disadvantage. This is a promising avenue for protecting high-stakes settings like public events and interviews, where the content on display is a likely target for manipulation, and while the illumination can be controlled, the cameras capturing video cannot. Peter F. Michael, Zekun Hao, Serge J. Belongie, Abe Davis |
ACM Trans. Graph. | 2 |
| 2022 | Implicit Neural Representations with Levels-of-ExpertsabstractCoordinate-based networks, usually in the forms of MLPs, have been successfully applied to the task of predicting high-frequency but low-dimensional signals using coordinate inputs. To scale them to model large-scale signals, previous works resort to hybrid representations, combining a coordinate-based network with a grid-based representation, such as sparse voxels. However, such approaches lack a compact global latent representation in its grid, making it difficult to model a distribution of signals, which is important for generalization tasks. To address the limitation, we propose the Levels-of-Experts (LoE) framework, which is a novel coordinate-based representation consisting of an MLP with periodic, position-dependent weights arranged hierarchically. For each linear layer of the MLP, multiple candidate values of its weight matrix are tiled and replicated across the input space, with different layers replicating at different frequencies. Based on the input, only one of the weight matrices is chosen for each layer. This greatly increases the model capacity without incurring extra computation or compromising generalization capability. We show that the new representation is an efficient and competitive drop-in replacement for a wide range of tasks, including signal fitting, novel view synthesis, and generative modeling. Zekun Hao, Arun Mallya, Serge J. Belongie, Ming-Yu Liu 0001 |
NeurIPS | 1 |
| 2021 | GANcraft: Unsupervised 3D Neural Rendering of Minecraft WorldsabstractWe present GANcraft, an unsupervised neural rendering framework for generating photorealistic images of large 3D block worlds such as those created in Minecraft. Our method takes a semantic block world as input, where each block is assigned a semantic label such as dirt, grass, or water. We represent the world as a continuous volumetric function and train our model to render view-consistent photorealistic images for a user-controlled camera. In the absence of paired ground truth real images for the block world, we devise a training technique based on pseudo-ground truth and adversarial training. This stands in contrast to prior work on neural rendering for view synthesis, which requires ground truth images to estimate scene geometry and view-dependent appearance. In addition to camera trajectory, GANcraft allows user control over both scene semantics and output style. Experimental results with comparison to strong baselines show the effectiveness of GANcraft on this novel task of photorealistic 3D block world synthesis. The project website is available at https://nvlabs.github.io/GANcraft/. Zekun Hao, Arun Mallya, Serge J. Belongie, Ming-Yu Liu 0001 |
ICCV | 1 |
| 2020 | DualSDF: Semantic Shape Manipulation Using a Two-Level RepresentationabstractWe are seeing a Cambrian explosion of 3D shape representations for use in machine learning. Some representations seek high expressive power in capturing high-resolution detail. Other approaches seek to represent shapes as compositions of simple parts, which are intuitive for people to understand and easy to edit and manipulate. However, it is difficult to achieve both fidelity and interpretability in the same representation. We propose DualSDF, a representation expressing shapes at two levels of granularity, one capturing fine details and the other representing an abstracted proxy shape using simple and semantically consistent shape primitives. To achieve a tight coupling between the two representations, we use a variational objective over a shared latent space. Our two-level model gives rise to a new shape manipulation technique in which a user can interactively manipulate the coarse proxy shape and see the changes instantly mirrored in the high-resolution shape. Moreover, our model actively augments and guides the manipulation towards producing semantically meaningful shapes, making complex manipulations possible with minimal user input. Zekun Hao, Hadar Averbuch-Elor, Noah Snavely, Serge J. Belongie |
CVPR | 1 |
| 2020 | Learning Gradient Fields for Shape Generation
Ruojin Cai, Guandao Yang, Hadar Averbuch-Elor, Zekun Hao, Serge J. Belongie, Noah Snavely, Bharath Hariharan |
ECCV (3) | 4 |
| 2019 | PointFlow: 3D Point Cloud Generation With Continuous Normalizing FlowsabstractAs 3D point clouds become the representation of choice for multiple vision and graphics applications, the ability to synthesize or reconstruct high-resolution, high-fidelity point clouds becomes crucial. Despite the recent success of deep learning models in discriminative tasks of point clouds, generating point clouds remains challenging. This paper proposes a principled probabilistic framework to generate 3D point clouds by modeling them as a distribution of distributions. Specifically, we learn a two-level hierarchy of distributions where the first level is the distribution of shapes and the second level is the distribution of points given a shape. This formulation allows us to both sample shapes and sample an arbitrary number of points from a shape. Our generative model, named PointFlow, learns each level of the distribution with a continuous normalizing flow. The invertibility of normalizing flows enables the computation of the likelihood during training and allows us to train our model in the variational inference framework. Empirically, we demonstrate that PointFlow achieves state-of-the-art performance in point cloud generation. We additionally show that our model can faithfully reconstruct point clouds and learn useful representations in an unsupervised manner. The code is available at https://github.com/stevenygd/PointFlow. Guandao Yang, Xun Huang 0002, Zekun Hao, Ming-Yu Liu 0001, Serge J. Belongie, Bharath Hariharan |
ICCV | 3 |
| 2018 | Controllable Video Generation With Sparse TrajectoriesabstractVideo generation and manipulation is an important yet challenging task in computer vision. Existing methods usually lack ways to explicitly control the synthesized motion. In this work, we present a conditional video generation model that allows detailed control over the motion of the generated video. Given the first frame and sparse motion trajectories specified by users, our model can synthesize a video with corresponding appearance and motion. We propose to combine the advantage of copying pixels from the given frame and hallucinating the lightness difference from scratch which help generate sharp video while keeping the model robust to occlusion and lightness change. We also propose a training paradigm that calculate trajectories from video clips, which eliminated the need of annotated training data. Experiments on several standard benchmarks demonstrate that our approach can generate realistic videos comparable to state-of-the-art video generation and video prediction methods while the motion of the generated videos can correspond well with user input. Zekun Hao, Xun Huang 0002, Serge J. Belongie |
CVPR | 1 |
| 2017 | Scale-Aware Face DetectionabstractConvolutional neural network (CNN) based face detectors are inefficient in handling faces of diverse scales. They rely on either fitting a large single model to faces across a large scale range or multi-scale testing. Both are computationally expensive. We propose Scale-aware Face Detection (SAFD) to handle scale explicitly using CNN, and achieve better performance with less computation cost. Prior to detection, an efficient CNN predicts the scale distribution histogram of the faces. Then the scale histogram guides the zoom-in and zoom-out of the image. Since the faces will be approximately in uniform scale after zoom, they can be detected accurately even with much smaller CNN. Actually, more than 99% of the faces in AFW can be covered with less than two zooms per image. Extensive experiments on FDDB, MALF and AFW show advantages of SAFD. Zekun Hao, Yu Liu 0015, Hongwei Qin, Xiu Li 0001, Xiaolin Hu 0001 |
CVPR | 1 |