VLDB 2026 Research / reviewers in the wild / expert
Chenghui Li
dblp:80/2898
· DBLP profile ↗
11ranked-venue papers
0as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Efficient and distributed learning · 33% 3D vision · 30% Robot navigation and mapping · 13% | |
| Computer graphics and multimedia
5 papers |
Virtual and augmented reality · 65% Computer animation and physical simulation · 32% Visual content generation and editing · 4% | |
| Theoretical computer science
2 papers |
Mathematical optimization · 60% Graph algorithms and graph theory · 35% Algorithms and data structures · 5% | |
| Computer networks
1 paper |
Edge and fog computing · 44% Physical-layer communications · 44% Wireless sensing and localization · 13% | |
| Databases, data mining, and information retrieval
2 papers |
Data mining · 100% |
Topics — the 30 heaviest of 32, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Virtual and augmented reality › avatar
avatar animation |
2.5 | 3 | 2025 | Generative Head-Mounted Camera Captures for Photorealistic Avatars · ACM Trans. Graph. 2025 Audio Driven Real-Time Facial Animation for Social Telepresence · SIGGRAPH Asia 2025 Universal Facial Encoding of Codec Avatars from VR Headsets · ACM Trans. Graph. 2024 |
Machine learning › Generative modeling
conditional generative model |
0.9 | 1 | 2025 | Generative Head-Mounted Camera Captures for Photorealistic Avatars · ACM Trans. Graph. 2025 |
Computer animation and physical simulation
facial animation |
0.9 | 1 | 2025 | Audio Driven Real-Time Facial Animation for Social Telepresence · SIGGRAPH Asia 2025 |
Computer animation and physical simulation › facial animation
speech-driven facial animation |
0.9 | 1 | 2025 | Audio Driven Real-Time Facial Animation for Social Telepresence · SIGGRAPH Asia 2025 |
Edge and fog computing
edge intelligence |
0.9 | 1 | 2025 | Enabling Over-the-Air AI for Edge Computing via Metasurface-Driven Physical Neural Networks · SIGCOMM 2025 |
Physical-layer communications
programmable metasurface |
0.9 | 1 | 2025 | Enabling Over-the-Air AI for Edge Computing via Metasurface-Driven Physical Neural Networks · SIGCOMM 2025 |
Computer vision › 3D vision › 3d human reconstruction
human avatar modeling |
0.8 | 1 | 2024 | Codec Avatar Studio: Paired Human Captures for Complete, Driveable, and Generalizable Avatars · NeurIPS 2024 |
Machine learning › Efficient and distributed learning
model compression |
0.8 | 1 | 2024 | POCA: Post-training Quantization with Temporal Alignment for Codec Avatars · ECCV (40) 2024 |
Computer vision › 3D vision › 3d reconstruction
multi-view reconstruction |
0.8 | 1 | 2024 | Universal Facial Encoding of Codec Avatars from VR Headsets · ACM Trans. Graph. 2024 |
Machine learning › Efficient and distributed learning › model compression › quantization
post-training quantization |
0.8 | 1 | 2024 | POCA: Post-training Quantization with Temporal Alignment for Codec Avatars · ECCV (40) 2024 |
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search |
0.7 | 1 | 2023 | Auto-CARD: Efficient and Robust Codec Avatar Driving for Real-time Mobile Telepresence · CVPR 2023 |
Data mining
clustering |
0.7 | 1 | 2023 | Large sample spectral analysis of graph-based multi-manifold clustering · J. Mach. Learn. Res. 2023 |
Data mining › clustering › high-dimensional clustering
multi-manifold clustering |
0.7 | 1 | 2023 | Large sample spectral analysis of graph-based multi-manifold clustering · J. Mach. Learn. Res. 2023 |
Graph algorithms and graph theory › spectral graph theory
graph laplacian |
0.7 | 1 | 2023 | Large sample spectral analysis of graph-based multi-manifold clustering · J. Mach. Learn. Res. 2023 |
Graph algorithms and graph theory › graph clustering
spectral clustering |
0.7 | 1 | 2023 | Large sample spectral analysis of graph-based multi-manifold clustering · J. Mach. Learn. Res. 2023 |
Mathematical optimization › continuous optimization › convex optimization › first-order methods › gradient-based optimization
accelerated gradient methods |
0.6 | 1 | 2022 | A Fast Scale-Invariant Algorithm for Non-negative Least Squares with Non-negative Data · NeurIPS 2022 |
Mathematical optimization › continuous optimization
convex optimization |
0.6 | 1 | 2022 | A Fast Scale-Invariant Algorithm for Non-negative Least Squares with Non-negative Data · NeurIPS 2022 |
Mathematical optimization › continuous optimization › convex optimization
first-order methods |
0.6 | 1 | 2022 | A Fast Scale-Invariant Algorithm for Non-negative Least Squares with Non-negative Data · NeurIPS 2022 |
Mathematical optimization › least squares
non-negative least squares |
0.6 | 1 | 2022 | A Fast Scale-Invariant Algorithm for Non-negative Least Squares with Non-negative Data · NeurIPS 2022 |
Virtual and augmented reality › telepresence
avatar-mediated telepresence |
0.5 | 2 | 2025 | Generative Head-Mounted Camera Captures for Photorealistic Avatars · ACM Trans. Graph. 2025 Universal Facial Encoding of Codec Avatars from VR Headsets · ACM Trans. Graph. 2024 |
Virtual and augmented reality
telepresence |
0.5 | 2 | 2025 | Generative Head-Mounted Camera Captures for Photorealistic Avatars · ACM Trans. Graph. 2025 Universal Facial Encoding of Codec Avatars from VR Headsets · ACM Trans. Graph. 2024 |
Robotics › Robot navigation and mapping › sensor calibration
LiDAR-camera extrinsic calibration |
0.4 | 1 | 2020 | Online Camera-LiDAR Calibration with Sensor Semantic Information · ICRA 2020 |
Computer vision › 3D vision
multimodal perception |
0.4 | 1 | 2020 | Online Camera-LiDAR Calibration with Sensor Semantic Information · ICRA 2020 |
Robotics › Autonomous driving
perception |
0.4 | 1 | 2020 | Online Camera-LiDAR Calibration with Sensor Semantic Information · ICRA 2020 |
Robotics › Robot navigation and mapping
sensor calibration |
0.4 | 1 | 2020 | Online Camera-LiDAR Calibration with Sensor Semantic Information · ICRA 2020 |
Computer animation and physical simulation › audio-driven animation
speech-driven animation |
0.3 | 1 | 2025 | Audio Driven Real-Time Facial Animation for Social Telepresence · SIGGRAPH Asia 2025 |
Wireless sensing and localization
multi-sensor fusion |
0.3 | 1 | 2025 | Enabling Over-the-Air AI for Edge Computing via Metasurface-Driven Physical Neural Networks · SIGCOMM 2025 |
Visual content generation and editing › avatar generation
codec avatars |
0.2 | 1 | 2024 | POCA: Post-training Quantization with Temporal Alignment for Codec Avatars · ECCV (40) 2024 |
Algorithms and data structures › numerical linear algebra › dimensionality reduction › nonlinear dimensionality reduction
manifold learning |
0.2 | 1 | 2023 | Large sample spectral analysis of graph-based multi-manifold clustering · J. Mach. Learn. Res. 2023 |
Data mining
direct marketing |
0.0 | 1 | 1998 | Data Mining for Direct Marketing: Problems and Solutions · KDD 1998 |
Methods — techniques the papers use, named apart from their topics
unpaired learning · 1.7generative model · 1.7disentanglement · 1.7temporal alignment · 1.5self-supervised learning · 1.5expression calibration · 1.5cross-view reconstruction · 1.5spectral analysis · 1.3graph laplacian · 1.3error bounds · 1.3transformer · 0.9metasurface · 0.9linear neural network · 0.9knowledge distillation · 0.9diffusion model · 0.9relightable capture · 0.8multi-view capture · 0.8avatar generation models · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enabling Over-the-Air AI for Edge Computing via Metasurface-Driven Physical Neural NetworksabstractWe present MetaAI, a novel wireless computing paradigm that integrates neural network computation directly into wireless signal propagation. Unlike traditional approaches that treat wireless channels as mere data conduits, MetaAI transforms them into active computing elements through programmable metasurfaces, enabling concurrent data transmission and neural network processing. By leveraging the inherent linearity of both wireless propagation and neural networks, our design resolves the fundamental mismatch between sequential wireless transmission and parallel neural computation, while supporting efficient multi-sensor late-stage data fusion. We implemented MetaAI using metasurfaces at both dual-band (2.4/5 GHz) and single-band (3.5 GHz) frequencies. Extensive experiments demonstrate robust performance across diverse classification tasks, achieving 82.8% average accuracy (up to 89.8%) even with a simple linear architecture. Multi-sensor fusion further improves accuracy by up to 27.06%. MetaAI represents a fundamental shift in Edge AI architecture, where wireless infrastructure becomes an integral part of the computing pipeline. Chao Feng 0004, Shuo Liang, Chenghui Li, Gaoteng Zhao, Beier Jing, Yaxiong Xie, Xiaojiang Chen |
SIGCOMM | 3 |
| 2025 | Audio Driven Real-Time Facial Animation for Social TelepresenceabstractWe present an audio-driven real-time system for animating photorealistic 3D facial avatars with minimal latency, designed for social interactions in virtual reality for anyone. Central to our approach is an encoder model that transforms audio signals into latent facial expression sequences in real time, which are then decoded as photorealistic 3D facial avatars. Leveraging the generative capabilities of diffusion models, we capture the rich spectrum of facial expressions necessary for natural communication while achieving real-time performance (<15ms GPU time). Our novel architecture minimizes latency through two key innovations: an online transformer that eliminates dependency on future inputs and a distillation pipeline that accelerates iterative denoising into a single step. We further address critical design challenges in live scenarios for processing continuous audio signals frame-by-frame while maintaining consistent animation quality. The versatility of our framework extends to multimodal applications, including semantic modalities such as emotion conditions and multimodal sensors with head-mounted eye cameras on VR headsets. Experimental results demonstrate significant improvements in facial animation accuracy over existing offline state-of-the-art baselines, achieving 100 to 1000 × faster inference speed. We validate our approach through live VR demonstrations and across various scenarios such as multilingual speeches. Jiye Lee 0001, Chenghui Li, Shih-En Wei, Jason M. Saragih, Alexander Richard, Hanbyul Joo, Shaojie Bai |
SIGGRAPH Asia | 2 |
| 2025 | Generative Head-Mounted Camera Captures for Photorealistic AvatarsabstractEnabling photorealistic avatar animations in virtual and augmented reality (VR/AR) has been challenging because of the difficulty of obtaining ground truth state of faces. It is physically impossible to obtain synchronized images from head-mounted cameras (HMC) sensing input, which has partial observations in infrared (IR), and an array of outside-in dome cameras, which have full observations that match avatars' appearance. Prior works relying on analysis-by-synthesis methods could generate accurate ground truth, but suffer from imperfect disentanglement between expression and style in their personalized training. The reliance of extensive paired captures (HMC and dome) for the same subject makes it operationally expensive to collect large-scale datasets, which cannot be reused for different HMC viewpoints and lighting. In this work, we propose a novel generative approach, Generative HMC (GenHMC), that leverages large unpaired HMC captures , which are much easier to collect, to directly generate high-quality synthetic HMC images given any conditioning avatar state from dome captures. We show that our method is able to properly disentangle the input conditioning signal that specifies facial expression and viewpoint, from facial appearance, leading to more accurate ground truth. Furthermore, our method can generalize to unseen identities, removing the reliance on the paired captures. We demonstrate these breakthroughs by both evaluating synthetic HMC images and universal face encoders trained from these new HMC-avatar correspondences, which achieve better data efficiency and state-of-the-art accuracy. Shaojie Bai, Seunghyeon Seo, Chenghui Li, Owen Wang, Te-Li Wang, Tianyang Ma, Jason M. Saragih, Shih-En Wei, Nojun Kwak, Hyung Jun(John) Kim |
ACM Trans. Graph. | 4 |
| 2024 | POCA: Post-training Quantization with Temporal Alignment for Codec Avatars
Jian Meng, Yuecheng Li, Chenghui Li, Syed Shakib Sarwar, Dilin Wang, Jae-sun Seo |
ECCV (40) | 3 |
| 2024 | Codec Avatar Studio: Paired Human Captures for Complete, Driveable, and Generalizable AvatarsabstractTo build photorealistic avatars that users can embody, human modelling must be complete (cover the full body), driveable (able to reproduce the current motion and appearance from the user), and generalizable (i.e., easily adaptable to novel identities).Towards these goals, paired captures, that is, captures of the same subject obtained from systems of diverse quality and availability, are crucial.However, paired captures are rarely available to researchers outside of dedicated industrial labs: Codec Avatar Studio is our proposal to close this gap.Towards generalization and driveability, we introduce a dataset of 256 subjects captured in two modalities: high resolution multi-view scans of their heads, and video from the internal cameras of a headset.Towards completeness, we introduce a dataset of 4 subjects captured in eight modalities: high quality relightable multi-view captures of heads and hands, full body multi-view captures with minimal and regular clothes, and corresponding head, hands and body phone captures.Together with our data, we also provide code and pre-trained models for different state-of-the-art human generation models.Our datasets and code are available at https://github.com/facebookresearch/ava-256 and https://github.com/facebookresearch/goliath. Julieta Martinez 0001, Emily Kim, Javier Romero 0002, Timur M. Bagautdinov, Shunsuke Saito, Shoou-I Yu, Michael Zollhöfer, Te-Li Wang, Shaojie Bai, Chenghui Li, Shih-En Wei, Rohan Joshi, Wyatt Borsos, Tomas Simon, Jason M. Saragih, Paul Theodosis, Alexander Greene, Anjani Josyula, Silvio Maeta, Andrew Jewett, Simion Venshtain, Christopher Heilman, Yueh-Tung Chen, Sidi Fu, Mohamed Elshaer, Tingfang Du, Longhua Wu, Shen-Chi Chen, Youssef Emad, Steven Longay, Ashley Brewer, Hitesh Shah, Taylor Koska, Kayla Haidle, Matthew Andromalos, Joanna Hsu, Thomas Dauer, Peter Selednik, Timothy Godisart, Scott Ardisson, Matthew Cipperly, Ben Humberston, Lon Farr, Bob Hansen, Peihong Guo, Dave Braun, Steven Krenn, He Wen 0001, Lucas Evans, Natalia Fadeeva, Matthew Stewart, Gabriel Schwartz, Divam Gupta, Gyeongsik Moon, Takaaki Shiratori, Fabian Prada, Bernardo Pires, Julia Buffalini, Autumn Trimble, Kevyn McPhail, Melissa Schoeller, Yaser Sheikh |
NeurIPS | 11 |
| 2024 | Universal Facial Encoding of Codec Avatars from VR HeadsetsabstractFaithful real-time facial animation is essential for avatar-mediated telepresence in Virtual Reality (VR). To emulate authentic communication, avatar animation needs to be efficient and accurate: able to capture both extreme and subtle expressions within a few milliseconds to sustain the rhythm of natural conversations. The oblique and incomplete views of the face, variability in the donning of headsets, and illumination variation due to the environment are some of the unique challenges in generalization to unseen faces. In this paper, we present a method that can animate a photorealistic avatar in realtime from head-mounted cameras (HMCs) on a consumer VR headset. We present a self-supervised learning approach, based on a cross-view reconstruction objective, that enables generalization to unseen users. We present a lightweight expression calibration mechanism that increases accuracy with minimal additional cost to run-time efficiency. We present an improved parameterization for precise ground-truth generation that provides robustness to environmental variation. The resulting system produces accurate facial animation for unseen users wearing VR headsets in realtime. We compare our approach to prior face-encoding methods demonstrating significant improvements in both quantitative metrics and qualitative results. Shaojie Bai, Te-Li Wang, Chenghui Li, Akshay Venkatesh, Tomas Simon, Chen Cao 0001, Gabriel Schwartz, Jason M. Saragih, Yaser Sheikh, Shih-En Wei |
ACM Trans. Graph. | 3 |
| 2023 | Auto-CARD: Efficient and Robust Codec Avatar Driving for Real-time Mobile TelepresenceabstractReal-time and robust photorealistic avatars for telepresence in AR/VR have been highly desired for enabling im-mersive photorealistic telepresence. However, there still exists one key bottleneck: the considerable computational expense needed to accurately infer facial expressions captured from headset-mounted cameras with a quality level that can match the realism of the avatar's human appearance. To this end, we propose a framework called Auto-CARD, which for the first time enables realtime and robust driving of Codec Avatars when exclusively using merely on-device computing resources. This is achieved by minimizing two sources of redundancy. First, we develop a dedicated neural architecture search technique called AVE-NAS for avatar encoding in AR/VR, which explicitly boosts both the searched architectures' robustness in the presence of extreme facial ex-pressions and hardware friendliness on fast evolving AR/VR headsets. Second, we leverage the temporal redundancy in consecutively captured images during continuous rendering and develop a mechanism dubbed LATEX to skip the computation of redundant frames. Specifically, we first identify an opportunity from the linearity of the latent space derived by the avatar decoder and then propose to perform adaptive latent extrapolation for redundant frames. For evaluation, we demonstrate the efficacy of our Auto-CARD framework in realtime Codec Avatar driving settings, where we achieve a$5.05\times$speedup on Meta Quest 2 while maintaining a compa-rable or even better animation quality than state-of-the-art avatar encoder designs. Yonggan Fu, Yuecheng Li, Chenghui Li, Jason M. Saragih, Peizhao Zhang, Xiaoliang Dai, Yingyan (Celine) Lin |
CVPR | 3 |
| 2023 | Large sample spectral analysis of graph-based multi-manifold clusteringabstractIn this work we study statistical properties of graph-based algorithms for multi-manifold clustering (MMC). In MMC the goal is to retrieve the multi-manifold structure underlying a given Euclidean data set when this one is assumed to be obtained by sampling a distribution on a union of manifolds $\M = \M_1 \cup\dots \cup \M_N$ that may intersect with each other and that may have different dimensions. We investigate sufficient conditions that similarity graphs on data sets must satisfy in order for their corresponding graph Laplacians to capture the right geometric information to solve the MMC problem. Precisely, we provide high probability error bounds for the spectral approximation of a tensorized Laplacian on $\M$ with a suitable graph Laplacian built from the observations; the recovered tensorized Laplacian contains all geometric information of all the individual underlying manifolds. We provide an example of a family of similarity graphs, which we call annular proximity graphs with angle constraints, satisfying these sufficient conditions. We contrast our family of graphs with other constructions in the literature based on the alignment of tangent planes. Extensive numerical experiments expand the insights that our theory provides on the MMC problem. Nicolás García Trillos, Pengfei He 0002, Chenghui Li |
J. Mach. Learn. Res. | 3 |
| 2022 | A Fast Scale-Invariant Algorithm for Non-negative Least Squares with Non-negative DataabstractNonnegative (linear) least square problems are a fundamental class of problems that is well-studied in statistical learning and for which solvers have been implemented in many of the standard programming languages used within the machine learning community. The existing off-the-shelf solvers view the non-negativity constraint in these problems as an obstacle and, compared to unconstrained least squares, perform additional effort to address it. However, in many of the typical applications, the data itself is nonnegative as well, and we show that the nonnegativity in this case makes the problem easier. In particular, while the worst-case dimension-independent oracle complexity of unconstrained least squares problems necessarily scales with one of the data matrix constants (typically the spectral norm) and these problems are solved to additive error, we show that nonnegative least squares problems with nonnegative data are solvable to multiplicative error and with complexity that is independent of any matrix constants. The algorithm we introduce is accelerated and based on a primal-dual perspective. We further show how to provably obtain linear convergence using adaptive restart coupled with our method and demonstrate its effectiveness on large-scale data via numerical experiments. Jelena Diakonikolas, Chenghui Li, Swati Padmanabhan, Chaobing Song |
NeurIPS | 2 |
| 2020 | Online Camera-LiDAR Calibration with Sensor Semantic InformationabstractAs a crucial step of sensor data fusion, sensor calibration plays a vital role in many cutting-edge machine vision applications, such as autonomous vehicles and AR/VR. Existing techniques either require quite amount of manual work and complex settings, or are unrobust and prone to produce suboptimal results. In this paper, we investigate the extrinsic calibration of an RGB camera and a light detection and ranging (LiDAR) sensor, which are two of the most widely used sensors in autonomous vehicles for perceiving the outdoor environment. Specifically, we introduce an online calibration technique that automatically computes the optimal rigid motion transformation between the aforementioned two sensors and maximizes their mutual information of perceived data, without the need of tuning environment settings. By formulating the calibration as an optimization problem with a novel calibration quality metric based on semantic features, we successfully and robustly align pairs of temporally synchronized camera and LiDAR frames in real time. Demonstrated on several autonomous driving tasks, our method outperforms state-of-the-art edge feature based auto-calibration approaches in terms of robustness and accuracy. Yufeng Zhu, Chenghui Li |
ICRA | 2 |
| 1998 | Data Mining for Direct Marketing: Problems and Solutions
Charles Ling 0001, Chenghui Li |
KDD | 2 |