EDBT 2026 Demo / reviewers in the wild / expert
Chao Liu 0064
dblp:15/5923-64
· DBLP profile ↗
15ranked-venue papers
7as first author
10since 2021 · last 2025
0009-0007-5751-8723ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Coherent 3D Portrait Video Reconstruction via Triplane FusionabstractRecent breakthroughs in single-image 3D portrait reconstruction have enabled telepresence systems to stream 3D portrait videos from a single camera in real-time, democratizing telepresence. However, per-frame 3D reconstruction exhibits temporal inconsistency and forgets the user’s appearance. On the other hand, self-reenactment methods can render coherent 3D portraits by driving a 3D avatar built from a single reference image but fail to faithfully preserve the user’s per-frame appearance (e.g., instantaneous facial expressions and lighting). As a result, neither of these two frameworks is an ideal solution for democratized 3D telepresence. In this work, we address this dilemma and propose a novel solution that maintains both coherent identity and dynamic per-frame appearance to enable the best possible realism. To this end, we propose a new fusion-based method that takes the best of both worlds by fusing a canonical 3D prior from a reference view with dynamic appearance from per-frame input views, producing temporally stable 3D videos with faithful reconstruction of the user’s per-frame appearance. Trained only using synthetic data produced by an expression-conditioned 3D GAN, our encoder-based method achieves both state-of-the-art 3D reconstruction and temporal consistency on in-studio and in-the-wild datasets. Shengze Wang 0002, Chao Liu 0064, Matthew A. Chan 0001, Michael Stengel, Henry Fuchs, Shalini De Mello, Koki Nagano |
CVPR | 3 |
| 2025 | BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video RepresentationsabstractExisting video generation models struggle to follow complex text prompts and synthesize multiple objects, raising the need for additional grounding input for improved controllability. In this work, we propose to decompose videos into visual primitives – blob video representation, a general representation for controllable video generation. Based on blob conditions, we develop a blob-grounded video diffusion model named BlobGEN-Vid that allows users to control object motions and fine-grained object appearance. In particular, we introduce a masked 3D attention module that effectively improves regional consistency across frames. In addition, we introduce a learnable module to interpolate text embeddings so that users can control semantics in specific frames and obtain smooth object transitions. We show that our framework is model-agnostic and can build BlobGEN-Vid on both U-Net and DiT-based video diffusion models. Extensive experimental results show that BlobGEN-Vid achieves superior zero-shot video generation ability and state-of-the-art layout controllability on multiple benchmarks. When combined with an LLM for layout planning, our framework even outperforms proprietary text-to-video generators regarding compositional accuracy. Our project page: blobgen-vid.github.io Weixi Feng, Chao Liu 0064, Sifei Liu, William Yang Wang, Arash Vahdat, Weili Nie |
CVPR | 2 |
| 2024 | Learning to Jointly Understand Visual and Tactile SignalsabstractModeling and analyzing object and shape has been well studied in the past. However, manipulation of these complex tools and articulated objects remains difficult for autonomous agents. Our human hands, however, are dexterous and adaptive. We can easily adapt a manipulation skill on one object to all objects in the class and to other similar classes. Our intuition comes from that there is a close connection between manipulations and topology and articulation of objects. The possible articulation of objects indicates the types of manipulation necessary to operate the object. In this work, we aim to take a manipulation perspective to understand everyday objects and tools. We collect a multi-modal visual-tactile dataset that contains paired full-hand force pressure maps and manipulation videos. We also propose a novel method to learn a cross-modal latent manifold that allow for cross-modal prediction and discovery of latent structure in different data modalities. We conduct extensive experiments to demonstrate the effectiveness of our method. Yichen Li 0004, Yilun Du, Chao Liu 0021, Chao Liu 0064, Francis Williams, Michael Foshey, Benjamin Eckart, Jan Kautz, Josh Tenenbaum, Antonio Torralba 0001, Wojciech Matusik |
ICLR | 4 |
| 2024 | Compositional Text-to-Image Generation with Dense Blob RepresentationsabstractExisting text-to-image models struggle to follow complex text prompts, raising the need for extra grounding inputs for better controllability. In this work, we propose to decompose a scene into visual primitives - denoted as dense blob representations - that contain fine-grained details of the scene while being modular, human-interpretable, and easy-to-construct. Based on blob representations, we develop a blob-grounded text-to-image diffusion model, termed BlobGEN, for compositional generation. Particularly, we introduce a new masked cross-attention module to disentangle the fusion between blob representations and visual features. To leverage the compositionality of large language models (LLMs), we introduce a new in-context learning approach to generate blob representations from text prompts. Our extensive experiments show that BlobGEN achieves superior zero-shot generation quality and better layout-guided controllability on MS-COCO. When augmented by LLMs, our method exhibits superior numerical and spatial correctness on compositional image generation benchmarks. Weili Nie, Sifei Liu, Morteza Mardani, Chao Liu 0064, Benjamin Eckart, Arash Vahdat |
ICML | 4 |
| 2024 | BlobGEN-3D: Compositional 3D-Consistent Freeview Image Generation with 3D Blobs
Chao Liu 0064, Weili Nie, Sifei Liu, Abhishek Badki, Hang Su 0005, Morteza Mardani, Benjamin Eckart, Arash Vahdat |
SIGGRAPH Asia | 1 |
| 2023 | Online Consistent Video Depth with Gaussian Mixture RepresentationabstractWe demonstrate how off-the-shelf single-image depth estimation methods can be augmented with guidance from optical flow to achieve consistent and accurate online depth estimation using video sequences of static scenes. While previous work has successfully leveraged the complementary nature of optical flow and depth estimation, these techniques use computationally expensive test time optimization strategies that do not generalize beyond a single video sequence and also require knowledge of the future. In contrast, we present a computationally efficient feed-forward design that runs in an online fashion by utilizing learned data priors from previously seen video sequences. To accomplish this, we propose a continuous geometric scene representation that parametrically and compositionally represents the scene as a Gaussian Mixture Model (GMM). Based on this representation, our pipeline learns to estimate consistent depths and associated camera poses from video sequences of static scenes without direct supervision. Our online method achieves state-of-the-art results compared against offline methods that require all sequence frames. Chao Liu 0064, Benjamin Eckart, Jan Kautz |
ICRA | 1 |
| 2023 | SMRD: SURE-Based Robust MRI Reconstruction with Diffusion Models
Batu Ozturkler, Chao Liu 0064, Benjamin Eckart, Morteza Mardani, Jiaming Song, Jan Kautz |
MICCAI (3) | 2 |
| 2023 | Real-Time Radiance Fields for Single-Image Portrait View SynthesisabstractWe present a one-shot method to infer and render a photorealistic 3D representation from a single unposed image (e.g., face portrait) in real-time. Given a single RGB input, our image encoder directly predicts a canonical triplane representation of a neural radiance field for 3D-aware novel view synthesis via volume rendering. Our method is fast (24 fps) on consumer hardware, and produces higher quality results than strong GAN-inversion baselines that require test-time optimization. To train our triplane encoder pipeline, we use only synthetic data, showing how to distill the knowledge from a pretrained 3D GAN into a feedforward encoder. Technical contributions include a Vision Transformer-based triplane encoder, a camera data augmentation strategy, and a well-designed loss function for synthetic data training. We benchmark against the state-of-the-art methods, demonstrating significant improvements in robustness and image quality in challenging real-world settings. We showcase our results on portraits of faces (FFHQ) and cats (AFHQ), but our algorithm can also be applied in the future to other categories with a 3D-aware image generator. Alex Trevithick, Matthew A. Chan 0001, Michael Stengel, Eric R. Chan, Chao Liu 0064, Zhiding Yu, Sameh Khamis, Manmohan Krishna Chandraker, Ravi Ramamoorthi, Koki Nagano |
ACM Trans. Graph. | 5 |
| 2022 | Neural Interferometry: Image Reconstruction from Astronomical Interferometers Using Transformer-Conditioned Neural FieldsabstractAstronomical interferometry enables a collection of telescopes to achieve angular resolutions comparable to that of a single, much larger telescope. This is achieved by combining simultaneous observations from pairs of telescopes such that the signal is mathematically equivalent to sampling the Fourier domain of the object. However, reconstructing images from such sparse sampling is a challenging and ill-posed problem, with current methods requiring precise tuning of parameters and manual, iterative cleaning by experts. We present a novel deep learning approach in which the representation in the Fourier domain of an astronomical source is learned implicitly using a neural field representation. Data-driven priors can be added through a transformer encoder. Results on synthetically observed galaxies show that transformer-conditioned neural fields can successfully reconstruct astronomical observations even when the number of visibilities is very sparse. Benjamin Wu, Chao Liu 0064, Benjamin Eckart, Jan Kautz |
AAAI | 2 |
| 2021 | Self-Supervised Learning on 3D Point Clouds by Learning Discrete Generative ModelsabstractWhile recent pre-training tasks on 2D images have proven very successful for transfer learning, pre-training for 3D data remains challenging. In this work, we introduce a general method for 3D self-supervised representation learning that 1) remains agnostic to the underlying neural network architecture, and 2) specifically leverages the geometric nature of 3D point cloud data. The proposed task softly segments 3D points into a discrete number of geometric partitions. A self-supervised loss is formed under the interpretation that these soft partitions implicitly parameterize a latent Gaussian Mixture Model (GMM), and that this generative model establishes a data likelihood function. Our pretext task can therefore be viewed in terms of an encoder-decoder paradigm that squeezes learned representations through an implicitly defined parametric discrete generative model bottleneck. We show that any existing neural network architecture designed for supervised point cloud segmentation can be repurposed for the proposed unsupervised pretext task. By maximizing data likelihood with respect to the soft partitions formed by the unsupervised point-wise segmentation network, learned representations are encouraged to contain compositionally rich geometric information. In tests, we show that our method naturally induces semantic separation in feature space, resulting in state-of-the-art performance on downstream applications like model classification and semantic segmentation. Benjamin Eckart, Chao Liu 0064, Jan Kautz |
CVPR | 3 |
| 2020 | High Resolution Diffuse Optical Tomography using Short Range Indirect Subsurface ImagingabstractDiffuse optical tomography (DOT) is an approach to recover subsurface structures beneath the skin by measuring light propagation beneath the surface. The method is based on optimizing the difference between the images collected and a forward model that accurately represents diffuse photon propagation within a heterogeneous scattering medium. However, to date, most works have used a few source-detector pairs and recover the medium at only a very low resolution. And increasing the resolution requires prohibitive computations/storage. In this work, we present a fast imaging and algorithm for high resolution diffuse optical tomography with a line imaging and illumination system. Key to our approach is a convolution approximation of the forward heterogeneous scattering model that can be inverted to produce deeper than ever before structured beneath the surface. We show that our proposed method can detect reasonably accurate boundaries and relative depth of heterogeneous structures up to a depth of 8 mm below highly scattering medium such as milk. This work can extend the potential of DOT to recover more intricate structures (vessels, tissue, tumors, etc.) beneath the skin for diagnosing many dermatological and cardio-vascular conditions. Chao Liu 0064, Akash K. Maity, Artur Dubrawski, Ashutosh Sabharwal, Srinivasa G. Narasimhan |
ICCP | 1 |
| 2019 | Neural RGB(r)D Sensing: Depth and Uncertainty From a Video CameraabstractDepth sensing is crucial for 3D reconstruction and scene understanding. Active depth sensors provide dense metric measurements, but often suffer from limitations such as restricted operating ranges, low spatial resolution, sensor interference, and high power consumption. In this paper, we propose a deep learning (DL) method to estimate per-pixel depth and its uncertainty continuously from a monocular video stream, with the goal of effectively turning an RGB camera into an RGB-D camera. Unlike prior DL-based methods, we estimate a depth probability distribution for each pixel rather than a single depth value, leading to an estimate of a 3D depth probability volume for each input frame. These depth probability volumes are accumulated over time under a Bayesian filtering framework as more incoming frames are processed sequentially, which effectively reduces depth uncertainty and improves accuracy, robustness, and temporal stability. Compared to prior work, the proposed approach achieves more accurate and stable results, and generalizes better to new datasets. Experimental results also show the output of our approach can be directly fed into classical RGB-D based 3D scanning methods for 3D scene reconstruction. Chao Liu 0064, Jinwei Gu, Srinivasa G. Narasimhan, Jan Kautz |
CVPR | 1 |
| 2018 | Near-light photometric stereo using circularly placed point light sourcesabstractMost photometric stereo approaches assume distant or directional lighting and orthographic imaging. However, when the source is divergent and is near the object and the camera is projective, the image intensity of a Lambertian object is a non-linear function of both the unknown surface normals and the unknown distances of the source to the surface points. The resulting non-linear optimization is non-convex and highly sensitive to the initial guess. In this paper, we propose a two-stage near-light photometric stereo method using circularly placed point light sources (commonly seen in recent consumer imaging devices like NESTcam, Amazon Cloudcam, etc). We represent the scene using a 3D mesh and directly optimize the vertices of the mesh. This reduces the complexity of the relationship between surface normals and depths in the image formation model. In the first stage, we optimize the vertex positions using the differential images induced by small changes in light source position. This procedure yields a strong initial guess for the second stage that refines the estimations using the raw captured images. We propose an accurate calibration approach to estimate the positions of the sources. Our approach performs better on simulations and on real Lambertian scenes with complex shapes than the state-of-the-art method with near-field lighting. Chao Liu 0064, Srinivasa G. Narasimhan, Artur Dubrawski |
ICCP | 1 |
| 2017 | Matting and Depth Recovery of Thin Structures Using a Focal StackabstractThin structures such as fence, grass and vessels are common in photography and scientific imaging. They exhibit complex 3D structures with sharp depth variations/discontinuities and mutual occlusions. In this paper, we develop a method to estimate the occlusion matte and depths of thin structures from a focal image stack, which is obtained either by varying the focus/aperture of the lens or computed from a one-shot light field image. We propose an image formation model that explicitly describes the spatially varying optical blur and mutual occlusions for structures located at different depths. Based on the model, we derive an efficient MCMC inference algorithm that enables direct and analytical computations of the iterative update for the model/images without re-rendering images in the sampling process. Then, the depths of the thin structures are recovered using gradient descent with the differential terms computed using the image formation model. We apply the proposed method to scenes at both macro and micro scales. For macro-scale, we evaluate our method on scenes with complex 3D thin structures such as tree branches and grass. For micro-scale, we apply our method to in-vivo microscopic images of micro-vessels with diameters less than 50 μm. To our knowledge, the proposed method is the first approach to reconstruct the 3D structures of micro-vessels from non-invasive in-vivo image measurements. Chao Liu 0064, Srinivasa G. Narasimhan, Artur Dubrawski |
CVPR | 1 |
| 2015 | Real-time visual analysis of microvascular blood flow for critical careabstractMicrocirculatory monitoring plays an important role in diagnosis and treatment of critical care patients. Sidestream Dark Field (SDF) imaging devices have been used to visualize and support interpretation of the micro-vascular blood flow. However, due to subsurface scattering within the tissue that embeds the capillaries, transparency of plasma, imaging noise and lack of features, it is difficult to obtain reliable physiological data from SDF videos. Therefore, thus far microcirculatory videos have been analyzed manually with significant input from expert clinicians. In this paper, we present a framework that automates the analysis process. It includes stages of video stabilization, enhancement, and micro-vessel extraction, in order to automatically estimate statistics of the micro blood flows from SDF videos. Our method has been validated in critical care experiments conducted carefully to record the microcirculatory blood flow in test animal subjects before, during and after induced bleeding episodes, as well as to study the effect of fluid resuscitation. Our method is able to extract microcirculatory measurements that are consistent with clinical intuition and it has a potential to become a useful tool in critical care medicine. Chao Liu 0064, Hernando Gómez, Srinivasa G. Narasimhan, Artur Dubrawski, Michael R. Pinsky, Brian Zuckerbraun |
CVPR | 1 |