EDBT 2026 Demo / reviewers in the wild / expert
Yuanzhen Li
dblp:97/371
· DBLP profile ↗
40ranked-venue papers
12as first author
31since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 4 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 8 first-author · 20 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CamCtrl3D: Single-Image Scene Exploration with Precise 3D Camera ControlabstractWe propose a method for generating fly-through videos of a scene, from a single image and a given camera trajectory. We build upon an image-to-video latent diffusion model [5]. We condition its UNet [25] denoiser on the camera trajectory, using four techniques. (1) We condition UNet's temporal blocks on raw camera extrinsics, similar to MotionCtrl [36]. (2) We use images containing camera ray parameters, similar to CameraCtrl [14]. (3) We re-project the initial image to subsequent frames and condition on the resulting video. (4) We introduce a global 3D representation using 2D ⇔ 3D transformers [32], which implicitly conditions on the camera poses. We combine all conditions in a ContolNet-style [42] architecture. We then propose a metric that evaluates overall video quality and the ability to preserve details with view changes, which we use to analyze the trade-offs of individual and combined conditions. Finally, we identify an optimal combination of conditions. We calibrate camera positions in our datasets for scale consistency across scenes, and we train our scene exploration model, CamCtrl3D, demonstrating state-of-the-art results. Stefan Popov, Amit Raj, Michael Krainin, Yuanzhen Li, William T. Freeman, Michael Rubinstein |
3DV | 4 |
| 2025 | Magic Insert: Style-Aware Drag-And-DropabstractWe present Magic Insert, a method for dragging-and-dropping subjects from a user-provided image into a target image of a different style in a physically plausible manner while matching the style of the target image. This work formalizes the problem of style-aware drag-and-drop and presents a method for tackling it by addressing two sub-problems: style-aware personalization and realistic object insertion in stylized images. For style-aware personalization, our method first fine-tunes a pretrained text-to-image diffusion model using LoRA and learned text tokens on the subject image, and then infuses it with a CLIP representation of the target style. For object insertion, we use Bootstrapped Domain Adaption to adapt a domain-specific photorealistic object insertion model to the domain of diverse artistic styles. Overall, the method significantly outperforms traditional approaches such as inpainting. Finally, we present a dataset, SubjectPlop, to facilitate evaluation and future progress in this area. Project page: https://magicinsert.github.io/ Nataniel Ruiz, Yuanzhen Li, Neal Wadhwa, Yael Pritch, Michael Rubinstein, David E. Jacobs, Shlomi Fruchter |
ICCV | 2 |
| 2025 | Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous TokensabstractScaling up autoregressive models in vision has not proven as beneficial as in large language models.
In this work, we investigate this scaling problem in the context of text-to-image generation, focusing on two critical factors: whether models use discrete or continuous tokens, and whether tokens are generated in a random or fixed raster order using BERT- or GPT-like transformer architectures. Our empirical results show that, while all models scale effectively in terms of validation loss, their evaluation performance -- measured by FID, GenEval score, and visual quality -- follows different trends. Models based on continuous tokens achieves significantly better visual quality than those using discrete tokens. Furthermore, the generation order and attention mechanisms significantly affect the GenEval score: random-order models achieve notably better GenEval scores compared to raster-order models.
Inspired by these findings, we train Fluid, a random-order autoregressive model on continuous tokens. Fluid 10.5B model achieves a new state-of-the-art zeor-shot FID of 6.16 on MS-COCO 30K, and 0.69 overall score on the GenEval benchmark. We hope our findings and results will encourage future efforts to further bridge the scaling gap between vision and language models. Lijie Fan, Tianhong Li, Siyang Qin, Yuanzhen Li, Chen Sun 0002, Michael Rubinstein, Deqing Sun, Kaiming He, Yonglong Tian |
ICLR | 4 |
| 2025 | Unbounded: A Generative Infinite Game of Character Life SimulationabstractWe introduce the concept of a generative infinite game, a video game that transcends the traditional boundaries of finite, hard-coded systems by using generative models. Inspired by James P. Carse's distinction between finite and infinite games, we leverage recent advances in generative AI to create Unbounded: a game of character life simulation that is fully encapsulated in generative models. Specifically, Unbounded draws inspiration from sandbox life simulations and allows you to interact with your autonomous virtual character in a virtual world by feeding, playing with and guiding it - with open-ended mechanics generated by an LLM, some of which can be emergent. In order to develop Unbounded, we propose technical innovations in both the LLM and visual generation domains. Specifically, we present: (1) a specialized, distilled large language model (LLM) that dynamically generates game mechanics, narratives, and character interactions in real-time, and (2) a new dynamic regional image prompt Adapter (IP-Adapter) for vision models that ensures consistent yet flexible visual generation of a character across multiple environments. We evaluate our system through both qualitative and quantitative analysis, showing significant improvements in character life simulation, user instruction following, narrative coherence, and visual consistency for both characters and the environments compared to traditional related approaches. Jialu Li 0001, Yuanzhen Li, Neal Wadhwa, Yael Pritch, David E. Jacobs, Michael Rubinstein, Mohit Bansal, Nataniel Ruiz |
ICLR | 2 |
| 2024 | Probing the 3D Awareness of Visual Foundation ModelsabstractRecent advances in large-scale pretraining have yielded visual foundation models with strong capabilities. Not only can recent models generalize to arbitrary images for their training task, their intermediate representations are useful for other visual tasks such as detection and segmentation. Given that such models can classify, delineate, and local- ize objects in 2D, we ask whether they also represent their 3D structure? In this work, we analyze the 3D awareness of visual foundation models. We posit that 3D awareness implies that representations (1) encode the 3D structure of the scene and (2) consistently represent the surface across views. We conduct a series of experiments using task-specific probes and zero-shot inference procedures on frozen fea- tures. Our experiments reveal several limitations of the current models. Our code and analysis can be found at https://github.com/mbanani/probe3d. Mohamed El Banani, Amit Raj, Kevis-Kokitsi Maninis, Abhishek Kar, Yuanzhen Li, Michael Rubinstein, Deqing Sun, Leonidas J. Guibas, Justin Johnson 0001, Varun Jampani |
CVPR | 5 |
| 2024 | SHINOBI: Shape and Illumination using Neural Object Decomposition via BRDF Optimization In-the-wildabstractWe present SHINOBI, an end-to-end frameworkfor the re-construction of shape, material, and illumination from object images captured with varying lighting, pose, and background. Inverse rendering of an object based on unconstrained image collections is a long-standing challenge in computer vision and graphics and requires a joint optimization over shape, radiance, and pose. We show that an implicit shape repre-sentation based on a multi-resolution hash encoding enables faster and robust shape reconstruction with joint camera alignment optimization that outperforms prior work. Further, to enable the editing of illumination and object reflectance (i.e. material) we jointly optimize BRDF and illumination to-gether with the object's shape. Our method is class-agnostic and works on in-the-wild image collections of objects to produce relightable 3D assets for several use cases such as AR/VR, movies, games, etc. Andreas Engelhardt, Amit Raj, Mark Boss, Abhishek Kar, Yuanzhen Li, Deqing Sun, Ricardo Martin-Brualla, Jonathan T. Barron, Hendrik P. A. Lensch, Varun Jampani |
CVPR | 6 |
| 2024 | Diffusion-FOF: Single-View Clothed Human Reconstruction via Diffusion-Based Fourier Occupancy FieldabstractReconstructing a clothed human from a single-view image has several challenging issues, including flexibly representing various body shapes and poses, estimating complete 3D geometry and consistent texture, and achieving more fine-grained details. To address them, we propose a new diffusion-based Fourier occupancy field method to improve the human representing ability and the geometry generating ability. First, we estimate the back-view image from the given reference image by incorporating a style consistency constraint. Then, we extract multi-scale features of the two images as conditional and design a diffusion model to generate the Fourier occupancy field in the wavelet domain. We refine the initial estimated Fourier occupancy field with image features as conditions to improve the geometric accuracy. Finally, the reference and estimated back-view images are mapped onto the human model, creating a textured clothed human model. Substantial experiments are conducted, and the experimental results show that our method outperforms the state-of-the-art methods in geometry and texture reconstruction performance. Yuanzhen Li, Fei Luo 0004, Chunxia Xiao |
CVPR | 1 |
| 2024 | HyperDreamBooth: HyperNetworks for Fast Personalization of Text-to-Image ModelsabstractPersonalization has emerged as a prominent aspect within the field of generative AI, enabling the synthesis of individuals in diverse contexts and styles, while retaining high-fidelity to their identities. However, the process of personalization presents inherent challenges in terms of time and memory requirements. Fine-tuning each personalized model needs considerable GPU time investment, and storing a personalized model per subject can be demanding in terms of storage capacity. To overcome these challenges, we propose HyperDreamBooth—a hypernetwork capable of efficiently generating a small set of personalized weights from a single image of a person. By composing these weights into the diffusion model, coupled with fast finetuning, HyperDreamBooth can generate a person's face in various contexts and styles, with high subject details while also preserving the model's crucial knowledge of diverse styles and semantic modifications. Our method achieves personalization on faces in roughly 20 seconds, 25x faster than DreamBooth and 125x faster than Textual Inversion, using as few as one reference image, with the same quality and style diversity as DreamBooth. Also our method yields a model that is 10,000x smaller than a normal DreamBooth model. Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Tingbo Hou, Yael Pritch, Neal Wadhwa, Michael Rubinstein, Kfir Aberman |
CVPR | 2 |
| 2024 | Alchemist: Parametric Control of Material Properties with Diffusion ModelsabstractWe propose a method to control material attributes of objects like roughness, metallic, albedo, and transparency in real images. Our method capitalizes on the generative prior of text-to-image models known for photorealism, employing a scalar value and instructions to alter low-level material properties. Addressing the lack of datasets with controlled material attributes, we generated an object-centric synthetic dataset with physically-based materials. Finetuning a modified pretrained text-to-image model on this synthetic dataset enables us to edit material properties in real-world images while preserving all other attributes. We show the potential application of our model to material edited NeRFs. Prafull Sharma, Varun Jampani, Yuanzhen Li, Xuhui Jia, Dmitry Lagun, Frédo Durand, William T. Freeman, Mark J. Matthews |
CVPR | 3 |
| 2024 | ZipLoRA: Any Subject in Any Style by Effectively Merging LoRAs
Viraj Shah, Nataniel Ruiz, Forrester Cole, Erika Lu, Svetlana Lazebnik, Yuanzhen Li, Varun Jampani |
ECCV (1) | 6 |
| 2024 | 3D Congealing: 3D-Aware Image Alignment in the Wild
Zizhang Li, Amit Raj, Andreas Engelhardt, Yuanzhen Li, Tingbo Hou, Jiajun Wu 0001, Varun Jampani |
ECCV (1) | 5 |
| 2024 | Lumiere: A Space-Time Diffusion Model for Video Generation
Omer Bar-Tal, Hila Chefer, Omer Tov, Charles Herrmann, Roni Paiss, Shiran Zada, Ariel Ephrat, Junhwa Hur, Amit Raj, Yuanzhen Li, Michael Rubinstein, Tomer Michaeli, Oliver Wang, Deqing Sun, Tali Dekel, Inbar Mosseri |
SIGGRAPH Asia | 11 |
| 2024 | RealFill: Reference-Driven Generation for Authentic Image CompletionabstractRecent advances in generative imagery have brought forth outpainting and inpainting models that can produce high-quality, plausible image content in unknown regions. However, the content these models hallucinate is necessarily inauthentic, since they are unaware of the true scene. In this work, we propose RealFill, a novel generative approach for image completion that fills in missing regions of an image with the content that should have been there. RealFill is a generative inpainting model that is personalized using only a few reference images of a scene. These reference images do not have to be aligned with the target image, and can be taken with drastically varying viewpoints, lighting conditions, camera apertures, or image styles. Once personalized, RealFill is able to complete a target image with visually compelling contents that are faithful to the original scene. We evaluate RealFill on a new image completion benchmark that covers a set of diverse and challenging scenarios, and find that it outperforms existing approaches by a large margin. Project page: https://realfill.github.io. Luming Tang, Nataniel Ruiz, Qinghao Chu, Yuanzhen Li, Aleksander Holynski, David E. Jacobs, Bharath Hariharan, Yael Pritch, Neal Wadhwa, Kfir Aberman, Michael Rubinstein |
ACM Trans. Graph. | 4 |
| 2023 | MetaCLUE: Towards Comprehensive Visual Metaphors ResearchabstractCreativity is an indispensable part of human cognition and also an inherent part of how we make sense of the world. Metaphorical abstraction is fundamental in communicating creative ideas through nuanced relationships between abstract concepts such as feelings. While computer vision benchmarks and approaches predominantly focus on understanding and generating literal interpretations of images, metaphorical comprehension of images remains relatively unexplored. Towards this goal, we introduce Meta-CLUE, a set of vision tasks on visual metaphor. We also collect high-quality and rich metaphor annotations (abstract objects, concepts, relationships along with their corresponding object boxes) as there do not exist any datasets that facilitate the evaluation of these tasks. We perform a comprehensive analysis of state-of-the-art models in vision and language based on our annotations, highlighting strengths and weaknesses of current approaches in visual metaphor classification, localization, understanding (retrieval, question answering, captioning) and generation (text-to-image synthesis) tasks. We hope this work provides a concrete step towards developing AI systems with human-like creative capabilities. Project page: https://metaclue.github.io Arjun R. Akula, Brendan Driscoll, Pradyumna Narayana, Soravit Changpinyo, Zhiwei Jia, Suyash Damle, Garima Pruthi, Sugato Basu, Leonidas J. Guibas, William T. Freeman, Yuanzhen Li, Varun Jampani |
CVPR | 11 |
| 2023 | ShapeClipper: Scalable 3D Shape Learning from Single-View Images via Geometric and CLIP-Based ConsistencyabstractWe present ShapeClipper, a novel method that reconstructs 3D object shapes from real-world single-view RGB images. Instead of relying on laborious 3D, multi-view or camera pose annotation, ShapeClipper learns shape reconstruction from a set of single-view segmented images. The key idea is to facilitate shape learning via CLIP-based shape consistency, where we encourage objects with similar CLIP encodings to share similar shapes. We also leverage off-the-shelf normals as an additional geometric constraint so the model can learn better bottom-up reasoning of detailed surface geometry. These two novel consistency constraints, when used to regularize our model, improve its ability to learn both global shape structure and local geometric details. We evaluate our method over three challenging real-world datasets, Pix3D, Pascal3D+, and Open-Images, where we achieve superior performance over state-of-the-art methods.11project website at: https://zixuanh.com/projects/shapeclipper.html Zixuan Huang 0001, Varun Jampani, Ngoc Anh Thai, Yuanzhen Li, Stefan Stojanov, James M. Rehg |
CVPR | 4 |
| 2023 | DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven GenerationabstractLarge text-to-image models achieved a remarkable leap in the evolution of AI, enabling high-quality and diverse synthesis of images from a given text prompt. However, these models lack the ability to mimic the appearance of subjects in a given reference set and synthesize novel renditions of them in different contexts. In this work, we present a new approach for “personalization” of text-to-image diffusion models. Given as input just a few images of a subject, we fine-tune a pretrained text-to-image model such that it learns to bind a unique identifier with that specific subject. Once the subject is embedded in the output domain of the model, the unique identifier can be used to synthesize novel photorealistic images of the subject contextualized in different scenes. By leveraging the semantic prior embedded in the model with a new autogenous class-specific prior preservation loss, our technique enables synthesizing the subject in diverse scenes, poses, views and lighting conditions that do not appear in the reference images. We apply our technique to several previously-unassailable tasks, including subject recontextualization, text-guided view synthesis, and artistic rendering, all while preserving the subject's key features. We also provide a new dataset and evaluation protocol for this new task of subject-driven generation. Project page: https://dreambooth.github.io/ Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, Kfir Aberman |
CVPR | 2 |
| 2023 | Hi-LASSIE: High-Fidelity Articulated Shape and Skeleton Discovery from Sparse Image EnsembleabstractAutomatically estimating 3D skeleton, shape, camera viewpoints, and part articulation from sparse in-the-wild image ensembles is a severely under-constrained and challenging problem. Most prior methods rely on large-scale image datasets, dense temporal correspondence, or human annotations like camera pose, 2D keypoints, and shape templates. We propose Hi-LASSIE, which performs 3D articulated reconstruction from only 20–30 online images in the wild without any user-defined shape or skeleton templates. We follow the recent work of LASSIE that tackles a similar problem setting and make two significant advances. First, instead of relying on a manually annotated 3D skeleton, we automatically estimate a class-specific skeleton from the selected reference image. Second, we improve the shape reconstructions with novel instance-specific optimization strategies that allow reconstructions to faithful fit on each instance while preserving the class-specific priors learned across all images. Experiments on in-the-wild image ensembles show that Hi-LASSIE obtains higher fidelity state-of-the-art 3D reconstructions despite requiring minimum user input. Project page: chhankyao.github.io/hi-lassie/ Chun-Han Yao, Wei-Chih Hung, Yuanzhen Li, Michael Rubinstein, Ming-Hsuan Yang 0001, Varun Jampani |
CVPR | 3 |
| 2023 | DreamBooth3D: Subject-Driven Text-to-3D GenerationabstractWe present DreamBooth3D, an approach to personalize text-to-3D generative models from as few as 3-6 casually captured images of a subject. Our approach combines recent advances in personalizing text-to-image models (DreamBooth) with text-to-3D generation (DreamFusion). We find that naïvely combining these methods fails to yield satisfactory subject-specific 3D assets due to personalized text-to-image models overfitting to the input viewpoints of the subject. We overcome this through a 3-stage optimization strategy where we jointly leverage the 3D consistency of neural radiance fields together with the personalization capability of text-to-image models. Our method can produce high-quality, subject-specific 3D assets with text-driven modifications such as novel poses, colors and attributes that are not seen in any of the input images of the subject. More results are available at our project page: https://dreambooth3d.github.io Amit Raj, Srinivas Kaza, Ben Poole, Michael Niemeyer, Nataniel Ruiz, Ben Mildenhall, Shiran Zada, Kfir Aberman, Michael Rubinstein, Jonathan T. Barron, Yuanzhen Li, Varun Jampani |
ICCV | 11 |
| 2023 | Muse: Text-To-Image Generation via Masked Generative TransformersabstractWe present Muse, a text-to-image Transformermodel that achieves state-of-the-art image genera-tion performance while being significantly moreefficient than diffusion or autoregressive models.Muse is trained on a masked modeling task indiscrete token space: given the text embeddingextracted from a pre-trained large language model(LLM), Muse learns to predict randomly maskedimage tokens. Compared to pixel-space diffusionmodels, such as Imagen and DALL-E 2, Muse issignificantly more efficient due to the use of dis-crete tokens and requires fewer sampling itera-tions; compared to autoregressive models such asParti, Muse is more efficient due to the use of par-allel decoding. The use of a pre-trained LLM en-ables fine-grained language understanding, whichtranslates to high-fidelity image generation andthe understanding of visual concepts such as ob-jects, their spatial relationships, pose, cardinalityetc. Our 900M parameter model achieves a newSOTA on CC3M, with an FID score of 6.06. TheMuse 3B parameter model achieves an FID of7.88 on zero-shot COCO evaluation, along with aCLIP score of 0.32. Muse also directly enables anumber of image editing applications without theneed to fine-tune or invert the model: inpainting,outpainting, and mask-free editing. More resultsand videos demonstrating editing are available at https://muse-icml.github.io/ Huiwen Chang, Han Zhang 0010, Jarred Barber, Aaron Maschinot, José Lezama, Lu Jiang 0004, Ming-Hsuan Yang 0001, Kevin Murphy 0002, William T. Freeman, Michael Rubinstein, Yuanzhen Li, Dilip Krishnan |
ICML | 11 |
| 2023 | NAVI: Category-Agnostic Image Collections with High-Quality 3D Shape and Pose AnnotationsabstractRecent advances in neural reconstruction enable high-quality 3D object reconstruction from casually captured image collections. Current techniques mostly analyze their progress on relatively simple image collections where SfM techniques can provide ground-truth (GT) camera poses. We note that SfM techniques tend to fail on in-the-wild image collections such as image search results with varying backgrounds and illuminations. To enable systematic research progress on 3D reconstruction from casual image captures, we propose `NAVI': a new dataset of category-agnostic image collections of objects with high-quality 3D scans along with per-image 2D-3D alignments providing near-perfect GT camera parameters. These 2D-3D alignments allow us to extract accurate derivative annotations such as dense pixel correspondences, depth and segmentation maps. We demonstrate the use of NAVI image collections on different problem settings and show that NAVI enables more thorough evaluations that were not possible with existing datasets. We believe NAVI is beneficial for systematic research progress on 3D reconstruction and correspondence estimation. Varun Jampani, Kevis-Kokitsi Maninis, Andreas Engelhardt, Arjun Karpur, Karen Truong, Kyle Sargent, Stefan Popov, André Araújo 0001, Ricardo Martin-Brualla, Kaushal Patel, Daniel Vlasic, Vittorio Ferrari, Ameesh Makadia, Ce Liu 0001, Yuanzhen Li, Howard Zhou |
NeurIPS | 15 |
| 2023 | StyleDrop: Text-to-Image Synthesis of Any StyleabstractPre-trained large text-to-image models synthesize impressive images with an appropriate use of text prompts. However, ambiguities inherent in natural language, and out-of-distribution effects make it hard to synthesize arbitrary image styles, leveraging a specific design pattern, texture or material. In this paper, we introduce *StyleDrop*, a method that enables the synthesis of images that faithfully follow a specific style using a text-to-image model. StyleDrop is extremely versatile and captures nuances and details of a user-provided style, such as color schemes, shading, design patterns, and local and global effects. StyleDrop works by efficiently learning a new style by fine-tuning very few trainable parameters (less than 1\% of total model parameters), and improving the quality via iterative training with either human or automated feedback. Better yet, StyleDrop is able to deliver impressive results even when the user supplies only a *single* image specifying the desired style. An extensive study shows that, for the task of style tuning text-to-image models, StyleDrop on Muse convincingly outperforms other methods, including DreamBooth and textual inversion on Imagen or Stable Diffusion. More results are available at our project website: [https://styledrop.github.io](https://styledrop.github.io). Kihyuk Sohn, Lu Jiang 0004, Jarred Barber, Kimin Lee, Nataniel Ruiz, Dilip Krishnan, Huiwen Chang, Yuanzhen Li, Irfan A. Essa, Michael Rubinstein, Yuan Hao, Glenn Entis, Irina Blok, Daniel Castro Chin |
NeurIPS | 8 |
| 2023 | ARTIC3D: Learning Robust Articulated 3D Shapes from Noisy Web Image CollectionsabstractEstimating 3D articulated shapes like animal bodies from monocular images is inherently challenging due to the ambiguities of camera viewpoint, pose, texture, lighting, etc. We propose ARTIC3D, a self-supervised framework to reconstruct per-instance 3D shapes from a sparse image collection in-the-wild. Specifically, ARTIC3D is built upon a skeleton-based surface representation and is further guided by 2D diffusion priors from Stable Diffusion. First, we enhance the input images with occlusions/truncation via 2D diffusion to obtain cleaner mask estimates and semantic features. Second, we perform diffusion-guided 3D optimization to estimate shape and texture that are of high-fidelity and faithful to input images. We also propose a novel technique to calculate more stable image-level gradients via diffusion models compared to existing alternatives. Finally, we produce realistic animations by fine-tuning the rendered shape and texture under rigid part transformations. Extensive evaluations on multiple existing datasets as well as newly introduced noisy web image collections with occlusions and truncation demonstrate that ARTIC3D outputs are more robust to noisy images, higher quality in terms of shape and texture details, and more realistic when animated. Chun-Han Yao, Amit Raj, Wei-Chih Hung, Michael Rubinstein, Yuanzhen Li, Ming-Hsuan Yang 0001, Varun Jampani |
NeurIPS | 5 |
| 2023 | A problem-specific knowledge based artificial bee colony algorithm for scheduling distributed permutation flowshop problems with peak power consumption
Yuanzhen Li, Kai-Zhou Gao, Leilei Meng, Ponnuthurai N. Suganthan |
Eng. Appl. Artif. Intell. | 1 |
| 2023 | Self-Supervised Monocular Depth Estimation by Digging into Uncertainty Quantification
Yuanzhen Li, Shengjie Zheng, Zi-Xin Tan, Tuo Cao, Fei Luo 0004, Chunxia Xiao |
J. Comput. Sci. Technol. | 1 |
| 2023 | Monocular human depth estimation with 3D motion flow and surface normals
Yuanzhen Li, Fei Luo 0004, Chunxia Xiao |
Vis. Comput. | 1 |
| 2022 | Deep Image-based Illumination HarmonizationabstractIntegrating a foreground object into a background scene with illumination harmonization is an important but challenging task in computer vision and augmented reality community. Existing methods mainly focus on foreground and background appearance consistency or the foreground object shadow generation, which rarely consider global appearance and illumination harmonization. In this paper, we formulate seamless illumination harmonization as an illumination exchange and aggregation problem. Specifically, we firstly apply a physically-based rendering method to construct a large-scale, high-quality dataset (named IH) for our task, which contains various types of foreground objects and background scenes with different lighting conditions. Then, we propose a deep image-based illumination harmonization GAN framework named DIH-GAN, which makes full use of a multi-scale attention mechanism and illumination exchange strategy to directly infer mapping relationship between the inserted foreground object and the corresponding background scene. Meanwhile, we also use adversarial learning strategy to further refine the illumination harmonization result. Our method can not only achieve harmonious appearance and illumination for the foreground object but also can generate compelling shadow cast by the foreground object. Comprehensive experiments on both our IH dataset and real-world images show that our proposed DIH-GAN provides a practical and effective solution for image-based object illumination harmonization editing, and validate the superiority of our method against state-of-the-art methods. Our IH dataset is available at https://github.com/zhongyunbao/Dataset. Zhongyun Bao, Chengjiang Long, Gang Fu 0003, Daquan Liu, Yuanzhen Li, Chunxia Xiao |
CVPR | 5 |
| 2022 | SAMURAI: Shape And Material from Unconstrained Real-world Arbitrary Image collectionsabstractInverse rendering of an object under entirely unknown capture conditions is a fundamental challenge in computer vision and graphics. Neural approaches such as NeRF have achieved photorealistic results on novel view synthesis, but they require known camera poses. Solving this problem with unknown camera poses is highly challenging as it requires joint optimization over shape, radiance, and pose. This problem is exacerbated when the input images are captured in the wild with varying backgrounds and illuminations. Standard pose estimation techniques fail in such image collections in the wild due to very few estimated correspondences across images. Furthermore, NeRF cannot relight a scene under any illumination, as it operates on radiance (the product of reflectance and illumination). We propose a joint optimization framework to estimate the shape, BRDF, and per-image camera pose and illumination. Our method works on in-the-wild online image collections of an object and produces relightable 3D assets for several use-cases such as AR/VR. To our knowledge, our method is the first to tackle this severely unconstrained task with minimal user interaction. Mark Boss, Andreas Engelhardt, Abhishek Kar, Yuanzhen Li, Deqing Sun, Jonathan T. Barron, Hendrik P. A. Lensch, Varun Jampani |
NeurIPS | 4 |
| 2022 | LASSIE: Learning Articulated Shapes from Sparse Image Ensemble via 3D Part DiscoveryabstractCreating high-quality articulated 3D models of animals is challenging either via manual creation or using 3D scanning tools. Therefore, techniques to reconstruct articulated 3D objects from 2D images are crucial and highly useful. In this work, we propose a practical problem setting to estimate 3D pose and shape of animals given only a few (10-30) in-the-wild images of a particular animal species (say, horse). Contrary to existing works that rely on pre-defined template shapes, we do not assume any form of 2D or 3D ground-truth annotations, nor do we leverage any multi-view or temporal information. Moreover, each input image ensemble can contain animal instances with varying poses, backgrounds, illuminations, and textures. Our key insight is that 3D parts have much simpler shape compared to the overall animal and that they are robust w.r.t. animal pose articulations. Following these insights, we propose LASSIE, a novel optimization framework which discovers 3D parts in a self-supervised manner with minimal user intervention. A key driving force behind LASSIE is the enforcing of 2D-3D part consistency using self-supervisory deep features. Experiments on Pascal-Part and self-collected in-the-wild animal datasets demonstrate considerably better 3D reconstructions as well as both 2D and 3D part discovery compared to prior arts. Project page: https://chhankyao.github.io/lassie/ Chun-Han Yao, Wei-Chih Hung, Yuanzhen Li, Michael Rubinstein, Ming-Hsuan Yang 0001, Varun Jampani |
NeurIPS | 3 |
| 2022 | Self-supervised coarse-to-fine monocular depth estimation using a lightweight attention moduleabstractSelf-supervised monocular depth estimation has been widely investigated and applied in previous works. However, existing methods suffer from texture-copy, depth drift, and incomplete structure. It is difficult for normal CNN networks to completely understand the relationship between the object and its surrounding environment. Moreover, it is hard to design the depth smoothness loss to balance depth smoothness and sharpness. To address these issues, we propose a coarse-to-fine method with a normalized convolutional block attention module (NCBAM). In the coarse estimation stage, we incorporate the NCBAM into depth and pose networks to overcome the texture-copy and depth drift problems. Then, we use a new network to refine the coarse depth guided by the color image and produce a structure-preserving depth result in the refinement stage. Our method can produce results competitive with state-of-the-art methods. Comprehensive experiments prove the effectiveness of our two-stage method using the NCBAM. Yuanzhen Li, Fei Luo 0004, Chunxia Xiao |
Comput. Vis. Media | 1 |
| 2022 | A referenced iterated greedy algorithm for the distributed assembly mixed no-idle permutation flowshop scheduling problem with the total tardiness criterion
Yuanzhen Li, Quan-Ke Pan, Rubén Ruiz, Hongyan Sang |
Knowl. Based Syst. | 1 |
| 2021 | Self-supervised monocular depth estimation based on image texture detail enhancement
Yuanzhen Li, Fei Luo 0004, Shenjie Zheng, Huanhuan Wu, Chunxia Xiao |
Vis. Comput. | 1 |
| 2019 | EPVNE: An Efficient Parallelizable Virtual Network Embedding AlgorithmabstractVirtual network embedding (VNE) problem is a key issue in network virtualization technology, and much attention has been paid to the virtual network embedding. However, very little research work focuses on parallelized virtual network embedding problems which assumes that the substrate infrastructure supports parallel computing and allows one virtual node to be mapped to multiple substrate nodes. Based on the work of Liang and Zhang, we extend the well-known VNE to parallelizable virtual network embedding (PVNE) in this paper. Furthermore, to the best of our knowledge, we give the first formulation of the PVNE problem. A new heuristic algorithm named efficient parallelizable virtual network embedding (EPVNE) is proposed to reduce the cost of embedding the VN request and increase the VN request acceptance ratio. EPVNE is a two-stage mapping algorithm, which first performs node mapping and then performs link mapping. In the node mapping phase, we present a simple and efficient virtual node and physical node sorting formula and perform the virtual node mapping in order. When mapping virtual nodes, we map virtual nodes to physical nodes that just meet the CPU requirements. Substrate nodes with more CPU resources will be retained for subsequent virtual network mapping requests. In the link mapping phase, Dijkstra’s algorithm is used to find a substrate path for each virtual link. Finally, simulations are carried out and simulation results show that our algorithm performs better than the existing heuristic algorithms. Yuanzhen Li, Yingyu Zhang |
Wirel. Commun. Mob. Comput. | 1 |
| 2011 | Flexible Job Shop Scheduling Problem by Chemical-Reaction Optimization Algorithm
Junqing Li 0001, Yuanzhen Li, Huaqing Yang, Kai-Zhou Gao, Yuting Wang 0003 |
ICIC (1) | 2 |
| 2008 | ScribbleBoost: Adding Classification to Edge-Aware Interpolation of Local Image and Video AdjustmentsabstractAbstract One of the most common tasks in image and video editing is the local adjustment of various properties (e.g., saturation or brightness) of regions within an image or video. Edge‐aware interpolation of user‐drawn scribbles offers a less effort‐intensive approach to this problem than traditional region selection and matting. However, the technique suffers a number of limitations, such as reduced performance in the presence of texture contrast, and the inability to handle fragmented appearances. We significantly improve the performance of edge‐aware interpolation for this problem by adding a boosting‐based classification step that learns to discriminate between the appearance of scribbled pixels. We show that this novel data term in combination with an existing edge‐aware optimization technique achieves substantially better results for the local image and video adjustment problem than edge‐aware interpolation techniques without classification, or related methods such as matting techniques or graph cut segmentation. Yuanzhen Li, Edward H. Adelson, Aseem Agarwala |
Comput. Graph. Forum | 1 |
| 2005 | Feature congestion: a measure of display clutterabstractManagement of clutter is an important factor in the design of user interfaces and information visualizations, allowing improved usability and aesthetics. However, clutter is not a well defined concept. In this paper, we present the Feature Congestion measure of display clutter. This measure is based upon extensive modeling of the saliency of elements of a display, and upon a new operational definition of clutter. The current implementation is based upon two features: color and luminance contrast. We have tested this measure on maps that observers ranked by perceived clutter. Results show good agreement between the observers' rankings and our measure of clutter. Furthermore, our measure can be used to make design suggestions in an automated UI critiquing tool. Ruth Rosenholtz, Yuanzhen Li, Jonathan Mansfield, Zhenlan Jin |
CHI | 2 |
| 2005 | Removing photography artifacts using gradient projection and flash-exposure samplingabstractFlash images are known to suffer from several problems: saturation of nearby objects, poor illumination of distant objects, reflections of objects strongly lit by the flash and strong highlights due to the reflection of flash itself by glossy surfaces. We propose to use a flash and no-flash (ambient) image pair to produce better flash images. We present a novel gradient projection scheme based on a gradient coherence model that allows removal of reflections and highlights from flash images. We also present a brightness-ratio based algorithm that allows us to compensate for the falloff in the flash image brightness due to depth. In several practical scenarios, the quality of flash/no-flash images may be limited in terms of dynamic range. In such cases, we advocate using several images taken under different flash intensities and exposures. We analyze the flash intensity-exposure space and propose a method for adaptively sampling this space so as to minimize the number of captured images for any given scene. We present several experimental results that demonstrate the ability of our algorithms to produce improved flash images. Amit K. Agrawal, Ramesh Raskar, Shree K. Nayar, Yuanzhen Li |
ACM Trans. Graph. | 4 |
| 2005 | Compressing and companding high dynamic range images with subband architecturesabstractHigh dynamic range (HDR) imaging is an area of increasing importance, but most display devices still have limited dynamic range (LDR). Various techniques have been proposed for compressing the dynamic range while retaining important visual information. Multi-scale image processing techniques, which are widely used for many image processing tasks, have a reputation of causing halo artifacts when used for range compression. However, we demonstrate that they can work when properly implemented. We use a symmetrical analysis-synthesis filter bank, and apply local gain control to the subbands. We also show that the technique can be adapted for the related problem of "companding", in which an HDR image is converted to an LDR image, and later expanded back to high dynamic range. Yuanzhen Li, Lavanya Sharan, Edward H. Adelson |
ACM Trans. Graph. | 1 |
| 2003 | Multiple-cue Illumination Estimation in Textured ScenesabstractIn this paper, we present a method that integrates cues from shading, shadow and specular reflections for estimating directional illumination in a textured scene. Texture poses a problem for lighting estimation, since texture edges can be mistaken for changes in illumination condition, and unknown variations in albedo make reflectance model fitting impractical. Unlike previous works which all assume known or uniform reflectance, our method can deal with the effects of textures by capitalizing on physical consistencies that exist among the lighting cues. Since scene textures do not exhibit such coherence, we use this property to minimize the influence of texture on illumination direction estimation. For the recovered light source directions, a technique for estimating their intensities in the presence of texture is also proposed. Yuanzhen Li, Stephen Lin 0001, Hanqing Lu, Harry Shum |
ICCV | 1 |
| 2002 | Diffuse-Specular Separation and Depth Recovery from Image Sequences
Stephen Lin 0001, Yuanzhen Li, Sing Bing Kang, Xin Tong 0001, Harry Shum |
ECCV (3) | 2 |
| 2002 | Single-Image Reflectance Estimation for Relighting by Iterative Soft GroupingabstractReflectance values for image-based relighting are often estimated from grouped pixels with similar reflectance, but such groupings are difficult to compute with certainty for sparse image data. To address this problem, we propose an iterative method that aggregates BRDF data in a single image with known geometry and lighting by soft grouping, where pixels contribute to one another's estimate according to their degree of reflectance similarity. Estimation of specular reflectance is further improved by albedo-independent soft grouping of pixels based on shape continuity. With recovered reflectances, we demonstrate realistic relighting for synthetic and real scenes, including surfaces with spatially-varying reflectance. Yuanzhen Li, Stephen Lin 0001, Sing Bing Kang, Hanqing Lu, Harry Shum |
PG | 1 |