VLDB 2026 Research / reviewers in the wild / expert
Yuxuan Zhang 0001
dblp:126/5240-1
· DBLP profile ↗
16ranked-venue papers
7as first author
16since 2021 · last 2025
0009-0004-3255-0901ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 14 since 2021Artificial intelligence and machine learning · 12 · 6 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Stable-Hair: Real-World Hair Transfer via Diffusion ModelabstractCurrent hair transfer methods struggle to handle diverse and intricate hairstyles, limiting their applicability in real-world scenarios. In this paper, we propose a novel diffusion-based hair transfer framework, named Stable-Hair, which robustly transfers a wide range of real-world hairstyles to user-provided faces for virtual hair try-on. To achieve this goal, our Stable-Hair framework is designed as a two-stage pipeline. In the first stage, we train a Bald Converter alongside stable diffusion to remove hair from the user-provided face images, resulting in bald images. In the second stage, we specifically designed a Hair Extractor and a Latent IdentityNet to transfer the target hairstyle with highly detailed and high-fidelity to the bald image. The Hair Extractor is trained to encode reference images with the desired hairstyles, while the Latent IdentityNet ensures consistency in identity and background. To minimize color deviations between source images and transfer results, we introduce a novel Latent ControlNet architecture, which functions as both the Bald Converter and Latent IdentityNet. After training on our curated triplet dataset, our method accurately transfers highly detailed and high-fidelity hairstyles to the source images. Extensive experiments demonstrate that our approach achieves state-of-the-art performance compared to existing hair transfer methods. Yuxuan Zhang 0001, Yiren Song, Jichao Zhang, Hao Tang 0005 |
AAAI | 1 |
| 2025 | DIFIX3D+: Improving 3D Reconstructions with Single-Step Diffusion ModelsabstractNeural Radiance Fields and 3D Gaussian Splatting have revolutionized 3D reconstruction and novel-view synthesis task. However, achieving photorealistic rendering from extreme novel viewpoints remains challenging, as artifacts persist across representations. In this work, we introduce Difix3D+, a novel pipeline designed to enhance 3D reconstruction and novel-view synthesis through single-step diffusion models. At the core of our approach is Difix, a single-step image diffusion model trained to enhance and remove artifacts in rendered novel views caused by under-constrained regions of the 3D representation. Difix serves two critical roles in our pipeline. First, it is used during the reconstruction phase to clean up pseudo-training views that are rendered from the reconstruction and then distilled back into 3D. This greatly enhances underconstrained regions and improves the overall 3D representation quality. More importantly, Difix also acts as a neural enhancer during inference, effectively removing residual artifacts arising from imperfect 3D supervision and the limited capacity of current reconstruction models. Difix3D+ is a general solution, a single model compatible with both NeRF and 3DGS representations, and it achieves an average 2× improvement in FID score over baselines while maintaining 3D consistency. Jay Zhangjie Wu, Yuxuan Zhang 0001, Haithem Turki, Xuanchi Ren, Jun Gao 0004, Zheng Shou 0001, Sanja Fidler, Zan Gojcic, Huan Ling |
CVPR | 2 |
| 2025 | ArtEditor: Learning Customized Instructional Image Editor From Few-Shot Examples
Yiren Song, Yuxuan Zhang 0001, Hailong Guo, Xueyin Wang |
ICCV | 3 |
| 2025 | Controllable Weather Synthesis and Removal with Video Diffusion Models
Chih-Hao Lin, Ruofan Liang, Yuxuan Zhang 0001, Sanja Fidler, Shenlong Wang, Zan Gojcic |
ICCV | 4 |
| 2025 | EasyControl: Adding Efficient and Flexible Control for Diffusion Transformer
Yuxuan Zhang 0001, Yirui Yuan, Yiren Song |
ICCV | 1 |
| 2024 | SSR-Encoder: Encoding Selective Subject Representation for Subject-Driven GenerationabstractRecent advancements in subject-driven image generation have led to zero-shot generation, yet precise selection and focus on crucial subject representations remain challenging. Addressing this, we introduce the SSR-Encoder, a novel architecture designed for selectively capturing any subject from single or multiple reference images. It responds to various query modalities including text and masks, without necessitating test-time fine-tuning. The SSR-Encoder combines a Token-to-Patch Aligner that aligns query inputs with image patches and a Detail-Preserving Subject Encoder for extracting and preserving fine features of the subjects, thereby generating subject embeddings. These embeddings, used in conjunction with original text embeddings, condition the generation process. Characterized by its model generalizability and efficiency, the SSR-Encoder adapts to a range of custom models and control modules. Enhanced by the Embedding Consistency Regularization Loss for improved training, our extensive experiments demonstrate its effectiveness in versatile and high-quality image generation, indicating its broad applicability. Project page: ssr-encoder.github.io Yuxuan Zhang 0001, Yiren Song, Rui Wang 0124, Jinpeng Yu 0002, Hao Tang 0005, Huaxia Li, Xu Tang 0007, Yao Hu 0002, Han Pan, Zhongliang Jing |
CVPR | 1 |
| 2024 | Fast Personalized Text to Image Synthesis with Attention InjectionabstractCurrently, personalized image generation methods mostly require considerable time to finetune and often overfit the concept resulting in generated images that are similar to custom concepts but difficult to edit by prompts. We propose an effective and fast approach that could balance the text-image consistency and identity consistency of the generated image and reference image. Our method can generate personalized images without any fine-tuning while maintaining the inherent text-to-image generation ability of diffusion models. Given a prompt and a reference image, we merge the custom concept into generated images by manipulating cross-attention and self-attention layers of the original diffusion model to generate personalized images that match the text description. Comprehensive experiments highlight the superiority of our method. Yuxuan Zhang 0001, Yiren Song, Jinpeng Yu 0002, Han Pan, Zhongliang Jing |
ICASSP | 1 |
| 2024 | GeoFormer: Learning Point Cloud Completion with Tri-Plane Integrated TransformerabstractPoint cloud completion aims to recover accurate global geometry and preserve fine-grained local details from partial point clouds. Conventional methods typically predict unseen points directly from 3D point cloud coordinates or use self-projected multi-view depth maps to ease this task. However, these gray-scale depth maps cannot reach multi-view consistency, consequently restricting the performance. In this paper, we introduce a GeoFormer that simultaneously enhances the global geometric structure of the points and improves the local details. Specifically, we design a CCM Feature Enhanced Point Generator to integrate image features from multi-view consistent canonical coordinate maps (CCMs) and align them with pure point features, thereby enhancing the global geometry feature. Additionally, we employ the Multi-scale Geometry-aware Upsampler module to progressively enhance local details. This is achieved through cross attention between the multi-scale features extracted from the partial input and the features derived from previously estimated points. Extensive experiments on the PCN, ShapeNet-55/34, and KITTI benchmarks demonstrate that our GeoFormer outperforms recent methods, achieving the state-of-the-art performance. Our code is available at https://github.com/Jinpeng-Yu/GeoFormer. Jinpeng Yu 0002, Binbin Huang 0004, Yuxuan Zhang 0001, Huaxia Li, Xu Tang 0007, Shenghua Gao |
ACM Multimedia | 3 |
| 2024 | ProcessPainter: Learning to draw from sequence data
Yiren Song, Hai Ci, Xiaojun Ye 0002, Yuxuan Zhang 0001, Zheng Shou 0001 |
SIGGRAPH Asia | 7 |
| 2023 | Shakes on a Plane: Unsupervised Depth Estimation from Unstabilized PhotographyabstractModern mobile burst photography pipelines capture and merge a short sequence of frames to recover an enhanced image, but often disregard the 3D nature of the scene they capture, treating pixel motion between images as a 2D aggregation problem. We show that in a “long-burst”, forty-two 12-megapixel RAW frames captured in a two-second sequence, there is enough parallax information from natural hand tremor alone to recover high-quality scene depth. To this end, we devise a test-time optimization approach that fits a neural RGB-D representation to long-burst data and simultaneously estimates scene depth and camera motion. Our plane plus depth model is trained end-to-end, and performs coarse-to-fine refinement by controlling which multi-resolution volume features the network has access to at what time during training. We validate the method experimentally, and demonstrate geometrically accurate depth reconstructions with no additional hardware or separate data pre-processing and pose-estimation steps. Ilya Chugunov, Yuxuan Zhang 0001, Felix Heide |
CVPR | 2 |
| 2022 | The Implicit Values of A Good Hand Shake: Handheld Multi-Frame Neural Depth RefinementabstractModern smartphones can continuously stream multi-megapixel RGB images at 60 Hz, synchronized with high-quality 3D pose information and low-resolution LiDAR-driven depth estimates. During a snapshot photograph, the natural unsteadiness of the photographer's hands offers millimeter-scale variation in camera pose, which we can capture along with RGB and depth in a circular buffer. In this work we explore how, from a bundle of these measurements acquired during viewfinding, we can combine dense micro-baseline parallax cues with kilopixel LiDAR depth to distill a high-fidelity depth map. We take a test-time optimization approach and train a coordinate MLP to output photometrically and geometrically consistent depth estimates at the continuous coordinates along the path traced by the photographer's natural hand shake. With no additional hardware, artificial hand motion, or user interaction beyond the press of a button, our proposed method brings high-resolution depth estimates to point-and-shoot “table-top” photography – textured objects at close range. Ilya Chugunov, Yuxuan Zhang 0001, Zhihao Xia, Xuaner Cecilia Zhang, Jiawen Chen 0001, Felix Heide |
CVPR | 2 |
| 2022 | All You Need Is RAW: Defending Against Adversarial Attacks with Camera Image Pipelines
Yuxuan Zhang 0001, Felix Heide |
ECCV (19) | 1 |
| 2022 | Neural Photo-FinishingabstractImage processing pipelines are ubiquitous and we rely on them either directly, by filtering or adjusting an image post-capture, or indirectly, as image signal processing (ISP) pipelines on broadly deployed camera systems. Used by artists, photographers, system engineers, and for downstream vision tasks, traditional image processing pipelines feature complex algorithmic branches developed over decades. Recently, image-to-image networks have made great strides in image processing, style transfer, and semantic understanding. The differentiable nature of these networks allows them to fit a large corpus of data; however, they do not allow for intuitive, fine-grained controls that photographers find in modern photo-finishing tools. This work closes that gap and presents an approach to making complex photo-finishing pipelines differentiable, allowing legacy algorithms to be trained akin to neural networks using first-order optimization methods. By concatenating tailored network proxy models of individual processing steps (e.g. white-balance, tone-mapping, color tuning), we can model a non-differentiable reference image finishing pipeline more faithfully than existing proxy image-to-image network models. We validate the method for several diverse applications, including photo and video style transfer, slider regression for commercial camera ISPs, photography-driven neural demosaicking, and adversarial photo-editing. Ethan Tseng, Yuxuan Zhang 0001, Lars Jebe, Xuaner Cecilia Zhang, Zhihao Xia, Felix Heide, Jiawen Chen 0001 |
ACM Trans. Graph. | 2 |
| 2021 | DatasetGAN: Efficient Labeled Data Factory With Minimal Human EffortabstractWe introduce DatasetGAN: an automatic procedure to generate massive datasets of high-quality semantically segmented images requiring minimal human effort. Current deep networks are extremely data-hungry, benefiting from training on large-scale datasets, which are time consuming to annotate. Our method relies on the power of recent GANs to generate realistic images. We show how the GAN latent code can be decoded to produce a semantic segmentation of the image. Training the decoder only needs a few labeled examples to generalize to the rest of the latent space, resulting in an infinite annotated dataset generator! These generated datasets can then be used for training any computer vision architecture just as real datasets are. As only a few images need to be manually segmented, it becomes possible to annotate images in extreme detail and generate datasets with rich object and part segmentations. To showcase the power of our approach, we generated datasets for 7 image segmentation tasks which include pixel-level labels for 34 human face parts, and 32 car parts. Our approach outperforms all semi-supervised baselines significantly and is on par with fully supervised methods, which in some cases require as much as 100x more annotated data as our method. Yuxuan Zhang 0001, Huan Ling, Jun Gao 0004, Kangxue Yin, Jean-Francois Lafleche, Adela Barriuso, Antonio Torralba 0001, Sanja Fidler |
CVPR | 1 |
| 2021 | Deep Neural Network Fingerprinting by Conferrable Adversarial Examples
Nils Lukas, Yuxuan Zhang 0001, Florian Kerschbaum |
ICLR | 2 |
| 2021 | Image GANs meet Differentiable Rendering for Inverse Graphics and Interpretable 3D Neural Rendering
Yuxuan Zhang 0001, Wenzheng Chen, Huan Ling, Jun Gao 0004, Antonio Torralba 0001, Sanja Fidler |
ICLR | 1 |