VLDB 2026 Research / reviewers in the wild / expert
He Zhang 0004
dblp:24/2058-4
· DBLP profile ↗
44ranked-venue papers
11as first author
31since 2021 · last 2025
0000-0002-7036-6820ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 6 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 9 first-author · 23 since 2021Security and privacy · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Text2Relight: Creative Portrait Relighting with Text GuidanceabstractWe present a lighting-aware image editing pipeline that, given a portrait image and a text prompt, performs single image relighting. Our model modifies the lighting and color of both the foreground and background to align with the provided text description. The unbounded nature in creativeness of a text allows us to describe the lighting of a scene with any sensory features including temperature, emotion, smell, time, and so on. However, the modeling of such mapping between the unbounded text and lighting is extremely challenging due to the lack of dataset where there exists no scalable data that provides large pairs of text and relighting, and therefore, current text-driven image editing models does not generalize to lighting-specific use cases. We overcome this problem by introducing a novel data synthesis pipeline: First, diverse and creative text prompts that describe the scenes with various lighting are automatically generated under a crafted hierarchy using a large language model (e.g., ChatGPT). A text-guided image generation model creates a lighting image that best matches the text. As a condition of the lighting images, we perform image-based relighting for both foreground and background using a single portrait image or a set of OLAT (One-Light-at-A-Time) images captured from lightstage system. Particularly for the background relighting, we represent the lighting image as a set of point lights and transfer them to other background images. A generative diffusion model learns the synthesized large-scale data with auxiliary task augmentation (e.g., portrait delighting and light positioning) to correlate the latent text and lighting distribution for text-guided portrait relighting. In our experiment, we demonstrate that our model outperforms existing text-guided image generation models, showing high-quality portrait relighting results with a strong generalization to unconstrained scenes. Junuk Cha, Mengwei Ren, Krishna Kumar Singh, He Zhang 0004, Yannick Hold-Geoffroy, Hyunjoon Jung, Jae Shin Yoon, Seungryul Baek |
AAAI | 4 |
| 2025 | UniReal: Universal Image Generation and Editing via Learning Real-world DynamicsabstractWe introduce UniReal, a unified framework designed to address various image generation and editing tasks. Existing solutions often vary by tasks, yet share fundamental principles: preserving consistency between inputs and outputs while capturing visual variations. Inspired by recent video generation models that effectively balance consistency and variation across frames, we propose a unifying approach that treats image-level tasks as discontinuous video generation. Specifically, we treat varying numbers of input and output images as frames, enabling seamless support for tasks such as image generation, editing, customization, composition, etc. Although designed for image-level tasks, we leverage videos as a scalable source for universal supervision. UniReal learns world dynamics from large-scale videos, demonstrating advanced capability in handling shadows, reflections, pose variation, and object interaction, while also exhibiting emergent capability for novel applications. Xi Chen 0119, He Zhang 0004, Yuqian Zhou, Soo Ye Kim, Qing Liu 0017, Yijun Li 0001, Jianming Zhang 0001, Nanxuan Zhao, Yilin Wang 0002, Zhe Lin 0001, Hengshuang Zhao |
CVPR | 3 |
| 2025 | Multitwine: Multi-Object Compositing with Text and Layout ControlabstractWe introduce the first generative model capable of simultaneous multi-object compositing, guided by both text and layout. Our model allows for the addition of multiple objects within a scene, capturing a range of interactions from simple positional relations (e.g., next to, in front of) to complex actions requiring reposing (e.g., hugging, playing guitar). When an interaction implies additional props, like ‘taking a selfie’, our model autonomously generates these supporting objects. By jointly training for compositing and subject-driven generation, also known as customization, we achieve a more balanced integration of textual and visual inputs for text-driven object compositing. As a result, we obtain a versatile model with state-of-the-art performance in both tasks. We further present a data generation pipeline leveraging visual and language models to effortlessly synthesize multimodal, aligned training data. Gemma Canet Tarrés, Zhe Lin 0001, He Zhang 0004, Andrew Gilbert, John P. Collomosse, Soo Ye Kim |
CVPR | 4 |
| 2025 | TransPixeler: Advancing Text-to-Video Generation with TransparencyabstractText-to-video generative models have made significant strides, enabling diverse applications in entertainment, advertising, and education. However, generating RGBA video, which includes alpha channels for transparency, remains a challenge due to limited datasets and the difficulty of adapting existing models. Alpha channels are crucial for visual effects (VFX), allowing transparent elements like smoke and reflections to blend seamlessly into scenes. We introduce TransPixeler, a method to extend pretrained video models for RGBA generation while retaining the original RGB capabilities. TransPixeler leverages a diffusion transformer (DiT) architecture, incorporating alpha-specific tokens and using LoRA-based fine-tuning to jointly generate RGB and alpha channels with high consistency. By optimizing attention mechanisms, TransPixeler preserves the strengths of the original RGB model and achieves strong alignment between RGB and alpha channels despite limited training data. Our approach effectively generates diverse and consistent RGBA videos, advancing the possibilities for VFX and interactive content creation. The code is available at https://wileewang.github.io/TransPixeler/. Luozhou Wang, Yijun Li 0001, Zhifei Chen, Jui-Hsien Wang, He Zhang 0004, Zhe Lin 0001, Ying-Cong Chen |
CVPR | 6 |
| 2025 | Comprehensive Relighting: Generalizable and Consistent Monocular Human Relighting and HarmonizationabstractThis paper introduces Comprehensive Relighting, the first all-in-one approach that can both control and harmonize the lighting from an image or video of humans with arbitrary body parts from any scene. Building such a generalizable model is extremely challenging due to the lack of dataset, restricting existing image-based relighting models to a specific scenario (e.g., face or static human). To address this challenge, we repurpose a pre-trained diffusion model as a general image prior and jointly model the human relighting and background harmonization in the coarse-to-fine framework. To further enhance the temporal coherence of the relighting, we introduce an unsupervised temporal lighting model that learns the lighting cycle consistency from many real-world videos without any ground truth. In inference time, our temporal lighting module is combined with the diffusion models through the spatio-temporal feature blending algorithms without extra training; and we apply a new guided refinement as a post-processing to pre-serve the high-frequency details from the input image. In the experiments, Comprehensive Relighting shows a strong generalizability and lighting temporal coherence, outperforming existing image-based human relighting and harmonization methods. Xin Sun 0014, Krishna Kumar Singh, Zhixin Shu, He Zhang 0004, Jimei Yang, Nanxuan Zhao, Tuanfeng Y. Wang, Simon S. Chen, Ulrich Neumann, Jae Shin Yoon |
CVPR | 6 |
| 2025 | Baking Gaussian Splatting Into Diffusion Denoiser for Fast and Scalable Single-Stage Image-to-3D Generation and Reconstruction
Yuanhao Cai, He Zhang 0004, Kai Zhang 0045, Yixun Liang, Mengwei Ren, Fujun Luan, Qing Liu 0017, Soo Ye Kim, Jianming Zhang 0001, Yuqian Zhou, Yulun Zhang 0001, Xiaokang Yang 0001, Zhe Lin 0001, Alan L. Yuille |
ICCV | 2 |
| 2025 | DIVE: Taming DINO for Subject-Driven Video EditingabstractBuilding on the success of diffusion models in image generation and editing, video editing has recently gained substantial attention. However, maintaining temporal consistency and motion alignment still remains challenging. To address these issues, this paper proposes DINO-guided Video Editing (DIVE), a framework designed to facilitate subject-driven editing in source videos conditioned on either target text prompts or reference images with specific identities. The core of DIVE lies in leveraging the powerful semantic features extracted from a pretrained DINOv2 model as implicit correspondences to guide the editing process. Specifically, to ensure temporal motion consistency, DIVE employs DINO features to align with the motion trajectory of the source video. For precise subject editing, DIVE incorporates the DINO features of reference images into a pretrained text-to-image model to learn Low-Rank Adaptations (LoRAs), effectively registering the target subject's identity. Extensive experiments on diverse real-world videos demonstrate that our framework can achieve high-quality editing results with robust motion consistency, highlighting the potential of DINO to contribute to video editing. Project page: https://dino-video-editing.github.io Yi Huang 0035, Wei Xiong 0008, He Zhang 0004, Chaoqi Chen, Jianzhuang Liu, Mingfu Yan, Shifeng Chen |
ICCV | 3 |
| 2025 | Refine-by-Align: Reference-Guided Artifacts Refinement through Semantic AlignmentabstractPersonalized image generation has emerged from the recent advancements in generative models. However, these generated personalized images often suffer from localized artifacts such as incorrect logos, reducing fidelity and fine-grained identity details of the generated results. Furthermore, there is little prior work tackling this problem. To help improve these identity details in the personalized image generation, we introduce a new task: reference-guided artifacts refinement. We present Refine-by-Align, a first-of-its-kind model that employs a diffusion-based framework to address this challenge. Our model consists of two stages: Alignment Stage and Refinement Stage, which share weights of a unified neural network model. Given a generated image, a masked artifact region, and a reference image, the alignment stage identifies and extracts the corresponding regional features in the reference, which are then used by the refinement stage to fix the artifacts. Our model-agnostic pipeline requires no test-time tuning or optimization. It automatically enhances image fidelity and reference identity in the generated image, generalizing well to existing models on various tasks including but not limited to customization, generative compositing, view synthesis, and virtual try-on. Extensive experiments and comparisons demonstrate that our pipeline greatly pushes the boundary of fine details in the image synthesis models. Soo Ye Kim, He Zhang 0004, Wei Xiong 0008, Zhe Lin 0001, Brian L. Price, Scott Cohen, Jianming Zhang 0001, Daniel G. Aliaga |
ICLR | 5 |
| 2025 | RelitLRM: Generative Relightable Radiance for Large Reconstruction ModelsabstractWe propose RelitLRM, a Large Reconstruction Model (LRM) for generating high-quality Gaussian splatting representations of 3D objects under novel illuminations from sparse (4-8) posed images captured under unknown static lighting. Unlike prior inverse rendering methods requiring dense captures and slow optimization, often causing artifacts like incorrect highlights or shadow baking, RelitLRM adopts a feed-forward transformer-based model with a novel combination of a geometry reconstructor and a relightable appearance generator based on diffusion. The model is trained end-to-end on synthetic multi-view renderings of objects under varying known illuminations. This architecture design enables to effectively decompose geometry and appearance, resolve the ambiguity between material and lighting, and capture the multi-modal distribution of shadows and specularity in the relit appearance. We show our sparse-view feed-forward RelitLRM offers competitive relighting results to state-of-the-art dense-view optimization-based baselines while being significantly faster. Our project page is available at: https://relit-lrm.github.io/. Zhengfei Kuang, Haian Jin, Zexiang Xu, Sai Bi, Hao Tan 0002, He Zhang 0004, Milos Hasan, William T. Freeman, Kai Zhang 0045, Fujun Luan |
ICLR | 7 |
| 2025 | OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control ConditionsabstractExisting feedforward subject-driven video customization methods mainly study single-subject scenarios due to the difficulty of constructing multi-subject training data pairs. Another challenging problem that how to use the signals such as depth, mask, camera, and text prompts to control and edit the subject in the customized video is still less explored. In this paper, we first propose a data construction pipeline, VideoCus-Factory, to produce training data pairs for multi-subject customization from raw videos without labels and control signals such as depth-to-video and mask-to-video pairs. Based on our constructed data, we develop an Image-Video Transfer Mixed (IVTM) training with image editing data to enable instructive editing for the subject in the customized video. Then we propose a diffusion Transformer framework, OmniVCus, with two embedding mechanisms, Lottery Embedding (LE) and Temporally Aligned Embedding (TAE). LE enables inference with more subjects by using the training subjects to activate more frame embeddings. TAE encourages the generation process to extract guidance from temporally aligned control signals by assigning the same frame embeddings to the control and noise tokens. Experiments demonstrate that our method significantly surpasses state-of-the-art methods in both quantitative and qualitative evaluations. Project page is at https://caiyuanhao1998.github.io/project/OmniVCus/ Yuanhao Cai, He Zhang 0004, Jinbo Xing, Yuqian Zhou, Soo Ye Kim, Yulun Zhang 0001, Xiaokang Yang 0001, Alan L. Yuille |
NeurIPS | 2 |
| 2025 | Diffusion Model-Based Image Editing: A SurveyabstractDenoising diffusion models have emerged as a powerful tool for various image generation and editing tasks, facilitating the synthesis of visual content in an unconditional or input-conditional manner. The core idea behind them is learning to reverse the process of gradually adding noise to images, allowing them to generate high-quality samples from a complex distribution. In this survey, we provide an exhaustive overview of existing methods using diffusion models for image editing, covering both theoretical and practical aspects in the field. We delve into a thorough analysis and categorization of these works from multiple perspectives, including learning strategies, user-input conditions, and the array of specific editing tasks that can be accomplished. In addition, we pay special attention to image inpainting and outpainting, and explore both earlier traditional context-driven and current multimodal conditional methods, offering a comprehensive analysis of their methodologies. To further evaluate the performance of text-guided image editing algorithms, we propose a systematic benchmark, EditEval, featuring an innovative metric, LMM Score. Finally, we address current limitations and envision some potential directions for future research. Yi Huang 0035, Jiancheng Huang, Yifan Liu 0001, Mingfu Yan, Jiaxi Lv, Jianzhuang Liu, Wei Xiong 0008, He Zhang 0004, Liangliang Cao, Shifeng Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2024 | Holo-Relighting: Controllable Volumetric Portrait Relighting from a Single ImageabstractAt the core of portrait photography is the search for ideal lighting and viewpoint. The process often requires advanced knowledge in photography and an elaborate studio setup. In this work, we propose Holo-Relighting, a volumetric relighting method that is capable of synthesizing novel viewpoints, and novel lighting from a single image. Holo-Relighting leverages the pretrained 3D GAN (EG3D) to reconstruct geometry and appearance from an input portrait as a set of 3D-aware features. We design a relighting module conditioned on a given lighting to process these features, and predict a relit 3D representation in the form of a tri-plane, which can render to an arbitrary viewpoint through volume rendering. Besides viewpoint and lighting control, Holo-Relighting also takes the head pose as a condition to enable head-pose-dependent lighting effects. With these novel designs, Holo-Relighting can generate complex non-Lambertian lighting effects (e.g., specular highlights and cast shadows) without using any explicit physical lighting priors. We train Holo-Relighting with data captured with a light stage, and propose two data-rendering techniques to improve the data quality for training the volumetric relighting system. Through quantitative and qualitative experiments, we demonstrate Holo-Relighting can achieve state-of-the-arts relighting quality with better photorealism, 3D consistency and controllability. Yiqun Mei, Yu Zeng 0001, He Zhang 0004, Zhixin Shu, Xuaner Cecilia Zhang, Sai Bi, Jianming Zhang 0001, Hyunjoon Jung, Vishal M. Patel |
CVPR | 3 |
| 2024 | Relightful Harmonization: Lighting-Aware Portrait Background ReplacementabstractPortrait harmonization aims to composite a subject into a new background, adjusting its lighting and color to ensure harmony with the background scene. Existing harmo-nization techniques often only focus on adjusting the global color and brightness of the foreground and ignore crucial illumination cues from the background such as apparent lighting direction, leading to unrealistic compositions. We introduce Relightful Harmonization, a lighting-aware diffusion model designed to seamlessly harmonize sophisticated lighting effect for the foreground portrait using any back-ground image. Our approach unfolds in three stages. First, we introduce a lighting representation module that allows our diffusion model to encode lighting information from target image background. Second, we introduce an alignment network that aligns lighting features learned from image background with lighting features learned from panorama environment maps, which is a complete representation for scene illumination. Last, to further boost the photorealism of the proposed method, we introduce a novel data simulation pipeline that generates synthetic training pairs from a diverse range of natural images, which are used to refine the model. Our method outperforms existing benchmarks in visual fidelity and lighting coherence, showing superior generalization in real-world testing scenarios, highlighting its versatility and practicality. Mengwei Ren, Wei Xiong 0008, Jae Shin Yoon, Zhixin Shu, Jianming Zhang 0001, Hyunjoon Jung, Guido Gerig, He Zhang 0004 |
CVPR | 8 |
| 2024 | IMPRINT: Generative Object Compositing by Learning Identity-Preserving RepresentationabstractGenerative object compositing emerges as a promising new avenue for compositional image editing. However, the requirement of object identity preservation poses a significant challenge, limiting practical usage of most existing methods. In response, this paper introduces IMPRINT, a novel diffusion-based generative model trained with a two-stage learning framework that decouples learning of identity preservation from that of compositing. The first stage is targeted for context-agnostic, identity-preserving pretraining of the object encoder, enabling the encoder to learn an embedding that is both view-invariant and conducive to enhanced detail preservation. The subsequent stage leverages this representation to learn seamless harmonization of the object composited to the background. In addition, IMPRINT incorporates a shape-guidance mechanism offering user-directed control over the compositing process. Extensive experiments demonstrate that IMPRINT significantly outperforms existing methods and various baselines on identity preservation and composition quality. Project page: https://song630.github.io/IMPRINT-Project-Page/ Zhe Lin 0001, Scott Cohen, Brian L. Price, Jianming Zhang 0001, Soo Ye Kim, He Zhang 0004, Wei Xiong 0008, Daniel G. Aliaga |
CVPR | 8 |
| 2024 | SwapAnything: Enabling Arbitrary Object Swapping in Personalized Image Editing
Nanxuan Zhao, Wei Xiong 0008, Qing Liu 0017, He Zhang 0004, Jianming Zhang 0001, Hyunjoon Jung, Yilin Wang 0002, Xin Wang 0061 |
ECCV (32) | 6 |
| 2024 | COMPOSE: Comprehensive Portrait Shadow Editing
Andrew Hou, Zhixin Shu, Xuaner Cecilia Zhang, He Zhang 0004, Yannick Hold-Geoffroy, Jae Shin Yoon, Xiaoming Liu 0002 |
ECCV (61) | 4 |
| 2024 | Generative Portrait Shadow RemovalabstractWe introduce a high-fidelity portrait shadow removal model that can effectively enhance the image of a portrait by predicting its appearance under disturbing shadows and highlights. Portrait shadow removal is a highly ill-posed problem where multiple plausible solutions can be found based on a single image. For example, disentangling complex environmental lighting from original skin color is a non-trivial problem. While existing works have solved this problem by predicting the appearance residuals that can propagate local shadow distribution, such methods are often incomplete and lead to unnatural predictions, especially for portraits with hard shadows. We overcome the limitations of existing local propagation methods by formulating the removal problem as a generation task where a diffusion model learns to globally rebuild the human appearance from scratch as a condition of an input portrait image. For robust and natural shadow removal, we propose to train the diffusion model with a compositional repurposing framework: a pre-trained text-guided image generation model is first fine-tuned to harmonize the lighting and color of the foreground with a background scene by using a background harmonization dataset; and then the model is further fine-tuned to generate a shadow-free portrait image via a shadow-paired dataset. To overcome the limitation of losing fine details in the latent diffusion model, we propose a guided-upsampling network to restore the original high-frequency details (e.g. , wrinkles and dots) from the input image. To enable our compositional training framework, we construct a high-fidelity and large-scale dataset using a lightstage capturing system and synthetic graphics simulation. Our generative framework effectively removes shadows caused by both self and external occlusions while maintaining original lighting distribution and high-frequency details. Our method also demonstrates robustness to diverse subjects captured in real environments. Jae Shin Yoon, Zhixin Shu, Mengwei Ren, Cecilia Zhang, Yannick Hold-Geoffroy, Krishna Kumar Singh, He Zhang 0004 |
ACM Trans. Graph. | 7 |
| 2023 | LightPainter: Interactive Portrait Relighting with Freehand ScribbleabstractRecent portrait relighting methods have achieved realistic results of portrait lighting effects given a desired lighting representation such as an environment map. However, these methods are not intuitive for user interaction and lack precise lighting control. We introduce LightPainter, a scribble-based relighting system that allows users to interactively manipulate portrait lighting effect with ease. This is achieved by two conditional neural networks, a delighting module that recovers geometry and albedo optionally conditioned on skin tone, and a scribble-based module for re-lighting. To train the relighting module, we propose a novel scribble simulation procedure to mimic real user scribbles, which allows our pipeline to be trained without any human annotations. We demonstrate high-quality and flexible portrait lighting editing capability with both quantitative and qualitative experiments. User study comparisons with commercial lighting editing tools also demonstrate consistent user preference for our method. Yiqun Mei, He Zhang 0004, Xuaner Cecilia Zhang, Jianming Zhang 0001, Zhixin Shu, Yilin Wang 0002, Zijun Wei, Hyunjoon Jung, Vishal M. Patel |
CVPR | 2 |
| 2023 | PixHt-Lab: Pixel Height Based Light Effect Generation for Image CompositingabstractLighting effects such as shadows or reflections are key in making synthetic images realistic and visually appealing. To generate such effects, traditional computer graphics uses a physically-based renderer along with 3D geometry. To compensate for the lack of geometry in 2D Image compositing, recent deep learning-based approaches introduced a pixel height representation to generate soft shadows and reflections. However, the lack of geometry limits the quality of the generated soft shadows and constrains reflections to pure specular ones. We introduce PixHt-Lab, a system leveraging an explicit mapping from pixel height representation to 3D space. Using this mapping, PixHt- Lab reconstructs both the cutout and background geometry and renders realistic, diverse lighting effects for image compositing. Given a surface with physically-based materials, we can render reflections with varying glossiness. To generate more realistic soft shadows, we further propose using 3D-aware buffer channels to guide a neural renderer. Both quantitative and qualitative evaluations demonstrate that PixHt-Lab significantly improves soft shadow generation. Project: https://shengcn.github.io/PixHtLab/ Yichen Sheng, Jianming Zhang 0001, Julien Philip, Yannick Hold-Geoffroy, Xin Sun 0014, He Zhang 0004, Lu Ling, Bedrich Benes |
CVPR | 6 |
| 2023 | Semi-Supervised Parametric Real-World Image HarmonizationabstractLearning-based image harmonization techniques are usually trained to undo synthetic random global transformations applied to a masked foreground in a single ground truth photo. This simulated data does not model many of the important appearance mismatches (illumination, object boundaries, etc.) between foreground and background in real composites, leading to models that do not generalize well and cannot model complex local changes. We propose a new semi-supervised training strategy that addresses this problem and lets us learn complex local appearance harmonization from unpaired real composites, where foreground and background come from different images. Our model is fully parametric. It uses RGB curves to correct the global colors and tone and a shading map to model local variations. Our method outperforms previous work on established benchmarks and real composites, as shown in a user study, and processes high-resolution images interactively. Code, and project page available at: https://kewang0622.github.io/sprih/ Michaël Gharbi, He Zhang 0004, Zhihao Xia, Eli Shechtman |
CVPR | 3 |
| 2023 | Perceptual Artifacts Localization for Image Synthesis TasksabstractRecent advancements in deep generative models have facilitated the creation of photo-realistic images across various tasks. However, these generated images often exhibit perceptual artifacts in specific regions, necessitating manual correction. In this study, we present a comprehensive empirical examination of Perceptual Artifacts Localization (PAL) spanning diverse image synthesis endeavors. We introduce a novel dataset comprising 10, 168 generated images, each annotated with per-pixel perceptual artifact labels across ten synthesis tasks. A segmentation model, trained on our proposed dataset, effectively localizes artifacts across a range of tasks. Additionally, we illustrate its proficiency in adapting to previously unseen models using minimal training samples. We further propose an innovative zoom-in inpainting pipeline that seamlessly rectifies perceptual artifacts in the generated images. Through our experimental analyses, we elucidate several invaluable downstream applications, such as automated artifact rectification, non-referential image quality evaluation, and abnormal region detection in images. The dataset and code are released here: https://owenzlz.github.io/PAL4VST Lingzhi Zhang, Zhengjie Xu, Connelly Barnes, Yuqian Zhou, Qing Liu 0017, He Zhang 0004, Sohrab Amirghodsi, Zhe Lin 0001, Eli Shechtman, Jianbo Shi |
ICCV | 6 |
| 2023 | Interactive Portrait Harmonization
Jeya Maria Jose Valanarasu, He Zhang 0004, Jianming Zhang 0001, Yilin Wang 0002, Zhe Lin 0001, Jose Echevarria, Yinglan Ma, Zijun Wei, Kalyan Sunkavalli, Vishal M. Patel |
ICLR | 2 |
| 2023 | PHOTOSWAP: Personalized Subject Swapping in ImagesabstractIn an era where images and visual content dominate our digital landscape, the ability to manipulate and personalize these images has become a necessity.
Envision seamlessly substituting a tabby cat lounging on a sunlit window sill in a photograph with your own playful puppy, all while preserving the original charm and composition of the image.
We present \emph{Photoswap}, a novel approach that enables this immersive image editing experience through personalized subject swapping in existing images.
\emph{Photoswap} first learns the visual concept of the subject from reference images and then swaps it into the target image using pre-trained diffusion models in a training-free manner. We establish that a well-conceptualized visual subject can be seamlessly transferred to any image with appropriate self-attention and cross-attention manipulation, maintaining the pose of the swapped subject and the overall coherence of the image.
Comprehensive experiments underscore the efficacy and controllability of \emph{Photoswap} in personalized subject swapping. Furthermore, \emph{Photoswap} significantly outperforms baseline methods in human ratings across subject swapping, background preservation, and overall quality, revealing its vast application potential, from entertainment to professional editing. Yilin Wang 0002, Nanxuan Zhao, Tsu-Jui Fu, Wei Xiong 0008, Qing Liu 0017, He Zhang 0004, Jianming Zhang 0001, Hyunjoon Jung, Xin Wang 0061 |
NeurIPS | 8 |
| 2023 | Semi-Supervised Domain Alignment Learning for Single Image DehazingabstractConvolutional neural networks (CNNs) have attracted much research attention and achieved great improvements in single-image dehazing. However, previous learning-based dehazing methods are mainly trained on synthetic data, which greatly degrades their generalization capability on natural hazy images. To address this issue, this article proposes a semi-supervised learning approach for single-image dehazing, where both synthetic and realistic images are leveraged during training. Considering the situation that it is hard to obtain the realistic pairs of hazy and haze-free images, how to utilize the realistic data is not a trivial work. In this article, a domain alignment module is introduced to narrow the distribution distance between synthetic data and realistic hazy images in a latent feature space. Meanwhile, a haze-aware attention module is designed to describe haze densities of different regions in the image, thus adaptively responds for different hazy areas. Furthermore, the dark channel prior is introduced to the framework to improve the quality of the unsupervised learning results by considering the statistical characters of haze-free images. Such a semi-supervised design can significantly address the domain shift issue between the synthetic and realistic data, and improve generalization performance in the real world. Experiments indicate that the proposed method obtains state-of-the-art performance on both public synthetic and realistic hazy images with better visual results. Yunan Li 0001, He Zhang 0004, Shifeng Chen |
IEEE Trans. Cybern. | 4 |
| 2022 | Boosting Robustness of Image Matting with Context Assembling and Strong Data AugmentationabstractDeep image matting methods have achieved increasingly better results on benchmarks (e.g., Composition-1k/alphamatting.com). However, the robustness, including robustness to trimaps and generalization to images from different domains, is still underexplored. Although some works propose to either refine the trimaps or adapt the algorithms to real-world images via extra data augmentation, none of them has taken both into consideration, not to mention the significant performance deterioration on benchmarks while using those data augmentation. To fill this gap, we propose an image matting method which achieves higher robustness (RMat) via multilevel context assembling and strong data augmentation targeting matting. Specifically, we first build a strong matting framework by modeling ample global information with transformer blocks in the encoder, and focusing on details in combination with convolution layers as well as a low-level feature assembling attention block in the decoder. Then, based on this strong baseline, we analyze current data augmentation and explore simple but effective strong data augmentation to boost the baseline model and contribute a more generalizable matting method. Compared with previous methods, the proposed method not only achieves state-of-the-art results on the Composition-1k benchmark (11 % improvement on SAD and 27% improvement on Grad) with smaller model size, but also shows more robust generalization results on other benchmarks, on real-world images, and also on varying coarse-to-fine trimaps with our extensive experiments.11This work was in part done when YD was an intern at Adobe and CS was with The University of Adelaide. CS is the corresponding author. Project page: https://dongdong93.github.io/RMat/. Yutong Dai 0001, Brian L. Price, He Zhang 0004, Chunhua Shen |
CVPR | 3 |
| 2022 | Lite Vision Transformer with Enhanced Self-AttentionabstractDespite the impressive representation capacity of vision transformer models, current light-weight vision transformer models still suffer from inconsistent and incorrect dense predictions at local regions. We suspect that the power of their self-attention mechanism is limited in shallower and thinner networks. We propose Lite Vision Transformer (LVT), a novel light-weight transformer network with two enhanced self-attention mechanisms to improve the model performances for mobile deployment. For the low-level features, we introduce Convolutional Self-Attention (CSA). Unlike previous approaches of merging convolution and self-attention, CSA introduces local self-attention into the convolution within a kernel of size$3\times 3$to enrich low-level features in the first stage of LVT. For the high-level features, we propose Recursive Atrous Self-Attention (RASA), which utilizes the multi-scale context when calculating the similarity map and a recursive mechanism to increase the representation capability with marginal extra parameter cost. The superiority of LVT is demonstrated on ImageNet recognition, ADE20K semantic segmentation, and COCO panoptic segmentation. The code is made publicly available11https://github.com/Chenglin-Yang/LVT. Yilin Wang 0002, Jianming Zhang 0001, He Zhang 0004, Zijun Wei, Zhe Lin 0001, Alan L. Yuille |
CVPR | 4 |
| 2022 | Controllable Shadow Generation Using Pixel Height Maps
Yichen Sheng, Yifan Liu 0001, Jianming Zhang 0001, Wei Yin 0006, A. Cengiz Öztireli, He Zhang 0004, Zhe Lin 0001, Eli Shechtman, Bedrich Benes |
ECCV (23) | 6 |
| 2022 | Hierarchical Density-Aware Dehazing NetworkabstractThe commonly used atmospheric model in image dehazing cannot hold in real cases. Although deep end-to-end networks were presented to solve this problem by disregarding the physical model, the transmission map in the atmospheric model contains significant haze density information, which cannot simply be ignored. In this article, we propose a novel hierarchical density-aware dehazing network, which consists of a the densely connected pyramid encoder, a density generator, and a Laplacian pyramid decoder. The proposed network incorporates density estimation but alleviates the constraint of the atmospheric model. The predicted haze density then guides the Laplacian pyramid decoder to generate a haze-free image in a coarse-to-fine fashion. In addition, we introduce a multiscale discriminator to preserve global and local consistency for dehazing. We conduct extensive experiments on natural and synthetic hazy images, which prove that the proposed model performs favorably against the state-of-the-art dehazing approaches. Jingang Zhang, Wenqi Ren, Shengdong Zhang, He Zhang 0004, Yunfeng Nie, Zhe Xue, Xiaochun Cao |
IEEE Trans. Cybern. | 4 |
| 2021 | Mask Guided Matting via Progressive Refinement NetworkabstractWe propose Mask Guided (MG) Matting, a robust matting framework that takes a general coarse mask as guidance. MG Matting leverages a network (PRN) design which encourages the matting model to provide self-guidance to progressively refine the uncertain regions through the decoding process. A series of guidance mask perturbation operations are also introduced in the training to further enhance its robustness to external guidance. We show that PRN can generalize to unseen types of guidance masks such as trimap and low-quality alpha matte, making it suitable for various application pipelines. In addition, we revisit the foreground color prediction problem for matting and propose a surprisingly simple improvement to address the dataset issue. Evaluation on real and synthetic benchmarks shows that MG Matting achieves state-of-the-art performance using various types of guidance inputs. Code and models are available at https://github.com/yucornetto/MGMatting. Qihang Yu, Jianming Zhang 0001, He Zhang 0004, Yilin Wang 0002, Zhe Lin 0001, Ning Xu 0007, Yutong Bai, Alan L. Yuille |
CVPR | 3 |
| 2021 | SSH: A Self-Supervised Framework for Image HarmonizationabstractImage harmonization aims to improve the quality of image compositing by matching the "appearance" (e.g., color tone, brightness and contrast) between foreground and background images. However, collecting large-scale annotated datasets for this task requires complex professional retouching. Instead, we propose a novel Self-Supervised Harmonization framework (SSH) that can be trained using just "free" natural images without being edited. We reformulate the image harmonization problem from a representation fusion perspective, which separately processes the foreground and background examples, to address the background occlusion issue. This framework design allows for a dual data augmentation method, where diverse [foreground, background, pseudo GT] triplets can be generated by cropping an image with perturbations using 3D color lookup tables (LUTs). In addition, we build a real-world harmonization dataset as carefully created by expert users, for evaluation and benchmarking purposes. Our results show that the proposed self-supervised method outperforms previous state-of-the-art methods in terms of reference metrics, visual quality, and subject user study. Code and dataset are available at https://github.com/VITA-Group/SSHarmonization. Yifan Jiang 0001, He Zhang 0004, Jianming Zhang 0001, Yilin Wang 0002, Zhe Lin 0001, Kalyan Sunkavalli, Simon Chen, Sohrab Amirghodsi, Sarah Kong, Zhangyang Wang |
ICCV | 2 |
| 2021 | Deep Image CompositingabstractImage compositing is a task of combining regions from different images to compose a new image. A common use case is background replacement of portrait images. To obtain high quality composites, professionals typically manually perform multiple editing steps such as segmentation, matting and foreground color decontamination, which is very time consuming even with sophisticated photo editing tools. In this paper, we propose a new method which can automatically generate high-quality image compositing with-out any user input. Our method can be trained end-to-end to optimize exploitation of contextual and color information of both foreground and background images, where the com-positing quality is considered in the optimization. Specifically, inspired by Laplacian pyramid blending, a dense-connected multi-stream fusion network is proposed to effectively fuse the information from the foreground and back-ground images at different scales. In addition, we intro-duce a self-taught strategy to progressively train from easy to complex cases to mitigate the lack of training data. Experiments show that the proposed method can automatically generate high-quality composites and outperforms existing methods both qualitatively and quantitatively. He Zhang 0004, Jianming Zhang 0001, Federico Perazzi, Zhe Lin 0001, Vishal M. Patel |
WACV | 1 |
| 2020 | FD-GAN: Generative Adversarial Networks with Fusion-Discriminator for Single Image DehazingabstractRecently, convolutional neural networks (CNNs) have achieved great improvements in single image dehazing and attained much attention in research. Most existing learning-based dehazing methods are not fully end-to-end, which still follow the traditional dehazing procedure: first estimate the medium transmission and the atmospheric light, then recover the haze-free image based on the atmospheric scattering model. However, in practice, due to lack of priors and constraints, it is hard to precisely estimate these intermediate parameters. Inaccurate estimation further degrades the performance of dehazing, resulting in artifacts, color distortion and insufficient haze removal. To address this, we propose a fully end-to-end Generative Adversarial Networks with Fusion-discriminator (FD-GAN) for image dehazing. With the proposed Fusion-discriminator which takes frequency information as additional priors, our model can generator more natural and realistic dehazed images with less color distortion and fewer artifacts. Moreover, we synthesize a large-scale training dataset including various indoor and outdoor hazy images to boost the performance and we reveal that for learning-based dehazing methods, the performance is strictly influenced by the training data. Experiments have shown that our method reaches state-of-the-art performance on both public synthetic datasets and real-world images with more visually pleasing dehazed results. Yihao Liu 0001, He Zhang 0004, Shifeng Chen, Yu Qiao 0001 |
AAAI | 3 |
| 2020 | Joint Transmission Map Estimation and Dehazing Using Deep NetworksabstractSingle image haze removal is an extremely challenging problem due to its inherent ill-posed nature. Several prior-based and learning-based methods have been proposed in the literature to solve this problem and they have achieved visually appealing results. However, most of the existing methods assume constant atmospheric light model and tend to follow a two-step procedure involving prior-based methods for estimating transmission map followed by calculation of dehazed image using the closed form solution. In this paper, we relax the constant atmospheric light assumption and propose a novel unified single image dehazing network that jointly estimates the transmission map and performs dehazing. In other words, our new approach provides an end-to-end learning framework, where the inherent transmission map and dehazed result are learned jointly from the loss function. The extensive experiments evaluated on synthetic and real datasets with challenging hazy images demonstrate that the proposed method achieves significant improvements over the state-of-the-art methods. He Zhang 0004, Vishwanath A. Sindagi, Vishal M. Patel |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Image De-Raining Using a Conditional Generative Adversarial NetworkabstractSevere weather conditions, such as rain and snow, adversely affect the visual quality of images captured under such conditions, thus rendering them useless for further usage and sharing. In addition, such degraded images drastically affect the performance of vision systems. Hence, it is important to address the problem of single image de-raining. However, the inherent ill-posed nature of the problem presents several challenges. We attempt to leverage powerful generative modeling capabilities of the recently introduced conditional generative adversarial networks (CGAN) by enforcing an additional constraint that the de-rained image must be indistinguishable from its corresponding ground truth clean image. The adversarial loss from GAN provides additional regularization and helps to achieve superior results. In addition to presenting a new approach to de-rain images, we introduce a new refined loss function and architectural novelties in the generator-discriminator pair for achieving improved results. The loss function is aimed at reducing artifacts introduced by GANs and ensure better visual quality. The generator sub-network is constructed using the recently introduced densely connected networks, whereas the discriminator is designed to leverage global and local information to decide if an image is real/fake. Based on this, we propose a novel single image de-raining method called image de-raining conditional generative adversarial network (ID-CGAN) that considers quantitative, visual, and also discriminative performance into the objective function. The experiments evaluated on synthetic and real images show that the proposed method outperforms many recent state-of-the-art single image de-raining methods in terms of quantitative and visual performances. Furthermore, the experimental results evaluated on object detection datasets using the Faster-RCNN also demonstrate the effectiveness of proposed method in improving the detection performance on images degraded by rain. He Zhang 0004, Vishwanath A. Sindagi, Vishal M. Patel |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Synthesis of High-Quality Visible Faces from Polarimetric Thermal Faces using Generative Adversarial Networks
He Zhang 0004, Benjamin S. Riggan, Shuowen Hu, Nathan J. Short, Vishal M. Patel |
Int. J. Comput. Vis. | 1 |
| 2019 | Convolutional Sparse Coding for Compressed Sensing CT ReconstructionabstractOver the past few years, dictionary learning (DL)-based methods have been successfully used in various image reconstruction problems. However, the traditional DL-based computed tomography (CT) reconstruction methods are patch-based and ignore the consistency of pixels in overlapped patches. In addition, the features learned by these methods always contain shifted versions of the same features. In recent years, convolutional sparse coding (CSC) has been developed to address these problems. In this paper, inspired by several successful applications of CSC in the field of signal processing, we explore the potential of CSC in sparse-view CT reconstruction. By directly working on the whole image, without the necessity of dividing the image into overlapped patches in DL-based methods, the proposed methods can maintain more details and avoid artifacts caused by patch aggregation. With predetermined filters, an alternating scheme is developed to optimize the objective function. Extensive experiments with simulated and real CT data were performed to validate the effectiveness of the proposed methods. The qualitative and quantitative results demonstrate that the proposed methods achieve better performance than the several existing state-of-the-art methods. Peng Bao 0001, Huaiqiang Sun, Zhangyang Wang, Yi Zhang 0018, Wenjun Xia, Mianyi Chen, Yan Xi, Shanzhou Niu, Jiliu Zhou, He Zhang 0004 |
IEEE Trans. Medical Imaging | 12 |
| 2018 | Density-Aware Single Image De-Raining Using a Multi-Stream Dense NetworkabstractSingle image rain streak removal is an extremely challenging problem due to the presence of non-uniform rain densities in images. We present a novel density-aware multi-stream densely connected convolutional neural network-based algorithm, called DID-MDN, for joint rain density estimation and de-raining. The proposed method enables the network itself to automatically determine the rain-density information and then efficiently remove the corresponding rain-streaks guided by the estimated rain-density label. To better characterize rain-streaks with different scales and shapes, a multi-stream densely connected de-raining network is proposed which efficiently leverages features from different scales. Furthermore, a new dataset containing images with rain-density labels is created and used to train the proposed density-aware network. Extensive experiments on synthetic and real datasets demonstrate that the proposed method achieves significant improvements over the recent state-of-the-art methods. In addition, an ablation study is performed to demonstrate the improvements obtained by different modules in the proposed method. The code can be downloaded at https://github.com/hezhangsprinter/DID-MDN. He Zhang 0004, Vishal M. Patel |
CVPR | 1 |
| 2018 | Densely Connected Pyramid Dehazing NetworkabstractWe propose a new end-to-end single image dehazing method, called Densely Connected Pyramid Dehazing Network (DCPDN), which can jointly learn the transmission map, atmospheric light and dehazing all together. The end-to-end learning is achieved by directly embedding the atmospheric scattering model into the network, thereby ensuring that the proposed method strictly follows the physics-driven scattering model for dehazing. Inspired by the dense network that can maximize the information flow along features from different levels, we propose a new edge-preserving densely connected encoder-decoder structure with multi-level pyramid pooling module for estimating the transmission map. This network is optimized using a newly introduced edge-preserving loss function. To further incorporate the mutual structural information between the estimated transmission map and the dehazed result, we propose a joint-discriminator based on generative adversarial network framework to decide whether the corresponding dehazed image and the estimated transmission map are real or fake. An ablation study is conducted to demonstrate the effectiveness of each module evaluated at both estimated transmission map and dehazed result. Extensive experiments demonstrate that the proposed method achieves significant improvements over the state-of-the-art methods. Code and dataset is made available at: https://github.com/hezhangsprinter/DCPDN. He Zhang 0004, Vishal M. Patel |
CVPR | 1 |
| 2018 | Convolutional Sparse and Low-Rank Coding-Based Image DecompositionabstractWe propose novel convolutional sparse and low-rank coding-based methods for cartoon and texture decomposition. In our method, we first learn a set of generic filters that can efficiently represent cartoon-and texture-type images. Then, using these learned filters, we propose two optimization frameworks to decompose a given image into cartoon and texture components: convolutional sparse coding-based image decomposition; and convolutional low-rank coding-based image decomposition. By working directly on the whole image, the proposed image separation algorithms do not need to divide the image into overlapping patches for leaning local dictionaries. The shift-invariance property is directly modeled into the objective function for learning filters. Extensive experiments show that the proposed methods perform favorably compared with state-of-the-art image separation methods. He Zhang 0004, Vishal M. Patel |
IEEE Trans. Image Process. | 1 |
| 2017 | Generative adversarial network-based synthesis of visible faces from polarimetrie thermal facesabstractThe large domain discrepancy between faces captured in polarimetric (or conventional) thermal and visible domain makes cross-domain face recognition quite a challenging problem for both human-examiners and computer vision algorithms. Previous approaches utilize a two-step procedure (visible feature estimation and visible image reconstruction) to synthesize the visible image given the corresponding polarimetric thermal image. However, these are regarded as two disjoint steps and hence may hinder the performance of visible face reconstruction. We argue that joint optimization would be a better way to reconstruct more photo-realistic images for both computer vision algorithms and human-examiners to examine. To this end, this paper proposes a Generative Adversarial Network-based Visible Face Synthesis (GAN-VFS) method to synthesize more photo-realistic visible face images from their corresponding polarimetric images. To ensure that the encoded visible-features contain more semantically meaningful information in reconstructing the visible face image, a guidance sub-network is involved into the training procedure. To achieve photo realistic property while preserving discriminative characteristics for the reconstructed outputs, an identity loss combined with the perceptual loss are optimized in the framework. Multiple experiments evaluated on different experimental protocols demonstrate that the proposed method achieves state-of-the-art performance. He Zhang 0004, Vishal M. Patel, Benjamin S. Riggan, Shuowen Hu |
IJCB | 1 |
| 2017 | Convolutional Sparse and Low-Rank Coding-Based Rain Streak RemovalabstractWe propose a novel Convolutional Coding-based Rain Removal (CCRR) algorithm for automatically removing rain streaks from a single rainy image. Our method first learns a set of generic sparsity-based and low-rank representation-based convolutional filters for efficiently representing background clear image and rain streaks, respectively. To this end, we first develop a new method for learning a set of convolutional low-rank filters. Then, using these learned filter, we propose an optimization problem to decompose a rainy image into a clear background image and a rain streak image. By working directly on the whole image, the proposed rain streak removal algorithm does not need to divide the image into overlapping patches for leaning local dictionaries. Extensive experiments on synthetic and real images show that the proposed method performs favorably compared to state-of-the-art rain streak removal algorithms. He Zhang 0004, Vishal M. Patel |
WACV | 1 |
| 2017 | Sparse Representation-Based Open Set RecognitionabstractWe propose a generalized Sparse Representation-based Classification (SRC) algorithm for open set recognition where not all classes presented during testing are known during training. The SRC algorithm uses class reconstruction errors for classification. As most of the discriminative information for open set recognition is hidden in the tail part of the matched and sum of non-matched reconstruction error distributions, we model the tail of those two error distributions using the statistical Extreme Value Theory (EVT). Then we simplify the open set recognition problem into a set of hypothesis testing problems. The confidence scores corresponding to the tail distributions of a novel test sample are then fused to determine its identity. The effectiveness of the proposed method is demonstrated using four publicly available image and object classification datasets and it is shown that this method can perform significantly better than many competitive open set recognition algorithms. He Zhang 0004, Vishal M. Patel |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2017 | SAR Image Despeckling Using a Convolutional Neural NetworkabstractSynthetic aperture radar (SAR) images are often contaminated by a multiplicative noise known as speckle. Speckle makes the processing and interpretation of SAR images difficult. We propose a deep-learning-based approach called, image despeckling convolutional neural network (ID-CNN), for automatically removing speckle from the input noisy images. In particular, ID-CNN uses a set of convolutional layers along with batch normalization and rectified linear unit activation function and a componentwise division residual layer to estimate speckle and it is trained in an end-to-end fashion using a combination of Euclidean loss and total variation loss. Extensive experiments on synthetic and real SAR images show that the proposed method achieves significant improvements over the state-of-the-art speckle reduction methods. Puyang Wang, He Zhang 0004, Vishal M. Patel |
IEEE Signal Process. Lett. | 2 |
| 2016 | Convolutional Sparse Coding-based Image Decomposition
He Zhang 0004, Vishal M. Patel |
BMVC | 1 |