VLDB 2026 Research / reviewers in the wild / expert
Jianming Zhang 0001
dblp:12/3933-1
· DBLP profile ↗
100ranked-venue papers
10as first author
54since 2021 · last 2025
0000-0002-9954-6294ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 86 · 9 first-author · 48 since 2021Graphics, computer vision, multimedia, augmented reality and games · 77 · 7 first-author · 41 since 2021Human-computer interaction and ubiquitous computing · 2Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | UniReal: Universal Image Generation and Editing via Learning Real-world DynamicsabstractWe introduce UniReal, a unified framework designed to address various image generation and editing tasks. Existing solutions often vary by tasks, yet share fundamental principles: preserving consistency between inputs and outputs while capturing visual variations. Inspired by recent video generation models that effectively balance consistency and variation across frames, we propose a unifying approach that treats image-level tasks as discontinuous video generation. Specifically, we treat varying numbers of input and output images as frames, enabling seamless support for tasks such as image generation, editing, customization, composition, etc. Although designed for image-level tasks, we leverage videos as a scalable source for universal supervision. UniReal learns world dynamics from large-scale videos, demonstrating advanced capability in handling shadows, reflections, pose variation, and object interaction, while also exhibiting emergent capability for novel applications. Xi Chen 0119, He Zhang 0004, Yuqian Zhou, Soo Ye Kim, Qing Liu 0017, Yijun Li 0001, Jianming Zhang 0001, Nanxuan Zhao, Yilin Wang 0002, Zhe Lin 0001, Hengshuang Zhao |
CVPR | 8 |
| 2025 | FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any GranularityabstractThe advent of large Vision-Language Models (VLMs) has significantly advanced multimodal tasks, enabling more sophisticated and accurate reasoning across various applications, including image and video captioning, visual question answering, and cross-modal retrieval. Despite their superior capabilities, VLMs struggle with fine-grained image regional composition information perception. Specifically, they have difficulty accurately aligning the segmentation masks with the corresponding semantics and precisely describing the compositional aspects of the referred regions. However, compositionality – the ability to understand and generate novel combinations of known visual and textual components – is critical for facilitating coherent reasoning and understanding across modalities by VLMs. To address this issue, we propose FineCaption, a novel VLM that can recognize arbitrary masks as referential inputs and process high-resolution images for compositional image captioning at different granularity levels. To support this endeavor, we introduce CompositionCap, a new dataset for multi-grained region compositional image captioning, which introduces the task of compositional attribute-aware regional image captioning. Empirical results demonstrate the effectiveness of our proposed model compared to other state-of-the-art VLMs. Additionally, we analyze the capabilities of current VLMs in recognizing various visual prompts for compositional region image captioning, highlighting areas for improvement in VLM design and training. https://hanghuacs.github.io/FineCaption/ Hang Hua, Qing Liu 0017, Lingzhi Zhang, Jing Shi 0005, Soo Ye Kim, Yilin Wang 0002, Jianming Zhang 0001, Zhe Lin 0001, Jiebo Luo 0001 |
CVPR | 8 |
| 2025 | MetaShadow: Object-Centered Shadow Detection, Removal, and SynthesisabstractShadows are often under-considered or even ignored in image editing applications, limiting the realism of the edited results. In this paper, we introduce MetaShadow, a three-in-one versatile framework that enables detection, removal, and controllable synthesis of shadows in natural images in an object-centered fashion. MetaShadow combines the strengths of two cooperative components: Shadow Analyzer, for object-centered shadow detection and removal, and Shadow Synthesizer, for reference-based controllable shadow synthesis. Notably, we optimize the learning of the intermediate features from Shadow Analyzer to guide Shadow Synthesizer to generate more realistic shadows that blend seamlessly with the scene. Extensive evaluations on multiple shadow benchmark datasets show significant improvements of MetaShadow over the existing state-of-the-art methods on object-centered shadow detection, removal, and synthesis. MetaShadow excels in image-editing tasks such as object removal, relocation, and insertion, pushing the boundaries of object-centered image editing. Tianyu Wang 0003, Jianming Zhang 0001, Haitian Zheng, Zhihong Ding, Scott Cohen, Zhe Lin 0001, Wei Xiong 0008, Chi-Wing Fu, Luis Figueroa, Soo Ye Kim |
CVPR | 2 |
| 2025 | Generative Image Layer Decomposition with Visual EffectsabstractRecent advancements in large generative models, particularly diffusion-based methods, have significantly enhanced the capabilities of image editing. However, achieving precise control over image composition tasks remains a challenge. Layered representations, which allow for independent editing of image components, are essential for user-driven content creation, yet existing approaches often struggle to decompose an image into plausible layers with accurately retained transparent visual effects such as shadows and reflections. We propose LayerDecomp, a generative framework for image layer decomposition which outputs photorealistic clean backgrounds and high-quality transparent foregrounds with faithfully preserved visual effects. To enable effective training, we first introduce a dataset preparation pipeline that automatically scales up simulated multi-layer data with synthesized visual effects. To further enhance real-world applicability, we supplement this simulated dataset with camera-captured images containing natural visual effects. Additionally, we propose a consistency loss which enforces the model to learn accurate representations for the transparent foreground layer when ground-truth annotations are not available. Our method achieves superior quality in layer decomposition, outperforming existing approaches in object removal and spatial editing tasks across several benchmarks and multiple user studies, unlocking various creative possibilities for layer-wise image editing. Jinrui Yang, Qing Liu 0017, Yijun Li 0001, Soo Ye Kim, Daniil Pakhomov, Mengwei Ren, Jianming Zhang 0001, Zhe Lin 0001, Cihang Xie, Yuyin Zhou |
CVPR | 7 |
| 2025 | Baking Gaussian Splatting Into Diffusion Denoiser for Fast and Scalable Single-Stage Image-to-3D Generation and Reconstruction
Yuanhao Cai, He Zhang 0004, Kai Zhang 0045, Yixun Liang, Mengwei Ren, Fujun Luan, Qing Liu 0017, Soo Ye Kim, Jianming Zhang 0001, Yuqian Zhou, Yulun Zhang 0001, Xiaokang Yang 0001, Zhe Lin 0001, Alan L. Yuille |
ICCV | 9 |
| 2025 | Refine-by-Align: Reference-Guided Artifacts Refinement through Semantic AlignmentabstractPersonalized image generation has emerged from the recent advancements in generative models. However, these generated personalized images often suffer from localized artifacts such as incorrect logos, reducing fidelity and fine-grained identity details of the generated results. Furthermore, there is little prior work tackling this problem. To help improve these identity details in the personalized image generation, we introduce a new task: reference-guided artifacts refinement. We present Refine-by-Align, a first-of-its-kind model that employs a diffusion-based framework to address this challenge. Our model consists of two stages: Alignment Stage and Refinement Stage, which share weights of a unified neural network model. Given a generated image, a masked artifact region, and a reference image, the alignment stage identifies and extracts the corresponding regional features in the reference, which are then used by the refinement stage to fix the artifacts. Our model-agnostic pipeline requires no test-time tuning or optimization. It automatically enhances image fidelity and reference identity in the generated image, generalizing well to existing models on various tasks including but not limited to customization, generative compositing, view synthesis, and virtual try-on. Extensive experiments and comparisons demonstrate that our pipeline greatly pushes the boundary of fine details in the image synthesis models. Soo Ye Kim, He Zhang 0004, Wei Xiong 0008, Zhe Lin 0001, Brian L. Price, Scott Cohen, Jianming Zhang 0001, Daniel G. Aliaga |
ICLR | 10 |
| 2025 | BokehMe++: Harmonious Fusion of Classical and Neural Rendering for Versatile Bokeh CreationabstractDespite significant advancements in simulating the bokeh effect of Digital Single Lens Reflex Camera (DSLR) from an all-in-focus image, challenges remain in processing highlight points, preserving boundary details for in-focus objects and processing high-resolution images efficiently. To tackle these issues, we first develop a ray-tracing-based bokeh simulator. An innovative pipeline with weight redistribution is introduced to handle highlight rendering. By considering the front length of lens barrel, we can simulate realistic cat-eye effect. This bokeh simulator serves as the foundation for creating our training dataset. Building on this dataset, we introduce a hybrid framework BokehMe++, combining a classical renderer and a neural renderer. The classical renderer is implemented by a hierarchical scattering-based method, which suffers from boundary inaccuracies. These erroneous areas will be identified by an error map generator and be corrected by a two-stage neural renderer. Adaptive resizing and iterative upsampling are introduced in the neural renderer to process arbitrary blur size efficiently. Extensive experiments demonstrate that BokehMe++ outperforms existing methods and provides highly customizable rendering features, such as adjustable blur amount, focal plane, highlight mode and cat-eye effect. Furthermore, BokehMe++ can maintain the sharpness of hair details in portraits through an auxiliary alpha map input. Juewen Peng, Zhiguo Cao 0001, Xianrui Luo, Ke Xian, Wenfeng Tang, Jianming Zhang 0001, Guosheng Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | NVDS$^{\mathbf{+}}$+: Towards Efficient and Versatile Neural Stabilizer for Video Depth EstimationabstractVideo depth estimation aims to infer temporally consistent depth. One approach is to finetune a single-image model on each video with geometry constraints, which proves inefficient and lacks robustness. An alternative is learning to enforce consistency from data, which requires well-designed models and sufficient video depth data. To address both challenges, we introduce NVDS that stabilizes inconsistent depth estimated by various single-image models in a plug-and-play manner. We also elaborate a large-scale Video Depth in the Wild (VDW) dataset, which contains 14,203 videos with over two million frames, making it the largest natural-scene video depth dataset. Additionally, a bidirectional inference strategy is designed to improve consistency by adaptively fusing forward and backward predictions. We instantiate a model family ranging from small to large scales for different applications. The method is evaluated on VDW dataset and three public benchmarks. To further prove the versatility, we extend NVDS to video semantic segmentation and several downstream applications like bokeh rendering, novel view synthesis, and 3D reconstruction. Experimental results show that our method achieves significant improvements in consistency, accuracy, and efficiency. Our work serves as a solid baseline and data foundation for learning-based video depth estimation. Yiran Wang 0005, Min Shi 0004, Jiaqi Li 0007, Chaoyi Hong, Zihao Huang 0001, Juewen Peng, Zhiguo Cao 0001, Jianming Zhang 0001, Ke Xian, Guosheng Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2024 | Amodal Scene Analysis via Holistic Occlusion Relation Inference and Generative Mask CompletionabstractAmodal scene analysis entails interpreting the occlusion relationship among scene elements and inferring the possible shapes of the invisible parts. Existing methods typically frame this task as an extended instance segmentation or a pair-wise object de-occlusion problem. In this work, we propose a new framework, which comprises a Holistic Occlusion Relation Inference (HORI) module followed by an instance-level Generative Mask Completion (GMC) module. Unlike previous approaches, which rely on mask completion results for occlusion reasoning, our HORI module directly predicts an occlusion relation matrix in a single pass. This approach is much more efficient than the pair-wise de-occlusion process and it naturally handles mutual occlusion, a common but often neglected situation. Moreover, we formulate the mask completion task as a generative process and use a diffusion-based GMC module for instance-level mask completion. This improves mask completion quality and provides multiple plausible solutions. We further introduce a large-scale amodal segmentation dataset with high-quality human annotations, including mutual occlusions. Experiments on our dataset and two public benchmarks demonstrate the advantages of our method. code public available at https://github.com/zbwxp/Amodal-AAAI. Bowen Zhang 0009, Qing Liu 0017, Jianming Zhang 0001, Yilin Wang 0002, Liyang Liu, Zhe Lin 0001, Yifan Liu 0001 |
AAAI | 3 |
| 2024 | UniHuman: A Unified Model For Editing Human Images in the WildabstractHuman image editing includes tasks like changing a person's pose, their clothing, or editing the image according to a text prompt. However, prior work often tackles these tasks separately, overlooking the benefit of mutual reinforcement from learning them jointly. In this paper, we propose UniHuman, a unified model that addresses multiple facets of human image editing in real-world settings. To enhance the model's generation quality and generalization capacity, we leverage guidance from human visual encoders and introduce a lightweight pose-warping module that can exploit different pose representations, accommodating unseen textures and patterns. Furthermore, to bridge the disparity between existing human editing benchmarks with real-world data, we curated 400K high-quality human image-text pairs for training and collected 2K human images for out-of-domain testing, both encompassing diverse clothing styles, backgrounds, and age groups. Experiments on both in-domain and out-of-domain test sets demonstrate that UniHuman outperforms task-specific models by a significant margin. In user studies, UniHuman is preferred by the users in an average of 77% of cases. Our project is available at this link. Nannan Li 0004, Qing Liu 0017, Krishna Kumar Singh, Yilin Wang 0002, Jianming Zhang 0001, Bryan A. Plummer, Zhe Lin 0001 |
CVPR | 5 |
| 2024 | Holo-Relighting: Controllable Volumetric Portrait Relighting from a Single ImageabstractAt the core of portrait photography is the search for ideal lighting and viewpoint. The process often requires advanced knowledge in photography and an elaborate studio setup. In this work, we propose Holo-Relighting, a volumetric relighting method that is capable of synthesizing novel viewpoints, and novel lighting from a single image. Holo-Relighting leverages the pretrained 3D GAN (EG3D) to reconstruct geometry and appearance from an input portrait as a set of 3D-aware features. We design a relighting module conditioned on a given lighting to process these features, and predict a relit 3D representation in the form of a tri-plane, which can render to an arbitrary viewpoint through volume rendering. Besides viewpoint and lighting control, Holo-Relighting also takes the head pose as a condition to enable head-pose-dependent lighting effects. With these novel designs, Holo-Relighting can generate complex non-Lambertian lighting effects (e.g., specular highlights and cast shadows) without using any explicit physical lighting priors. We train Holo-Relighting with data captured with a light stage, and propose two data-rendering techniques to improve the data quality for training the volumetric relighting system. Through quantitative and qualitative experiments, we demonstrate Holo-Relighting can achieve state-of-the-arts relighting quality with better photorealism, 3D consistency and controllability. Yiqun Mei, Yu Zeng 0001, He Zhang 0004, Zhixin Shu, Xuaner Cecilia Zhang, Sai Bi, Jianming Zhang 0001, Hyunjoon Jung, Vishal M. Patel |
CVPR | 7 |
| 2024 | Relightful Harmonization: Lighting-Aware Portrait Background ReplacementabstractPortrait harmonization aims to composite a subject into a new background, adjusting its lighting and color to ensure harmony with the background scene. Existing harmo-nization techniques often only focus on adjusting the global color and brightness of the foreground and ignore crucial illumination cues from the background such as apparent lighting direction, leading to unrealistic compositions. We introduce Relightful Harmonization, a lighting-aware diffusion model designed to seamlessly harmonize sophisticated lighting effect for the foreground portrait using any back-ground image. Our approach unfolds in three stages. First, we introduce a lighting representation module that allows our diffusion model to encode lighting information from target image background. Second, we introduce an alignment network that aligns lighting features learned from image background with lighting features learned from panorama environment maps, which is a complete representation for scene illumination. Last, to further boost the photorealism of the proposed method, we introduce a novel data simulation pipeline that generates synthetic training pairs from a diverse range of natural images, which are used to refine the model. Our method outperforms existing benchmarks in visual fidelity and lighting coherence, showing superior generalization in real-world testing scenarios, highlighting its versatility and practicality. Mengwei Ren, Wei Xiong 0008, Jae Shin Yoon, Zhixin Shu, Jianming Zhang 0001, Hyunjoon Jung, Guido Gerig, He Zhang 0004 |
CVPR | 5 |
| 2024 | IMPRINT: Generative Object Compositing by Learning Identity-Preserving RepresentationabstractGenerative object compositing emerges as a promising new avenue for compositional image editing. However, the requirement of object identity preservation poses a significant challenge, limiting practical usage of most existing methods. In response, this paper introduces IMPRINT, a novel diffusion-based generative model trained with a two-stage learning framework that decouples learning of identity preservation from that of compositing. The first stage is targeted for context-agnostic, identity-preserving pretraining of the object encoder, enabling the encoder to learn an embedding that is both view-invariant and conducive to enhanced detail preservation. The subsequent stage leverages this representation to learn seamless harmonization of the object composited to the background. In addition, IMPRINT incorporates a shape-guidance mechanism offering user-directed control over the compositing process. Extensive experiments demonstrate that IMPRINT significantly outperforms existing methods and various baselines on identity preservation and composition quality. Project page: https://song630.github.io/IMPRINT-Project-Page/ Zhe Lin 0001, Scott Cohen, Brian L. Price, Jianming Zhang 0001, Soo Ye Kim, He Zhang 0004, Wei Xiong 0008, Daniel G. Aliaga |
CVPR | 6 |
| 2024 | SwapAnything: Enabling Arbitrary Object Swapping in Personalized Image Editing
Nanxuan Zhao, Wei Xiong 0008, Qing Liu 0017, He Zhang 0004, Jianming Zhang 0001, Hyunjoon Jung, Yilin Wang 0002, Xin Wang 0061 |
ECCV (32) | 7 |
| 2024 | Fast View Synthesis of Casual Videos with Soup-of-Planes
Yao-Chih Lee, Zhoutong Zhang, Kevin Matzen, Simon Niklaus, Jianming Zhang 0001, Jia-Bin Huang 0001, Feng Liu 0015 |
ECCV (38) | 5 |
| 2024 | DreamMover: Leveraging the Prior of Diffusion Models for Image Interpolation with Large Motion
Liao Shen, Tianqi Liu 0003, Huiqiang Sun, Baopu Li, Jianming Zhang 0001, Zhiguo Cao 0001 |
ECCV (15) | 6 |
| 2024 | Thinking Outside the BBox: Unconstrained Generative Object Compositing
Gemma Canet Tarrés, Zhe Lin 0001, Jianming Zhang 0001, Dan Ruta, Andrew Gilbert, John P. Collomosse, Soo Ye Kim |
ECCV (62) | 4 |
| 2024 | Self-Distilled Depth Refinement with Noisy Poisson FusionabstractDepth refinement aims to infer high-resolution depth with fine-grained edges and details, refining low-resolution results of depth estimation models. The prevailing methods adopt tile-based manners by merging numerous patches, which lacks efficiency and produces inconsistency. Besides, prior arts suffer from fuzzy depth boundaries and limited generalizability. Analyzing the fundamental reasons for these limitations, we model depth refinement as a noisy Poisson fusion problem with local inconsistency and edge deformation noises. We propose the Self-distilled Depth Refinement (SDDR) framework to enforce robustness against the noises, which mainly consists of depth edge representation and edge-based guidance. With noisy depth predictions as input, SDDR generates low-noise depth edge representations as pseudo-labels by coarse-to-fine self-distillation. Edge-based guidance with edge-guided gradient loss and edge-based fusion loss serves as the optimization objective equivalent to Poisson fusion. When depth maps are better refined, the labels also become more noise-free. Our model can acquire strong robustness to the noises, achieving significant improvements in accuracy, edge quality, efficiency, and generalizability on five different benchmarks. Moreover, directly training another model with edge labels produced by SDDR brings improvements, suggesting that our method could help with training robust refinement models in future works. Jiaqi Li 0007, Yiran Wang 0005, Jinghong Zheng 0002, Zihao Huang 0001, Ke Xian, Zhiguo Cao 0001, Jianming Zhang 0001 |
NeurIPS | 7 |
| 2024 | Structure-Guided Image Completion With Image-Level and Object-Level Semantic DiscriminatorsabstractStructure-guided image completion aims to inpaint a local region of an image according to an input guidance map from users. While such a task enables many practical applications for interactive editing, existing methods often struggle to hallucinate realistic object instances in complex natural scenes. Such a limitation is partially due to the lack of semantic-level constraints inside the hole region as well as the lack of a mechanism to enforce realistic object generation. In this work, we propose a learning paradigm that consists of semantic discriminators and object-level discriminators for improving the generation of complex semantics and objects. Specifically, the semantic discriminators leverage pretrained visual features to improve the realism of the generated visual concepts. Moreover, the object-level discriminators take aligned instances as inputs to enforce the realism of individual objects. Our proposed scheme significantly improves the generation quality and achieves state-of-the-art results on various tasks, including segmentation-guided completion, edge-guided manipulation and panoptically-guided manipulation on Places2 datasets. Furthermore, our trained model is flexible and can support multiple editing use cases, such as object insertion, replacement, removal and standard inpainting. In particular, our trained model combined with a novel automatic image completion pipeline achieves state-of-the-art results on the standard inpainting task. Haitian Zheng, Zhe Lin 0001, Jingwan Lu, Scott Cohen, Eli Shechtman, Connelly Barnes, Jianming Zhang 0001, Qing Liu 0017, Sohrab Amirghodsi, Yuqian Zhou, Jiebo Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | ViTA: Video Transformer Adaptor for Robust Video Depth EstimationabstractDepth information plays a pivotal role in numerous computer vision applications, including autonomous driving, 3D reconstruction, and 3D content generation. When deploying depth estimation models in practical applications, it is essential to ensure that the models have strong generalization capabilities. However, existing depth estimation methods primarily concentrate on robust single-image depth estimation, leading to the occurrence of flickering artifacts when applied to video inputs. On the other hand, video depth estimation methods either consume excessive computational resources or lack robustness. To address the above issues, we propose ViTA, a video transformer adaptor, to estimate temporally consistent video depth in the wild. In particular, we leverage a pre-trained image transformer (i.e., DPT) and introduce additional temporal embeddings in the transformer blocks. Such designs enable our ViTA to output reliable results given an unconstrained video. Besides, we present a spatio-temporal consistency loss for supervision. The spatial loss computes the per-pixel discrepancy between the prediction and the ground truth in space, while the temporal loss regularizes the inconsistent outputs of the same point in consecutive frames. To find the correspondences between consecutive frames, we design a bi-directional warping strategy based on the forward and backward optical flow. During inference, our ViTA no longer requires optical flow estimation, which enables it to estimate spatially accurate and temporally consistent video depth maps with fine-grained details in real time. We conduct a detailed ablation study to verify the effectiveness of the proposed components. Extensive experiments on the zero-shot cross-dataset evaluation demonstrate that the proposed method is superior to previous methods. Ke Xian, Juewen Peng, Zhiguo Cao 0001, Jianming Zhang 0001, Guosheng Lin |
IEEE Trans. Multim. | 4 |
| 2023 | Perspective Fields for Single Image Camera CalibrationabstractGeometric camera calibration is often required for applications that understand the perspective of the image. We propose Perspective Fields as a representation that models the local perspective properties of an image. Perspective Fields contain per-pixel information about the camera view, parameterized as an Up-vector and a Latitude value. This representation has a number of advantages; it makes minimal assumptions about the camera model and is invariant or equivariant to common image editing operations like cropping, warping, and rotation. It is also more interpretable and aligned with human perception. We train a neural network to predict Perspective Fields and the predicted Perspective Fields can be converted to calibration parameters easily. We demonstrate the robustness of our approach under various scenarios compared with camera calibration-based methods and show example applications in image compositing. Project page: https://jinlinyi.github.io/PerspectiveFields/. Linyi Jin, Jianming Zhang 0001, Yannick Hold-Geoffroy, Oliver Wang, Kevin Matzen, Matthew Sticha, David F. Fouhey |
CVPR | 2 |
| 2023 | Single View Scene Scale Estimation using Scale FieldabstractIn this paper, we propose a single image scale estimation method based on a novel scale field representation. A scale field defines the local pixel-to-metric conversion ratio along the gravity direction on all the ground pixels. This representation resolves the ambiguity in camera parameters, allowing us to use a simple yet effective way to collect scale annotations on arbitrary images from human annotators. By training our model on calibrated panoramic image data and the in-the-wild human annotated data, our single image scene scale estimation network generates robust scale field on a variety of image, which can be utilized in various 3D understanding and scale-aware image editing applications. Byeong-Uk Lee, Jianming Zhang 0001, Yannick Hold-Geoffroy, In-So Kweon |
CVPR | 2 |
| 2023 | 3D Cinemagraphy from a Single ImageabstractWe present 3D Cinemagraphy, a new technique that mar-ries 2D image animation with 3D photography. Given a single still image as input, our goal is to generate a video that contains both visual content animation and camera motion. We empirically find that naively combining existing 2D image animation and 3D photography methods leads to obvious artifacts or inconsistent animation. Our key insight is that representing and animating the scene in 3D space offers a natural solution to this task. To this end, we first convert the input image into feature-based layered depth images using predicted depth values, followed by unprojecting them to a feature point cloud. To animate the scene, we perform motion estimation and lift the 2D motion into the 3D scene flow. Finally, to resolve the problem of hole emer-gence as points move forward, we propose to bidirectionally displace the point cloud as per the scene flow and synthe-size novel views by separately projecting them into target image planes and blending the results. Extensive experiments demonstrate the effectiveness of our method. A user study is also conducted to validate the compelling rendering results of our method. Xingyi Li 0005, Zhiguo Cao 0001, Huiqiang Sun, Jianming Zhang 0001, Ke Xian, Guosheng Lin |
CVPR | 4 |
| 2023 | LightPainter: Interactive Portrait Relighting with Freehand ScribbleabstractRecent portrait relighting methods have achieved realistic results of portrait lighting effects given a desired lighting representation such as an environment map. However, these methods are not intuitive for user interaction and lack precise lighting control. We introduce LightPainter, a scribble-based relighting system that allows users to interactively manipulate portrait lighting effect with ease. This is achieved by two conditional neural networks, a delighting module that recovers geometry and albedo optionally conditioned on skin tone, and a scribble-based module for re-lighting. To train the relighting module, we propose a novel scribble simulation procedure to mimic real user scribbles, which allows our pipeline to be trained without any human annotations. We demonstrate high-quality and flexible portrait lighting editing capability with both quantitative and qualitative experiments. User study comparisons with commercial lighting editing tools also demonstrate consistent user preference for our method. Yiqun Mei, He Zhang 0004, Xuaner Cecilia Zhang, Jianming Zhang 0001, Zhixin Shu, Yilin Wang 0002, Zijun Wei, Hyunjoon Jung, Vishal M. Patel |
CVPR | 4 |
| 2023 | PixHt-Lab: Pixel Height Based Light Effect Generation for Image CompositingabstractLighting effects such as shadows or reflections are key in making synthetic images realistic and visually appealing. To generate such effects, traditional computer graphics uses a physically-based renderer along with 3D geometry. To compensate for the lack of geometry in 2D Image compositing, recent deep learning-based approaches introduced a pixel height representation to generate soft shadows and reflections. However, the lack of geometry limits the quality of the generated soft shadows and constrains reflections to pure specular ones. We introduce PixHt-Lab, a system leveraging an explicit mapping from pixel height representation to 3D space. Using this mapping, PixHt- Lab reconstructs both the cutout and background geometry and renders realistic, diverse lighting effects for image compositing. Given a surface with physically-based materials, we can render reflections with varying glossiness. To generate more realistic soft shadows, we further propose using 3D-aware buffer channels to guide a neural renderer. Both quantitative and qualitative evaluations demonstrate that PixHt-Lab significantly improves soft shadow generation. Project: https://shengcn.github.io/PixHtLab/ Yichen Sheng, Jianming Zhang 0001, Julien Philip, Yannick Hold-Geoffroy, Xin Sun 0014, He Zhang 0004, Lu Ling, Bedrich Benes |
CVPR | 2 |
| 2023 | ObjectStitch: Object Compositing with Diffusion ModelabstractObject compositing based on 2D images is a challenging problem since it typically involves multiple processing stages such as color harmonization, geometry correction and shadow generation to generate realistic results. Furthermore, annotating training data pairs for compositing requires substantial manual effort from professionals, and is hardly scalable. Thus, with the recent advances in generative models, in this work, we propose a selfsupervised framework for object compositing by leveraging the power of conditional diffusion models. Our framework can hollistically address the object compositing task in a unified model, transforming the viewpoint, geometry, color and shadow of the generated object while requiring no manual labeling. To preserve the input object's characteristics, we introduce a content adaptor that helps to maintain categori-cal semantics and object appearance. A data augmentation method is further adopted to improve the fidelity of the generator. Our method outperforms relevant baselines in both realism and faithfulness of the synthesized result images in a user study on various real-world images. Zhe Lin 0001, Scott Cohen, Brian L. Price, Jianming Zhang 0001, Soo Ye Kim, Daniel G. Aliaga |
CVPR | 6 |
| 2023 | SceneComposer: Any-Level Semantic Image SynthesisabstractWe propose a new framework for conditional image synthesis from semantic layouts of any precision levels, ranging from pure text to a 2D semantic canvas with precise shapes. More specifically, the input layout consists of one or more semantic regions with free-form text descriptions and adjustable precision levels, which can be set based on the desired controllability. The framework naturally reduces to text-to-image (T2I) at the lowest level with no shape information, and it becomes segmentation-to-image (S2I) at the highest level. By supporting the levels in-between, our framework is flexible in assisting users of different drawing expertise and at different stages of their creative workflow. We introduce several novel techniques to address the challenges coming with this new setup, including a pipeline for collecting training data; a precision-encoded mask pyramid and a text feature map representation to jointly encode precision level, semantics, and composition information; and a multi-scale guided diffusion model to synthesize images. To evaluate the proposed method, we collect a test dataset containing user-drawn layouts with diverse scenes and styles. Experimental results show that the proposed method can generate high-quality images following the layout at given precision, and compares favorably against existing methods. Project page https://zengxianyu.github.io/scenec/ Yu Zeng 0001, Zhe Lin 0001, Jianming Zhang 0001, Qing Liu 0017, John P. Collomosse, Jason Kuen, Vishal M. Patel |
CVPR | 3 |
| 2023 | Neural Video Depth StabilizerabstractVideo depth estimation aims to infer temporally consistent depth. Some methods achieve temporal consistency by finetuning a single-image depth model during test time using geometry and re-projection constraints, which is inefficient and not robust. An alternative approach is to learn how to enforce temporal consistency from data, but this requires well-designed models and sufficient video depth data. To address these challenges, we propose a plug-and-play framework called Neural Video Depth Stabilizer (NVDS) that stabilizes inconsistent depth estimations and can be applied to different single-image depth models without extra effort. We also introduce a large-scale dataset, Video Depth in the Wild (VDW), which consists of 14,203 videos with over two million frames, making it the largest natural-scene video depth dataset to our knowledge. We evaluate our method on the VDW dataset as well as two public benchmarks and demonstrate significant improvements in consistency, accuracy, and efficiency compared to previous approaches. Our work serves as a solid baseline and provides a data foundation for learning-based video depth models. We will release our dataset and code for future research. Yiran Wang 0005, Min Shi 0004, Jiaqi Li 0007, Zihao Huang 0001, Zhiguo Cao 0001, Jianming Zhang 0001, Ke Xian, Guosheng Lin |
ICCV | 6 |
| 2023 | Lens Parameter Estimation for Realistic Depth of Field ModelingabstractWe present a method to estimate the depth of field effect from a single image. Most existing methods related to this task provide either a per-pixel estimation of blur and/or depth. Instead, we go further and propose to use a lens-based representation that models the depth of field using two parameters: the blur factor and focus disparity. Those two parameters, along with the signed defocus representation, result in a more intuitive and linear representation which we solve using a novel weighting network. Furthermore, our method explicitly enforces consistency between the estimated defocus blur, the lens parameters, and the depth map. Finally, we train our deep-learning-based model on a mix of real images with synthetic depth of field and fully synthetic images. These improvements result in a more robust and accurate method, as demonstrated by our state-of-the-art results. In particular, our lens parametrization enables several applications, such as 3D staging for AR environments and seamless object compositing. Dominique Piché-Meunier, Yannick Hold-Geoffroy, Jianming Zhang 0001, Jean-François Lalonde |
ICCV | 3 |
| 2023 | GAIT: Generating Aesthetic Indoor Tours with Deep Reinforcement LearningabstractPlacing and orienting a camera to compose aesthetically meaningful shots of a scene is not only a key objective in real-world photography and cinematography but also for virtual content creation. The framing of a camera often significantly contributes to the story telling in movies, games, and mixed reality applications. Generating single camera poses or even contiguous trajectories either requires a significant amount of manual labor or requires solving high-dimensional optimization problems, which can be computationally demanding and error-prone. In this paper, we introduce GAIT, a framework for training a Deep Reinforcement Learning (DRL) agent, that learns to automatically control a camera to generate a sequence of aesthetically meaningful views for synthetic 3D indoor scenes. To generate sequences of frames with high aesthetic value, GAIT relies on a neural aesthetics estimator, which is trained on a crowed-sourced dataset. Additionally, we introduce regularization techniques for diversity and smoothness to generate visually interesting trajectories for a 3D environment, and to constrain agent acceleration in the reward function to generate a smooth sequence of camera frames. We validated our method by comparing it to baseline algorithms, based on a perceptual user study, and through ablation studies. Code and visual results are available on the project website: https://desaixie.github.io/gait-rl Desai Xie, Ping Hu 0003, Xin Sun 0014, Sören Pirk, Jianming Zhang 0001, Radomír Mech, Arie E. Kaufman |
ICCV | 5 |
| 2023 | Interactive Portrait Harmonization
Jeya Maria Jose Valanarasu, He Zhang 0004, Jianming Zhang 0001, Yilin Wang 0002, Zhe Lin 0001, Jose Echevarria, Yinglan Ma, Zijun Wei, Kalyan Sunkavalli, Vishal M. Patel |
ICLR | 3 |
| 2023 | Diffusion-Augmented Depth Prediction with Sparse AnnotationsabstractDepth estimation aims to predict dense depth maps. In autonomous driving scenes, sparsity of annotations makes the task challenging. Supervised models produce concave objects due to insufficient structural information. They overfit to valid pixels and fail to restore spatial structures. Self-supervised methods are proposed for the problem. Their robustness is limited by pose estimation, leading to erroneous results in natural scenes. In this paper, we propose a supervised framework termed Diffusion-Augmented Depth Prediction (DADP). We leverage the structural characteristics of diffusion model to enforce depth structures of depth models in a plug-and-play manner. An object-guided integrality loss is also proposed to further enhance regional structure integrality by fetching objective information. We evaluate DADP on three driving benchmarks and achieve significant improvements in depth structures and robustness. Our work provides a new perspective on depth estimation with sparse annotations in autonomous driving scenes. Jiaqi Li 0007, Yiran Wang 0005, Zihao Huang 0001, Jinghong Zheng 0002, Ke Xian, Zhiguo Cao 0001, Jianming Zhang 0001 |
ACM Multimedia | 7 |
| 2023 | PHOTOSWAP: Personalized Subject Swapping in ImagesabstractIn an era where images and visual content dominate our digital landscape, the ability to manipulate and personalize these images has become a necessity.
Envision seamlessly substituting a tabby cat lounging on a sunlit window sill in a photograph with your own playful puppy, all while preserving the original charm and composition of the image.
We present \emph{Photoswap}, a novel approach that enables this immersive image editing experience through personalized subject swapping in existing images.
\emph{Photoswap} first learns the visual concept of the subject from reference images and then swaps it into the target image using pre-trained diffusion models in a training-free manner. We establish that a well-conceptualized visual subject can be seamlessly transferred to any image with appropriate self-attention and cross-attention manipulation, maintaining the pose of the swapped subject and the overall coherence of the image.
Comprehensive experiments underscore the efficacy and controllability of \emph{Photoswap} in personalized subject swapping. Furthermore, \emph{Photoswap} significantly outperforms baseline methods in human ratings across subject swapping, background preservation, and overall quality, revealing its vast application potential, from entertainment to professional editing. Yilin Wang 0002, Nanxuan Zhao, Tsu-Jui Fu, Wei Xiong 0008, Qing Liu 0017, He Zhang 0004, Jianming Zhang 0001, Hyunjoon Jung, Xin Wang 0061 |
NeurIPS | 9 |
| 2023 | Towards Accurate Reconstruction of 3D Scene Shape From A Single Monocular ImageabstractDespite significant progress made in the past few years, challenges remain for depth estimation using a single monocular image. First, it is nontrivial to train a metric-depth prediction model that can generalize well to diverse scenes mainly due to limited training data. Thus, researchers have built large-scale relative depth datasets that are much easier to collect. However, existing relative depth estimation models often fail to recover accurate 3D scene shapes due to the unknown depth shift caused by training with the relative depth data. We tackle this problem here and attempt to estimate accurate scene shapes by training on large-scale relative depth data, and estimating the depth shift. To do so, we propose a two-stage framework that first predicts depth up to an unknown scale and shift from a single monocular image, and then exploits 3D point cloud data to predict the depth shift and the camera's focal length that allow us to recover 3D scene shapes. As the two modules are trained separately, we do not need strictly paired training data. In addition, we propose an image-level normalized regression loss and a normal-based geometry loss to improve training with relative depth annotation. We test our depth model on nine unseen datasets and achieve state-of-the-art performance on zero-shot evaluation. Code is available at: https://github.com/aim-uofa/depth/. Wei Yin 0006, Jianming Zhang 0001, Oliver Wang, Simon Niklaus, Simon Chen, Yifan Liu 0001, Chunhua Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Semantic Layout Manipulation With High-Resolution Sparse AttentionabstractWe tackle the problem of semantic image layout manipulation, which aims to manipulate an input image by editing its semantic label map. A core problem of this task is how to transfer visual details from the input images to the new semantic layout while making the resulting image visually realistic. Recent work on learning cross-domain correspondence has shown promising results for global layout transfer with dense attention-based warping. However, this method tends to lose texture details due to the resolution limitation and the lack of smoothness constraint on correspondence. To adapt this paradigm for the layout manipulation task, we propose a high-resolution sparse attention module that effectively transfers visual details to new layouts at a resolution up to 512x512. To further improve visual quality, we introduce a novel generator architecture consisting of a semantic encoder and a two-stage decoder for coarse-to-fine synthesis. Experiments on the ADE20k and Places365 datasets demonstrate that our proposed approach achieves substantial improvements over the existing inpainting and layout manipulation methods. Haitian Zheng, Zhe Lin 0001, Jingwan Lu, Scott Cohen, Jianming Zhang 0001, Ning Xu 0007, Jiebo Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Layered Depth Refinement with Mask GuidanceabstractDepth maps are used in a wide range of applications from 3D rendering to 2D image effects such as Bokeh. However, those predicted by single image depth estimation (SIDE) models often fail to capture isolated holes in objects and/or have inaccurate boundary regions. Meanwhile, high-quality masks are much easier to obtain, using commercial auto-masking tools or off-the-shelf methods of segmentation and matting or even by manual editing. Hence, in this paper, we formulate a novel problem of mask-guided depth refinement that utilizes a generic mask to refine the depth prediction of SIDE models. Our framework performs layered refinement and inpainting/outpainting, decomposing the depth map into two separate layers signified by the mask and the inverse mask. As datasets with both depth and mask annotations are scarce, we propose a self-supervised learning scheme that uses arbitrary masks and RGB-D datasets. We empirically show that our method is robust to different types of masks and initial depth predictions, accurately refining depth values in inner and outer mask boundary regions. We further analyze our model with an ablation study and demonstrate results on real applications. More information can be found on our project page.11https://sooyekim.github.io/MaskDepth/ Soo Ye Kim, Jianming Zhang 0001, Simon Niklaus, Simon Chen, Zhe Lin 0001, Munchurl Kim |
CVPR | 2 |
| 2022 | BokehMe: When Neural Rendering Meets Classical RenderingabstractWe propose BokehMe, a hybrid bokeh rendering framework that marries a neural renderer with a classical physically motivated renderer. Given a single image and a potentially imperfect disparity map, BokehMe generates high-resolution photo-realistic bokeh effects with adjustable blur size, focal plane, and aperture shape. To this end, we analyze the errors from the classical scattering-based method and derive a formulation to calculate an error map. Based on this formulation, we implement the classical renderer by a scattering-based method and propose a two-stage neural renderer to fix the erroneous areas from the classical renderer. The neural renderer employs a dynamic multi-scale scheme to efficiently handle arbitrary blur sizes, and it is trained to handle imperfect disparity input. Experiments show that our method compares favorably against previous methods on both synthetic image data and real image data with predicted disparity. A user study is further conducted to validate the advantage of our method. Juewen Peng, Zhiguo Cao 0001, Xianrui Luo, Hao Lu 0003, Ke Xian, Jianming Zhang 0001 |
CVPR | 6 |
| 2022 | Lite Vision Transformer with Enhanced Self-AttentionabstractDespite the impressive representation capacity of vision transformer models, current light-weight vision transformer models still suffer from inconsistent and incorrect dense predictions at local regions. We suspect that the power of their self-attention mechanism is limited in shallower and thinner networks. We propose Lite Vision Transformer (LVT), a novel light-weight transformer network with two enhanced self-attention mechanisms to improve the model performances for mobile deployment. For the low-level features, we introduce Convolutional Self-Attention (CSA). Unlike previous approaches of merging convolution and self-attention, CSA introduces local self-attention into the convolution within a kernel of size$3\times 3$to enrich low-level features in the first stage of LVT. For the high-level features, we propose Recursive Atrous Self-Attention (RASA), which utilizes the multi-scale context when calculating the similarity map and a recursive mechanism to increase the representation capability with marginal extra parameter cost. The superiority of LVT is demonstrated on ImageNet recognition, ADE20K semantic segmentation, and COCO panoptic segmentation. The code is made publicly available11https://github.com/Chenglin-Yang/LVT. Yilin Wang 0002, Jianming Zhang 0001, He Zhang 0004, Zijun Wei, Zhe Lin 0001, Alan L. Yuille |
CVPR | 3 |
| 2022 | MPIB: An MPI-Based Bokeh Rendering Framework for Realistic Partial Occlusion Effects
Juewen Peng, Jianming Zhang 0001, Xianrui Luo, Hao Lu 0003, Ke Xian, Zhiguo Cao 0001 |
ECCV (6) | 2 |
| 2022 | Controllable Shadow Generation Using Pixel Height Maps
Yichen Sheng, Yifan Liu 0001, Jianming Zhang 0001, Wei Yin 0006, A. Cengiz Öztireli, He Zhang 0004, Zhe Lin 0001, Eli Shechtman, Bedrich Benes |
ECCV (23) | 3 |
| 2022 | Image Inpainting with Cascaded Modulation GAN and Object-Aware Training
Haitian Zheng, Zhe Lin 0001, Jingwan Lu, Scott Cohen, Eli Shechtman, Connelly Barnes, Jianming Zhang 0001, Ning Xu 0007, Sohrab Amirghodsi, Jiebo Luo 0001 |
ECCV (16) | 7 |
| 2022 | Less is More: Consistent Video Depth Estimation with Masked Frames ModelingabstractTemporal consistency is the key challenge of video depth estimation. Previous works are based on additional optical flow or camera poses, which is time-consuming. By contrast, we derive consistency with less information. Since videos inherently exist with heavy temporal redundancy, a missing frame could be recovered from neighboring ones. Inspired by this, we propose the frame masking network (FMNet), a spatial-temporal transformer network predicting the depth of masked frames based on their neighboring frames. By reconstructing masked temporal features, the FMNet can learn intrinsic inter-frame correlations, which leads to consistency. Compared with prior arts, experimental results demonstrate that our approach achieves comparable spatial accuracy and higher temporal consistency without any additional information. Our work provides a new perspective on consistent video depth estimation. Yiran Wang 0005, Xingyi Li 0005, Zhiguo Cao 0001, Ke Xian, Jianming Zhang 0001 |
ACM Multimedia | 6 |
| 2021 | Learning To Recover 3D Scene Shape From a Single ImageabstractDespite significant progress in monocular depth estimation in the wild, recent state-of-the-art methods cannot be used to recover accurate 3D scene shape due to an unknown depth shift induced by shift-invariant reconstruction losses used in mixed-data depth prediction training, and possible unknown camera focal length. We investigate this problem in detail, and propose a two-stage framework that first predicts depth up to an unknown scale and shift from a single monocular image, and then use 3D point cloud encoders to predict the missing depth shift and focal length that allow us to recover a realistic 3D scene shape. In addition, we propose an image-level normalized regression loss and a normal-based geometry loss to enhance depth prediction models trained on mixed datasets. We test our depth model on nine unseen datasets and achieve state-of-the-art performance on zero-shot dataset generalization. Code is available at: https://git.io/Depth Wei Yin 0006, Jianming Zhang 0001, Oliver Wang, Simon Niklaus, Long Mai, Simon Chen, Chunhua Shen |
CVPR | 2 |
| 2021 | SSN: Soft Shadow Network for Image CompositingabstractWe introduce an interactive Soft Shadow Network (SSN) to generates controllable soft shadows for image compositing. SSN takes a 2D object mask as input and thus is agnostic to image types such as painting and vector art. An environment light map is used to control the shadow’s characteristics, such as angle and softness. SSN employs an Ambient Occlusion Prediction module to predict an intermediate ambient occlusion map, which can be further refined by the user to provides geometric cues to modulate the shadow generation. To train our model, we design an efficient pipeline to produce diverse soft shadow training data using 3D object models. In addition, we propose an inverse shadow map representation to improve model training. We demonstrate that our model produces realistic soft shadows in real-time. Our user studies show that the generated shadows are often indistinguishable from shadows calculated by a physics-based renderer and users can easily use SSN through an interactive application to generate specific shadow effects in minutes. Yichen Sheng, Jianming Zhang 0001, Bedrich Benes |
CVPR | 2 |
| 2021 | Mask Guided Matting via Progressive Refinement NetworkabstractWe propose Mask Guided (MG) Matting, a robust matting framework that takes a general coarse mask as guidance. MG Matting leverages a network (PRN) design which encourages the matting model to provide self-guidance to progressively refine the uncertain regions through the decoding process. A series of guidance mask perturbation operations are also introduced in the training to further enhance its robustness to external guidance. We show that PRN can generalize to unseen types of guidance masks such as trimap and low-quality alpha matte, making it suitable for various application pipelines. In addition, we revisit the foreground color prediction problem for matting and propose a surprisingly simple improvement to address the dataset issue. Evaluation on real and synthetic benchmarks shows that MG Matting achieves state-of-the-art performance using various types of guidance inputs. Code and models are available at https://github.com/yucornetto/MGMatting. Qihang Yu, Jianming Zhang 0001, He Zhang 0004, Yilin Wang 0002, Zhe Lin 0001, Ning Xu 0007, Yutong Bai, Alan L. Yuille |
CVPR | 2 |
| 2021 | Multimodal Contrastive Training for Visual Representation LearningabstractWe develop an approach to learning visual representations that embraces multimodal data, driven by a combination of intra- and inter-modal similarity preservation objectives. Unlike existing visual pre-training methods, which solve a proxy prediction task in a single domain, our method exploits intrinsic data properties within each modality and semantic information from cross-modal correlation simultaneously, hence improving the quality of learned visual representations. By including multimodal training in a unified framework with different types of contrastive losses, our method can learn more powerful and generic visual features. We first train our model on COCO and evaluate the learned visual representations on various downstream tasks including image classification, object detection, and instance segmentation. For example, the visual representations pre-trained on COCO by our method achieve state-of-the-art top-1 validation accuracy of 55.3% on ImageNet classification, under the common transfer protocol. We also evaluate our method on the large-scale Stock images dataset and show its effectiveness on multi-label image tagging, and cross-modal retrieval tasks. Zhe Lin 0001, Jason Kuen, Jianming Zhang 0001, Yilin Wang 0002, Michael Maire, Ajinkya Kale, Baldo Faieta |
CVPR | 4 |
| 2021 | SSH: A Self-Supervised Framework for Image HarmonizationabstractImage harmonization aims to improve the quality of image compositing by matching the "appearance" (e.g., color tone, brightness and contrast) between foreground and background images. However, collecting large-scale annotated datasets for this task requires complex professional retouching. Instead, we propose a novel Self-Supervised Harmonization framework (SSH) that can be trained using just "free" natural images without being edited. We reformulate the image harmonization problem from a representation fusion perspective, which separately processes the foreground and background examples, to address the background occlusion issue. This framework design allows for a dual data augmentation method, where diverse [foreground, background, pseudo GT] triplets can be generated by cropping an image with perturbations using 3D color lookup tables (LUTs). In addition, we build a real-world harmonization dataset as carefully created by expert users, for evaluation and benchmarking purposes. Our results show that the proposed self-supervised method outperforms previous state-of-the-art methods in terms of reference metrics, visual quality, and subject user study. Code and dataset are available at https://github.com/VITA-Group/SSHarmonization. Yifan Jiang 0001, He Zhang 0004, Jianming Zhang 0001, Yilin Wang 0002, Zhe Lin 0001, Kalyan Sunkavalli, Simon Chen, Sohrab Amirghodsi, Sarah Kong, Zhangyang Wang |
ICCV | 3 |
| 2021 | Deep Image CompositingabstractImage compositing is a task of combining regions from different images to compose a new image. A common use case is background replacement of portrait images. To obtain high quality composites, professionals typically manually perform multiple editing steps such as segmentation, matting and foreground color decontamination, which is very time consuming even with sophisticated photo editing tools. In this paper, we propose a new method which can automatically generate high-quality image compositing with-out any user input. Our method can be trained end-to-end to optimize exploitation of contextual and color information of both foreground and background images, where the com-positing quality is considered in the optimization. Specifically, inspired by Laplacian pyramid blending, a dense-connected multi-stream fusion network is proposed to effectively fuse the information from the foreground and back-ground images at different scales. In addition, we intro-duce a self-taught strategy to progressively train from easy to complex cases to mitigate the lack of training data. Experiments show that the proposed method can automatically generate high-quality composites and outperforms existing methods both qualitatively and quantitatively. He Zhang 0004, Jianming Zhang 0001, Federico Perazzi, Zhe Lin 0001, Vishal M. Patel |
WACV | 2 |
| 2021 | Excitation Dropout: Encouraging Plasticity in Deep Neural Networks
Andrea Zunino, Sarah Adel Bargal, Pietro Morerio, Jianming Zhang 0001, Stan Sclaroff, Vittorio Murino |
Int. J. Comput. Vis. | 4 |
| 2021 | Guided Zoom: Zooming into Network Evidence to Refine Fine-Grained Model DecisionsabstractIn state-of-the-art deep single-label classification models, the top- k (k=2,3,4, ...) accuracy is usually significantly higher than the top-1 accuracy. This is more evident in fine-grained datasets, where differences between classes are quite subtle. Exploiting the information provided in the top k predicted classes boosts the final prediction of a model. We propose Guided Zoom, a novel way in which explainability could be used to improve model performance. We do so by making sure the model has "the right reasons" for a prediction. The reason/evidence upon which a deep neural network makes a prediction is defined to be the grounding, in the pixel space, for a specific class conditional probability in the model output. Guided Zoom examines how reasonable the evidence used to make each of the top- k predictions is. Test time evidence is deemed reasonable if it is coherent with evidence used to make similar correct decisions at training time. This leads to better informed predictions. We explore a variety of grounding techniques and study their complementarity for computing evidence. We show that Guided Zoom results in an improvement of a model's classification accuracy and achieves state-of-the-art classification performance on four fine-grained classification datasets. Our code is available at https://github.com/andreazuna89/Guided-Zoom. Sarah Adel Bargal, Andrea Zunino, Vitali Petsiuk, Jianming Zhang 0001, Kate Saenko, Vittorio Murino, Stan Sclaroff |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | LayoutGAN: Synthesizing Graphic Layouts With Vector-Wireframe Adversarial NetworksabstractLayout is important for graphic design and scene generation. We propose a novel Generative Adversarial Network, called LayoutGAN, that synthesizes layouts by modeling geometric relations of different types of 2D elements. The generator of LayoutGAN takes as input a set of randomly-placed 2D graphic elements, represented by vectors and uses self-attention modules to refine their labels and geometric parameters jointly to produce a realistic layout. Accurate alignment is critical for good layouts. We, thus, propose a novel differentiable wireframe rendering layer that maps the generated layout to a wireframe image, upon which a CNN-based discriminator is used to optimize the layouts in image space. We validate the effectiveness of LayoutGAN in various experiments including MNIST digit generation, document layout generation, clipart abstract scene generation, tangram graphic design, mobile app layout design, and webpage layout optimization from hand-drawn sketches. Jianan Li 0001, Jimei Yang, Aaron Hertzmann, Jianming Zhang 0001, Tingfa Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Sequence-to-Segments Networks for Detecting Segments in VideosabstractDetecting segments of interest from videos is a common problem for many applications. And yet it is a challenging problem as it often requires not only knowledge of individual target segments, but also contextual understanding of the entire video and the relationships between the target segments. To address this problem, we propose the Sequence-to-Segments Network (S2N), a novel and general end-to-end sequential encoder-decoder architecture. S2N first encodes the input video into a sequence of hidden states that capture information progressively, as it appears in the video. It then employs the Segment Detection Unit (SDU), a novel decoding architecture, that sequentially detects segments. At each decoding step, the SDU integrates the decoder state and encoder hidden states to detect a target segment. During training, we address the problem of finding the best assignment of predicted segments to ground truth using the Hungarian Matching Algorithm with Lexicographic Cost. Additionally we propose to use the squared Earth Mover's Distance to optimize the localization errors of the segments. We show the state-of-the-art performance of S2N across numerous tasks, including video highlighting, video summarization, and human action proposal generation. Zijun Wei, Boyu Wang 0001, Minh Hoai, Jianming Zhang 0001, Xiaohui Shen, Zhe Lin 0001, Radomír Mech, Dimitris Samaras |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Black-Box Diagnosis and Calibration on GAN Intra-Mode Collapse: A Pilot StudyabstractGenerative adversarial networks (GANs) nowadays are capable of producing images of incredible realism. Two concerns raised are whether the state-of-the-art GAN’s learned distribution still suffers from mode collapse and what to do if so. Existing diversity tests of samples from GANs are usually conducted qualitatively on a small scale and/or depend on the access to original training data as well as the trained model parameters. This article explores GAN intra-mode collapse and calibrates that in a novel black-box setting: access to neither training data nor the trained model parameters is assumed. The new setting is practically demanded yet rarely explored and significantly more challenging. As a first stab, we devise a set of statistical tools based on sampling that can visualize, quantify, and rectify intra-mode collapse . We demonstrate the effectiveness of our proposed diagnosis and calibration techniques, via extensive simulations and experiments, on unconditional GAN image generation (e.g., face and vehicle). Our study reveals that the intra-mode collapse is still a prevailing problem in state-of-the-art GANs and the mode collapse is diagnosable and calibratable in black-box settings. Our codes are available at https://github.com/VITA-Group/BlackBoxGANCollapse . Zhenyu Wu 0002, Ye Yuan 0012, Jianming Zhang 0001, Zhangyang Wang, Hailin Jin |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2021 | Attribute-Conditioned Layout GAN for Automatic Graphic DesignabstractModeling layout is an important first step for graphic design. Recently, methods for generating graphic layouts have progressed, particularly with Generative Adversarial Networks (GANs). However, the problem of specifying the locations and sizes of design elements usually involves constraints with respect to element attributes, such as area, aspect ratio and reading-order. Automating attribute conditional graphic layouts remains a complex and unsolved problem. In this article, we introduce Attribute-conditioned Layout GAN to incorporate the attributes of design elements for graphic layout generation by forcing both the generator and the discriminator to meet attribute conditions. Due to the complexity of graphic designs, we further propose an element dropout method to make the discriminator look at partial lists of elements and learn their local patterns. In addition, we introduce various loss designs following different design principles for layout optimization. We demonstrate that the proposed method can synthesize graphic layouts conditioned on different element attributes. It can also adjust well-designed layouts to new sizes while retaining elements' original reading-orders. The effectiveness of our method is validated through a user study. Jianan Li 0001, Jimei Yang, Jianming Zhang 0001, Christina Wang, Tingfa Xu |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2020 | M2KD: Incremental Learning via Multi-model and Multi-level Knowledge Distillation
Peng Zhou 0009, Long Mai, Jianming Zhang 0001, Ning Xu 0007, Zuxuan Wu, Larry Davis 0001 |
BMVC | 3 |
| 2020 | Adaptive Photographic Composition GuidanceabstractPhotographic composition is often taught as alignment with composition grids-most commonly, the rule of thirds. Professional photographers use more complex grids, like the harmonic armature, to achieve more diverse dynamic compositions. We are interested in understanding whether these complex grids are helpful to amateurs. Jane E, Ohad Fried, Jingwan Lu, Jianming Zhang 0001, Radomír Mech, Jose Echevarria, Pat Hanrahan, James A. Landay |
CHI | 4 |
| 2020 | SDC-Depth: Semantic Divide-and-Conquer Network for Monocular Depth EstimationabstractMonocular depth estimation is an ill-posed problem, and as such critically relies on scene priors and semantics. Due to its complexity, we propose a deep neural network model based on a semantic divide-and-conquer approach. Our model decomposes a scene into semantic segments, such as object instances and background stuff classes, and then predicts a scale and shift invariant depth map for each semantic segment in a canonical space. Semantic segments of the same category share the same depth decoder, so the global depth prediction task is decomposed into a series of category-specific ones, which are simpler to learn and easier to generalize to new scene types. Finally, our model stitches each local depth segment by predicting its scale and shift based on the global context of the image. The model is trained end-to-end using a multi-task loss for panoptic segmentation and depth prediction, and is therefore able to leverage large-scale panoptic segmentation datasets to boost its semantic understanding. We validate the effectiveness of our approach and show state-of-the-art performance on three benchmark datasets. Lijun Wang 0001, Jianming Zhang 0001, Oliver Wang, Zhe Lin 0001, Huchuan Lu |
CVPR | 2 |
| 2020 | Learning Visual Emotion Representations From Web DataabstractWe present a scalable approach for learning powerful visual features for emotion recognition. A critical bottleneck in emotion recognition is the lack of large scale datasets that can be used for learning visual emotion features. To this end, we curate a webly derived large scale dataset, StockEmotion, which has more than a million images. StockEmotion uses 690 emotion related tags as labels giving us a fine-grained and diverse set of emotion labels, circumventing the difficulty in manually obtaining emotion annotations. We use this dataset to train a feature extraction network, EmotionNet, which we further regularize using joint text and visual embedding and text distillation. Our experimental results establish that EmotionNet trained on the StockEmotion dataset outperforms SOTA models on four different visual emotion tasks. An aded benefit of our joint embedding training approach is that EmotionNet achieves competitive zero-shot recognition performance against fully supervised baselines on a challenging visual emotion dataset, EMOTIC, which further highlights the generalizability of the learned emotion features. Zijun Wei, Jianming Zhang 0001, Zhe Lin 0001, Joon-Young Lee, Niranjan Balasubramanian, Minh Hoai, Dimitris Samaras |
CVPR | 2 |
| 2020 | Structure-Guided Ranking Loss for Single Image Depth PredictionabstractSingle image depth prediction is a challenging task due to its ill-posed nature and challenges with capturing ground truth for supervision. Large-scale disparity data generated from stereo photos and 3D videos is a promising source of supervision, however, such disparity data can only approximate the inverse ground truth depth up to an affine transformation. To more effectively learn from such pseudo-depth data, we propose to use a simple pair-wise ranking loss with a novel sampling strategy. Instead of randomly sampling point pairs, we guide the sampling to better characterize structure of important regions based on the low-level edge maps and high-level object instance masks. We show that the pair-wise ranking loss, combined with our structure-guided sampling strategies, can significantly improve the quality of depth map prediction. In addition, we introduce a new relative depth dataset of about 21K diverse high-resolution web stereo photos to enhance the generalization ability of our model. In experiments, we conduct cross-dataset evaluation on six benchmark datasets and show that our method consistently improves over the baselines, leading to superior quantitative and qualitative results. Ke Xian, Jianming Zhang 0001, Oliver Wang, Long Mai, Zhe Lin 0001, Zhiguo Cao 0001 |
CVPR | 2 |
| 2020 | Shape Adaptor: A Learnable Resizing Module
Shikun Liu, Zhe Lin 0001, Yilin Wang 0002, Jianming Zhang 0001, Federico Perazzi, Edward Johns |
ECCV (12) | 4 |
| 2020 | Open-Edit: Open-Domain Image Manipulation with Open-Vocabulary Instructions
Xihui Liu, Zhe Lin 0001, Jianming Zhang 0001, Handong Zhao, Quan Tran, Xiaogang Wang 0001, Hongsheng Li 0001 |
ECCV (11) | 3 |
| 2020 | CLIFFNet for Monocular Depth Estimation with Hierarchical Embedding Loss
Lijun Wang 0001, Jianming Zhang 0001, Yifan Wang 0004, Huchuan Lu, Xiang Ruan |
ECCV (5) | 2 |
| 2020 | High-Resolution Image Inpainting with Iterative Confidence Feedback and Guided Upsampling
Yu Zeng 0001, Zhe Lin 0001, Jimei Yang, Jianming Zhang 0001, Eli Shechtman, Huchuan Lu |
ECCV (19) | 4 |
| 2020 | Unsupervised Video Object Segmentation with Joint Hotspot Tracking
Lu Zhang 0053, Jianming Zhang 0001, Zhe Lin 0001, Radomír Mech, Huchuan Lu, You He 0002 |
ECCV (14) | 2 |
| 2020 | Adversarial Knowledge Transfer from Unlabeled DataabstractWhile machine learning approaches to visual recognition offer great promise, most of the existing methods rely heavily on the availability of large quantities of labeled training data. However, in the vast majority of real-world settings, manually collecting such large labeled datasets is infeasible due to the cost of labeling data or the paucity of data in a given domain. In this paper, we present a novel Adversarial Knowledge Transfer (AKT) framework for transferring knowledge from internet-scale unlabeled data to improve the performance of a classifier on a given visual recognition task. The proposed adversarial learning framework aligns the feature space of the unlabeled source data with the labeled target data such that the target classifier can be used to predict pseudo labels on the source data. An important novel aspect of our method is that the unlabeled source data can be of different classes from those of the labeled target data, and there is no need to define a separate pretext task, unlike some existing approaches. Extensive experiments well demonstrate that models learned using our approach hold a lot of promise across a variety of visual recognition tasks on multiple standard datasets. Project page is at \texttthttps://agupt013.github.io/akt.html. Akash Gupta 0001, Rameswar Panda, Sujoy Paul, Jianming Zhang 0001, Amit K. Roy-Chowdhury |
ACM Multimedia | 4 |
| 2020 | Multi-way Encoding for RobustnessabstractDeep models are state-of-the-art for many computer vision tasks including image classification and object detection. However, it has been shown that deep models are vulnerable to adversarial examples. We highlight how one-hot encoding directly contributes to this vulnerability and propose breaking away from this widely-used, but highly-vulnerable mapping. We demonstrate that by leveraging a different output encoding, multi-way encoding, we decorre-late source and target models, making target models more secure. Our approach makes it more difficult for adversaries to find useful gradients for generating adversarial attacks. We present robustness for black-box and white-box attacks on four benchmark datasets: MNIST, CIFAR-10, CIFAR-100, and SVHN. The strength of our approach is also presented in the form of an attack for model watermarking, raising challenges in detecting stolen models. Donghyun Kim 0006, Sarah Adel Bargal, Jianming Zhang 0001, Stan Sclaroff |
WACV | 3 |
| 2020 | Reducing Footskate in Human Motion Reconstruction with Ground Contact ConstraintsabstractIn this paper, we aim to reduce the footskate artifacts when reconstructing human dynamics from monocular RGB videos. Recent work has made substantial progress in improving the temporal smoothness of the reconstructed motion trajectories. Their results, however, still suffer from severe foot skating and slippage artifacts. To tackle this issue, we present a neural network based detector for localizing ground contact events of human feet and use it to impose a physical constraint for optimization of the whole human dynamics in a video. We present a detailed study on the proposed ground contact detector and demonstrate high-quality human motion reconstruction results in various videos. Yuliang Zou, Jimei Yang, Duygu Ceylan, Jianming Zhang 0001, Federico Perazzi, Jia-Bin Huang 0001 |
WACV | 4 |
| 2019 | Guided Zoom: Questioning Network Evidence for Fine-grained Classification
Sarah Adel Bargal, Andrea Zunino, Vitali Petsiuk, Jianming Zhang 0001, Kate Saenko, Vittorio Murino, Stan Sclaroff |
BMVC | 4 |
| 2019 | SmartEye: Assisting Instant Photo Taking via Integrating User Preference with Deep View Proposal NetworkabstractInstant photo taking and sharing has become one of the most popular forms of social networking. However, taking high-quality photos is difficult as it requires knowledge and skill in photography that most non-expert users lack. In this paper we present SmartEye, a novel mobile system to help users take photos with good compositions in-situ. The back-end of SmartEye integrates the View Proposal Network (VPN), a deep learning based model that outputs composition suggestions in real time, and a novel, interactively updated module (P-Module) that adjusts the VPN outputs to account for personalized composition preferences. We also design a novel interface with functions at the front-end to enable real-time and informative interactions for photo taking. We conduct two user studies to investigate SmartEye qualitatively and quantitatively. Results show that SmartEye effectively models and predicts personalized composition preferences, provides instant high-quality compositions in-situ, and outperforms the non-personalized systems significantly. Shuai Ma 0005, Zijun Wei, Feng Tian 0001, Xiangmin Fan, Jianming Zhang 0001, Xiaohui Shen, Zhe Lin 0001, Jin Huang 0009, Radomír Mech, Dimitris Samaras, Hongan Wang |
CHI | 5 |
| 2019 | Neural Rejuvenation: Improving Deep Network Training by Enhancing Computational Resource UtilizationabstractIn this paper, we study the problem of improving computational resource utilization of neural networks. Deep neural networks are usually over-parameterized for their tasks in order to achieve good performances, thus are likely to have underutilized computational resources. This observation motivates a lot of research topics, e.g. network pruning, architecture search, etc. As models with higher computational costs (e.g. more parameters or more computations) usually have better performances, we study the problem of improving the resource utilization of neural networks so that their potentials can be further realized. To this end, we propose a novel optimization method named Neural Rejuvenation. As its name suggests, our method detects dead neurons and computes resource utilization in real time, rejuvenates dead neurons by resource reallocation and reinitialization, and trains them with new training schemes. By simply replacing standard optimizers with Neural Rejuvenation, we are able to improve the performances of neural networks by a very large margin while using similar training efforts and maintaining their original resource usages. The code is available here: https://github.com/joe-siyuan-qiao/NeuralRejuvenation-CVPR19 Siyuan Qiao, Zhe Lin 0001, Jianming Zhang 0001, Alan L. Yuille |
CVPR | 3 |
| 2019 | CapSal: Leveraging Captioning to Boost Semantics for Salient Object DetectionabstractDetecting salient objects in cluttered scenes is a big challenge. To address this problem, we argue that the model needs to learn discriminative semantic features for salient objects. To this end, we propose to leverage captioning as an auxiliary semantic task to boost salient object detection in complex scenarios. Specifically, we develop a CapSal model which consists of two sub-networks, the Image Captioning Network (ICN) and the Local-Global Perception Network (LGPN). ICN encodes the embedding of a generated caption to capture the semantic information of major objects in the scene, while LGPN incorporates the captioning embedding with local-global visual contexts for predicting the saliency map. ICN and LGPN are jointly trained to model high-level semantics as well as visual saliency. Extensive experiments demonstrate the effectiveness of image captioning in boosting the performance of salient object detection. In particular, our model performs significantly better than the state-of-the-art methods on several challenging datasets of complex scenarios. Lu Zhang 0053, Jianming Zhang 0001, Zhe Lin 0001, Huchuan Lu, You He 0002 |
CVPR | 2 |
| 2019 | Scaling Object Detection by Transferring Classification WeightsabstractLarge scale object detection datasets are constantly increasing their size in terms of the number of classes and annotations count. Yet, the number of object-level categories annotated in detection datasets is an order of magnitude smaller than image-level classification labels. State-of-the art object detection models are trained in a supervised fashion and this limits the number of object classes they can detect. In this paper, we propose a novel weight transfer network (WTN) to effectively and efficiently transfer knowledge from classification network's weights to detection network's weights to allow detection of novel classes without box supervision. We first introduce input and feature normalization schemes to curb the under-fitting during training of a vanilla WTN. We then propose autoencoder-WTN (AE-WTN) which uses reconstruction loss to preserve classification network's information over all classes in the target latent space to ensure generalization to novel classes. Compared to vanilla WTN, AE-WTN obtains absolute performance gains of 6% on two Open Images evaluation sets with 500 seen and 57 novel classes respectively, and 25% on a Visual Genome evaluation set with 200 novel classes. Jason Kuen, Federico Perazzi, Zhe Lin 0001, Jianming Zhang 0001, Yap-Peng Tan |
ICCV | 4 |
| 2019 | Towards High-Resolution Salient Object DetectionabstractDeep neural network based methods have made a significant breakthrough in salient object detection. However, they are typically limited to input images with low resolutions (400×400 pixels or less). Little effort has been made to train neural networks to directly handle salient object segmentation in high-resolution images. This paper pushes forward high-resolution saliency detection, and contributes a new dataset, named High-Resolution Salient Object Detection (HRSOD) dataset. To our best knowledge, HRSOD is the first high-resolution saliency detection dataset to date. As another contribution, we also propose a novel approach, which incorporates both global semantic information and local high-resolution details, to address this challenging task. More specifically, our approach consists of a Global Semantic Network (GSN), a Local Refinement Network (LRN) and a Global-Local Fusion Network (GLFN). The GSN extracts the global semantic information based on downsampled entire image. Guided by the results of GSN, the LRN focuses on some local regions and progressively produces high-resolution predictions. The GLFN is further proposed to enforce spatial consistency and boost performance. Experiments illustrate that our method outperforms existing state-of-the-art methods on high-resolution saliency datasets by a large margin, and achieves comparable or even better performance than them on some widely used saliency benchmarks. Yi Zeng 0006, Zhe Lin 0001, Jianming Zhang 0001, Huchuan Lu |
ICCV | 4 |
| 2019 | Fast Video Object Segmentation via Dynamic Targeting NetworkabstractWe propose a new model for fast and accurate video object segmentation. It consists of two convolutional neural networks, a Dynamic Targeting Network (DTN) and a Mask Refinement Network (MRN). DTN locates the object by dynamically focusing on regions of interest surrounding the target object. The target region is predicted by DTN via two sub-streams, Box Propagation (BP) and Box Re-identification (BR). The BP stream is faster but less effective at objects with large deformation or occlusion. The BR stream performs better in difficult scenarios at a higher computation cost. We propose a Decision Module (DM) to adaptively determine which sub-stream to use for each frame. Finally, MRN is exploited to predict segmentation within the target region. Experimental results on two public datasets demonstrate that the proposed model significantly outperforms existing methods without online training in both accuracy and efficiency, and is comparable to online training-based methods in accuracy with an order of magnitude faster speed. Lu Zhang 0053, Zhe Lin 0001, Jianming Zhang 0001, Huchuan Lu, You He 0002 |
ICCV | 3 |
| 2019 | LayoutGAN: Generating Graphic Layouts with Wireframe Discriminators
Jianan Li 0001, Jimei Yang, Aaron Hertzmann, Jianming Zhang 0001, Tingfa Xu |
ICLR (Poster) | 4 |
| 2018 | Excitation Backprop for RNNs
Sarah Adel Bargal, Andrea Zunino, Donghyun Kim 0006, Jianming Zhang 0001, Vittorio Murino, Stan Sclaroff |
CVPR | 4 |
| 2018 | Good View Hunting: Learning Photo Composition From Dense View PairsabstractFinding views with good photo composition is a challenging task for machine learning methods. A key difficulty is the lack of well annotated large scale datasets. Most existing datasets only provide a limited number of annotations for good views, while ignoring the comparative nature of view selection. In this work, we present the first large scale Comparative Photo Composition dataset, which contains over one million comparative view pairs annotated using a cost-effective crowdsourcing workflow. We show that these comparative view annotations are essential for training a robust neural network model for composition. In addition, we propose a novel knowledge transfer framework to train a fast view proposal network, which runs at 75+ FPS and achieves state-of-the-art performance in image cropping and thumbnail generation tasks on three benchmark datasets. The superiority of our method is also demonstrated in a user study on a challenging experiment, where our method significantly outperforms the baseline methods in producing diversified well-composed views. Zijun Wei, Jianming Zhang 0001, Xiaohui Shen, Zhe Lin 0001, Radomír Mech, Minh Hoai, Dimitris Samaras |
CVPR | 2 |
| 2018 | Learning to Blend Photos
Wei-Chih Hung, Jianming Zhang 0001, Xiaohui Shen, Zhe Lin 0001, Joon-Young Lee, Ming-Hsuan Yang 0001 |
ECCV (7) | 2 |
| 2018 | Contemplating Visual Emotions: Understanding and Overcoming Dataset Bias
Rameswar Panda, Jianming Zhang 0001, Joon-Young Lee, Xin Lu 0006, Amit K. Roy-Chowdhury |
ECCV (2) | 2 |
| 2018 | Concept Mask: Large-Scale Segmentation from Semantic Concepts
Yufei Wang 0001, Zhe Lin 0001, Xiaohui Shen, Jianming Zhang 0001, Scott Cohen |
ECCV (12) | 4 |
| 2018 | Sequence-to-Segment Networks for Segment DetectionabstractDetecting segments of interest from an input sequence is a challenging problem which often requires not only good knowledge of individual target segments, but also contextual understanding of the entire input sequence and the relationships between the target segments. To address this problem, we propose the Sequence-to-Segment Network (S$^2$N), a novel end-to-end sequential encoder-decoder architecture. S$^2$N first encodes the input into a sequence of hidden states that progressively capture both local and holistic information. It then employs a novel decoding architecture, called Segment Detection Unit (SDU), that integrates the decoder state and encoder hidden states to detect segments sequentially. During training, we formulate the assignment of predicted segments to ground truth as bipartite matching and use the Earth Mover's Distance to calculate the localization errors. We experiment with S$^2$N on temporal action proposal generation and video summarization and show that S$^2$N achieves state-of-the-art performance on both tasks. Zijun Wei, Boyu Wang 0001, Minh Hoai, Jianming Zhang 0001, Zhe Lin 0001, Xiaohui Shen, Radomír Mech, Dimitris Samaras |
NeurIPS | 4 |
| 2018 | Predicting Foreground Object Ambiguity and Efficiently Crowdsourcing the Segmentation(s)
Danna Gurari, Kun He 0003, Jianming Zhang 0001, Mehrnoosh Sameki, Suyog Dutt Jain, Stan Sclaroff, Margrit Betke, Kristen Grauman |
Int. J. Comput. Vis. | 4 |
| 2018 | Space-Time Tree Ensemble for Action Recognition and Localization
Shugao Ma, Jianming Zhang 0001, Stan Sclaroff, Nazli Ikizler-Cinbis, Leonid Sigal |
Int. J. Comput. Vis. | 2 |
| 2018 | Top-Down Neural Attention by Excitation Backprop
Jianming Zhang 0001, Sarah Adel Bargal, Zhe Lin 0001, Jonathan Brandt, Xiaohui Shen, Stan Sclaroff |
Int. J. Comput. Vis. | 1 |
| 2018 | Exemplar-Aided Salient Object Detection via Joint Latent Space EmbeddingabstractTraditional unsupervised salient object detection methods majorly rely on pre-defined assumptions about saliency. However, these assumptions may not be sufficient for handling test images of varied content and context. Meanwhile, supervised models learn saliency knowledge from thousands of annotated images, which are usually expensive to obtain. In this paper, we propose an exemplar-aided salient object detection method, which can complement heuristic saliency assumptions by leveraging only a few exemplar images. This is a challenging task since the appearances between the query images and the exemplars can be quite different. We handle it by learning the matching relationship of the intra-class instances in a latent embedding space in an online fashion. Given a test image and an annotated reference image (retrieved from several exemplar images), our method transfers the foreground and background information of the reference image to the test image via a joint latent embedding of image superpixels. Extensive experiments show that our method can easily improve the performance of existing unsupervised methods even when a very small reference image dataset (e.g. one image) is used. In addition, our method is able to attain competitive performance against fully supervised methods. Yuqiu Kong, Jianming Zhang 0001, Huchuan Lu, Xiuping Liu |
IEEE Trans. Image Process. | 2 |
| 2018 | DeepLens: shallow depth of field from a single imageabstractWe aim to generate high resolution shallow depth-of-field (DoF) images from a single all-in-focus image with controllable focal distance and aperture size. To achieve this, we propose a novel neural network model comprised of a depth prediction module, a lens blur module, and a guided upsampling module. All modules are differentiable and are learned from data. To train our depth prediction module, we collect a dataset of 2462 RGB-D images captured by mobile phones with a dual-lens camera, and use existing segmentation datasets to improve border prediction. We further leverage a synthetic dataset with known depth to supervise the lens blur and guided upsampling modules. The effectiveness of our system and training strategies are verified in the experiments. Our method can generate high-quality shallow DoF images at high resolution, and produces significantly fewer artifacts than the baselines and existing solutions for single image shallow DoF synthesis. Compared with the iPhone portrait mode, which is a state-of-the-art shallow DoF solution based on a dual-lens depth camera, our method generates comparable results, while allowing for greater flexibility to choose focal points and aperture size, and is not limited to one capture setup. Lijun Wang 0001, Xiaohui Shen, Jianming Zhang 0001, Oliver Wang, Zhe Lin 0001, Chih-Yao Hsieh, Sarah Kong, Huchuan Lu |
ACM Trans. Graph. | 3 |
| 2017 | Top-Down Visual Saliency Guided by CaptionsabstractNeural image/video captioning models can generate accurate descriptions, but their internal process of mapping regions to words is a black box and therefore difficult to explain. Top-down neural saliency methods can find important regions given a high-level semantic task such as object classification, but cannot use a natural language sentence as the top-down input for the task. In this paper, we propose Caption-Guided Visual Saliency to expose the region-to-word mapping in modern encoder-decoder networks and demonstrate that it is learned implicitly from caption training data, without any pixel-level annotations. Our approach can produce spatial or spatiotemporal heatmaps for both predicted captions, and for arbitrary query sentences. It recovers saliency without the overhead of introducing explicit attention layers, and can be used to analyze a variety of existing model architectures and improve their design. Evaluation on large-scale video and image datasets demonstrates that our approach achieves comparable captioning performance with existing methods while providing more accurate saliency heatmaps. Our code is available at visionlearninggroup.github.io/caption-guided-saliency/. Vasili Ramanishka, Abir Das, Jianming Zhang 0001, Kate Saenko |
CVPR | 3 |
| 2017 | Salient Object Subitizing
Jianming Zhang 0001, Shugao Ma, Mehrnoosh Sameki, Stan Sclaroff, Margrit Betke, Zhe Lin 0001, Xiaohui Shen, Brian L. Price, Radomír Mech |
Int. J. Comput. Vis. | 1 |
| 2017 | Do less and achieve more: Training CNNs for action recognition utilizing action images from the Web
Shugao Ma, Sarah Adel Bargal, Jianming Zhang 0001, Leonid Sigal, Stan Sclaroff |
Pattern Recognit. | 3 |
| 2016 | Unconstrained Salient Object Detection via Proposal Subset OptimizationabstractWe aim at detecting salient objects in unconstrained images. In unconstrained images, the number of salient objects (if any) varies from image to image, and is not given. We present a salient object detection system that directly outputs a compact set of detection windows, if any, for an input image. Our system leverages a Convolutional-Neural-Network model to generate location proposals of salient objects. Location proposals tend to be highly overlapping and noisy. Based on the Maximum a Posteriori principle, we propose a novel subset optimization framework to generate a compact set of detection windows out of noisy proposals. In experiments, we show that our subset optimization formulation greatly enhances the performance of our system, and our system attains 16-34% relative improvement in Average Precision compared with the state-of-the-art on three challenging salient object datasets. Jianming Zhang 0001, Stan Sclaroff, Zhe Lin 0001, Xiaohui Shen, Brian L. Price, Radomír Mech |
CVPR | 1 |
| 2016 | Top-Down Neural Attention by Excitation Backprop
Jianming Zhang 0001, Zhe Lin 0001, Jonathan Brandt, Xiaohui Shen, Stan Sclaroff |
ECCV (4) | 1 |
| 2016 | Exploiting Surroundedness for Saliency Detection: A Boolean Map ApproachabstractWe demonstrate the usefulness of surroundedness for eye fixation prediction by proposing a Boolean Map based Saliency model (BMS). In our formulation, an image is characterized by a set of binary images, which are generated by randomly thresholding the image's feature maps in a whitened feature space. Based on a Gestalt principle of figure-ground segregation, BMS computes a saliency map by discovering surrounded regions via topological analysis of Boolean maps. Furthermore, we draw a connection between BMS and the Minimum Barrier Distance to provide insight into why and how BMS can properly captures the surroundedness cue via Boolean maps. The strength of BMS is verified by its simplicity, efficiency and superior performance compared with 10 state-of-the-art methods on seven eye tracking benchmark datasets. Jianming Zhang 0001, Stan Sclaroff |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Salient Object SubitizingabstractPeople can immediately and precisely identify that an image contains 1, 2, 3 or 4 items by a simple glance. The phenomenon, known as Subitizing, inspires us to pursue the task of Salient Object Subitizing (SOS), i.e. predicting the existence and the number of salient objects in a scene using holistic cues. To study this problem, we propose a new image dataset annotated using an online crowdsourcing marketplace. We show that a proposed subitizing technique using an end-to-end Convolutional Neural Network (CNN) model achieves significantly better than chance performance in matching human labels on our dataset. It attains 94% accuracy in detecting the existence of salient objects, and 42–82% accuracy (chance is 20%) in predicting the number of salient objects (1, 2, 3, and 4+), without resorting to any object localization process. Finally, we demonstrate the usefulness of the proposed subitizing technique in two computer vision applications: salient object detection and object proposal. Jianming Zhang 0001, Shugao Ma, Mehrnoosh Sameki, Stan Sclaroff, Margrit Betke, Zhe Lin 0001, Xiaohui Shen, Brian L. Price, Radomír Mech |
CVPR | 1 |
| 2015 | Minimum Barrier Salient Object Detection at 80 FPSabstractWe propose a highly efficient, yet powerful, salient object detection method based on the Minimum Barrier Distance (MBD) Transform. The MBD transform is robust to pixel-value fluctuation, and thus can be effectively applied on raw pixels without region abstraction. We present an approximate MBD transform algorithm with 100X speedup over the exact algorithm. An error bound analysis is also provided. Powered by this fast MBD transform algorithm, the proposed salient object detection method runs at 80 FPS, and significantly outperforms previous methods with similar speed on four large benchmark datasets, and achieves comparable or better performance than state-of-the-art methods. Furthermore, a technique based on color whitening is proposed to extend our method to leverage the appearance-based backgroundness cue. This extended version further improves the performance, while still being one order of magnitude faster than all the other leading methods. Jianming Zhang 0001, Stan Sclaroff, Zhe Lin 0001, Xiaohui Shen, Brian L. Price, Radomír Mech |
ICCV | 1 |
| 2014 | MEEM: Robust Tracking via Multiple Experts Using Entropy Minimization
Jianming Zhang 0001, Shugao Ma, Stan Sclaroff |
ECCV (6) | 1 |
| 2013 | Online Motion Agreement TrackingabstractThis paper proposes a fast online multi-target tracking method, called motion agreement algorithm, which dynamically selects stable object regions to track. The appearance of each object, here pedestrians, is represented by multiple local patches. For each patch, the algorithm computes a local estimate of the direction of motion. By fusion of the agreements between a global estimate of the object motion and each local estimate, the algorithm identifies the object stable regions and enables robust tracking. The proposed patch-based appearance model was integrated into an efficient online tracking system that uses bipartite matching for data association. The experiments on recent pedestrian tracking benchmark sequences show that the proposed method achieves competitive results compared to state-of-the-art methods, including several offline tracking techniques. Zheng Wu 0003, Jianming Zhang 0001, Margrit Betke |
BMVC | 2 |
| 2013 | Decoding Children's Social BehaviorabstractWe introduce a new problem domain for activity recognition: the analysis of children's social and communicative behaviors based on video and audio data. We specifically target interactions between children aged 1-2 years and an adult. Such interactions arise naturally in the diagnosis and treatment of developmental disorders such as autism. We introduce a new publicly-available dataset containing over 160 sessions of a 3-5 minute child-adult interaction. In each session, the adult examiner followed a semi-structured play interaction protocol which was designed to elicit a broad range of social behaviors. We identify the key technical challenges in analyzing these behaviors, and describe methods for decoding the interactions. We present experimental results that demonstrate the potential of the dataset to drive interesting research questions, and show preliminary results for multi-modal activity recognition. James M. Rehg, Gregory D. Abowd, Agata Rozga, Mario Romero, Mark A. Clements, Stan Sclaroff, Irfan A. Essa, Opal Y. Ousley, Yin Li 0003, Chanho Kim, Hrishikesh Rao 0001, Jonathan C. Kim, Liliana Lo Presti, Jianming Zhang 0001, Denis Lantsman, Jonathan Bidwell, Zhefan Ye |
CVPR | 14 |
| 2013 | Action Recognition and Localization by Hierarchical Space-Time SegmentsabstractWe propose Hierarchical Space-Time Segments as a new representation for action recognition and localization. This representation has a two-level hierarchy. The first level comprises the root space-time segments that may contain a human body. The second level comprises multi-grained space-time segments that contain parts of the root. We present an unsupervised method to generate this representation from video, which extracts both static and non-static relevant space-time segments, and also preserves their hierarchical and temporal relationships. Using simple linear SVM on the resultant bag of hierarchical space-time segments representation, we attain better than, or comparable to, state-of-the-art action recognition performance on two challenging benchmark datasets and at the same time produce good action localization results. Shugao Ma, Jianming Zhang 0001, Nazli Ikizler-Cinbis, Stan Sclaroff |
ICCV | 2 |
| 2013 | Saliency Detection: A Boolean Map ApproachabstractA novel Boolean Map based Saliency (BMS) model is proposed. An image is characterized by a set of binary images, which are generated by randomly thresholding the image's color channels. Based on a Gestalt principle of figure-ground segregation, BMS computes saliency maps by analyzing the topological structure of Boolean maps. BMS is simple to implement and efficient to run. Despite its simplicity, BMS consistently achieves state-of-the-art performance compared with ten leading methods on five eye tracking datasets. Furthermore, BMS is also shown to be advantageous in salient object detection. Jianming Zhang 0001, Stan Sclaroff |
ICCV | 1 |
| 2012 | Online Multi-person Tracking by Tracker HierarchyabstractTracking-by-detection is a widely used paradigm for multi-person tracking but is affected by variations in crowd density, obstacles in the scene, varying illumination, human pose variation, scale changes, etc. We propose an improved tracking-by-detection framework for multi-person tracking where the appearance model is formulated as a template ensemble updated online given detections provided by a pedestrian detector. We employ a hierarchy of trackers to select the most effective tracking strategy and an algorithm to adapt the conditions for trackers' initialization and termination. Our formulation is online and does not require calibration information. In experiments with four pedestrian tracking benchmark datasets, our formulation attains accuracy that is comparable to, or better than, the state-of-the-art pedestrian trackers that must exploit calibration information and operate offline. Jianming Zhang 0001, Liliana Lo Presti, Stan Sclaroff |
AVSS | 1 |