EDBT 2026 Demo / reviewers in the wild / expert
Qing Liu 0017
dblp:53/4481-17
· DBLP profile ↗
27ranked-venue papers
3as first author
22since 2021 · last 2026
0000-0003-0879-7440ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 3 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 3 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LVM-Lite: Training Large Vision Models with Efficient Sequential ModelingabstractLarge Vision Models (LVMs) have demonstrated impressive capabilities by leveraging large-scale generative pre-training on visual sequences. However, training these models on massive datasets of single images and image sequences can be computationally expensive and limit their accessibility to researchers without substantial computational resources. This paper introduces LVM-Lite, a new two-stage learning pipeline for more efficient and effective LVM training. In the first stage, the model is pre-trained on a large corpus of single images. Subsequently, in the second stage, the model is fine-tuned on curated long image/video sequences. This decoupled training approach substantially accelerates the training process, achieving up to 2.7× speedup compared to baseline LVM training. Extensive experiments demonstrate that LVM-Lite achieves competitive performance on various generative and discriminative benchmarks while maintaining high training efficiency and strong scalability. https://github.com/UCSC-VLAA/LVM-Lite Xianhang Li, Hongru Zhu, Sucheng Ren, Peng Wang 0001, Xiaohui Shen, Qing Liu 0017, Cihang Xie |
WACV | 8 |
| 2025 | UniReal: Universal Image Generation and Editing via Learning Real-world DynamicsabstractWe introduce UniReal, a unified framework designed to address various image generation and editing tasks. Existing solutions often vary by tasks, yet share fundamental principles: preserving consistency between inputs and outputs while capturing visual variations. Inspired by recent video generation models that effectively balance consistency and variation across frames, we propose a unifying approach that treats image-level tasks as discontinuous video generation. Specifically, we treat varying numbers of input and output images as frames, enabling seamless support for tasks such as image generation, editing, customization, composition, etc. Although designed for image-level tasks, we leverage videos as a scalable source for universal supervision. UniReal learns world dynamics from large-scale videos, demonstrating advanced capability in handling shadows, reflections, pose variation, and object interaction, while also exhibiting emergent capability for novel applications. Xi Chen 0119, He Zhang 0004, Yuqian Zhou, Soo Ye Kim, Qing Liu 0017, Yijun Li 0001, Jianming Zhang 0001, Nanxuan Zhao, Yilin Wang 0002, Zhe Lin 0001, Hengshuang Zhao |
CVPR | 6 |
| 2025 | FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any GranularityabstractThe advent of large Vision-Language Models (VLMs) has significantly advanced multimodal tasks, enabling more sophisticated and accurate reasoning across various applications, including image and video captioning, visual question answering, and cross-modal retrieval. Despite their superior capabilities, VLMs struggle with fine-grained image regional composition information perception. Specifically, they have difficulty accurately aligning the segmentation masks with the corresponding semantics and precisely describing the compositional aspects of the referred regions. However, compositionality – the ability to understand and generate novel combinations of known visual and textual components – is critical for facilitating coherent reasoning and understanding across modalities by VLMs. To address this issue, we propose FineCaption, a novel VLM that can recognize arbitrary masks as referential inputs and process high-resolution images for compositional image captioning at different granularity levels. To support this endeavor, we introduce CompositionCap, a new dataset for multi-grained region compositional image captioning, which introduces the task of compositional attribute-aware regional image captioning. Empirical results demonstrate the effectiveness of our proposed model compared to other state-of-the-art VLMs. Additionally, we analyze the capabilities of current VLMs in recognizing various visual prompts for compositional region image captioning, highlighting areas for improvement in VLM design and training. https://hanghuacs.github.io/FineCaption/ Hang Hua, Qing Liu 0017, Lingzhi Zhang, Jing Shi 0005, Soo Ye Kim, Yilin Wang 0002, Jianming Zhang 0001, Zhe Lin 0001, Jiebo Luo 0001 |
CVPR | 2 |
| 2025 | Generative Video PropagationabstractLarge-scale video generation models have the inherent ability to realistically model natural scenes. In this paper, we demonstrate that through a careful design of a generative video propagation framework, various video tasks can be addressed in a unified way by leveraging the generative power of such models. Specifically, our framework, Gen-Prop, encodes the original video with a selective content encoder and propagates the changes made to the first frame using an image-to-video generation model. We propose a data generation scheme to cover multiple video tasks based on instance-level video segmentation datasets. Our model is trained by incorporating a mask prediction decoder head and optimizing a region-aware loss to aid the encoder to preserve the original content while the generation model propagates the modified region. This novel design opens up new possibilities: In editing scenarios, GenProp allows substantial changes to an object’s shape; for insertion, the inserted objects can exhibit independent motion; for removal, GenProp effectively removes effects like shadows and reflections from the whole video; for tracking, GenProp is capable of tracking objects and their associated effects together. Experiment results demonstrate the leading performance of our model in various video tasks, and we further provide in-depth analyses of the proposed framework. Shaoteng Liu, Tianyu Wang 0003, Jui-Hsien Wang, Qing Liu 0017, Joon-Young Lee, Yijun Li 0001, Bei Yu 0001, Zhe Lin 0001, Soo Ye Kim, Jiaya Jia |
CVPR | 4 |
| 2025 | Generative Image Layer Decomposition with Visual EffectsabstractRecent advancements in large generative models, particularly diffusion-based methods, have significantly enhanced the capabilities of image editing. However, achieving precise control over image composition tasks remains a challenge. Layered representations, which allow for independent editing of image components, are essential for user-driven content creation, yet existing approaches often struggle to decompose an image into plausible layers with accurately retained transparent visual effects such as shadows and reflections. We propose LayerDecomp, a generative framework for image layer decomposition which outputs photorealistic clean backgrounds and high-quality transparent foregrounds with faithfully preserved visual effects. To enable effective training, we first introduce a dataset preparation pipeline that automatically scales up simulated multi-layer data with synthesized visual effects. To further enhance real-world applicability, we supplement this simulated dataset with camera-captured images containing natural visual effects. Additionally, we propose a consistency loss which enforces the model to learn accurate representations for the transparent foreground layer when ground-truth annotations are not available. Our method achieves superior quality in layer decomposition, outperforming existing approaches in object removal and spatial editing tasks across several benchmarks and multiple user studies, unlocking various creative possibilities for layer-wise image editing. Jinrui Yang, Qing Liu 0017, Yijun Li 0001, Soo Ye Kim, Daniil Pakhomov, Mengwei Ren, Jianming Zhang 0001, Zhe Lin 0001, Cihang Xie, Yuyin Zhou |
CVPR | 2 |
| 2025 | ObjectMover: Generative Object Movement with Video PriorabstractSimple as it seems, moving an object to another location within an image is, in fact, a challenging image-editing task that requires re-harmonizing the lighting, adjusting the pose based on perspective, accurately filling occluded regions, and ensuring coherent synchronization of shadows and reflections while maintaining the object identity. In this paper, we present ObjectMover, a generative model that can perform object movement in highly challenging scenes. Our key insight is that we model this task as a sequence-to-sequence problem and fine-tune a video generation model to leverage its knowledge of consistent object generation across video frames. We show that with this approach, our model is able to adjust to complex real-world scenarios, handling extreme lighting harmonization and object effect movement. As large-scale data for object movement are unavailable, we construct a data generation pipeline using a modern game engine to synthesize high-quality data pairs. We further propose a multi-task learning strategy that enables training on real-world video data to improve the model generalization. Through extensive experiments, we demonstrate that ObjectMover achieves outstanding results and adapts well to real-world scenarios. Xin Yu 0004, Tianyu Wang 0003, Soo Ye Kim, Paul Guerrero 0001, Xi Chen 0119, Qing Liu 0017, Zhe Lin 0001, Xiaojuan Qi 0001 |
CVPR | 6 |
| 2025 | Baking Gaussian Splatting Into Diffusion Denoiser for Fast and Scalable Single-Stage Image-to-3D Generation and Reconstruction
Yuanhao Cai, He Zhang 0004, Kai Zhang 0045, Yixun Liang, Mengwei Ren, Fujun Luan, Qing Liu 0017, Soo Ye Kim, Jianming Zhang 0001, Yuqian Zhou, Yulun Zhang 0001, Xiaokang Yang 0001, Zhe Lin 0001, Alan L. Yuille |
ICCV | 7 |
| 2025 | What If We Recaption Billions of Web Images with LLaMA-3?abstractWeb-crawled image-text pairs are inherently noisy. Prior studies demonstrate that semantically aligning and enriching textual descriptions of these pairs can significantly enhance model training across various vision-language tasks, particularly text-to-image generation. However, large-scale investigations in this area remain predominantly closed-source. Our paper aims to bridge this community effort, leveraging the powerful and $\textit{open-sourced}$ LLaMA-3, a GPT-4 level LLM. Our recaptioning pipeline is simple: first, we fine-tune a LLaMA-3-8B powered LLaVA-1.5 and then employ it to recaption ~1.3 billion images from the DataComp-1B dataset. Our empirical results confirm that this enhanced dataset, Recap-DataComp-1B, offers substantial benefits in training advanced vision-language models. For discriminative models like CLIP, we observe an average of 3.1% enhanced zero-shot performance cross four cross-modal retrieval tasks using a mixed set of the original and our captions. For generative models like text-to-image Diffusion Transformers, the generated images exhibit a significant improvement in alignment with users' text instructions, especially in following complex queries. Our project page is https://www.haqtu.me/Recap-Datacomp-1B/. Xianhang Li, Haoqin Tu, Mude Hui, Zeyu Wang 0008, Bingchen Zhao, Junfei Xiao, Sucheng Ren, Jieru Mei, Qing Liu 0017, Huangjie Zheng, Yuyin Zhou, Cihang Xie |
ICML | 9 |
| 2024 | Amodal Scene Analysis via Holistic Occlusion Relation Inference and Generative Mask CompletionabstractAmodal scene analysis entails interpreting the occlusion relationship among scene elements and inferring the possible shapes of the invisible parts. Existing methods typically frame this task as an extended instance segmentation or a pair-wise object de-occlusion problem. In this work, we propose a new framework, which comprises a Holistic Occlusion Relation Inference (HORI) module followed by an instance-level Generative Mask Completion (GMC) module. Unlike previous approaches, which rely on mask completion results for occlusion reasoning, our HORI module directly predicts an occlusion relation matrix in a single pass. This approach is much more efficient than the pair-wise de-occlusion process and it naturally handles mutual occlusion, a common but often neglected situation. Moreover, we formulate the mask completion task as a generative process and use a diffusion-based GMC module for instance-level mask completion. This improves mask completion quality and provides multiple plausible solutions. We further introduce a large-scale amodal segmentation dataset with high-quality human annotations, including mutual occlusions. Experiments on our dataset and two public benchmarks demonstrate the advantages of our method. code public available at https://github.com/zbwxp/Amodal-AAAI. Bowen Zhang 0009, Qing Liu 0017, Jianming Zhang 0001, Yilin Wang 0002, Liyang Liu, Zhe Lin 0001, Yifan Liu 0001 |
AAAI | 2 |
| 2024 | UniHuman: A Unified Model For Editing Human Images in the WildabstractHuman image editing includes tasks like changing a person's pose, their clothing, or editing the image according to a text prompt. However, prior work often tackles these tasks separately, overlooking the benefit of mutual reinforcement from learning them jointly. In this paper, we propose UniHuman, a unified model that addresses multiple facets of human image editing in real-world settings. To enhance the model's generation quality and generalization capacity, we leverage guidance from human visual encoders and introduce a lightweight pose-warping module that can exploit different pose representations, accommodating unseen textures and patterns. Furthermore, to bridge the disparity between existing human editing benchmarks with real-world data, we curated 400K high-quality human image-text pairs for training and collected 2K human images for out-of-domain testing, both encompassing diverse clothing styles, backgrounds, and age groups. Experiments on both in-domain and out-of-domain test sets demonstrate that UniHuman outperforms task-specific models by a significant margin. In user studies, UniHuman is preferred by the users in an average of 77% of cases. Our project is available at this link. Nannan Li 0004, Qing Liu 0017, Krishna Kumar Singh, Yilin Wang 0002, Jianming Zhang 0001, Bryan A. Plummer, Zhe Lin 0001 |
CVPR | 2 |
| 2024 | SwapAnything: Enabling Arbitrary Object Swapping in Personalized Image Editing
Nanxuan Zhao, Wei Xiong 0008, Qing Liu 0017, He Zhang 0004, Jianming Zhang 0001, Hyunjoon Jung, Yilin Wang 0002, Xin Wang 0061 |
ECCV (32) | 4 |
| 2024 | GroupDiff: Diffusion-Based Group Portrait Editing
Yuming Jiang 0003, Nanxuan Zhao, Qing Liu 0017, Krishna Kumar Singh, Shuai Yang 0001, Chen Change Loy, Ziwei Liu 0002 |
ECCV (34) | 3 |
| 2024 | SegGen: Supercharging Segmentation Models with Text2Mask and Mask2Img Synthesis
Hanrong Ye, Jason Kuen, Qing Liu 0017, Zhe Lin 0001, Brian L. Price, Dan Xu 0002 |
ECCV (8) | 3 |
| 2024 | Structure-Guided Image Completion With Image-Level and Object-Level Semantic DiscriminatorsabstractStructure-guided image completion aims to inpaint a local region of an image according to an input guidance map from users. While such a task enables many practical applications for interactive editing, existing methods often struggle to hallucinate realistic object instances in complex natural scenes. Such a limitation is partially due to the lack of semantic-level constraints inside the hole region as well as the lack of a mechanism to enforce realistic object generation. In this work, we propose a learning paradigm that consists of semantic discriminators and object-level discriminators for improving the generation of complex semantics and objects. Specifically, the semantic discriminators leverage pretrained visual features to improve the realism of the generated visual concepts. Moreover, the object-level discriminators take aligned instances as inputs to enforce the realism of individual objects. Our proposed scheme significantly improves the generation quality and achieves state-of-the-art results on various tasks, including segmentation-guided completion, edge-guided manipulation and panoptically-guided manipulation on Places2 datasets. Furthermore, our trained model is flexible and can support multiple editing use cases, such as object insertion, replacement, removal and standard inpainting. In particular, our trained model combined with a novel automatic image completion pipeline achieves state-of-the-art results on the standard inpainting task. Haitian Zheng, Zhe Lin 0001, Jingwan Lu, Scott Cohen, Eli Shechtman, Connelly Barnes, Jianming Zhang 0001, Qing Liu 0017, Sohrab Amirghodsi, Yuqian Zhou, Jiebo Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2023 | Towards Open-World Segmentation of PartsabstractSegmenting object parts such as cup handles and animal bodies is important in many real-world applications but requires more annotation effort. The largest dataset nowadays contains merely two hundred object categories, implying the difficulty to scale up part segmentation to an unconstrained setting. To address this, we propose to explore a seemingly simplified but empirically useful and scalable task, class-agnostic part segmentation. In this problem, we disregard the part class labels in training and instead treat all of them as a single part class. We argue and demonstrate that models trained without part classes can better localize parts and segment them on objects unseen in training. We then present two further improvements. First, we propose to make the model object-aware, leveraging the fact that parts are “compositions”, whose extents are bounded by the corresponding objects and whose appearances are by nature not independent but bundled. Second, we introduce a novel approach to improve part segmentation on unseen objects, inspired by an interesting finding - for unseen objects, the pixel-wise features extracted by the model often reveal high-quality part segments. To this end, we propose a novel self-supervised procedure that iterates between pixel clustering and supervised contrastive learning that pulls pixels closer or pushes them away. Via extensive experiments on PartImageNet and Pascal-Part, we show notable and consistent gains by our approach, essentially a critical step towards open-world part segmentation. Tai-Yu Pan, Qing Liu 0017, Wei-Lun Chao, Brian L. Price |
CVPR | 2 |
| 2023 | SceneComposer: Any-Level Semantic Image SynthesisabstractWe propose a new framework for conditional image synthesis from semantic layouts of any precision levels, ranging from pure text to a 2D semantic canvas with precise shapes. More specifically, the input layout consists of one or more semantic regions with free-form text descriptions and adjustable precision levels, which can be set based on the desired controllability. The framework naturally reduces to text-to-image (T2I) at the lowest level with no shape information, and it becomes segmentation-to-image (S2I) at the highest level. By supporting the levels in-between, our framework is flexible in assisting users of different drawing expertise and at different stages of their creative workflow. We introduce several novel techniques to address the challenges coming with this new setup, including a pipeline for collecting training data; a precision-encoded mask pyramid and a text feature map representation to jointly encode precision level, semantics, and composition information; and a multi-scale guided diffusion model to synthesize images. To evaluate the proposed method, we collect a test dataset containing user-drawn layouts with diverse scenes and styles. Experimental results show that the proposed method can generate high-quality images following the layout at given precision, and compares favorably against existing methods. Project page https://zengxianyu.github.io/scenec/ Yu Zeng 0001, Zhe Lin 0001, Jianming Zhang 0001, Qing Liu 0017, John P. Collomosse, Jason Kuen, Vishal M. Patel |
CVPR | 4 |
| 2023 | Perceptual Artifacts Localization for Image Synthesis TasksabstractRecent advancements in deep generative models have facilitated the creation of photo-realistic images across various tasks. However, these generated images often exhibit perceptual artifacts in specific regions, necessitating manual correction. In this study, we present a comprehensive empirical examination of Perceptual Artifacts Localization (PAL) spanning diverse image synthesis endeavors. We introduce a novel dataset comprising 10, 168 generated images, each annotated with per-pixel perceptual artifact labels across ten synthesis tasks. A segmentation model, trained on our proposed dataset, effectively localizes artifacts across a range of tasks. Additionally, we illustrate its proficiency in adapting to previously unseen models using minimal training samples. We further propose an innovative zoom-in inpainting pipeline that seamlessly rectifies perceptual artifacts in the generated images. Through our experimental analyses, we elucidate several invaluable downstream applications, such as automated artifact rectification, non-referential image quality evaluation, and abnormal region detection in images. The dataset and code are released here: https://owenzlz.github.io/PAL4VST Lingzhi Zhang, Zhengjie Xu, Connelly Barnes, Yuqian Zhou, Qing Liu 0017, He Zhang 0004, Sohrab Amirghodsi, Zhe Lin 0001, Eli Shechtman, Jianbo Shi |
ICCV | 5 |
| 2023 | PHOTOSWAP: Personalized Subject Swapping in ImagesabstractIn an era where images and visual content dominate our digital landscape, the ability to manipulate and personalize these images has become a necessity.
Envision seamlessly substituting a tabby cat lounging on a sunlit window sill in a photograph with your own playful puppy, all while preserving the original charm and composition of the image.
We present \emph{Photoswap}, a novel approach that enables this immersive image editing experience through personalized subject swapping in existing images.
\emph{Photoswap} first learns the visual concept of the subject from reference images and then swaps it into the target image using pre-trained diffusion models in a training-free manner. We establish that a well-conceptualized visual subject can be seamlessly transferred to any image with appropriate self-attention and cross-attention manipulation, maintaining the pose of the swapped subject and the overall coherence of the image.
Comprehensive experiments underscore the efficacy and controllability of \emph{Photoswap} in personalized subject swapping. Furthermore, \emph{Photoswap} significantly outperforms baseline methods in human ratings across subject swapping, background preservation, and overall quality, revealing its vast application potential, from entertainment to professional editing. Yilin Wang 0002, Nanxuan Zhao, Tsu-Jui Fu, Wei Xiong 0008, Qing Liu 0017, He Zhang 0004, Jianming Zhang 0001, Hyunjoon Jung, Xin Wang 0061 |
NeurIPS | 6 |
| 2022 | Learning Part Segmentation through Unsupervised Domain Adaptation from Synthetic VehiclesabstractPart segmentations provide a rich and detailed part-level description of objects. However, their annotation requires an enormous amount of work, which makes it difficult to apply standard deep learning methods. In this paper, we propose the idea of learning part segmentation through unsupervised domain adaptation (UDA) from synthetic data. We first introduce UDA-Part, a comprehensive part segmentation dataset for vehicles that can serve as an adequate benchmark for UDA11https://qliu24.github.io/udapart/. In UDA-Part, we label parts on 3D CAD models which enables us to generate a large set of annotated synthetic images. We also annotate parts on a number of real images to provide a real test set. Secondly, to advance the adaptation of part models trained from the synthetic data to the real images, we introduce a new UDA algorithm that leverages the object's spatial structure to guide the adaptation process. Our experimental results on two real test datasets confirm the superiority of our approach over existing works, and demonstrate the promise of learning part segmentation for general objects from synthetic data. We believe our dataset provides a rich testbed to study UDA for part segmentation and will help to significantly push forward research in this area. Qing Liu 0017, Adam Kortylewski, Zhishuai Zhang, Zizhang Li, Mengqi Guo, Qihao Liu, Xiaoding Yuan, Jiteng Mu, Weichao Qiu, Alan L. Yuille |
CVPR | 1 |
| 2021 | Visual Analogy: Deep Learning Versus Compositional Models
Nicholas Ichien, Qing Liu 0017, Shuhao Fu, Keith J. Holyoak, Alan L. Yuille, Hongjing Lu |
CogSci | 2 |
| 2021 | Weakly Supervised Instance Segmentation for Videos With Temporal Mask ConsistencyabstractWeakly supervised instance segmentation reduces the cost of annotations required to train models. However, existing approaches which rely only on image-level class labels predominantly suffer from errors due to (a) partial segmentation of objects and (b) missing object predictions. We show that these issues can be better addressed by training with weakly labeled videos instead of images. In videos, motion and temporal consistency of predictions across frames provide complementary signals which can help segmentation. We are the first to explore the use of these video signals to tackle weakly supervised instance segmentation. We propose two ways to leverage this information in our model. First, we adapt inter-pixel relation network (IRN) [1] to effectively incorporate motion information during training. Second, we introduce a new MaskConsist module, which addresses the problem of missing object instances by transferring stable predictions between neighboring frames during training. We demonstrate that both approaches together improve the instance segmentation metric AP50on video frames of two datasets: Youtube-VIS and Cityscapes by 5% and 3% respectively. Qing Liu 0017, Vignesh Ramanathan, Dhruv Mahajan 0001, Alan L. Yuille, Zhenheng Yang |
CVPR | 1 |
| 2021 | Compositional Convolutional Neural Networks: A Robust and Interpretable Model for Object Recognition Under Occlusion
Adam Kortylewski, Qing Liu 0017, Angtian Wang, Yihong Sun, Alan L. Yuille |
Int. J. Comput. Vis. | 2 |
| 2020 | Compositional Convolutional Neural Networks: A Deep Architecture With Innate Robustness to Partial OcclusionabstractRecent work has shown that deep convolutional neural networks (DCNNs) do not generalize well under partial occlusion. Inspired by the success of compositional models at classifying partially occluded objects, we propose to integrate compositional models and DCNNs into a unified deep model with innate robustness to partial occlusion. We term this architecture Compositional Convolutional Neural Network. In particular, we propose to replace the fully connected classification head of a DCNN with a differentiable compositional model. The generative nature of the compositional model enables it to localize occluders and subsequently focus on the non-occluded parts of the object. We conduct classification experiments on artificially occluded images as well as real images of partially occluded objects from the MS-COCO dataset. The results show that DCNNs do not classify occluded objects robustly, even when trained with data that is strongly augmented with partial occlusions. Our proposed model outperforms standard DCNNs by a large margin at classifying partially occluded objects, even when it has not been exposed to occluded objects during training. Additional experiments demonstrate that CompositionalNets can also localize the occluders accurately, despite being trained with class labels only. The code and data used in this work are publicly available. Adam Kortylewski, Ju He, Qing Liu 0017, Alan L. Yuille |
CVPR | 3 |
| 2020 | Combining Compositional Models and Deep Networks For Robust Object Classification under OcclusionabstractDeep convolutional neural networks (DCNNs) are powerful models that yield impressive results at object classification. However, recent work has shown that they do not generalize well to partially occluded objects and to mask attacks. In contrast to DCNNs, compositional models are robust to partial occlusion, however, they are not as discriminative as deep models. In this work, we combine DC-NNs and compositional object models to retain the best of both approaches: a discriminative model that is robust to partial occlusion and mask attacks. Our model is learned in two steps. First, a standard DCNN is trained for image classification. Subsequently, we cluster the DCNN features into dictionaries. We show that the dictionary components resemble object part detectors and learn the spatial distribution of parts for each object class. We propose mixtures of compositional models to account for large changes in the spatial activation patterns (e.g. due to changes in the 3D pose of an object). At runtime, an image is first classified by the DCNN in a feedforward manner. The prediction uncertainty is used to detect partially occluded objects, which in turn are classified by the compositional model. Our experimental results demonstrate that combining compositional models and DCNNs resolves a fundamental problem of current deep learning approaches to computer vision: The combined model recognizes occluded objects, even when it has not been exposed to occluded objects during training, while at the same time maintaining high discriminative performance for non-occluded objects. Adam Kortylewski, Qing Liu 0017, Zhishuai Zhang, Alan L. Yuille |
WACV | 2 |
| 2019 | Seeing the Meaning: Vision Meets Semantics in Solving Pictorial Analogy Problems
Hongjing Lu, Qing Liu 0017, Nicholas Ichien, Alan L. Yuille, Keith J. Holyoak |
CogSci | 2 |
| 2019 | Semantic-Aware Knowledge Preservation for Zero-Shot Sketch-Based Image RetrievalabstractSketch-based image retrieval (SBIR) is widely recognized as an important vision problem which implies a wide range of real-world applications. Recently, research interests arise in solving this problem under the more realistic and challenging setting of zero-shot learning. In this paper, we investigate this problem from the viewpoint of domain adaptation which we show is critical in improving feature embedding in the zero-shot scenario. Based on a framework which starts with a pre-trained model on ImageNet and fine-tunes it on the training set of SBIR benchmark, we advocate the importance of preserving previously acquired knowledge, e.g., the rich discriminative features learned from ImageNet, to improve the model's transfer ability. For this purpose, we design an approach named Semantic-Aware Knowledge prEservation (SAKE), which fine-tunes the pre-trained model in an economical way and leverages semantic information, e.g., inter-class relationship, to achieve the goal of knowledge preservation. Zero-shot experiments on two extended SBIR datasets, TU-Berlin and Sketchy, verify the superior performance of our approach. Extensive diagnostic experiments validate that knowledge preserved benefits SBIR in zero-shot settings, as a large fraction of the performance gain is from the more properly structured feature embedding for photo images. Qing Liu 0017, Lingxi Xie, Alan L. Yuille |
ICCV | 1 |
| 2019 | Semantic Part Detection via Matching: Learning to Generalize to Novel Viewpoints From Limited Training DataabstractDetecting semantic parts of an object is a challenging task, particularly because it is hard to annotate semantic parts and construct large datasets. In this paper, we present an approach which can learn from a small annotated dataset containing a limited range of viewpoints and generalize to detect semantic parts for a much larger range of viewpoints. The approach is based on our matching algorithm, which is used for finding accurate spatial correspondence between two images and transplanting semantic parts annotated on one image to the other. Images in the training set are matched to synthetic images rendered from a 3D CAD model, following which a clustering algorithm is used to automatically annotate semantic parts of the CAD model. During the testing period, this CAD model can synthesize annotated images under every viewpoint. These synthesized images are matched to images in the testing set to detect semantic parts in novel viewpoints. Our algorithm is simple, intuitive, and contains very few parameters. Experiments show our method outperforms standard deep learning approaches and, in particular, performs much better on novel viewpoints. For facilitating the future research, code is available: https://github.com/ytongbai/SemanticPartDetection. Yutong Bai, Qing Liu 0017, Lingxi Xie, Weichao Qiu, Alan L. Yuille |
ICCV | 2 |