EDBT 2026 Demo / reviewers in the wild / expert
Li Niu 0002
dblp:02/3166-2
· DBLP profile ↗
118ranked-venue papers
16as first author
85since 2021 · last 2026
0000-0003-1970-8634ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 94 · 11 first-author · 68 since 2021Artificial intelligence and machine learning · 84 · 15 first-author · 62 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | D3ToM: Decider-Guided Dynamic Token Merging for Accelerating Diffusion MLLMsabstractDiffusion-based multimodal large language models (Diffusion MLLMs) have recently demonstrated impressive non-autoregressive generative capabilities across vision-and-language tasks. However, Diffusion MLLMs exhibit substantially slower inference than autoregressive models: Each denoising step employs full bidirectional self-attention over the entire sequence, resulting in cubic decoding complexity that becomes computationally impractical with thousands of visual tokens. To address this challenge, we propose D³ToM, a Decider-guided dynamic token merging method that dynamically merges redundant visual tokens at different denoising steps to accelerate inference in Diffusion MLLMs. At each denoising step, D³ToM uses decider tokens—the tokens generated in the previous denoising step—to build an importance map over all visual tokens. Then it maintains a proportion of the most salient tokens and merges the remainder through similarity-based aggregation. This plug-and-play module integrates into a single transformer layer, physically shortening the visual token sequence for all subsequent layers without altering model parameters. Moreover, D³ToM employs a merge ratio that dynamically varies with each denoising step, aligns with the native decoding process of Diffusion MLLMs, achieving superior performance under equivalent computational budgets. Extensive experiments show that D³ToM accelerates inference while preserving competitive performance. Shuochen Chang, Xiaofeng Zhang 0006, Qingyang Liu 0008, Li Niu 0002 |
AAAI | 4 |
| 2026 | CareCom: Generative Image Composition with Calibrated Reference FeaturesabstractImage composition aims to seamlessly insert foreground object into background. Despite the huge progress in generative image composition, the existing methods are still struggling with simultaneous detail preservation and foreground pose/view adjustment. To address this issue, we extend the existing generative composition model to multi-reference version, which allows using arbitrary number of foreground reference images. Furthermore, we propose to calibrate the global and local features of foreground reference images to make them compatible with the background information. The calibrated reference features can supplement the original reference features with useful global and local information of proper pose/view. Extensive experiments on MVImgNet and MureCom demonstrate that the generative model can greatly benefit from the calibrated reference features. Bo Zhang 0075, Qingdong He, Jinlong Peng, Li Niu 0002 |
AAAI | 5 |
| 2026 | Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to InterventionabstractShuochen Chang, Tong Bai, Xiaofeng Zhang, Qianli Ma, Qingyang Liu, Zhaohe Liao, Yibo Miao, Li Niu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shuochen Chang, Tong Bai, Xiaofeng Zhang 0006, Qianli Ma 0008, Qingyang Liu 0008, Zhaohe Liao, Yibo Miao, Li Niu 0002 |
ACL (1) | 8 |
| 2026 | Bridging Visual Dynamics and Narrative Reasoning: Multimodal Large Language Models for Short Drama Quality Assessment
Qingyang Liu 0008, Jiangtong Li, Zelin Peng, Shaobo Wang 0001, Zhaohe Liao, Shuochen Chang, Bingjie Gao, Mu Liu, Jidong Jiang, Li Niu 0002 |
WWW | 11 |
| 2026 | Parse, Align and Aggregate: Graph-Driven Compositional Reasoning for Video Question AnsweringabstractVideo Question-Answering (VideoQA) enables machines to interpret and respond to complex video content, advancing human-computer interaction. However, existing multimodal large language models (MLLMs) often provide incomplete or opaque explanations and existing benchmarks mainly focus on the correction of final answers, limiting insight into their reasoning processes and hindering both transparency and verifiability. To address this gap, we propose the Question Parsing, Video Alignment and Answer Aggregation framework (QPVA$^{3}$3), which leverages a compositional graph to drive visual and logical reasoning in VideoQA. Specifically, QPVA$^{3}$3 consists of three core components, the planner, executor, and reasoner to generate the compositional graph and conduct graph-driven reasoning. For the original question, the planner parses it into the compositional graph, capturing the underlying reasoning logic and structuring it into a series of interconnected questions. For each question in compositional graph, the executor aligns the video by selecting relevant video clips and generates answers, ensuring accurate, context-specific responses. For each question with its first-order descents, the reasoner aggregates answers by integrating reasoning logic with visual evidence, resolving conflicts to produce a coherent and accurate response. Moreover, to assess the performance of existing MLLMs in the reasoning processes of VideoQA, we introduce novel compositional consistency metrics and construct a VideoQA benchmark (QPVA$^{3}$3 Bench) with 3,492 question-video tuples, each annotated with detailed compositional graphs and fine-grained answers. We evaluate the QPVA$^{3}$3 framework on QPVA$^{3}$3 Bench and 5 other VideoQA benchmarks. Experimental results demonstrate that our framework improves both consistency and accuracy compared to baselines, leading to a more transparent and verifiable VideoQA system. This approach has the potential to advance the field, as supported by our comprehensive evaluation and benchmarking efforts. Jiangtong Li, Zhaohe Liao, Fengshun Xiao, Tianjiao Li 0001, Qiang Zhang 0055, Haohua Zhao 0001, Li Niu 0002, Guang Chen 0001, Liqing Zhang 0001, Changjun Jiang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video GenerationabstractThe evolution of Text-To-video (T2V) generative models, trained on large-scale datasets, has been marked by significant progress. However, the sensitivity of T2V generative models to input prompts highlights the critical role of prompt design in influencing generative outcomes. Prior research has predominantly relied on Large Language Models (LLMs) to align user-provided prompts with the distribution of training prompts, albeit without tailored guidance encompassing prompt vocabulary and sentence structure nuances. To this end, we introduce RAPO, a novel Retrieval-Augmented Prompt Optimization framework. In order to address potential inaccuracies and ambiguous details generated by LLM-generated prompts. RAPO refines the naive prompts through dual optimization branches, selecting the superior prompt for T2V generation. The first branch augments user prompts with diverse modifiers extracted from a learned relational graph, refining them to align with the format of training prompts via a fine-tuned LLM. Conversely, the second branch rewrites the naive prompt using a pre-trained LLM following a well-defined instruction set. Extensive experiments demonstrate that RAPO can effectively enhance both the static and dynamic dimensions of generated videos, demonstrating the significance of prompt optimization for user-provided prompts. Project website: GitHub. Bingjie Gao, Yu Qiao 0001, Li Niu 0002, Yaohui Wang 0001 |
CVPR | 6 |
| 2025 | Decouple-Then-Merge: Finetune Diffusion Models as Multi-Task LearningabstractDiffusion models are trained by learning a sequence of models that reverse each step of noise corruption. Typically, the model parameters are fully shared across multiple timesteps to enhance training efficiency. However, since the denoising tasks differ at each timestep, the gradients computed at different timesteps may conflict, potentially degrading the overall performance of image generation. To solve this issue, this work proposes a Decouple-then-Merge (DeMe) framework, which begins with a pretrained model and finetunes separate models tailored to specific timesteps. We introduce several improved techniques during the fine-tuning stage to promote effective knowledge sharing while minimizing training interference across timesteps. Finally, after finetuning, these separate models can be merged into a single model in the parameter space, ensuring efficient and practical inference. Experimental results show significant generation quality improvements upon 6 benchmarks including Stable Diffusion on COCO30K, ImageNet1K, PartiPrompts, and DDPM on LSUN Church, LSUN Bedroom, and CIFAR10. Code is available at GitHub. Qianli Ma 0008, Xuefei Ning, Dongrui Liu, Li Niu 0002, Linfeng Zhang 0001 |
CVPR | 4 |
| 2025 | Shadow Generation Using Diffusion Model with Geometry PriorabstractImage composition involves integrating foreground object into background image to obtain a composite image. One of the key challenges is to produce realistic shadow for the inserted foreground object. Recently, diffusion-based methods have shown superior performance compared to GAN-based methods in shadow generation. However, they are still struggling to generate shadows with plausible geometry in complex cases. In this paper, we focus on promoting diffusion-based methods by leveraging geometry priors. Specifically, we first predict the rotated bounding box and matched shadow shapes for the foreground shadow. Then, the geometry information of rotated bounding box and matched shadow shapes is injected into ControlNet to facilitate shadow generation. Extensive experiments on both DESOBAv2 dataset and real composite images validate the effectiveness of our proposed method. The code and model are released at https://github.com/bcmi/GPSDiffusion-Object-Shadow-Generation. Qingyang Liu 0002, Xinhao Tao, Li Niu 0002, Guangtao Zhai |
CVPR | 4 |
| 2025 | Light-a-Video: Training-Free Video Relighting via Progressive Light FusionabstractRecent advancements in image relighting models, driven by large-scale datasets and pre-trained diffusion models, have enabled the imposition of consistent lighting. However, video relighting still lags, primarily due to the excessive training costs and the scarcity of diverse, high-quality video relighting datasets. A simple application of image relighting models on a frame-by-frame basis leads to several issues: lighting source inconsistency and relighted appearance inconsistency, resulting in flickers in the generated videos. In this work, we propose Light-A-Video, a training-free approach to achieve temporally smooth video relighting. Adapted from image relighting models, Light-A-Video introduces two key techniques to enhance lighting consistency. First, we design a Consistent Light Attention (CLA) module, which enhances cross-frame interactions within the self-attention layers of the image relight model to stabilize the generation of the background lighting source. Second, leveraging the physical principle of light transport independence, we apply linear blending between the source video's appearance and the relighted appearance, using a Progressive Light Fusion (PLF) strategy to ensure smooth temporal transitions in illumination. Experiments show that Light-A-Video improves the temporal consistency of relighted video while maintaining the relighted image quality, ensuring coherent lighting transitions across frames. Project page: https://bujiazi.github.io/light-a-video.github.io/. Jiazi Bu, Pengyang Ling, Pan Zhang 0001, Qidong Huang, Jinsong Li 0001, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Anyi Rao, Jiaqi Wang 0003, Li Niu 0002 |
ICCV | 13 |
| 2025 | Pedestrian Motion Reconstruction: A Large-scale Benchmark via Mixed Reality Rendering with Multiple Perspectives and ModalitiesabstractReconstructing pedestrian motion from dynamic sensors, with a focus on pedestrian intention, is crucial for advancing autonomous driving safety. However, this task is challenging due to data limitations arising from technical complexities, safety, and cost concerns. We introduce the Pedestrian Motion Reconstruction (PMR) dataset, which focuses on pedestrian intention to reconstruct behavior using multiple perspectives and modalities. PMR is developed from a mixed reality platform that combines real-world realism with the extensive, accurate labels of simulations, thereby reducing costs and risks. It captures the intricate dynamics of pedestrian interactions with objects and vehicles, using different modalities for a comprehensive understanding of human-vehicle interaction. Analyses show that PMR can naturally exhibit pedestrian intent and simulate extreme cases. PMR features a vast collection of data from 54 subjects interacting across 12 urban settings with 7 objects, encompassing 12,138 sequences with diverse weather conditions and vehicle speeds. This data provides a rich foundation for modeling pedestrian intent through multi-view and multi-modal insights. We also conduct comprehensive benchmark assessments across different modalities to thoroughly evaluate pedestrian motion reconstruction methods. Yiyi Zhang 0002, Xinhao Hu, Li Niu 0002, Jianfu Zhang 0003, Yasushi Makihara, Yasushi Yagi, Wenlong Liao, Junchi Yan, Liqing Zhang 0001 |
ICLR | 4 |
| 2025 | Object Placement for AnythingabstractObject placement aims to determine the appropriate placement (e.g., location and size) of a foreground object when placing it on the background image. Most previous works are limited by small-scale labeled dataset, which hinders the real-world application of object placement. In this work, we devise a semi-supervised framework which can exploit large-scale unlabeled dataset to promote the generalization ability of discriminative object placement models. The discriminative models predict the rationality label for each foreground placement given a foreground-background pair. To better leverage the labeled data, under the semi-supervised framework, we further propose to transfer the knowledge of rationality variation, i.e., whether the change of foreground placement would result in the change of rationality label, from labeled data to unlabeled data. Extensive experiments demonstrate that our framework can effectively enhance the generalization ability of discriminative object placement models. Bingjie Gao, Bo Zhang 0075, Li Niu 0002 |
ICME | 3 |
| 2025 | Divide and Conquer: Exploring Language-centric Tree Reasoning for Video Question-AnsweringabstractVideo Question-Answering (VideoQA) remains challenging in achieving advanced cognitive reasoning due to the uncontrollable and opaque reasoning processes in existing Multimodal Large Language Models (MLLMs). To address this issue, we propose a novel Language-centric Tree Reasoning (LTR) framework that targets on enhancing the reasoning ability of models. In detail, it recursively divides the original question into logically manageable parts and conquers them piece by piece, enhancing the reasoning capabilities and interpretability of existing MLLMs. Specifically, in the first stage, the LTR focuses on language to recursively generate a language-centric logical tree, which gradually breaks down the complex cognitive question into simple perceptual ones and plans the reasoning path through a RAG-based few-shot approach. In the second stage, with the aid of video content, the LTR performs bottom-up logical reasoning within the tree to derive the final answer along with the traceable reasoning path. Experiments across 11 VideoQA benchmarks demonstrate that our LTR framework significantly improves both accuracy and interpretability compared to state-of-the-art MLLMs. To our knowledge, this is the first work to implement a language-centric logical tree to guide MLLM reasoning in VideoQA, paving the way for language-centric video understanding from perception to cognition. Zhaohe Liao, Jiangtong Li, Qingyang Liu 0002, Fengshun Xiao, Tianjiao Li 0001, Qiang Zhang 0055, Guang Chen 0001, Li Niu 0002, Changjun Jiang 0002, Liqing Zhang 0001 |
ICML | 9 |
| 2025 | Dual-Space Adaptive Fusion for Self-supervised Text-guided Image EditingabstractDespite the recent progress in text-guided image editing, achieving coherent image transformation while preserving the untouched content remains challenging for diffusion models. Many methods attempt to blend source and target elements, but they often struggle to adaptively modify the fusion strategy for different editing cases, limiting editing capabilities or requiring manual parameter tuning. In response, we propose a dual-space adaptive editing framework that appropriately fuses the source and target elements in two spaces in a self-supervised manner. Specifically, for each editing case, we optimize a set of coefficients to balance the fusion of source and target elements in the inner space, while continuously optimizing the fusion mask in the latent space. The entire process does not require additional data or manual parameter tuning. Extensive experiments demonstrate the robustness and superior performance of our method across various editing scenarios. Qingyang Liu 0008, Li Niu 0002 |
MMAsia | 3 |
| 2025 | Weak-shot Keypoint Estimation via Keyness and Correspondence TransferabstractKeypoint estimation is a fundamental task in computer vision, but generally requires large-scale annotated data for training. Few-shot and unsupervised keypoint estimation are prevalent economical paradigms, but the former still requires annotations for extensive novel classes while the latter only supports for single class. In this paper, we focus on the task of weak-shot keypoint estimation, where multiple novel classes are learned from unlabeled images with the help of labeled base classes. The key problem is what to transfer from base classes to novel classes, and we propose to transfer keyness and correspondence, which essentially belong to comparing entities and thus are class-agnostic and class-wise transferable. The keyness compares which pixel in the local region is more key, which can guide the keypoints of novel classes to move towards the local maximum (i.e., obtaining keypoints). The correspondence compares whether the two pixels belongs to the same semantic part, which can activate the keypoints of novel classes by reinforcing the consistency between corresponding points on two paired images. By transferring keyness and correspondence, our framework achieves favourable performance for weak-shot keypoint estimation. Extensive experiments and analyses on large-scale benchmark MP-100 demonstrate our effectiveness. Junjie Chen 0008, Zeyu Luo, Zezheng Liu, Wenhui Jiang 0001, Li Niu 0002, Yuming Fang 0001 |
NeurIPS | 5 |
| 2025 | Webly Supervised Fine-Grained Classification by Integrally Tackling Noises and Subtle DifferencesabstractWebly-supervised fine-grained visual classification (WSL-FGVC) aims to learn similar sub-classes from cheap web images, which suffers from two major issues: label noises in web images and subtle differences among fine-grained classes. However, existing methods for WSL-FGVC only focus on suppressing noise at image-level, but neglect to mine cues at pixel-level to distinguish the subtle differences among fine-grained classes. In this paper, we propose a bag-level top-down attention framework, which could tackle label noises and mine subtle cues simultaneously and integrally. Specifically, our method first extracts high-level semantic information from a bag of images belonging to the same class, and then uses the bag-level information to mine discriminative regions in various scales of each image. Besides, we propose to derive attention weights from attention maps to weight the bag-level fusion for a robust supervision. We also propose an attention loss on self-bag attention and cross-bag attention to facilitate the learning of valid attention. Extensive experiments on four WSL-FGVC datasets, i.e., Web-Aircraft, Web-Bird, Web-Car, and WebiNat-5089, demonstrate the effectiveness of our method against the state-of-the-art methods. Junjie Chen 0008, Jiebin Yan, Yuming Fang 0001, Li Niu 0002 |
IEEE Trans. Image Process. | 4 |
| 2024 | WeditGAN: Few-Shot Image Generation via Latent Space RelocationabstractIn few-shot image generation, directly training GAN models on just a handful of images faces the risk of overfitting. A popular solution is to transfer the models pretrained on large source domains to small target ones. In this work, we introduce WeditGAN, which realizes model transfer by editing the intermediate latent codes w in StyleGANs with learned constant offsets (delta w), discovering and constructing target latent spaces via simply relocating the distribution of source latent spaces. The established one-to-one mapping between latent spaces can naturally prevents mode collapse and overfitting. Besides, we also propose variants of WeditGAN to further enhance the relocation process by regularizing the direction or finetuning the intensity of delta w. Experiments on a collection of widely used source/target datasets manifest the capability of WeditGAN in generating realistic and diverse images, which is simple yet highly effective in the research area of few-shot image generation. Codes are available at https://github.com/Ldhlwh/WeditGAN. Yuxuan Duan, Li Niu 0002, Yan Hong 0001, Liqing Zhang 0001 |
AAAI | 2 |
| 2024 | Painterly Image Harmonization by Learning from Painterly ObjectsabstractGiven a composite image with photographic object and painterly background, painterly image harmonization targets at stylizing the composite object to be compatible with the background. Despite the competitive performance of existing painterly harmonization works, they did not fully leverage the painterly objects in artistic paintings. In this work, we explore learning from painterly objects for painterly image harmonization. In particular, we learn a mapping from background style and object information to object style based on painterly objects in artistic paintings. With the learnt mapping, we can hallucinate the target style of composite object, which is used to harmonize encoder feature maps to produce the harmonized image. Extensive experiments on the benchmark dataset demonstrate the effectiveness of our proposed method. Li Niu 0002, Junyan Cao, Yan Hong 0001, Liqing Zhang 0001 |
AAAI | 1 |
| 2024 | Progressive Painterly Image Harmonization from Low-Level Styles to High-Level StylesabstractPainterly image harmonization aims to harmonize a photographic foreground object on the painterly background. Different from previous auto-encoder based harmonization networks, we develop a progressive multi-stage harmonization network, which harmonizes the composite foreground from low-level styles (e.g., color, simple texture) to high-level styles (e.g., complex texture). Our network has better interpretability and harmonization performance. Moreover, we design an early-exit strategy to automatically decide the proper stage to exit, which can skip the unnecessary and even harmful late stages. Extensive experiments on the benchmark dataset demonstrate the effectiveness of our progressive harmonization network. Li Niu 0002, Yan Hong 0001, Junyan Cao, Liqing Zhang 0001 |
AAAI | 1 |
| 2024 | Shadow Generation with Decomposed Mask Prediction and Attentive Shadow FillingabstractImage composition refers to inserting a foreground object into a background image to obtain a composite image. In this work, we focus on generating plausible shadows for the inserted foreground object to make the composite image more realistic. To supplement the existing small-scale dataset, we create a large-scale dataset called RdSOBA with rendering techniques. Moreover, we design a two-stage network named DMASNet with decomposed mask prediction and attentive shadow filling. Specifically, in the first stage, we decompose shadow mask prediction into box prediction and shape prediction. In the second stage, we attend to reference background shadow pixels to fill the foreground shadow. Abundant experiments prove that our DMASNet achieves better visual effects and generalizes well to real composite images. Xinhao Tao, Junyan Cao, Yan Hong 0001, Li Niu 0002 |
AAAI | 4 |
| 2024 | Meta-Point Learning and Refining for Category-Agnostic Pose EstimationabstractCategory-agnostic pose estimation (CAPE) aims to predict keypoints for arbitrary classes given a few support images annotated with keypoints. Existing methods only rely on the features extracted at support keypoints to predict or refine the keypoints on query image, but a few support feature vectors are local and inadequate for CAPE. Considering that human can quickly perceive potential keypoints of arbitrary objects, we propose a novel framework for CAPE based on such potential keypoints (named as meta-points). Specifically, we maintain learnable embeddings to capture inherent information of various keypoints, which interact with image feature maps to produce meta-points without any support. The produced meta-points could serve as meaningful potential keypoints for CAPE. Due to the inevitable gap between inherency and annotation, we finally utilize the identities and details offered by support key-points to assign and refine meta-points to desired keypoints in query image. In addition, we propose a progressive deformable point decoder and a slacked regression loss for better prediction and supervision. Our novel framework not only reveals the inherency of key points but also outperforms existing methods of CAPE. Comprehensive experiments and in-depth studies on large-scale MP-100 dataset demon-strate the effectiveness of our framework. Code is avaiable at https://github.com/chenbys/MetaPoint Junjie Chen 0008, Jiebin Yan, Yuming Fang 0001, Li Niu 0002 |
CVPR | 4 |
| 2024 | Align and Aggregate: Compositional Reasoning with Video Alignment and Answer Aggregation for Video Question-AnsweringabstractDespite the recent progress made in Video Question-Answering (VideoQA), these methods typically function as black-boxes, making it difficult to understand their reasoning processes and perform consistent compositional reasoning. To address these challenges, we propose a model-agnostic Video Alignment and Answer Aggregation (VA3) framework, which is capable of enhancing both compositional consistency and accuracy of existing VidQA methods by integrating video aligner and answer aggregator modules. The video aligner hierarchically selects the relevant video clips based on the question, while the answer ag-gregator deduces the answer to the question based on its sub-questions, with compositional consistency ensured by the information flow along question decomposition graph and the contrastive learning strategy. We evaluate our framework on three settings of the AGQA-Decomp dataset with three baseline methods, and propose new metrics to measure the compositional consistency of VidQA methods more comprehensively. Moreover, we propose a large language model (LLM) based automatic question decomposition pipeline to apply our framework to any VidQA dataset. We extend MSVD and NExT-QA datasets with it to evaluate our VA3framework on broader scenarios. Extensive experiments show that our framework improves both compositional consistency and accuracy of existing methods, leading to more interpretable real-world VidQA models. Zhaohe Liao, Jiangtong Li, Li Niu 0002, Liqing Zhang 0001 |
CVPR | 3 |
| 2024 | Shadow Generation for Composite Image Using Diffusion ModelabstractIn the realm of image composition, generating realistic shadow for the inserted foreground remains a formidable challenge. Previous works have developed image-to-image translation models which are trained on paired training data. However, they are struggling to generate shadows with accurate shapes and intensities, hindered by data scarcity and inherent task complexity. In this paper, we resort to foundation model with rich prior knowledge of natural shadow images. Specifically, we first adapt ControlNet to our task and then propose intensity modulation modules to improve the shadow intensity. Moreover, we extend the small-scale DESOBA dataset to DESOBAv2 using a novel data acquisition pipeline. Experimental results on both DESOBA and DESOBAv2 datasets as well as real composite images demonstrate the superior capability of our model for shadow generation task. The dataset, code, and model are released at https://github.com/bcmi/Object-Shadow-Generation-Dataset-DESOBAv2. Qingyang Liu 0002, Junqi You, Jianting Wang, Xinhao Tao, Bo Zhang 0075, Li Niu 0002 |
CVPR | 6 |
| 2024 | Unsupervised Exposure Correction
Ruodai Cui, Li Niu 0002, Guosheng Hu |
ECCV (6) | 2 |
| 2024 | COIN-Matting: Confounder Intervention for Image Matting
Zhaohe Liao, Jiangtong Li, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002, Li Niu 0002, Liqing Zhang 0001 |
ECCV (19) | 6 |
| 2024 | Arbitrary Style Transfer with Prototype-Based Channel AlignmentabstractStyle transfer aims to migrate the "style" from a style image to a content image. Despite the appealing results achieved by existing methods, few studies have considered the alignment of semantics or structures between the style image and the content image. To overcome this problem, we propose a novel network with two parallel branches: coarse-grained stylization branch and fine-grained decoration branch. In the stylization branch, we perform conventional AdaIN to produce globally stylized feature. In the decoration branch, we propose a ProtoType-based Channel Alignment module to align the channels between style feature and content feature, followed by Adaptive Group Transfer to produce locally stylized feature. Extensive experiments demonstrate that our proposed method outperforms state-of-the-art methods in terms of visual quality and efficiency. Yan Hong 0001, Li Niu 0002, Jianfu Zhang 0003 |
ICASSP | 2 |
| 2024 | DomainGallery: Few-shot Domain-driven Image Generation by Attribute-centric FinetuningabstractThe recent progress in text-to-image models pretrained on large-scale datasets has enabled us to generate various images as long as we provide a text prompt describing what we want. Nevertheless, the availability of these models is still limited when we expect to generate images that fall into a specific domain either hard to describe or just unseen to the models. In this work, we propose DomainGallery, a few-shot domain-driven image generation method which aims at finetuning pretrained Stable Diffusion on few-shot target datasets in an attribute-centric manner. Specifically, DomainGallery features prior attribute erasure, attribute disentanglement, regularization and enhancement. These techniques are tailored to few-shot domain-driven generation in order to solve key issues that previous works have failed to settle. Extensive experiments are given to validate the superior performance of DomainGallery on a variety of domain-driven generation scenarios. Yuxuan Duan, Yan Hong 0001, Bo Zhang 0075, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002, Jianfu Zhang 0003, Li Niu 0002, Liqing Zhang 0001 |
NeurIPS | 8 |
| 2024 | Painterly Image Harmonization via Adversarial Residual LearningabstractImage compositing plays a vital role in photo editing. After inserting a foreground object into another background image, the composite image may look unnatural and inharmonious. When the foreground is photorealistic and the background is an artistic painting, painterly image harmonization aims to transfer the style of background painting to the foreground object, which is a challenging task due to the large domain gap between foreground and background. In this work, we employ adversarial learning to bridge the domain gap between foreground feature map and background feature map. Specifically, we design a dual-encoder generator, in which the residual encoder produces the residual features added to the foreground feature map from main encoder. Then, a pixel-wise discriminator plays against the generator, encouraging the refined foreground feature map to be indistinguishable from background feature map. Extensive experiments demonstrate that our method could achieve more harmonious and visually appealing results than previous methods. Xudong Wang 0001, Li Niu 0002, Junyan Cao, Yan Hong 0001, Liqing Zhang 0001 |
WACV | 2 |
| 2023 | Painterly Image Harmonization in Dual DomainsabstractImage harmonization aims to produce visually harmonious composite images by adjusting the foreground appearance to be compatible with the background. When the composite image has photographic foreground and painterly background, the task is called painterly image harmonization. There are only few works on this task, which are either time-consuming or weak in generating well-harmonized results. In this work, we propose a novel painterly harmonization network consisting of a dual-domain generator and a dual-domain discriminator, which harmonizes the composite image in both spatial domain and frequency domain. The dual-domain generator performs harmonization by using AdaIN modules in the spatial domain and our proposed ResFFT modules in the frequency domain. The dual-domain discriminator attempts to distinguish the inharmonious patches based on the spatial feature and frequency feature of each patch, which can enhance the ability of generator in an adversarial manner. Extensive experiments on the benchmark dataset show the effectiveness of our method. Our code and model are available at https://github.com/bcmi/PHDNet-Painterly-Image-Harmonization. Junyan Cao, Yan Hong 0001, Li Niu 0002 |
AAAI | 3 |
| 2023 | Amodal Instance Segmentation via Prior-Guided ExpansionabstractAmodal instance segmentation aims to infer the amodal mask, including both the visible part and occluded part of each object instance. Predicting the occluded parts is challenging. Existing methods often produce incomplete amodal boxes and amodal masks, probably due to lacking visual evidences to expand the boxes and masks. To this end, we propose a prior-guided expansion framework, which builds on a two-stage segmentation model (i.e., Mask R-CNN) and performs box-level (resp., pixel-level) expansion for amodal box (resp., mask) prediction, by retrieving regression (resp., flow) transformations from a memory bank of expansion prior. We conduct extensive experiments on KINS, D2SA, and COCOA cls datasets, which show the effectiveness of our method. Junjie Chen 0008, Li Niu 0002, Jianfu Zhang 0003, Jianlou Si, Chen Qian 0006, Liqing Zhang 0001 |
AAAI | 2 |
| 2023 | Few-Shot Defect Image Generation via Defect-Aware Feature ManipulationabstractThe performances of defect inspection have been severely hindered by insufficient defect images in industries, which can be alleviated by generating more samples as data augmentation. We propose the first defect image generation method in the challenging few-shot cases. Given just a handful of defect images and relatively more defect-free ones, our goal is to augment the dataset with new defect images. Our method consists of two training stages. First, we train a data-efficient StyleGAN2 on defect-free images as the backbone. Second, we attach defect-aware residual blocks to the backbone, which learn to produce reasonable defect masks and accordingly manipulate the features within the masked regions by training the added modules on limited defect images. Extensive experiments on MVTec AD dataset not only validate the effectiveness of our method in generating realistic and diverse defect images, but also manifest the benefits it brings to downstream defect inspection tasks. Codes are available at https://github.com/Ldhlwh/DFMGAN. Yuxuan Duan, Yan Hong 0001, Li Niu 0002, Liqing Zhang 0001 |
AAAI | 3 |
| 2023 | Isometric Manifold Learning Using Hierarchical FlowabstractWe propose the Hierarchical Flow (HF) model constrained by isometric regularizations for manifold learning that combines manifold learning goals such as dimensionality reduction, inference, sampling, projection and density estimation into one unified framework. Our proposed HF model is regularized to not only produce embeddings preserving the geometric structure of the manifold, but also project samples onto the manifold in a manner conforming to the rigorous definition of projection. Theoretical guarantees are provided for our HF model to satisfy the two desired properties. In order to detect the real dimensionality of the manifold, we also propose a two-stage dimensionality reduction algorithm, which is a time-efficient algorithm thanks to the hierarchical architecture design of our HF model. Experimental results justify our theoretical analysis, demonstrate the superiority of our dimensionality reduction algorithm in cost of training time, and verify the effect of the aforementioned properties in improving performances on downstream tasks such as anomaly detection. Jianfu Zhang 0003, Li Niu 0002, Liqing Zhang 0001 |
AAAI | 3 |
| 2023 | Geometric Inductive Biases for Identifiable Unsupervised Learning of Disentangled Representations
Li Niu 0002, Liqing Zhang 0001 |
AAAI | 2 |
| 2023 | Video Object of Interest SegmentationabstractIn this work, we present a new computer vision task named video object of interest segmentation (VOIS). Given a video and a target image of interest, our objective is to simultaneously segment and track all objects in the video that are relevant to the target image. This problem combines the traditional video object segmentation task with an additional image indicating the content that users are concerned with. Since no existing dataset is perfectly suitable for this new task, we specifically construct a large-scale dataset called LiveVideos, which contains 2418 pairs of target images and live videos with instance-level annotations. In addition, we propose a transformer-based method for this task. We revisit Swin Transformer and design a dual-path structure to fuse video and image features. Then, a transformer decoder is employed to generate object proposals for segmentation and tracking from the fused features. Extensive experiments on LiveVideos dataset show the superiority of our proposed method. Chunru Zhan, Tiezheng Ge, Yuning Jiang 0001, Li Niu 0002 |
AAAI | 6 |
| 2023 | Image Cropping with Spatial-aware Feature and Rank ConsistencyabstractImage cropping aims to find visually appealing crops in an image. Despite the great progress made by previous methods, they are weak in capturing the spatial relationship between crops and aesthetic elements (e.g., salient objects, semantic edges). Besides, due to the high annotation cost of labeled data, the potential of unlabeled data awaits to be excavated. To address the first issue, we propose spatial-aware feature to encode the spatial relationship between candidate crops and aesthetic elements, by feeding the concatenation of crop mask and selectively aggregated feature maps to a light-weighted encoder. To address the second issue, we train a pair-wise ranking classifier on labeled images and transfer such knowledge to unlabeled images to enforce rank consistency. Experimental results on the benchmark datasets show that our proposed method performs favorably against state-of-the-art methods. Li Niu 0002, Bo Zhang 0075, Liqing Zhang 0001 |
CVPR | 2 |
| 2023 | Deep Image Harmonization with Learnable AugmentationabstractThe goal of image harmonization is adjusting the foreground appearance in a composite image to make the whole image harmonious. To construct paired training images, existing datasets adopt different ways to adjust the illumination statistics of foregrounds of real images to produce synthetic composite images. However, different datasets have considerable domain gap and the performances on small-scale datasets are limited by insufficient training data. In this work, we explore learnable augmentation to enrich the illumination diversity of small-scale datasets for better harmonization performance. In particular, our designed SYthetic COmposite Network (SycoNet) takes in a real image with foreground mask and a random vector to learn suitable color transformation, which is applied to the foreground of this real image to produce a synthetic composite image. Comprehensive experiments demonstrate the effectiveness of our proposed learnable augmentation for image harmonization. The code of SycoNet is released at https://github.com/bcmi/SycoNet-Adaptive-Image-Harmonization. Li Niu 0002, Junyan Cao, Wenyan Cong, Liqing Zhang 0001 |
ICCV | 1 |
| 2023 | Knowledge Proxy Intervention for Deconfounded Video Question AnsweringabstractRecently, Video Question-Answering (VideoQA) has drawn more and more attention from both the industry and the research community. Despite all the success achieved by recent works, dataset bias always harmfully misleads current methods focusing on spurious correlations in training data. To analyze the effects of dataset bias, we frame the VideoQA pipeline into a causal graph, which shows the causalities among video, question, aligned feature between video and question, answer, and underlying confounder. Through the causal graph, we prove that the confounder and the backdoor path lead to spurious causality. To tackle the challenge that the confounder in VideoQA is unobserved and non-enumerable in general, we propose a model-agnostic framework called Knowledge Proxy Intervention (KPI), which introduces an extra knowledge proxy variable in the causal graph to cut the backdoor path and remove the effect of confounder. Our KPI framework exploits the front-door adjustment, which requires no prior knowledge about the confounder. The effectiveness of our KPI framework is corroborated by three baseline methods on five benchmark datasets, including MSVD-QA, MSRVTT-QA, TGIF-QA, NExT-QA, and Causal-VidQA. Jiangtong Li, Li Niu 0002, Liqing Zhang 0001 |
ICCV | 2 |
| 2023 | Deep Image Harmonization with Globally Guided Feature Transformation and Relation DistillationabstractGiven a composite image, image harmonization aims to adjust the foreground illumination to be consistent with background. Previous methods have explored transforming foreground features to achieve competitive performance. In this work, we show that using global information to guide foreground feature transformation could achieve significant improvement. Besides, we propose to transfer the foreground-background relation from real images to composite images, which can provide intermediate supervision for the transformed encoder features. Additionally, considering the drawbacks of existing harmonization datasets, we also contribute a ccHarmony dataset which simulates the natural illumination variation. Extensive experiments on iHarmony4 and our contributed dataset demonstrate the superiority of our method. Our ccHarmony dataset is released at https://github.com/bcmi/Image-HarmonizationDataset-ccHarmony. Li Niu 0002, Linfeng Tan, Xinhao Tao, Junyan Cao, Fengjun Guo, Liqing Zhang 0001 |
ICCV | 1 |
| 2023 | Fine-grained Visible Watermark RemovalabstractVisible watermark removal aims to erase the watermark from watermarked image and recover the background image, which is a challenging task due to the diverse watermarks. Previous works have designed dynamic network to handle various types of watermarks adaptively, but they ignore that even the watermarked region in a single image can be divided into multiple local parts with distinct visual appearances. In this work, we advance image-specific dynamic network towards part-specific dynamic network, which discovers multiple local parts within the watermarked region and handle them adaptively. Specifically, we propose a query-based multi-task framework, in which part query embeddings are jointly used in two branches to predict part masks and restore watermarked parts. Extensive experiments demonstrate the effectiveness of our fine-grained watermark removal network. Li Niu 0002, Xing Zhao 0010, Bo Zhang 0075, Liqing Zhang 0001 |
ICCV | 1 |
| 2023 | Foreground Object Search by Distilling Composite Image FeatureabstractForeground object search (FOS) aims to find compatible foreground objects for a given background image, producing realistic composite image. We observe that competitive retrieval performance could be achieved by using a discriminator to predict the compatibility of composite image, but this approach has unaffordable time cost. To this end, we propose a novel FOS method via distilling composite feature (DiscoFOS). Specifically, the abovementioned discriminator serves as teacher network. The student network employs two encoders to extract foreground feature and background feature. Their interaction output is enforced to match the composite image feature from the teacher network. Additionally, previous works did not release their datasets, so we contribute two datasets for FOS task: S-FOSD dataset with synthetic composite images and R-FOSD dataset with real composite images. Extensive experiments on our two datasets demonstrate the superiority of the proposed method over previous approaches. The dataset and code are available at https://github.com/bcmi/Foreground-Object-Search-Dataset-FOSD. Bo Zhang 0075, Jiacheng Sui, Li Niu 0002 |
ICCV | 3 |
| 2023 | Object Part Parsing with Hierarchical Dual TransformerabstractObject part parsing involves segmenting objects into semantic parts, which has drawn great attention recently. The current methods ignore the specific hierarchical structure of the object, which can be used as strong prior knowledge. To address this, we propose the Hierarchical Dual Transformer (HDTR) to explore the contribution of the typical structural priors of the object parts. HDTR first generates the pyramid multi-granularity pixel representations under the supervision of the object part parsing maps at different semantic levels and then assigns each region an initial part embedding. Moreover, HDTR generates an edge pixel representation to extend the capability of the network to capture detailed information. Afterward, we design a Hierarchical Part Transformer to upgrade the part embeddings to their hierarchical counterparts with the assistance of the multi-granularity pixel representations. Next, we propose a Hierarchical Pixel Transformer to infer the hierarchical information from the part embeddings to enrich the pixel representations. Note that both transformer decoders rely on the structural relations between object parts, i.e., dependency, composition, and decomposition relations. The experiments on five large-scale datasets, i.e., LaPa, CelebAMask-HQ, CIHP, LIP and Pascal Animal, demonstrate that our method sets a new state-of-the-art performance for object part parsing. Jianlou Si, Naihao Liu, Li Niu 0002, Chen Qian 0006 |
ACM Multimedia | 5 |
| 2023 | Painterly Image Harmonization using Diffusion ModelabstractPainterly image harmonization aims to insert photographic objects into paintings and obtain artistically coherent composite images. Previous methods for this task mainly rely on inference optimization or generative adversarial network, but they are either very time-consuming or struggling at fine control of the foreground objects (e.g., texture and content details). To address these issues, we propose a novel Painterly Harmonization stable Diffusion model (PHDiffusion), which includes a lightweight adaptive encoder and a Dual Encoder Fusion (DEF) module. Specifically, the adaptive encoder and the DEF module first stylize foreground features within each encoder. Then, the stylized foreground features from both encoders are combined to guide the harmonization process. During training, besides the noise loss in diffusion model, we additionally employ content loss and two style losses, i.e., AdaIN style loss and contrastive style loss, aiming to balance the trade-off between style migration and content preservation. Compared with the state-of-the-art models from related fields, our PHDiffusion can stylize the foreground more sufficiently and simultaneously retain finer content. Our code and model are available at https://github.com/bcmi/PHDiffusion-Painterly-Image-Harmonization Lingxiao Lu, Jiangtong Li, Junyan Cao, Li Niu 0002, Liqing Zhang 0001 |
ACM Multimedia | 4 |
| 2023 | Deep Image Harmonization in Dual Color SpacesabstractImage harmonization is an essential step in image composition that adjusts the appearance of composite foreground to address the inconsistency between foreground and background. Existing methods primarily operate in correlated RGB color space, leading to entangled features and limited representation ability. In contrast, decorrelated color space (e.g., Lab) has decorrelated channels that provide disentangled color and illumination statistics. In this paper, we explore image harmonization in dual color spaces, which supplements entangled RGB features with disentangled L, a, b features to alleviate the workload in harmonization process. The network comprises a RGB harmonization backbone, an Lab encoding module, and an Lab control module. The backbone is a U-Net network translating composite image to harmonized image. Three encoders in Lab encoding module extract three control codes independently from L, a, b channels, which are used to manipulate the decoder features in harmonization backbone via Lab control module. Our code and model are available at https://github.com/bcmi/DucoNet-Image-Harmonization. Linfeng Tan, Jiangtong Li, Li Niu 0002, Liqing Zhang 0001 |
ACM Multimedia | 3 |
| 2023 | Scene-aware Human Pose Generation using TransformerabstractAffordance learning considers the interaction opportunities for an actor in the scene and thus has wide application in scene understanding and intelligent robotics. In this paper, we focus on contextual affordance learning, i.e., using affordance as context to generate a reasonable human pose in a scene. Existing scene-aware human pose generation methods could be divided into two categories depending on whether using pose templates. Our proposed method belongs to the template-based category, which benefits from the representative pose templates. Moreover, inspired by recent transformer-based methods, we associate each query embedding with a pose template, and use the interaction between query embeddings and scene feature map to effectively predict the scale and offsets for each pose template. In addition, we employ knowledge distillation to facilitate the offset learning given the predicted scale. Comprehensive experiments on Sitcom dataset demonstrate the effectiveness of our method. Jieteng Yao, Junjie Chen 0008, Li Niu 0002, Bin Sheng 0001 |
ACM Multimedia | 3 |
| 2023 | Natural Image Matting with Attended Global Context
Yiyi Zhang 0002, Li Niu 0002, Yasushi Makihara, Jianfu Zhang 0003, Weijie Zhao 0003, Yasushi Yagi, Liqing Zhang 0001 |
J. Comput. Sci. Technol. | 2 |
| 2023 | Diverse image inpainting with disentangled uncertainty
Wentao Wang 0009, Li Niu 0002, Jianfu Zhang 0003, Haoyu Ling, Liqing Zhang 0001 |
Pattern Recognit. | 3 |
| 2023 | From Pixel to Patch: Synthesize Context-Aware Features for Zero-Shot Semantic SegmentationabstractZero-shot learning (ZSL) has been actively studied for image classification tasks to relieve the burden of annotating image labels. Interestingly, the semantic segmentation task requires more labor-intensive pixel-wise annotation, but zero-shot semantic segmentation has not attracted extensive research interest. Thus, we focus on zero-shot semantic segmentation that aims to segment unseen objects with only category-level semantic representations provided for unseen categories. In this article, we propose a novel context-aware feature generation network (CaGNet) that can synthesize context-aware pixel-wise visual features for unseen categories based on category-level semantic representations and pixel-wise contextual information. The synthesized features are used to fine-tune the classifier to enable segmenting of unseen objects. Furthermore, we extend pixel-wise feature generation and fine-tuning to patch-wise feature generation and fine-tuning, which additionally considers the interpixel relationship. Experimental results on Pascal-VOC, Pascal-context, and COCO-stuff show that our method significantly outperforms the existing zero-shot semantic segmentation methods. Zhangxuan Gu, Li Niu 0002, Zihan Zhao 0001, Liqing Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Inharmonious Region Localization by Magnifying Domain DiscrepancyabstractInharmonious region localization aims to localize the region in a synthetic image which is incompatible with surrounding background. The inharmony issue is mainly attributed to the color and illumination inconsistency produced by image editing techniques. In this work, we tend to transform the input image to another color space to magnify the domain discrepancy between inharmonious region and background, so that the model can identify the inharmonious region more easily. To this end, we present a novel framework consisting of a color mapping module and an inharmonious region localization network, in which the former is equipped with a novel domain discrepancy magnification loss and the latter could be an arbitrary localization network. Extensive experiments on image harmonization dataset show the superiority of our designed framework. Jing Liang 0007, Li Niu 0002, Penghao Wu, Fengjun Guo |
AAAI | 2 |
| 2022 | Shadow Generation for Composite Image in Real-World ScenesabstractImage composition targets at inserting a foreground object into a background image. Most previous image composition methods focus on adjusting the foreground to make it compatible with background while ignoring the shadow effect of foreground on the background. In this work, we focus on generating plausible shadow for the foreground object in the composite image. First, we contribute a real-world shadow generation dataset DESOBA by generating synthetic composite images based on paired real images and deshadowed images. Then, we propose a novel shadow generation network SGRNet, which consists of a shadow mask prediction stage and a shadow filling stage. In the shadow mask prediction stage, foreground and background information are thoroughly interacted to generate foreground shadow mask. In the shadow filling stage, shadow parameters are predicted to fill the shadow area. Extensive experiments on our DESOBA dataset and real composite images demonstrate the effectiveness of our proposed method. Our dataset and code are available at https://github.com/bcmi/Object-Shadow-Generation- Dataset-DESOBA. Yan Hong 0001, Li Niu 0002, Jianfu Zhang 0003 |
AAAI | 2 |
| 2022 | Action-Aware Embedding Enhancement for Image-Text RetrievalabstractImage-text retrieval plays a central role in bridging vision and language, which aims to reduce the semantic discrepancy between images and texts. Most of existing works rely on refined words and objects representation through the data-oriented method to capture the word-object cooccurrence. Such approaches are prone to ignore the asymmetric action relation between images and texts, that is, the text has explicit action representation (i.e., verb phrase) while the image only contains implicit action information. In this paper, we propose Action-aware Memory-Enhanced embedding (AME) method for image-text retrieval, which aims to emphasize the action information when mapping the images and texts into a shared embedding space. Specifically, we integrate action prediction along with an action-aware memory bank to enrich the image and text features with action-similar text features. The effectiveness of our proposed AME method is verified by comprehensive experimental results on two benchmark datasets. Jiangtong Li, Li Niu 0002, Liqing Zhang 0001 |
AAAI | 2 |
| 2022 | Deep Image Harmonization by Bridging the Reality Gap
Junyan Cao, Wenyan Cong, Li Niu 0002, Jianfu Zhang 0003, Liqing Zhang 0001 |
BMVC | 3 |
| 2022 | Inharmonious Region Localization via Recurrent Self-Reasoning
Penghao Wu, Li Niu 0002, Jing Liang 0007, Liqing Zhang 0001 |
BMVC | 2 |
| 2022 | Inharmonious Region Localization with Auxiliary Style Feature
Penghao Wu, Li Niu 0002, Liqing Zhang 0001 |
BMVC | 2 |
| 2022 | Visible Watermark Removal with Dynamic Kernel and Semantic-aware Propagation
Xing Zhao 0010, Li Niu 0002, Liqing Zhang 0001 |
BMVC | 2 |
| 2022 | Weak-shot Semantic Segmentation by Transferring Semantic Affinity and Boundary
Li Niu 0002, Jianlou Si, Chen Qian 0006, Liqing Zhang 0001 |
BMVC | 2 |
| 2022 | High-Resolution Image Harmonization via Collaborative Dual TransformationsabstractGiven a composite image, image harmonization aims to adjust the foreground to make it compatible with the background. High-resolution image harmonization is in high demand, but still remains unexplored. Conventional image harmonization methods learn global RGB-to-RGB transformation which could effortlessly scale to high resolution, but ignore diverse local context. Recent deep learning methods learn the dense pixel-to-pixel transformation which could generate harmonious outputs, but are highly constrained in low resolution. In this work, we propose a high-resolution image harmonization network with Collaborative Dual Transformation (CDTNet) to combine pixel-to-pixel transformation and RGB-to-RGB transformation coherently in an end-to-end network. Our CDTNet consists of a low-resolution generator for pixel-to-pixel transformation, a color mapping module for RGB-to-RGB transformation, and a refinement module to take advantage of both. Extensive experiments on high-resolution bench-mark dataset and our created high-resolution real composite images demonstrate that our CDTNet strikes a good balance between efficiency and effectiveness. Our used datasets can be found in https://github.com/bcmi/CDTNet-High-Resolution-Image-Harmonization. Wenyan Cong, Xinhao Tao, Li Niu 0002, Jing Liang 0007, Xuesong Gao, Qihao Sun, Liqing Zhang 0001 |
CVPR | 3 |
| 2022 | From Representation to Reasoning: Towards both Evidence and Commonsense Reasoning for Video Question-AnsweringabstractVideo understanding has achieved great success in representation learning, such as video caption, video object grounding, and video descriptive question-answer. However, current methods still struggle on video reasoning, including evidence reasoning and commonsense reasoning. To facilitate deeper video understanding towards video reasoning, we present the task of Causal-VidQA, which includes four types of questions ranging from scene description (description) to evidence reasoning (explanation) and commonsense reasoning (prediction and counterfactual). For commonsense reasoning, we set up a two-step solution by answering the question and providing a proper reason. Through extensive experiments on existing VideoQA methods, we find that the state-of-the-art methods are strong in descriptions but weak in reasoning. We hope that Causal-VidQA can guide the research of video understanding from representation learning to deeper reasoning. The dataset and related resources are available at https://github.com/bcmi/Causal-VidQA.git. Jiangtong Li, Li Niu 0002, Liqing Zhang 0001 |
CVPR | 2 |
| 2022 | Dual-path Image Inpainting with Auxiliary GAN InversionabstractDeep image inpainting can inpaint a corrupted image using a feed-forward inference, but still fails to handle large missing area or complex semantics. Recently, GAN inversion based inpainting methods propose to leverage semantic information in pretrained generator (e.g., StyleGAN) to solve the above issues. Different from feed-forward methods, they seek for a closest latent code to the corrupted image and feed it to a pretrained generator. However, inferring the latent code is either time-consuming or inaccurate. In this paper, we develop a dual-path inpainting network with inversion path and feed-forward path, in which inversion path provides auxiliary information to help feed-forward path. We also design a novel deformable fusion module to align the feature maps in two paths. Experiments on FFHQ and LSUN demonstrate that our method is effective in solving the aforementioned problems while producing more realistic results than state-of-the-art methods. Wentao Wang 0009, Li Niu 0002, Jianfu Zhang 0003, Xue Yang 0005, Liqing Zhang 0001 |
CVPR | 2 |
| 2022 | DeltaGAN: Towards Diverse Few-Shot Image Generation with Sample-Specific Delta
Yan Hong 0001, Li Niu 0002, Jianfu Zhang 0003, Liqing Zhang 0001 |
ECCV (16) | 2 |
| 2022 | Human-Centric Image Cropping with Partition-Aware and Content-Preserving Features
Bo Zhang 0075, Li Niu 0002, Xing Zhao 0010, Liqing Zhang 0001 |
ECCV (7) | 2 |
| 2022 | Learning Object Placement via Dual-Path Graph Completion
Liu Liu 0022, Li Niu 0002, Liqing Zhang 0001 |
ECCV (17) | 3 |
| 2022 | Deep Video Harmonization With Color Mapping ConsistencyabstractVideo harmonization aims to adjust the foreground of a composite video to make it compatible with the background. So far, video harmonization has only received limited attention and there is no public dataset for video harmonization. In this work, we construct a new video harmonization dataset HYouTube by adjusting the foreground of real videos to create synthetic composite videos. Moreover, we consider the temporal consistency in video harmonization task. Unlike previous works which establish the spatial correspondence, we design a novel framework based on the assumption of color mapping consistency, which leverages the color mapping of neighboring frames to refine the current frame. Extensive experiments on our HYouTube dataset prove the effectiveness of our proposed framework. Our dataset and code are available at https://github.com/bcmi/Video-Harmonization-Dataset-HYouTube. Xinyuan Lu, Shengyuan Huang, Li Niu 0002, Wenyan Cong, Liqing Zhang 0001 |
IJCAI | 3 |
| 2022 | Few-shot Image Generation Using Discrete Content RepresentationabstractFew-shot image generation and few-shot image translation are two related tasks, both of which aim to generate new images for an unseen category with only a few images. In this work, we make the first attempt to adapt few-shot image translation method to few-shot image generation task. Few-shot image translation disentangles an image into style vector and content map. An unseen style vector can be combined with different seen content maps to produce different images. However, it needs to store seen images to provide content maps and the unseen style vector may be incompatible with seen content maps. To adapt it to few-shot image generation task, we learn a compact dictionary of local content vectors via quantizing continuous content maps into discrete content maps instead of storing seen images. Furthermore, we model the autoregressive distribution of discrete content map conditioned on style vector, which can alleviate the incompatibility between content map and style vector. Qualitative and quantitative results on three real datasets demonstrate that our model can produce images of higher diversity and fidelity for unseen categories than previous methods. Yan Hong 0001, Li Niu 0002, Jianfu Zhang 0003, Liqing Zhang 0001 |
ACM Multimedia | 2 |
| 2022 | Multi-Level Region Matching for Fine-Grained Sketch-Based Image RetrievalabstractFine-Grained Sketch-Based Image Retrieval (FG-SBIR) is to use free-hand sketches as queries to perform instance-level retrieval in an image gallery. Existing works usually leverage only high-level information and perform matching in a single region. However, both low-level and high-level information are helpful to establish fine-grained correspondence. Besides, we argue that matching different regions between each sketch-image pair can further boost model robustness. Therefore, we propose Multi-Level Region Matching (MLRM) for FG-SBIR, which consists of two modules: a Discriminative Region Extraction module (DRE) and a Region and Level Attention module (RLA). In DRE, we propose Light-weighted Attention Map Augmentation (LAMA) to extract local feature from different regions. In RLA, we propose a transformer-based attentive matching module to learn attention weights to explore different importance from different image/sketch regions and feature levels. Furthermore, to ensure that the geometrical and semantic distinctiveness is well modeled, we also explore a novel LAMA overlapping penalty and a local region-negative triplet loss in our proposed MLRM method. Comprehensive experiments conducted on five datasets (i.e., Sketchy, QMUL-ChairV2, QMUL-ShoeV2, QMUL-Chair, QMUL-Shoe) demonstrate effectiveness of our method. Zhixin Ling, Jiangtong Li, Li Niu 0002 |
ACM Multimedia | 4 |
| 2022 | Weak-shot Semantic Segmentation via Dual Similarity TransferabstractSemantic segmentation is a practical and active task, but severely suffers from the expensive cost of pixel-level labels when extending to more classes in wider applications. To this end, we focus on the problem named weak-shot semantic segmentation, where the novel classes are learnt from cheaper image-level labels with the support of base classes having off-the-shelf pixel-level labels. To tackle this problem, we propose a dual similarity transfer framework, which is built upon MaskFormer to disentangle the semantic segmentation task into single-label classification and binary segmentation for each proposal. Specifically, the binary segmentation sub-task allows proposal-pixel similarity transfer from base classes to novel classes, which enables the mask learning of novel classes. We also learn pixel-pixel similarity from base classes and distill such class-agnostic semantic similarity to the semantic masks of novel classes, which regularizes the segmentation model with pixel-level semantic relationship across images. In addition, we propose a complementary loss to facilitate the learning of novel classes. Comprehensive experiments on the challenging COCO-Stuff-10K and ADE20K datasets demonstrate the effectiveness of our method. Junjie Chen 0008, Li Niu 0002, Jianlou Si, Chen Qian 0006, Liqing Zhang 0001 |
NeurIPS | 2 |
| 2022 | UniGAN: Reducing Mode Collapse in GANs using a Uniform GeneratorabstractDespite the significant progress that has been made in the training of Generative Adversarial Networks (GANs), the mode collapse problem remains a major challenge in training GANs, which refers to a lack of diversity in generative samples. In this paper, we propose a new type of generative diversity named uniform diversity, which relates to a newly proposed type of mode collapse named $u$-mode collapse where the generative samples distribute nonuniformly over the data manifold. From a geometric perspective, we show that the uniform diversity is closely related with the generator uniformity property, and the maximum uniform diversity is achieved if the generator is uniform. To learn a uniform generator, we propose UniGAN, a generative framework with a Normalizing Flow based generator and a simple yet sample efficient generator uniformity regularization, which can be easily adapted to any other generative framework. A new type of diversity metric named udiv is also proposed to estimate the uniform diversity given a set of generative samples in practice. Experimental results verify the effectiveness of our UniGAN in learning a uniform generator and improving uniform diversity. Li Niu 0002, Liqing Zhang 0001 |
NeurIPS | 2 |
| 2022 | Zero-shot sketch-based image retrieval with structure-aware asymmetric disentanglement
Jiangtong Li, Zhixin Ling, Li Niu 0002, Liqing Zhang 0001 |
Comput. Vis. Image Underst. | 3 |
| 2022 | Hallucinating uncertain motion and future for static image action recognition
Li Niu 0002, Shengyuan Huang, Xing Zhao 0010, Liwei Kang, Yiyi Zhang 0002, Liqing Zhang 0001 |
Comput. Vis. Image Underst. | 1 |
| 2022 | FaceEngage: Robust Estimation of Gameplay Engagement from User-Contributed (YouTube) VideosabstractMeasuring user engagement in interactive tasks can facilitate numerous applications toward optimizing user experience, ranging from eLearning to gaming. However, a significant challenge is the lack of non-contact engagement estimation methods that are robust in unconstrained environments. We present FaceEngage, a non-intrusive engagement estimator leveraging user facial recordings during actual gameplay in naturalistic conditions. Our contributions are three-fold. First, we show the potential of using front-facing videos as training data to build the engagement estimator. We compile FaceEngage Dataset with over 700 picture-in-picture, realisitic, and user-contributed YouTube gaming videos (i.e., with both full-screen game scenes and time-synchronized user facial recordings in subwindows). Second, we develop FaceEngage system, that captures relevant gamer facial features from front-facing recordings to infer task engagement. We implement two FaceEngage pipelines: an estimator trained on user facial motion features inspired by prior psychological works, and a deep learning-enabled estimator. Lastly, we conduct extensive experiments and conclude: (i) certain user facial motion cues (e.g., blink rates, head movements) are engagement-indicative; (ii) our deep learning-enabled FaceEngage pipeline can automatically extract more informative features, outperforming the facial motion feature-based pipeline; (iii) FaceEngage is robust to various video lengths, users/game genres and interpretable. Despite the challenging nature of realistic videos, FaceEngage attains the accuracy of 83.8 percent and leave-one-user-out precision of 79.9 percent, both of which are superior to our face motion-based model. Xu Chen 0011, Li Niu 0002, Ashok Veeraraghavan, Ashutosh Sabharwal |
IEEE Trans. Affect. Comput. | 2 |
| 2021 | Activity Image-to-Video Retrieval by Disentangling Appearance and MotionabstractWith the rapid emergence of video data, image-to-video retrieval has attracted much attention. There are two types of image-to-video retrieval: instance-based and activity-based. The former task aims to retrieve videos containing the same main objects as the query image, while the latter focuses on finding the similar activity. Since dynamic information plays a significant role in the video, we pay attention to the latter task to explore the motion relation between images and videos. In this paper, we propose a Motion-assisted Activity Proposal-based Image-to-Video Retrieval (MAP-IVR) approach to disentangle the video features into motion features and appearance features and obtain appearance features from the images. Then, we perform image-to-video translation to improve the disentanglement quality. The retrieval is performed in both appearance and video feature spaces. Extensive experiments demonstrate that our MAP-IVR approach remarkably outperforms the state-of-the-art approaches on two benchmark activity-based video datasets. Liu Liu 0022, Jiangtong Li, Li Niu 0002, Ruicong Xu, Liqing Zhang 0001 |
AAAI | 3 |
| 2021 | Disentangled Information BottleneckabstractThe information bottleneck (IB) method is a technique for extracting information that is relevant for predicting the target random variable from the source random variable, which is typically implemented by optimizing the IB Lagrangian that balances the compression and prediction terms. However, the IB Lagrangian is hard to optimize, and multiple trials for tuning values of Lagrangian multiplier are required. Moreover, we show that the prediction performance strictly decreases as the compression gets stronger during optimizing the IB Lagrangian. In this paper, we implement the IB method from the perspective of supervised disentangling. Specifically, we introduce Disentangled Information Bottleneck (DisenIB) that is consistent on compressing source maximally without target prediction performance loss (maximum compression). Theoretical and experimental results demonstrate that our method is consistent on maximum compression, and performs well in terms of generalization, robustness to adversarial attack, out-of-distribution detection, and supervised disentangling. Li Niu 0002, Jianfu Zhang 0003, Liqing Zhang 0001 |
AAAI | 2 |
| 2021 | Depth Privileged Object Detection in Indoor Scenes via Deformation HallucinationabstractRGB-D object detection has achieved significant advance, because depth provides complementary geometric information to RGB images. Considering depth images are unavailable in some scenarios, we focus on depth privileged object detection in indoor scenes, where the depth images are only available in the training phase. Under this setting, one prevalent research line is modality hallucination, in which depth image and depth feature are the common choices for hallucinating. In contrast, we choose to hallucinate depth deformation, which is explicit geometric information and efficient to hallucinate. Specifically, we employ the deformable convolution layer with augmented offsets as our deformation module and regard the offsets as geometric deformation, because the offsets enable flexibly sampling over the object and transforming to a canonical shape for ease of detection. In addition, we design a quality-based mechanism to avoid negative transfer of depth deformation. Experimental results and analyses on NYUDv2 and SUN RGB-D demonstrate the effectiveness of our method against the state-of-the-art methods for depth privileged object detection. Junjie Chen 0008, Li Niu 0002, Liqing Zhang 0001 |
AAAI | 4 |
| 2021 | Image Composition Assessment with Saliency-augmented Multi-pattern Pooling
Bo Zhang 0075, Li Niu 0002, Liqing Zhang 0001 |
BMVC | 2 |
| 2021 | Parallel Multi-Resolution Fusion Network for Image InpaintingabstractConventional deep image inpainting methods are based on auto-encoder architecture, in which the spatial details of images will be lost in the down-sampling process, leading to the degradation of generated results. Also, the structure information in deep layers and texture information in shallow layers of the auto-encoder architecture can not be well integrated. Differing from the conventional image inpainting architecture, we design a parallel multi-resolution inpainting network with multi-resolution partial convolution, in which low-resolution branches focus on the global structure while high-resolution branches focus on the local texture details. All these high- and low-resolution streams are in parallel and fused repeatedly with multi-resolution masked representation fusion so that the reconstructed images are semantically robust and textually plausible. Experimental results show that our method can effectively fuse structure and texture information, producing more realistic results than state-of-the-art methods. Wentao Wang 0009, Jianfu Zhang 0003, Li Niu 0002, Haoyu Ling, Xue Yang 0005, Liqing Zhang 0001 |
ICCV | 3 |
| 2021 | Inharmonious Region LocalizationabstractThe advance of image editing techniques allows users to create artistic works, but the manipulated regions may be incompatible with the background. Localizing the inharmonious region is an appealing yet challenging task. Realizing that this task requires effective aggregation of multi-scale contextual information and suppression of redundant information, we design novel Bi-directional Feature Integration (BFI) block and Global-context Guided Decoder (GGD) block to fuse multi-scale features in the encoder and decoder respectively. We also employ Mask-guided Dual Attention (MDA) block between the encoder and decoder to suppress the redundant information. Experiments on the image harmonization dataset demonstrate that our method achieves competitive performance for inharmonious region localization. The source code is available at https://github.com/bcmi/DIRL. Jing Liang 0007, Li Niu 0002, Liqing Zhang 0001 |
ICME | 2 |
| 2021 | Bargainnet: Background-Guided Domain Translation for Image HarmonizationabstractGiven a composite image with inharmonious foreground and background, image harmonization aims to adjust the foreground to make it compatible with the background. Previous image harmonization methods mainly focus on learning the mapping from composite image to real image, while ignoring the crucial guidance role that background plays. In this work, we formulate image harmonization task as background-guided domain translation. Specifically, we use a domain code extractor to capture the background domain information to guide the foreground harmonization, which is regulated by well-tailored triplet losses. Extensive experiments on the benchmark dataset demonstrate the effectiveness of our proposed method. Code is available at https://github.com/bcmi/BargainNet. Wenyan Cong, Li Niu 0002, Jianfu Zhang 0003, Jing Liang 0007, Liqing Zhang 0001 |
ICME | 2 |
| 2021 | Static Image Action Recognition with Hallucinated Fine-Grained Motion InformationabstractStatic image action recognition is a challenging task due to the lack of motion information in a static image. Some previous works have attempted to hallucinate the motion information in a static image using a generator learnt from freely available unlabeled videos. However, their hallucinated motion information is either low-level or coarse-grained, which may contain lots of noise or lose motion details. In contrast, we propose to hallucinate fine-grained high-level motion information, which is more robust and detail-preserving. Specifically, we hallucinate motion feature map which encodes the motion details of human body parts. We also hallucinate motion attention map to focus on motion-related regions. Our hallucinated motion information can greatly facilitate static image action recognition, which is confirmed by the experiments on two static action image datasets and two video datasets. Shengyuan Huang, Xing Zhao 0010, Li Niu 0002, Liqing Zhang 0001 |
ICME | 3 |
| 2021 | End-to-End Video Object Detection with Spatial-Temporal TransformersabstractRecently, DETR and Deformable DETR have been proposed to eliminate the need for many hand-designed components in object detection while demonstrating good performance as previous complex hand-crafted detectors. However, their performance on Video Object Detection (VOD) has not been well explored. In this paper, we present TransVOD, an end-to-end video object detection model based on a spatial-temporal Transformer architecture. The goal of this paper is to streamline the pipeline of VOD, effectively removing the need for many hand-crafted components for feature aggregation, e.g., optical flow, recurrent neural networks, relation networks. Besides, benefited from the object query design in DETR, our method does not need complicated post-processing methods such as Seq-NMS or Tubelet rescoring, which keeps the pipeline simple and clean. In particular, we present temporal Transformer to aggregate both the spatial object queries and the feature memories of each frame. Our temporal Transformer consists of three components: Temporal Deformable Transformer Encoder (TDTE) to encode the multiple frame spatial details, Temporal Query Encoder (TQE) to fuse object queries, and Temporal Deformable Transformer Decoder (TDTD) to obtain current frame detection results. These designs boost the strong baseline deformable DETR by a significant margin (3%-4% mAP) on the ImageNet VID dataset. TransVOD yields comparable results performance on the benchmark of ImageNet VID. We hope our TransVOD can provide a new perspective for video object detection. Qianyu Zhou 0001, Xiangtai Li, Li Niu 0002, Yunhai Tong, Lizhuang Ma, Liqing Zhang 0001 |
ACM Multimedia | 4 |
| 2021 | Video Semantic Segmentation via Sparse Temporal TransformerabstractCurrently, video semantic segmentation mainly faces two challenges: 1) the demand of temporal consistency; 2) the balance between segmentation accuracy and inference efficiency. For the first challenge, existing methods usually use optical flow to capture the temporal relation in consecutive frames and maintain the temporal consistency, but the low inference speed by means of optical flow limits the real-time applications. For the second challenge, flow based key frame warping is one mainstream solution. However, the unbalanced inference latency of flow-based key frame warping makes it unsatisfactory for real-time applications. Considering the segmentation accuracy and inference efficiency, we propose a novel Sparse Temporal Transformer (STT) to bridge temporal relation among video frames adaptively, which is also equipped with query selection and key selection. The key selection and query selection strategies are separately applied to filter out temporal and spatial redundancy in our temporal transformer. Specifically, our STT can reduce the time complexity of temporal transformer by a large margin without harming the segmentation accuracy and temporal consistency. Experiments on two benchmark datasets, Cityscapes and Camvid, demonstrate that our method achieves the state-of-the-art segmentation accuracy and temporal consistency with comparable inference speed. Jiangtong Li, Wentao Wang 0009, Junjie Chen 0008, Li Niu 0002, Jianlou Si, Chen Qian 0006, Liqing Zhang 0001 |
ACM Multimedia | 4 |
| 2021 | Visible Watermark Removal via Self-calibrated Localization and Background RefinementabstractSuperimposing visible watermarks on images provides a powerful weapon to cope with the copyright issue. Watermark removal techniques, which can strengthen the robustness of visible watermarks in an adversarial way, have attracted increasing research interest. Modern watermark removal methods perform watermark localization and background restoration simultaneously, which could be viewed as a multi-task learning problem. However, existing approaches suffer from incomplete detected watermark and degraded texture quality of restored background. Therefore, we design a two-stage multi-task network to address the above issues. The coarse stage consists of a watermark branch and a background branch, in which the watermark branch self-calibrates the roughly estimated mask and passes the calibrated mask to background branch to reconstruct the watermarked area. In the refinement stage, we integrate multi-level features to improve the texture quality of watermarked area. Extensive experiments on two datasets demonstrate the effectiveness of our proposed method. Jing Liang 0007, Li Niu 0002, Fengjun Guo, Liqing Zhang 0001 |
ACM Multimedia | 2 |
| 2021 | Weak-shot Fine-grained Classification via Similarity TransferabstractRecognizing fine-grained categories remains a challenging task, due to the subtle distinctions among different subordinate categories, which results in the need of abundant annotated samples. To alleviate the data-hungry problem, we consider the problem of learning novel categories from web data with the support of a clean set of base categories, which is referred to as weak-shot learning. In this setting, we propose a method called SimTrans to transfer pairwise semantic similarity from base categories to novel categories. Specifically, we firstly train a similarity net on clean data, and then leverage the transferred similarity to denoise web training data using two simple yet effective strategies. In addition, we apply adversarial loss on similarity net to enhance the transferability of similarity. Comprehensive experiments demonstrate the effectiveness of our weak-shot setting and our SimTrans method. Junjie Chen 0008, Li Niu 0002, Liu Liu 0022, Liqing Zhang 0001 |
NeurIPS | 2 |
| 2021 | Mixed Supervised Object Detection by Transferring Mask Prior and Semantic SimilarityabstractObject detection has achieved promising success, but requires large-scale fully-annotated data, which is time-consuming and labor-extensive. Therefore, we consider object detection with mixed supervision, which learns novel object categories using weak annotations with the help of full annotations of existing base object categories. Previous works using mixed supervision mainly learn the class-agnostic objectness from fully-annotated categories, which can be transferred to upgrade the weak annotations to pseudo full annotations for novel categories. In this paper, we further transfer mask prior and semantic similarity to bridge the gap between novel categories and base categories. Specifically, the ability of using mask prior to help detect objects is learned from base categories and transferred to novel categories. Moreover, the semantic similarity between objects learned from base categories is transferred to denoise the pseudo full annotations for novel categories. Experimental results on three benchmark datasets demonstrate the effectiveness of our method over existing methods. Codes are available at https://github.com/bcmi/TraMaS-Weak-Shot-Object-Detection. Li Niu 0002, Junjie Chen 0008, Liqing Zhang 0001 |
NeurIPS | 3 |
| 2021 | Depth Privileged Scene Recognition via Dual Attention HallucinationabstractRGB-D scene recognition has achieved promising performance because depth could provide complementary geometric information to RGB images. However, the inaccessibility of depth sensors severely limits RGB-D applications. In this paper, we focus on depth privileged setting, in which depth information is only available during training but not available during testing. Considering that the information obtained from RGB and depth images are complementary while attention is informative and transferable, our idea is using RGB input to hallucinate depth attention. We build our model upon modulated deformable convolutional layer and hallucinate dual attention: post-hoc importance weight and trainable spatial transformation. Specifically, we use modulation (resp., offset) learned from RGB to mimic Grad-CAM (resp., offset) learned from depth, to combine the strength of dual attention. We also design a weighted loss to avoid negative transfer according to the quality of depth attention. Extensive experiments on two benchmarks, i.e., SUN RGB-D and NYUDv2, demonstrate that our method outperforms the state-of-the-art methods for depth privileged scene recognition. Junjie Chen 0008, Li Niu 0002, Liqing Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Memorize, Associate and Match: Embedding Enhancement via Fine-Grained Alignment for Image-Text RetrievalabstractImage-text retrieval aims to capture the semantic correlation between images and texts. Existing image-text retrieval methods can be roughly categorized into embedding learning paradigm and pair-wise learning paradigm. The former paradigm fails to capture the fine-grained correspondence between images and texts. The latter paradigm achieves fine-grained alignment between regions and words, but the high cost of pair-wise computation leads to slow retrieval speed. In this paper, we propose a novel method named MEMBER by using Memory-based EMBedding Enhancement for image-text Retrieval (MEMBER), which introduces global memory banks to enable fine-grained alignment and fusion in embedding learning paradigm. Specifically, we enrich image (resp., text) features with relevant text (resp., image) features stored in the text (resp., image) memory bank. In this way, our model not only accomplishes mutual embedding enhancement across two modalities, but also maintains the retrieval efficiency. Extensive experiments demonstrate that our MEMBER remarkably outperforms state-of-the-art approaches on two large-scale benchmark datasets. Jiangtong Li, Liu Liu 0022, Li Niu 0002, Liqing Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Person Re-Identification With Reinforced Attribute Attention SelectionabstractPerson re-identification (Re-ID) aims to match pedestrian images across various scenes in video surveillance. There are a few works using attribute information to boost Re-ID performance. Specifically, those methods leverage attribute information to boost Re-ID performance by introducing auxiliary tasks like verifying the image level attribute information of two pedestrian images or recognizing identity level attributes. Identity level attribute annotations cost less manpower and are well-fitted for person re-identification task compared with image-level attribute annotations. However, the identity attribute information may be very noisy due to incorrect attribute annotation or lack of discriminativeness to distinguish different persons, which is probably unhelpful for the Re-ID task. In this paper, we propose a novel Attribute Attentional Block (AAB), which can be integrated into any backbone network or framework. Our AAB adopts reinforcement learning to drop noisy attributes based on our designed reward and then utilizes aggregated attribute attention of the remaining attributes to facilitate the Re-ID task. Experimental results demonstrate that our proposed method achieves state-of-the-art results on three benchmark datasets. Jianfu Zhang 0003, Li Niu 0002, Liqing Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Hard Pixel Mining for Depth Privileged Semantic SegmentationabstractSemantic segmentation has achieved remarkable progress but remains challenging due to the complex scene, object occlusion, and so on. Some research works have attempted to use extra information such as a depth map to help RGB based semantic segmentation because the depth map could provide complementary geometric cues. However, due to the inaccessibility of depth sensors, depth information is usually unavailable for the test images. In this paper, we leverage only the depth of training images as the privileged information to mine the hard pixels in semantic segmentation, in which depth information is only available for training images but not available for test images. Specifically, we propose a novel Loss Weight Module, which outputs a loss weight map by employing two depth-related measurements of hard pixels: Depth Prediction Error and Depth-aware Segmentation Error. The loss weight map is then applied to segmentation loss, with the goal of learning a more robust model by paying more attention to the hard pixels. Besides, we also explore a curriculum learning strategy based on the loss weight map. Meanwhile, to fully mine the hard pixels on different scales, we apply our loss weight module to multi-scale side outputs. Our hard pixels mining method achieves the state-of-the-art results on three benchmark datasets, and even outperforms the methods which need depth input during testing. Zhangxuan Gu, Li Niu 0002, Haohua Zhao 0001, Liqing Zhang 0001 |
IEEE Trans. Multim. | 2 |
| 2020 | Image Cropping with Composition and Saliency Aware Aesthetic Score MapabstractAesthetic image cropping is a practical but challenging task which aims at finding the best crops with the highest aesthetic quality in an image. Recently, many deep learning methods have been proposed to address this problem, but they did not reveal the intrinsic mechanism of aesthetic evaluation. In this paper, we propose an interpretable image cropping model to unveil the mystery. For each image, we use a fully convolutional network to produce an aesthetic score map, which is shared among all candidate crops during crop-level aesthetic evaluation. Then, we require the aesthetic score map to be both composition-aware and saliency-aware. In particular, the same region is assigned with different aesthetic scores based on its relative positions in different crops. Moreover, a visually salient region is supposed to have more sensitive aesthetic scores so that our network can learn to place salient objects at more proper positions. Such an aesthetic score map can be used to localize aesthetically important regions in an image, which sheds light on the composition rules learned by our model. We show the competitive performance of our model in the image cropping task on several benchmark datasets, and also demonstrate its generality in real-world applications. Li Niu 0002, Weijie Zhao 0003, Dawei Cheng, Liqing Zhang 0001 |
AAAI | 2 |
| 2020 | A Proposal-Based Approach for Activity Image-to-Video RetrievalabstractActivity image-to-video retrieval task aims to retrieve videos containing the similar activity as the query image, which is a challenging task because videos generally have many background segments irrelevant to the activity. In this paper, we utilize R-C3D model to represent a video by a bag of activity proposals, which can filter out background segments to some extent. However, there are still noisy proposals in each bag. Thus, we propose an Activity Proposal-based Image-to-Video Retrieval (APIVR) approach, which incorporates multi-instance learning into cross-modal retrieval framework to address the proposal noise issue. Specifically, we propose a Graph Multi-Instance Learning (GMIL) module with graph convolutional layer, and integrate this module with classification loss, adversarial loss, and triplet loss in our cross-modal retrieval framework. Moreover, we propose geometry-aware triplet loss based on point-to-subspace distance to preserve the structural information of activity proposals. Extensive experiments on three widely-used datasets verify the effectiveness of our approach. Ruicong Xu, Li Niu 0002, Jianfu Zhang 0003, Liqing Zhang 0001 |
AAAI | 2 |
| 2020 | Exploiting Motion Information from Unlabeled Videos for Static Image Action RecognitionabstractStatic image action recognition, which aims to recognize action based on a single image, usually relies on expensive human labeling effort such as adequate labeled action images and large-scale labeled image dataset. In contrast, abundant unlabeled videos can be economically obtained. Therefore, several works have explored using unlabeled videos to facilitate image action recognition, which can be categorized into the following two groups: (a) enhance visual representations of action images with a designed proxy task on unlabeled videos, which falls into the scope of self-supervised learning; (b) generate auxiliary representations for action images with the generator learned from unlabeled videos. In this paper, we integrate the above two strategies in a unified framework, which consists of Visual Representation Enhancement (VRE) module and Motion Representation Augmentation (MRA) module. Specifically, the VRE module includes a proxy task which imposes pseudo motion label constraint and temporal coherence constraint on unlabeled videos, while the MRA module could predict the motion information of a static action image by exploiting unlabeled videos. We demonstrate the superiority of our framework based on four benchmark human action datasets with limited labeled data. Yiyi Zhang 0002, Li Niu 0002, Meichao Luo, Jianfu Zhang 0003, Dawei Cheng, Liqing Zhang 0001 |
AAAI | 2 |
| 2020 | DoveNet: Deep Image Harmonization via Domain VerificationabstractImage composition is an important operation in image processing, but the inconsistency between foreground and background significantly degrades the quality of composite image. Image harmonization, aiming to make the foreground compatible with the background, is a promising yet challenging task. However, the lack of high-quality publicly available dataset for image harmonization greatly hinders the development of image harmonization techniques. In this work, we contribute an image harmonization dataset iHarmony4 by generating synthesized composite images based on COCO (resp., Adobe5k, Flickr, day2night) dataset, leading to our HCOCO (resp., HAdobe5k, HFlickr, Hday2night) sub-dataset. Moreover, we propose a new deep image harmonization method DoveNet using a novel domain verification discriminator, with the insight that the foreground needs to be translated to the same domain as background. Extensive experiments on our constructed dataset demonstrate the effectiveness of our proposed method. Our dataset and code are available at https://github.com/bcmi/Image_Harmonization_Datasets. Wenyan Cong, Jianfu Zhang 0003, Li Niu 0002, Liu Liu 0022, Zhixin Ling, Weiyuan Li, Liqing Zhang 0001 |
CVPR | 3 |
| 2020 | Learning From Web Data With Self-Organizing Memory ModuleabstractLearning from web data has attracted lots of research interest in recent years. However, crawled web images usually have two types of noises, label noise and background noise, which induce extra difficulties in utilizing them effectively. Most existing methods either rely on human supervision or ignore the background noise. In this paper, we propose a novel method, which is capable of handling these two types of noises together, without the supervision of clean images in the training stage. Particularly, we formulate our method under the framework of multi-instance learning by grouping ROIs (i.e., images and their region proposals) from the same category into bags. ROIs in each bag are assigned with different weights based on the representative/discriminative scores of their nearest clusters, in which the clusters and their scores are obtained via our designed memory module. Our memory module could be naturally integrated with the classification module, leading to an end-to-end trainable system. Extensive experiments on four benchmark datasets demonstrate the effectiveness of our method. Li Niu 0002, Junjie Chen 0008, Dawei Cheng, Liqing Zhang 0001 |
CVPR | 2 |
| 2020 | Beyond Without Forgetting: Multi-Task Learning for Classification with Disjoint DatasetsabstractMulti-task Learning (MTL) for classification with disjoint datasets aims to explore MTL when one task only has one labeled dataset. In existing methods, for each task, the unlabeled datasets are not fully exploited to facilitate this task. Inspired by semi-supervised learning, we use unlabeled datasets with pseudo labels to facilitate each task. However, there are two major issues: 1) the pseudo labels are very noisy; 2) the unlabeled datasets and the labeled dataset for each task has considerable data distribution mismatch. To address these issues, we propose our MTL with Selective Augmentation (MTL-SA) method to select the training samples in unlabeled datasets with confident pseudo labels and close data distribution to the labeled dataset. Then, we use the selected training samples to add information and use the remaining training samples to preserve information. Extensive experiments on face-centric and human-centric applications demonstrate the effectiveness of our MTL-SA method. Yan Hong 0001, Li Niu 0002, Jianfu Zhang 0003, Liqing Zhang 0001 |
ICME | 2 |
| 2020 | Matchinggan: Matching-Based Few-Shot Image GenerationabstractTo generate new images for a given category, most deep generative models require abundant training images from this category, which are often too expensive to acquire. To achieve the goal of generation based on only a few images, we propose matching-based Generative Adversarial Network (GAN) for few-shot generation, which includes a matching generator and a matching discriminator. Matching generator can match random vectors with a few conditional images from the same category and generate new images for this category based on the fused features. The matching discriminator extends conventional GAN discriminator by matching the feature of generated image with the fused feature of conditional images. Extensive experiments on three datasets demonstrate the effectiveness of our proposed method. Yan Hong 0001, Li Niu 0002, Jianfu Zhang 0003, Liqing Zhang 0001 |
ICME | 2 |
| 2020 | Context-aware Feature Generation For Zero-shot Semantic SegmentationabstractExisting semantic segmentation models heavily rely on dense pixel-wise annotations. To reduce the annotation pressure, we focus on a challenging task named zero-shot semantic segmentation, which aims to segment unseen objects with zero annotations. This task can be accomplished by transferring knowledge across categories via semantic word embeddings. In this paper, we propose a novel context-aware feature generation method for zero-shot segmentation named CaGNet. In particular, with the observation that a pixel-wise feature highly depends on its contextual information, we insert a contextual module in a segmentation network to capture the pixel-wise contextual information, which guides the process of generating more diverse and context-aware features from semantic word embeddings. Our method achieves state-of-the-art results on three benchmark datasets for zero-shot segmentation. Zhangxuan Gu, Li Niu 0002, Zihan Zhao 0001, Liqing Zhang 0001 |
ACM Multimedia | 3 |
| 2020 | F2GAN: Fusing-and-Filling GAN for Few-shot Image GenerationabstractIn order to generate images for a given category, existing deep generative models generally rely on abundant training images. However, extensive data acquisition is expensive and fast learning ability from limited data is necessarily required in real-world applications. Also, these existing methods are not well-suited for fast adaptation to a new category. Few-shot image generation, aiming to generate images from only a few images for a new category, has attracted some research interest. In this paper, we propose a Fusing-and-Filling Generative Adversarial Network (F2GAN) to generate realistic and diverse images for a new category with only a few images. In our F2GAN, a fusion generator is designed to fuse the high-level features of conditional images with random interpolation coefficients, and then fills in attended low-level details with non-local attention module to produce a new image. Moreover, our discriminator can ensure the diversity of generated images by a mode seeking loss and an interpolation regression loss. Extensive experiments on five datasets demonstrate the effectiveness of our proposed method for few-shot image generation. Yan Hong 0001, Li Niu 0002, Jianfu Zhang 0003, Weijie Zhao 0003, Liqing Zhang 0001 |
ACM Multimedia | 2 |
| 2020 | Adversarial Query-by-Image Video Retrieval Based on Attention Mechanism
Ruicong Xu, Li Niu 0002, Liqing Zhang 0001 |
MMM (1) | 2 |
| 2020 | Multi-mode neural network for human action recognitionabstractVideo data are of two different intrinsic modes, in‐frame and temporal. It is beneficial to incorporate static in‐frame features to acquire dynamic features for video applications. However, some existing methods such as recurrent neural networks do not have a good performance, and some other such as 3D convolutional neural networks (CNNs) are both memory consuming and time consuming. This study proposes an effective framework that takes the advantage of deep learning on the static image feature extraction to tackle the video data. After extracting in‐frame feature vectors using a pretrained deep network, the authors integrate them and form a multi‐mode feature matrix, which preserves the multi‐mode structure and high‐level representation. They propose two models for follow‐up classification. The authors first introduce a temporal CNN, which directly feeds the multi‐mode feature matrix into a CNN. However, they show that characteristics of the multi‐mode features differ significantly in distinct modes. The authors therefore further propose the multi‐mode neural network (MMNN), in which different modes deploy different types of layers. They evaluate their algorithm with the task of human action recognition. The experimental results show that the MMNN achieves a much better performance than the existing long short‐term memory‐based methods and consumes far fewer resources than the existing 3D end‐to‐end models. Haohua Zhao 0001, Weichen Xue, Zhangxuan Gu, Li Niu 0002, Liqing Zhang 0001 |
IET Comput. Vis. | 5 |
| 2019 | Multi-person 3D Pose Estimation from Monocular Image Sequences
Nayun Xu, Xutong Lu, Yucheng Xing, Haohua Zhao 0001, Li Niu 0002, Liqing Zhang 0001 |
ICONIP (2) | 6 |
| 2019 | GAIN: Gradient Augmented Inpainting Network for Irregular HolesabstractImage inpainting, which aims to fill the missing holes of the images, is a challenging task because the holes may contain complicated structures or different possible layouts. Deep learning methods have shown promising performance in image inpainting but still, suffer from generating poor-structured artifacts when the holes are large and irregular. Some existing methods use edge inpainting to help image inpainting, with binary edge map obtained from image gradient. However, by only using the binary edge map, these methods discard the rich information in image gradient and thus leave some critical issues (e.g. , color discrepancy) unattended. In this paper, we propose Gradient Augmented Inpainting Network (GAIN), which uses image gradient information instead of edge information to facilitate image inpainting. Specifically, we formulate a multi-task learning framework which performs image inpainting and gradient inpainting simultaneously. A novel GAI-Block is designed to encourage the information fusion between the image feature map and the gradient feature map. Moreover, gradient information is also used to determine the filling priority, which can guide the network to construct more plausible semantic structures for the holes. Experimental results on public datasets CelebA-HQ and Places2 show that our proposed method outperforms state-of-the-art methods quantitatively and qualitatively. Jianfu Zhang 0003, Li Niu 0002, Dexin Yang, Liwei Kang, Yaoyi Li, Weijie Zhao 0003, Liqing Zhang 0001 |
ACM Multimedia | 2 |
| 2019 | Zero-Shot Learning via Category-Specific Visual-Semantic Mapping and Label RefinementabstractZero-Shot Learning (ZSL) aims to classify a test instance from an unseen category based on the training instances from seen categories, in which the gap between seen categories and unseen categories is generally bridged via visual-semantic mapping between the low-level visual feature space and the intermediate semantic space. However, the visual-semantic mapping (i.e., projection) learnt based on seen categories may not generalize well to unseen categories, which is known as the projection domain shift in ZSL. To address this projection domain shift issue, we propose a method named Adaptive Embedding ZSL (AEZSL) to learn an adaptive visual-semantic mapping for each unseen category, followed by progressive label refinement. Moreover, to avoid learning visual-semantic mapping for each unseen category in the large-scale classification task, we additionally propose a deep adaptive embedding model named Deep AEZSL (DAEZSL) sharing the similar idea (i.e., visual-semantic mapping should be category-specific and related to the semantic space) with AEZSL, which only needs to be trained once, but can be applied to arbitrary number of unseen categories. Extensive experiments demonstrate that our proposed methods achieve the state-of-theart results for image classification on three small-scale benchmark datasets and one large-scale benchmark dataset. Li Niu 0002, Jianfei Cai 0001, Ashok Veeraraghavan, Liqing Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Learning From Noisy Web Data With Category-Level SupervisionabstractLearning from web data is increasingly popular due to abundant free web resources. However, the performance gap between webly supervised learning and traditional supervised learning is still very large, due to the label noise of web data as well as the domain shift between web data and test data. To fill this gap, most existing methods propose to purify or augment web data using instance-level supervision, which generally requires heavy annotation. Instead, we propose to address the label noise and domain shift by using more accessible category-level supervision. In particular, we build our deep probabilistic framework upon variational autoencoder (VAE), in which classification network and VAE can jointly leverage category-level hybrid information. Then, we extend our method for domain adaptation followed by our low-rank refinement strategy. Extensive experiments on three benchmark datasets demonstrate the effectiveness of our proposed method. Li Niu 0002, Qingtao Tang, Ashok Veeraraghavan, Ashutosh Sabharwal |
CVPR | 1 |
| 2018 | Webly Supervised Learning Meets Zero-Shot Learning: A Hybrid Approach for Fine-Grained ClassificationabstractFine-grained image classification, which targets at distinguishing subtle distinctions among various subordinate categories, remains a very difficult task due to the high annotation cost of enormous fine-grained categories. To cope with the scarcity of well-labeled training images, existing works mainly follow two research directions: 1) utilize freely available web images without human annotation; 2) only annotate some fine-grained categories and transfer the knowledge to other fine-grained categories, which falls into the scope of zero-shot learning (ZSL). However, the above two directions have their own drawbacks. For the first direction, the labels of web images are very noisy and the data distribution between web images and test images are considerably different. For the second direction, the performance gap between ZSL and traditional supervised learning is still very large. The drawbacks of the above two directions motivate us to design a new framework which can jointly leverage both web data and auxiliary labeled categories to predict the test categories that are not associated with any well-labeled training images. Comprehensive experiments on three benchmark datasets demonstrate the effectiveness of our proposed framework. Li Niu 0002, Ashok Veeraraghavan, Ashutosh Sabharwal |
CVPR | 1 |
| 2018 | Look, Imagine and Match: Improving Textual-Visual Cross-Modal Retrieval With Generative ModelsabstractTextual-visual cross-modal retrieval has been a hot research topic in both computer vision and natural language processing communities. Learning appropriate representations for multi-modal data is crucial for the cross-modal retrieval performance. Unlike existing image-text retrieval approaches that embed image-text pairs as single feature vectors in a common representational space, we propose to incorporate generative processes into the cross-modal feature embedding, through which we are able to learn not only the global abstract features but also the local grounded features. Extensive experiments show that our framework can well match images and sentences with complex content, and achieve the state-of-the-art cross-modal retrieval results on MSCOCO dataset. Jiuxiang Gu, Jianfei Cai 0001, Shafiq R. Joty, Li Niu 0002, Gang Wang 0012 |
CVPR | 4 |
| 2018 | Sure-Based Dual Domain Image DenoisingabstractRecently developed Dual Domain Image Denoising (DDID) algorithm is a simple version of block-matching 3D filtering (BM3D) by combining bilateral filter and frequency-based method. DDID and its invariants have achieved competitive results compared with state-of-the-art methods. However, this kind of methods share a common drawback: there are a few parameters of the algorithms that are data- and noise-dependent, and difficult to tune. In this paper, we propose to use Stein's unbiased risk estimate (SURE) to measure the mean square error (MSE) of the DDID algorithm for restoration of an image contaminated with additive white Gaussian noise. We derive an explicit expression for SURE value to optimize parameters without access to the noise-free signal. Experimental results demonstrate the effectiveness of the proposed parameter selection in term of both quantitative and qualitative metrics. Zhiya Xu, Tao Dai 0001, Li Niu 0002, Jiawei Li 0006, Qingtao Tang, Shutao Xia |
ICASSP | 3 |
| 2018 | Self -Paced Mixture of T Distribution ModelabstractGaussian mixture model (GMM) is a powerful probabilistic model for representing the probability distribution of observations in the population. However, the fitness of Gaussian mixture model can be significantly degraded when the data contain a certain amount of outliers. Although there are certain variants of GMM (e.g., mixture of Laplace, mixture of t distribution) attempting to handle outliers, none of them can sufficiently mitigate the effect of outliers if the outliers are far from the centroids. Aiming to remove the effect of outliers further, this paper introduces a Self-Paced Learning mechanism into mixture of t distribution, which leads to Self-Paced Mixture of t distribution model (SPTMM). We derive an Expectation-Maximization based algorithm to train SPTMM and show SPTMM is able to screen the outliers. To demonstrate the effectiveness of SPTMM, we apply the model to density estimation and clustering. Finally, the results indicate that SPTMM outperforms other methods, especially on the data with outliers. Yang Zhang 0016, Qingtao Tang, Li Niu 0002, Tao Dai 0001, Xi Xiao 0001, Shutao Xia |
ICASSP | 3 |
| 2018 | Portrait-Aware Artistic Style TransferabstractThe goal of artistic style transfer is to transfer the style of artistic works into photos. However, the performances of existing style transfer algorithms on portraits are not very satisfactory, because the synthetic photo is either not sufficiently stylized or distorted severely in the portrait domain (i.e., foreground), which limits the use of style transfer for portraits. In this paper, we propose a novel portrait-aware artistic style transfer algorithm, which treats foreground and background differently. Particularly, we separate the foreground from the background, and apply fine-grained style transfer to the background and coarse-grained style transfer to the entire image at the same time, so that the artistic style of entire image can be transferred with the details of the portrait well preserved. Extensive experiments demonstrate the effectiveness of our proposed method. Yeli Xing, Jiawei Li 0006, Tao Dai 0001, Qingtao Tang, Li Niu 0002, Shutao Xia |
ICIP | 5 |
| 2018 | Referenceless quality metric of multiply-distorted images based on structural degradation
Tao Dai 0001, Ke Gu 0001, Li Niu 0002, Yongbing Zhang 0002, Weizhi Lu, Shutao Xia |
Neurocomputing | 3 |
| 2018 | An Exemplar-Based Multi-View Domain Generalization Framework for Visual RecognitionabstractIn this paper, we propose a new exemplar-based multi-view domain generalization (EMVDG) framework for visual recognition by learning robust classifier that are able to generalize well to arbitrary target domain based on the training samples with multiple types of features (i.e., multi-view features). In this framework, we aim to address two issues simultaneously. First, the distribution of training samples (i.e., the source domain) is often considerably different from that of testing samples (i.e., the target domain), so the performance of the classifiers learnt on the source domain may drop significantly on the target domain. Moreover, the testing data are often unseen during the training procedure. Second, when the training data are associated with multi-view features, the recognition performance can be further improved by exploiting the relation among multiple types of features. To address the first issue, considering that it has been shown that fusing multiple SVM classifiers can enhance the domain generalization ability, we build our EMVDG framework upon exemplar SVMs (ESVMs), in which a set of ESVM classifiers are learnt with each one trained based on one positive training sample and all the negative training samples. When the source domain contains multiple latent domains, the learnt ESVM classifiers are expected to be grouped into multiple clusters. To address the second issue, we propose two approaches under the EMVDG framework based on the consensus principle and the complementary principle, respectively. Specifically, we propose an EMVDG_CO method by adding a co-regularizer to enforce the cluster structures of ESVM classifiers on different views to be consistent based on the consensus principle. Inspired by multiple kernel learning, we also propose another EMVDG_MK method by fusing the ESVM classifiers from different views based on the complementary principle. In addition, we further extend our EMVDG framework to exemplar-based multi-view domain adaptation (EMVDA) framework when the unlabeled target domain data are available during the training procedure. The effectiveness of our EMVDG and EMVDA frameworks for visual recognition is clearly demonstrated by comprehensive experiments on three benchmark data sets. Li Niu 0002, Wen Li 0001, Dong Xu 0001, Jianfei Cai 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Robust Survey Aggregation with Student-t Distribution and Sparse RepresentationabstractMost existing survey aggregation methods assume that the sample data follow Gaussian distribution. However, these methods are sensitive to outliers, due to the thin-tailed property of the Gaussian distribution. To address this issue, we propose a robust survey aggregation method based on Student-t distribution and sparse representation. Specifically, we assume that the samples follow Student-$t$ distribution, instead of the common Gaussian distribution. Due to the Student-t distribution, our method is robust to outliers, which can be explained from both Bayesian point of view and non-Bayesian point of view. In addition, inspired by James-Stain estimator (JS) and Compressive Averaging (CAvg), we propose to sparsely represent the global mean vector by an adaptive basis comprising both data-specific basis and combined generic bases. Theoretically, we prove that JS and CAvg are special cases of our method. Extensive experiments demonstrate that our proposed method achieves significant improvement over the state-of-the-art methods on both synthetic and real datasets. Qingtao Tang, Tao Dai 0001, Li Niu 0002, Yisen Wang 0001, Shutao Xia, Jianfei Cai 0001 |
IJCAI | 3 |
| 2017 | Student-t Process Regression with Student-t LikelihoodabstractGaussian Process Regression (GPR) is a powerful Bayesian method. However, the performance of GPR can be significantly degraded when the training data are contaminated by outliers, including target outliers and input outliers. Although there are some variants of GPR (e.g., GPR with Student-t likelihood (GPRT)) aiming to handle outliers, most of the variants focus on handling the target outliers while little effort has been done to deal with the input outliers. In contrast, in this work, we aim to handle both the target outliers and the input outliers at the same time. Specifically, we replace the Gaussian noise in GPR with independent Student-t noise to cope with the target outliers. Moreover, to enhance the robustness w.r.t. the input outliers, we use a Student-t Process prior instead of the common Gaussian Process prior, leading to Student-t Process Regression with Student-t Likelihood (TPRT). We theoretically show that TPRT is more robust to both input and target outliers than GPR and GPRT, and prove that both GPR and GPRT are special cases of TPRT. Various experiments demonstrate that TPRT outperforms GPR and its variants on both synthetic and real datasets. Qingtao Tang, Li Niu 0002, Yisen Wang 0001, Tao Dai 0001, Wangpeng An, Jianfei Cai 0001, Shutao Xia |
IJCAI | 2 |
| 2017 | Action proposals using hierarchical clustering of super-trajectoriesabstractAction localization aims to determine the spatial and temporal location of certain action which appears in a video. To facilitate action localization, spatio-temporal proposals which are likely to contain the action of interest are extracted to reduce the search space of candidate locations in a video, inspired by the object proposals in images. In this paper, considering the effectiveness of spatio-temporal trajectories for video action recognition and action proposal generation, we build our unsupervised action proposal generation pipeline upon super-trajectories. Specifically, we first group trajectories into super-trajectories inspired by super-voxels, and then employ hierarchical clustering on super-trajectories by taking different aspect and temporal ratios into consideration. Comprehensive experiments on two benchmark datasets (i.e., UCF-sports and MSR-II) demonstrate that our action proposal generation pipeline not only achieves the state-of-the-art recall, but also achieves competitive results for the action localization task. Tianyi Zhang 0004, Li Niu 0002, Jianfei Cai 0001, Alex Chichung Kot |
VCIP | 2 |
| 2017 | Visual Recognition by Learning From Web Data via Weakly Supervised Domain GeneralizationabstractIn this paper, a weakly supervised domain generalization (WSDG) method is proposed for real-world visual recognition tasks, in which we train classifiers by using Web data (e.g., Web images and Web videos) with noisy labels. In particular, two challenging problems need to be solved when learning robust classifiers, in which the first issue is to cope with the label noise of training Web data from the source domain, while the second issue is to enhance the generalization capability of learned classifiers to an arbitrary target domain. In order to handle the first problem, the training samples within each category are partitioned into clusters, where we use one bag to denote each cluster and instances to denote the samples in each cluster. Then, we identify a proportion of good training samples in each bag and train robust classifiers by using the good training samples, which leads to a multi-instance learning (MIL) problem. In order to handle the second problem, we assume that the training samples possibly form a set of hidden domains, with each hidden domain associated with a distinctive data distribution. Then, for each category and each hidden latent domain, we propose to learn one classifier by extending our MIL formulation, which leads to our WSDG approach. In the testing stage, our approach can obtain better generalization capability by effectively integrating multiple classifiers from different latent domains in each category. Moreover, our WSDG approach is further extended to utilize additional textual descriptions associated with Web data as privileged information (PI), although testing data do not have such PI. Extensive experiments on three benchmark data sets indicate that our newly proposed methods are effective for real-world visual recognition tasks by learning from Web data. Li Niu 0002, Wen Li 0001, Dong Xu 0001, Jianfei Cai 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Action and Event Recognition in Videos by Learning From Heterogeneous Web SourcesabstractIn this paper, we propose new approaches for action and event recognition by leveraging a large number of freely available Web videos (e.g., from Flickr video search engine) and Web images (e.g., from Bing and Google image search engines). We address this problem by formulating it as a new multi-domain adaptation problem, in which heterogeneous Web sources are provided. Specifically, we are given different types of visual features (e.g., the DeCAF features from Bing/Google images and the trajectory-based features from Flickr videos) from heterogeneous source domains and all types of visual features from the target domain. Considering the target domain is more relevant to some source domains, we propose a new approach named multi-domain adaptation with heterogeneous sources (MDA-HS) to effectively make use of the heterogeneous sources. In MDA-HS, we simultaneously seek for the optimal weights of multiple source domains, infer the labels of target domain samples, and learn an optimal target classifier. Moreover, as textual descriptions are often available for both Web videos and images, we propose a novel approach called MDA-HS using privileged information (MDA-HS+) to effectively incorporate the valuable textual information into our MDA-HS method, based on the recent learning using privileged information paradigm. MDA-HS+ can be further extended by using a new elastic-net-like regularization. We solve our MDA-HS and MDA-HS+ methods by using the cutting-plane algorithm, in which a multiple kernel learning problem is derived and solved. Extensive experiments on three benchmark data sets demonstrate that our proposed approaches are effective for action and event recognition without requiring any labeled samples from the target domain. Li Niu 0002, Xinxing Xu, Lin Chen 0021, Lixin Duan, Dong Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2016 | Domain Adaptive Fisher Vector for Visual Recognition
Li Niu 0002, Jianfei Cai 0001, Dong Xu 0001 |
ECCV (6) | 1 |
| 2016 | Exploiting Privileged Information from Web Data for Action and Event Recognition
Li Niu 0002, Wen Li 0001, Dong Xu 0001 |
Int. J. Comput. Vis. | 1 |
| 2015 | Visual recognition by learning from web data: A weakly supervised domain generalization approachabstractIn this work, we formulate a new weakly supervised domain generalization approach for visual recognition by using loosely labeled web images/videos as training data. Specifically, we aim to address two challenging issues when learning robust classifiers: 1) coping with noise in the labels of training web images/videos in the source domain; and 2) enhancing generalization capability of learnt classifiers to any unseen target domain. To address the first issue, we partition the training samples in each class into multiple clusters. By treating each cluster as a “bag” and the samples in each cluster as “instances”, we formulate a multi-instance learning (MIL) problem by selecting a subset of training samples from each training bag and simultaneously learning the optimal classifiers based on the selected samples. To address the second issue, we assume the training web images/videos may come from multiple hidden domains with different data distributions. We then extend our MIL formulation to learn one classifier for each class and each latent domain such that multiple classifiers from each class can be effectively integrated to achieve better generalization capability. Extensive experiments on three benchmark datasets demonstrate the effectiveness of our new approach for visual recognition by learning from web data. Li Niu 0002, Wen Li 0001, Dong Xu 0001 |
CVPR | 1 |
| 2015 | Multi-view Domain Generalization for Visual RecognitionabstractIn this paper, we propose a new multi-view domain generalization (MVDG) approach for visual recognition, in which we aim to use the source domain samples with multiple types of features (i.e., multi-view features) to learn robust classifiers that can generalize well to any unseen target domain. Considering the recent works show the domain generalization capability can be enhanced by fusing multiple SVM classifiers, we build upon exemplar SVMs to learn a set of SVM classifiers by using one positive sample and all negative samples in the source domain each time. When the source domain samples come from multiple latent domains, we expect the weight vectors of exemplar SVM classifiers can be organized into multiple hidden clusters. To exploit such cluster structure, we organize the weight vectors learnt on each view as a weight matrix and seek the low-rank representation by reconstructing this weight matrix using itself as the dictionary. To enforce the consistency of inherent cluster structures discovered from the weight matrices learnt on different views, we introduce a new regularizer to minimize the mismatch between any two representation matrices on different views. We also develop an efficient alternating optimization algorithm and further extend our MVDG approach for domain adaptation by exploiting the manifold structure of unlabeled target domain samples. Comprehensive experiments for visual recognition clearly demonstrate the effectiveness of our approaches for domain generalization and domain adaptation. Li Niu 0002, Wen Li 0001, Dong Xu 0001 |
ICCV | 1 |
| 2014 | Exploiting Privileged Information from Web Data for Image Categorization
Wen Li 0001, Li Niu 0002, Dong Xu 0001 |
ECCV (5) | 2 |
| 2014 | Exploiting Low-Rank Structure from Latent Domains for Domain Generalization
Zheng Xu 0002, Wen Li 0001, Li Niu 0002, Dong Xu 0001 |
ECCV (3) | 3 |