VLDB 2026 Research / reviewers in the wild / expert
Jia Li 0003
dblp:23/6950-3
· DBLP profile ↗
130ranked-venue papers
20as first author
66since 2021 · last 2026
0000-0002-4346-8696ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 98 · 14 first-author · 48 since 2021Artificial intelligence and machine learning · 60 · 7 first-author · 34 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adapting Vision-Language Models from Iconic to Inclusive for Multi-label Recognition Without Labels
Jingyu Zhou, Yifan Zhao 0002, Jia Li 0003 |
Int. J. Comput. Vis. | 4 |
| 2026 | Content-Rhythm Awareness Contrastive Learning for Music-Driven 3D Dance Generation
Yifan Zhao 0002, Jia Li 0003 |
Int. J. Comput. Vis. | 3 |
| 2026 | Single Image to Textured 3D Object Generation in Frequency Domain: From Theory to Pipeline
Yifan Zhao 0002, Jia Li 0003 |
Int. J. Comput. Vis. | 3 |
| 2026 | Learning When and How to Update Memory for Video Object SegmentationabstractRecent progress in semi-supervised video object segmentation has largely hinged on memory-based methods. However, when faced with increasingly tough challenges emerging in complex scenarios, such as fundamental semantic transformations and severe spatial deformations, the fixed-interval memory update mechanism usually adopted in these memory-based methods is insufficient to align with the pivotal moments of object changes. This inflexible mechanism motivates us to design an adaptive memory update mechanism in response to the semantic-spatial changes of target objects. To this end, we propose a novel Change-Sensitive Network (CSNet) to learn when and how to update memory to effectively address intricate challenges in complex scenarios. Specifically, we first design an Adaptive Perception-Capture module with a hierarchical contrastive learning loss to determine when to update memory moments by measuring the extent of object changes, thus dividing entire videos into different object-change clips. To further extract and highlight object changes to assist in the segmentation of frames after changes occur, we construct Dynamic Memory Update modules to redefine how to update memory by smoothly retaining the object prototypes within clips and dynamically amplifying the object variations across clips. Extensive experiments demonstrate that our proposed CSNet exhibits clear superiority when evaluated on eight datasets covering three kinds: common, complex and long-video datasets. Shengye Qiao, Changqun Xia, Xiaowu Chen 0001, Jia Li 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Beyond Heat Dissipation: Optimizing Diffusion Models in Frequency DomainabstractThe majority of standard diffusion models employ pixel-wise degradations while neglecting multi-scale characteristics of images. Recently, generalized diffusion models with Positive Semi-definite Degradations (PSD), such as heat dissipation and blurring, have been proposed to solve it, but suffering from problems of low generation quality due to incomplete optimization analysis and non-adaptiveness to the training process and different data distributions with hand-crafted and fixed inductive biases. In this paper, we present a comprehensive theoretical analysis of the optimization process in frequency domain for PSD-based generalized diffusion models, which implies the forward process of PSD frequency domain non-isotropic degradation implicitly acting on the inductive biases of the Variational Lower Bound non-isotropic weighting in the optimization reverse process. Based on this insight, we propose the Frequency Inductive Biases Bootstrapping Optimization (FIBBO) method, which parameterizes the forward process and learns distinct frequency degradation-generation trajectories iteratively. To tackle the problem of PSD hand-crafted and fixed inductive biases, FIBBO dynamically modifies the non-isotropic Gaussian kernel of the forward degradation process so that the inductive biases introduced can be adjusted adaptively during training. Experiments on public datasets show that FIBBO makes significant improvements in the generation quality of PSD-based generalized diffusion models. Yifan Zhao 0002, Jia Li 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | PolarGS: Polarimetric Cues for Ambiguity-Free Gaussian Splatting With Accurate Geometry RecoveryabstractRecent advances in surface reconstruction for 3D Gaussian Splatting (3DGS) have enabled remarkable geometric accuracy. However, their performance degrades in photometrically ambiguous regions such as reflective and textureless surfaces, where unreliable cues disrupt photometric consistency and hinder accurate geometry estimation. Reflected light is often partially polarized in a manner that reveals surface orientation, making polarization an optical complement to photometric cues in resolving such ambiguities. Therefore, we propose PolarGS, an optics-aware extension of RGB-based 3DGS that leverages polarization as an optical prior to resolve photometric ambiguities and enhance reconstruction accuracy. Specifically, we introduce two complementary modules: a polarization-guided photometric correction strategy, which ensures photometric consistency by identifying reflective regions via the Degree of Linear Polarization (DoLP) and refining reflective Gaussians with Color Refinement Maps; and a polarization-enhanced Gaussian densification mechanism for textureless area geometry recovery, which integrates both Angle and Degree of Linear Polarization (A&DoLP) into a PatchMatch-based depth completion process. This enables the back-projection and fusion of new Gaussians, leading to a more complete reconstruction. PolarGS is framework-agnostic and achieves superior geometric accuracy compared to state-of-the-art methods. Sijia Wen, Yifan Zhao 0002, Jia Li 0003, Zhiming Zheng 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | Toward Unified Co-Speech Gesture Generation via Hierarchical Implicit Periodicity LearningabstractGenerating 3D-based body movements from speech shows great potential in extensive downstream applications, while it still suffers challenges in imitating realistic human movements. Predominant research efforts focus on end-to-end generation schemes to generate co-speech gestures, spanning GANs, VQ-VAE, and recent diffusion models. As an ill-posed problem, in this paper, we argue that these prevailing learning schemes fail to model crucial inter- and intra-correlations across different motion units, i.e. head, body, and hands, thus leading to unnatural movements and poor coordination. To delve into these intrinsic correlations, we propose a unified Hierarchical Implicit Periodicity (HIP) learning approach for audio-inspired 3D gesture generation. Different from predominant research, our approach models this multi-modal implicit relationship by two explicit technique insights: i) To disentangle the complicated gesture movements, we first explore the gesture motion phase manifolds with periodic autoencoders to imitate human natures from realistic distributions while incorporating non-period ones from current latent states for instance-level diversities. ii) To model the hierarchical relationship of face motions, body gestures, and hand movements, driving the animation with cascaded guidance during learning. We exhibit our proposed approach on 3D avatars and extensive experiments show our method outperforms the state-of-the-art co-speech gesture generation methods by both quantitative and qualitative evaluations. Code and models will be publicly available. Yifan Zhao 0002, Jia Li 0003 |
IEEE Trans. Image Process. | 3 |
| 2026 | Generation in Generation: Fluid Co-Speech Gesture Synthesis With Generative Continuous QuantizationabstractMotion quantization codebooks have been widely adopted to facilitate co-speech motion generation. However, the conventional quantization-based generation paradigm-which relies on probabilistic token sampling from limited discrete codebooks-suffers from two major limitations: crude, unreasonable motion representations and fixed, homogenized motion token sequences. To overcome these issues, we propose a novel explicit generation paradigm based on generative continuous quantization. Specifically, we first introduce a continuous quantization method to derive a set of generative motion units. This approach enables smoother and more accurate representation of human motion compared to classical methods. Building on these generative units, we further propose a compositional weight generation paradigm that replaces probabilistic sampling with deterministic, explicit motion synthesis. Moreover, as generalization capability is crucial for real-world deployment, we design a fully audio-aware encoder to extract style features that are decoupled from content. These features are integrated into the motion decoder via Adaptive Instance Normalization to enhance cross-speaker facial style generalization. Our method achieves state-of-the-art performance on two public datasets. Notably, owing to its concise and efficient architecture, our model attains an inference speed exceeding 4000 fps on the SHOW dataset, demonstrating strong potential for practical real-time applications. Yifan Zhao 0002, Jia Li 0003 |
IEEE Trans. Image Process. | 4 |
| 2026 | UncNeRF: Uncovering Heavily Occluded Object With Multi-View CluesabstractNeural Radiance Fields can achieve photo-realistic rendering results, but the occlusion in front of the target object is a common and extreme scenario in practice that cannot be neglected. The prevailing works attempt to remove the occlusions using external 2D visual priors, which are not constrained to provide 3D-consistent guidance for the specific scenarios. In this paper, we propose UncNeRF, which utilizes multi-view clues from captured defective images to uncover the heavily occluded object. Specifically, we provide additional multi-view complementary optimization supervisions using object-centric forward warping and enhance the target object reconstruction by sampling pseudo-training views and introducing external spatial-relation regularization. To evaluate the reconstruction performance of occluded objects, we present the challenging and diverse Heavy Occlusion Removal (HOR) dataset consisting of synthetic and real-world scenes, whose target objects to be reconstructed are heavily occluded. Experimental results show that our method achieves state-of-the-art performance in heavy occlusion removal compared to other methods. Jiawei Ma, Yifan Zhao 0002, Jia Li 0003 |
IEEE Trans. Image Process. | 4 |
| 2025 | Holistic Correction with Object Prototype for Video Object SegmentationabstractRecently, memory-based methods have achieved progress in semi-supervised video object segmentation. However, these methods still suffer from unstructured challenges, such as object transformations, occlusions and disappearance-reappearance. To this end, we propose a Holistic Correction Network (HCNet) to adaptively acquire concise object prototypes for holistic correction at semantic, spatial and temporal aspects. Specifically, an Adaptive Prototype Update module is firstly designed to construct multi-level core object representations by associating object variations in consecutive frames with segmentation quality assessment. Based on the updated object prototypes, Semantic, Spatial and Temporal Correction modules are respectively designed to enhance the object semantics in the entire frame, eliminate the incorrect semantic enhancement outside the object regions and calibrate the estimated object regions with temporal changes of objects. Through the holistic correction mechanism with effective object prototypes, our proposed HCNet can robustly and efficiently deal with diverse complex scenarios. Extensive and comprehensive experiments conducted on seven datasets demonstrate that our proposed HCNet can significantly improve the segmentation performance. Shengye Qiao, Changqun Xia, Gongjin Lan, Jia Li 0003 |
AAAI | 5 |
| 2025 | A Comprehensive Evaluation on Event Reasoning of Large Language ModelsabstractEvent reasoning is a fundamental ability that underlies many applications. It requires event schema knowledge to perform global reasoning and needs to deal with the diversity of the inter-event relations and the reasoning paradigms. The extent to which LLMs excel in event reasoning across various relations and reasoning paradigms has not been thoroughly investigated. Additionally, it is still unclear whether LLMs utilize event knowledge in the same way humans do. To mitigate this disparity, we comprehensively evaluate the abilities of event reasoning of LLMs on different relations, paradigms, and levels of abstraction. We introduce a novel benchmark EV2 for EValuation of EVent reasoning. EV2 consists of two levels of evaluation on schema and instance and is comprehensive in relations and reasoning paradigms. We conduct extensive experiments on EV2. We find that 1) LLMs have abilities to accomplish event reasoning but their performances are far from satisfactory. 2) There are imbalances of event reasoning abilities on different relations and paradigms. 3) LLMs have event schema knowledge, however, they're not aligned with humans on how to utilize the knowledge. Based on these findings, we guide the LLMs in utilizing the event schema knowledge as memory leading to improvements in event reasoning. Zhengwei Tao, Zhi Jin 0001, Yifan Zhang 0004, Xiancai Chen, Haiyan Zhao 0001, Jia Li 0003, Bin Liang 0004, Chongyang Tao, Qun Liu 0001, Kam-Fai Wong |
AAAI | 6 |
| 2025 | FreeGen: Bridging Visual-Linguistic Discrepancies Towards Diffusion-based Pixel-level Data SynthesisabstractText-to-image diffusion model has inspired research into text-to-data synthesis without human intervention, where spatial attentions correlated with semantic entities in text prompts are primarily interpreted as pseudo-masks. However, these vannila attentions often deliver visual-linguistic discrepancies, in which the associations between image features and entity-level tokens are unstable and divergent, yielding inferior masks for realistic applications, especially in more practical open-vocabulary settings. To tackle this issue, we propose a novel text-guided self-driven generative paradigm, termed FreeGen, which addresses the discrepancies by recalibrating intrinsic visual-linguistic correlations and serves as a real-data-free method to automatically synthesize open-vocabulary pixel-level data for arbitrary entities. Specifically, we first learn an Attention Self-Rectification mechanism to reproject the inherent attention matrices to achieve robust semantic alignment, thereby obtaining class-discriminative masks. A Temporal Fluctuation Factor is present to assess mask quality based on its variation over uniform sampling timesteps, enabling the selection of reliable masks. These masks are then employed as self-supervised signals to support the learning of an Entity-level Grounding Decoder in a self-training manner, thus producing open-vocabulary segmentation results. Extensive experiments show that the existing segmenters trained on FreeGen narrow the performance gap with real data counterparts and remarkably outperform the state-of-the-art methods. Wenzhuang Wang, Mingcan Ma, Changqun Xia, Zhenbao Liang, Jia Li 0003 |
AAAI | 6 |
| 2025 | Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context LearningabstractIn-context learning (ICL), a predominant trend in instruction learning, aims at enhancing the performance of large language models by providing clear task guidance and examples, improving their capability in task understanding and execution. This paper investigates ICL on Large Vision-Language Models (LVLMs) and explores the policies of multi-modal demonstration selection. Existing research efforts in ICL face significant challenges: First, they rely on pre-defined demonstrations or heuristic selecting strategies based on human intuition, which are usually inadequate for covering diverse task requirements, leading to sub-optimal solutions; Second, individually selecting each demonstration fails in modeling the interactions between them, resulting in information redundancy. Unlike these prevailing efforts, we propose a new exploration-exploitation reinforcement learning framework, which explores policies to fuse multi-modal information and adaptively select adequate demonstrations as an integrated whole. The framework allows LVLMs to optimize themselves by continually refining their demonstrations through self-exploration, enabling the ability to autonomously identify and generate the most effective selection policies for in-context learning. Experimental results verify the superior performance of our approach on four Visual Question-Answering (VQA) datasets, demonstrating its effectiveness in enhancing the generalization capability of few-shot LVLMs. Yunpeng Zhai, Yifan Zhao 0002, Jinyang Gao, Bolin Ding, Jia Li 0003 |
CVPR | 6 |
| 2025 | FICGen: Frequency-Inspired Contextual Disentanglement for Layout-driven Degraded Image Generation
Wenzhuang Wang, Yifan Zhao 0002, Mingcan Ma, Zhonglin Jiang, Jia Li 0003 |
ICCV | 7 |
| 2025 | Learning Yourself: Class-Incremental Semantic Segmentation with Language-Inspired Bootstrapped DisentanglementabstractClass-Incremental Semantic Segmentation (CISS) requires continuous learning of newly introduced classes while retaining knowledge of past classes. By abstracting mainstream methods into two stages (visual feature extraction and prototype-feature matching), we identify a more fundamental challenge termed catastrophic semantic entanglement. This phenomenon involves Prototype-Feature Entanglement caused by semantic misalignment during the incremental process, and Background-Increment Entanglement due to dynamic data evolution. Existing techniques, which rely on visual feature learning without sufficient cues to distinguish targets, introduce significant noise and errors. To address these issues, we introduce a Language-inspired Bootstrapped Disentanglement framework (LBD). We leverage the prior class semantics of pre-trained visual-language models (e.g., CLIP) to guide the model in autonomously disentangling features through Language-guided Prototypical Disentanglement and Manifold Mutual Background Disentanglement. The former guides the disentangling of new prototypes by treating hand-crafted text features as topological templates, while the latter employs multiple learnable prototypes and mask-pooling-based supervision for background-incremental class disentanglement. By incorporating soft prompt tuning and encoder adaptation modifications, we further bridge the capability gap of CLIP between dense and sparse tasks, achieving state-of-the-art performance on both Pascal VOC and ADE20k, particularly in multi-step scenarios. Ruitao Wu, Yifan Zhao 0002, Jia Li 0003 |
ICCV | 3 |
| 2025 | Re-coding for Uncertainties: Edge-awareness Semantic Concordance for Resilient Event-RGB SegmentationabstractSemantic segmentation has achieved great success in ideal conditions. However, when facing extreme conditions (e.g., insufficient light, fierce camera motion), most existing methods suffer from significant information loss of RGB, severely damaging segmentation results. Several researches exploit the high-speed and high-dynamic event modality as a complement, but event and RGB are naturally heterogeneous, which leads to feature-level mismatch and inferior optimization of existing multi-modality methods. Different from these researches, we delve into the edge secret of both modalities for resilient fusion and propose a novel Edge-awareness Semantic Concordance framework to unify the multi-modality heterogeneous features with latent edge cues. In this framework, we first propose Edge-awareness Latent Re-coding, which obtains uncertainty indicators while realigning event-RGB features into unified semantic space guided by re-coded distribution, and transfers event-RGB distributions into re-coded features by utilizing a pre-established edge dictionary as clues. We then propose Re-coded Consolidation and Uncertainty Optimization, which utilize re-coded edge features and uncertainty indicators to solve the heterogeneous event-RGB fusion issues under extreme conditions. We establish two synthetic and one real-world event-RGB semantic segmentation datasets for extreme scenario comparisons. Experimental results show that our method outperforms the state-of-the-art by a 2.55% mIoU on our proposed DERS-XS, and possesses superior resilience under spatial occlusion. Our code and datasets are publicly available at https://github.com/iCVTEAM/ESC. Yifan Zhao 0002, Lin Zhu 0012, Jia Li 0003 |
NeurIPS | 4 |
| 2025 | Diffusion-Classifier Synergy: Reward-Aligned Learning via Mutual Boosting Loop for FSCILabstractFew-Shot Class-Incremental Learning (FSCIL) challenges models to sequentially learn new classes from minimal examples without forgetting prior knowledge, a task complicated by the stability-plasticity dilemma and data scarcity. Current FSCIL methods often struggle with generalization due to their reliance on limited datasets. While diffusion models offer a path for data augmentation, their direct application can lead to semantic misalignment or ineffective guidance. This paper introduces Diffusion-Classifier Synergy (DCS), a novel framework that establishes a mutual boosting loop between diffusion model and FSCIL classifier. DCS utilizes a reward-aligned learning strategy, where a dynamic, multi-faceted reward function derived from the classifier's state directs the diffusion model. This reward system operates at two levels: the feature level ensures semantic coherence and diversity using prototype-anchored maximum mean discrepancy and dimension-wise variance matching, while the logits level promotes exploratory image generation and enhances inter-class discriminability through confidence recalibration and cross-session confusion-aware mechanisms. This co-evolutionary process, where generated images refine the classifier and an improved classifier state yields better reward signals, demonstrably achieves state-of-the-art performance on FSCIL benchmarks, significantly enhancing both knowledge retention and new class learning. Ruitao Wu, Yifan Zhao 0002, Jia Li 0003 |
NeurIPS | 4 |
| 2025 | Free Lunch to Meet the Gap: Intermediate Domain Reconstruction for Cross-Domain Few-Shot Learning
Yifan Zhao 0002, Jia Li 0003 |
Int. J. Comput. Vis. | 4 |
| 2025 | Language-Inspired Relation Transfer for Few-Shot Class-Incremental LearningabstractDepicting novel classes with language descriptions by observing few-shot samples is inherent in human-learning systems. This lifelong learning capability helps to distinguish new knowledge from old ones through the increase of open-world learning, namely Few-Shot Class-Incremental Learning (FSCIL). Existing works to solve this problem mainly rely on the careful tuning of visual encoders, which shows an evident trade-off between the base knowledge and incremental ones. Motivated by human learning systems, we propose a new Language-inspired Relation Transfer (LRT) paradigm to understand objects by joint visual clues and text depictions, composed of two major steps. We first transfer the pretrained text knowledge to the visual domains by proposing a graph relation transformation module and then fuse the visual and language embedding by a text-vision prototypical fusion module. Second, to mitigate the domain gap caused by visual finetuning, we propose context prompt learning for fast domain alignment and imagined contrastive learning to alleviate the insufficient text data during alignment. With collaborative learning of domain alignments and text-image transfer, our proposed LRT outperforms the state-of-the-art models by over 13% and 7% on the final session of miniImageNet and CIFAR-100 FSCIL benchmarks. Yifan Zhao 0002, Jia Li 0003, Zeyin Song, Yonghong Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Progressive Semantic-Visual Alignment and Refinement for Vision-Language TrackingabstractIn recent years, vision-language tracking has drawn emerging attention in the tracking field. The critical challenge for the task is to fuse semantic representations of language information and visual representations of vision information. For this purpose, several vision-language tracking methods perform early or late fusion to fuse visual and semantic features. However, these methods cannot take full advantage of the transformer architecture to excavate useful cross-modal context at various levels. To this end, we propose a new progressive joint vision-language transformer (PJVLT) to progressively align and refine visual embedding with semantic embedding for vision-language tracking. Specifically, to align visual signals with semantic signals, we propose to insert a semantic-aware instance encoder layer (SAIEL) into each intermediate layer of transformer encoder to perform progressive alignment of visual and semantic features. Furthermore, to highlight the multi-modal feature channels and patches corresponding to target objects, we propose a unified channel communication patch interaction layer (CCPIL), which is plugged into each intermediate layer of transformer encoder to progressively activate target-aware channels and patches of aligned multi-modal features for fine-grained tracking. In general, by progressively aligning and refining visual features with semantic features in the transformer encoder, our PJVLT can adaptively excavate well-aligned vision-language context at coarse-to-fine levels, therefore highlighting target objects at various levels for more discriminative tracking. Experiments on several tracking datasets show that the proposed PJVLT can achieve favorable performance in comparison with both conventional trackers and other vision-language trackers. Qiangqiang Wu, Changqun Xia, Jia Li 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Text-prompt Camouflaged Instance Segmentation with Graduated Camouflage LearningabstractCamouflaged instance segmentation (CIS) aims to detect and segment objects blending with their surroundings. While existing CIS methods rely heavily on fully-supervised training with massive precisely annotated data, consuming considerable annotation efforts yet struggling to segment highly camouflaged objects accurately. Despite their visual similarity to the background, camouflaged objects differ semantically. Since text associated with images offers explicit semantic cues to underscore this difference, we propose a novel approach: the first Text-Prompt based weakly-supervised camouflaged instance segmentation method named TPNet, leveraging semantic distinctions for effective segmentation. TPNet operates in two stages: pseudo mask generation and a self-training process. In the first stage, we align text prompts with images using a language-image model to obtain region proposals containing camouflaged instances. A Semantic-Spatial Iterative Fusion module is designed to assimilate spatial information with semantic insights, iteratively refining pseudo mask. In the second stage, Graduated Camouflage Learning, a self-training strategy, sequences training from simple to complex images based on camouflage levels, facilitating an effective learning gradient. Through the collaboration of the dual phases, our method offers a comprehensive experiment on two common benchmark and demonstrates a significant advancement, delivering a novel solution that bridges the gap between weak-supervised and high camouflaged instance segmentation. Changqun Xia, Shengye Qiao, Jia Li 0003 |
ACM Multimedia | 4 |
| 2024 | Deblurring Neural Radiance Fields with Event-driven Bundle AdjustmentabstractNeural Radiance Fields (NeRF) achieves impressive 3D representation learning and novel view synthesis results with high-quality multi-view images as input. However, motion blur in images often occurs in low-light and high-speed motion scenes, which significantly degrades the reconstruction quality of NeRF. Previous deblurring NeRF methods struggle to estimate pose and lighting changes during the exposure time, making them unable to accurately model the motion blur. The bio-inspired event camera measuring intensity changes with high temporal resolution makes up this information deficiency. In this paper, we propose Event-driven Bundle Adjustment for Deblurring Neural Radiance Fields (EBAD-NeRF) to jointly optimize the learnable poses and NeRF parameters by leveraging the hybrid event-RGB data. An intensity-change-metric event loss and a photo-metric blur loss are introduced to strengthen the explicit modeling of camera motion blur. Experiments on both synthetic and real-captured data demonstrate that EBAD-NeRF can obtain accurate camera trajectory during the exposure time and learn a sharper 3D representations compared to prior works. Yunshan Qi, Lin Zhu 0012, Yifan Zhao 0002, Jia Li 0003 |
ACM Multimedia | 5 |
| 2024 | How to Use Diffusion Priors under Sparse Views?abstractNovel view synthesis under sparse views has been a long-term important challenge in 3D reconstruction. Existing works mainly rely on introducing external semantic or depth priors to supervise the optimization of 3D representations. However, the diffusion model, as an external prior that can directly provide visual supervision, has always underperformed in sparse-view 3D reconstruction using Score Distillation Sampling (SDS) due to the low information entropy of sparse views compared to text, leading to optimization challenges caused by mode deviation. To this end, we present a thorough analysis of SDS from the mode-seeking perspective and propose Inline Prior Guided Score Matching (IPSM), which leverages visual inline priors provided by pose relationships between viewpoints to rectify the rendered image distribution and decomposes the original optimization objective of SDS, thereby offering effective diffusion visual guidance without any fine-tuning or pre-training. Furthermore, we propose the IPSM-Gaussian pipeline, which adopts 3D Gaussian Splatting as the backbone and supplements depth and geometry consistency regularization based on IPSM to further improve inline priors and rectified distribution. Experimental results on different public datasets show that our method achieves state-of-the-art reconstruction quality. The code is released at https://github.com/iCVTEAM/IPSM. Yifan Zhao 0002, Jiawei Ma, Jia Li 0003 |
NeurIPS | 4 |
| 2024 | Towards imbalanced motion: part-decoupling network for video portrait segmentation
Tianshu Yu 0003, Changqun Xia, Jia Li 0003 |
Sci. China Inf. Sci. | 3 |
| 2024 | Reliable Event Generation With Invertible Conditional Normalizing FlowabstractEvent streams provide a novel paradigm to describe visual scenes by capturing intensity variations above specific thresholds along with various types of noise. Existing event generation methods usually rely on one-way mappings using hand-crafted parameters and noise rates, which may not adequately suit diverse scenarios and event cameras. To address this limitation, we propose a novel approach to learn a bidirectional mapping between the feature space of event streams and their inherent parameters, enabling the generation of reliable event streams with enhanced generalization capabilities. We first randomly generate a vast number of parameters and synthesize massive event streams using an event simulator. Subsequently, an event-based normalizing flow network is proposed to learn the invertible mapping between the representation of a synthetic event stream and its parameters. The invertible mapping is implemented by incorporating an intensity-guided conditional affine simulation mechanism, facilitating better alignment between event features and parameter spaces. Additionally, we impose constraints on event sparsity, edge distribution, and noise distribution through novel event losses, further emphasizing event priors in the bidirectional mapping. Our framework surpasses state-of-the-art methods in video reconstruction, optical flow estimation, and parameter estimation tasks on synthetic and real-world datasets, exhibiting excellent generalization across diverse scenes and cameras. Daxin Gu, Jia Li 0003, Lin Zhu 0012, Yu Zhang 0035, Jimmy S. J. Ren |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Learning Adaptive Parameter Representation for Event-Based Video ReconstructionabstractEvent-based video reconstruction aims to generate images from asynchronous event streams, which record the intensity changes exceeding specific contrast thresholds. However, the contrast thresholds are varied among pixels with manufacturing imperfections and circumstancing interference, which causes undesirable events. It may cause the existing works to output blurry frames with unpleasing artifacts. To address this, we propose a novel two-stage framework to reconstruct images with learnable parameter representations. The learnable representation of the contrast threshold is extracted with a transformer network from corresponding asynchronous events in the first stage. Then a UNet architecture is utilized in the second stage to fuse the representations with the event encoding features to refine the decoding features in spatiotemporal dimensions. The representation learned from asynchronous events can adapt to the variety of contrast thresholds when processing event data in diverse scenes, motivating the proposed framework to generate high-quality frames. Quantitative and qualitative experimental results on the four public datasets show that our approach achieves better performance. Daxin Gu, Jia Li 0003, Lin Zhu 0012 |
IEEE Signal Process. Lett. | 2 |
| 2024 | Joint Spatio-Temporal Similarity and Discrimination Learning for Visual TrackingabstractVisual tracking is a task of localizing a target unceasingly in a video with an initial target state at the first frame. The limited target information makes this problem an extremely challenging task. Existing tracking methods either perform matching based similarity learning or optimization based discrimination reasoning. However, these two types of tracking methods suffer from the problem of ineffectiveness for distinguishing target objects from background distractors and the problem of insufficiency in maintaining spatio-temporal consistency among successive frames, respectively. In this paper, we design a joint spatio-temporal similarity and discrimination learning (STSDL) framework for accurate and robust tracking. The designed framework is composed of two complementary branches: a similarity learning branch and a discrimination learning branch. The similarity learning branch uses an effective transformer encoder-decoder to gather rich spatio-temporal context information to generate a similarity map. In parallel, the discrimination learning branch exploits an efficient model predictor to train a target model to produce a discriminative map. Finally, the similarity map and the discriminative map are adaptively fused for accurate and robust target localization. Experimental results on six prevalent datasets demonstrate that the proposed STSDL can obtain satisfactory results, while it retains a real-time tracking speed of 50 FPS on a single GPU. Haosheng Chen 0001, Qiangqiang Wu, Changqun Xia, Jia Li 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Stripe Sensitive Convolution for Omnidirectional Image Dehazingabstractvirtual reality (VR) experience. The recent single image dehazing methods, to date, have been only focused on plane images. In this work, we propose a novel neural network pipeline for single omnidirectional image dehazing. To create the pipeline, we build the first hazy omnidirectional image dataset, which contains both synthetic and real-world samples. Then, we propose a new stripe sensitive convolution (SSConv) to handle the distortion problems due to the equirectangular projections. The SSConv calibrates distortion in two steps: 1) extracting features using different rectangular filters and, 2) learning to select the optimal features by a weighting of the feature stripes (a series of rows in the feature maps). Subsequently, using SSConv, we design an end-to-end network that jointly learns haze removal and depth estimation from a single omnidirectional image. The estimated depth map is leveraged as the intermediate representation and provides global context and geometric information to the dehazing module. Extensive experiments on challenging synthetic and real-world omnidirectional image datasets demonstrate the effectiveness of SSConv, and our network attains superior dehazing performance. The experiments on practical applications also demonstrate that our method can significantly improve the 3-D object detection and 3-D layout performances for hazy omnidirectional images. Dong Zhao 0010, Jia Li 0003, Long Xu 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | TriDet: Temporal Action Detection with Relative Boundary ModelingabstractIn this paper, we present a one-stage framework TriDet for temporal action detection. Existing methods often suffer from imprecise boundary predictions due to the ambiguous action boundaries in videos. To alleviate this problem, we propose a novel Trident-head to model the action boundary via an estimated relative probability distribution around the boundary. In the feature pyramid of TriDet, we propose an efficient Scalable-Granularity Perception (SGP) layer to mitigate the rank loss problem of self-attention that takes place in the video features and aggregate information across different temporal granularities. Benefiting from the Trident-head and the SGP-based feature pyramid, TriDet achieves state-of-the-art performance on three challenging benchmarks: THUMOS14, HACS and EPIC-KITCHEN 100, with lower computational costs, compared to previous methods. For example, TriDet hits an average mAP of 69.3% on THUMOS14, outperforming the previous best by 2.5%, but with only 74.6% of its latency. The code is released to https://github.com/dingfengshi/TriDet. Dingfeng Shi, Qiong Cao, Lin Ma 0002, Jia Li 0003, Dacheng Tao |
CVPR | 5 |
| 2023 | E2NeRF: Event Enhanced Neural Radiance Fields from Blurry ImagesabstractNeural Radiance Fields (NeRF) achieves impressive rendering performance by learning volumetric 3D representation from several images of different views. However, it is difficult to reconstruct a sharp NeRF from blurry input as often occurred in the wild. To solve this problem, we propose a novel Event-Enhanced NeRF (E2NeRF) by utilizing the combination data of a bio-inspired event camera and a standard RGB camera. To effectively introduce event stream into the learning process of neural volumetric representation, we propose a blur rendering loss and an event rendering loss, which guide the network via modelling real blur process and event generation process, respectively. Moreover, a camera pose estimation framework for real-world data is built with the guidance of event stream to generalize the method to practical applications. In contrast to previous image-based or event-based NeRF, our framework effectively utilizes the internal relationship between events and images. As a result, E2NeRF not only achieves image deblurring but also achieves high-quality novel view image generation. Extensive experiments on both synthetic data and real-world data demonstrate that E2NeRF can effectively learn a sharp NeRF from blurry images, especially in complex and low-light scenes. Our code and datasets are publicly available at https://github.com/iCVTEAM/E2NeRF. Yunshan Qi, Lin Zhu 0012, Yu Zhang 0035, Jia Li 0003 |
ICCV | 4 |
| 2023 | Frequency Representation Integration for Camouflaged Object DetectionabstractRecent camouflaged object detection (COD) approaches have been proposed to accurately segment objects blended into surroundings. The most challenging and critical issue in COD is to find out the lines of demarcation between objects and background in the camouflage environment. Because of the similarity between the target object and the background, these lines are difficult to be found accurately. However, these are easy to be observed in different frequency components of the image. To this end, in this paper we rethink COD from the perspective of frequency components and propose a Frequency Representation Integration Network to mine informative cues from them. Specifically, we obtain high-frequency components from the original image by Laplacian pyramid-like decomposition, and then respectively send the image to a transformer-based encoder and frequency components to a tailored CNN-based Residual Frequency Array Encoder. Besides, we utilize the multi-head self-attention in transformer encoder to capture low-frequency signals, which can effectively parse the overall contextual information of camouflage scenes. We also design a Frequency Representation Reasoning Module, which progressively eliminates discrepancies between differentiated frequency representations and integrates them by modeling their point-wise relations. Moreover, to further bridge different frequency representations, we introduce the image reconstruction task to implicitly guide their integration. Sufficient experiments on three widely-used COD benchmark datasets demonstrate that our method surpasses existing state-of-the-art methods by a large margin. Changqun Xia, Tianshu Yu 0003, Jia Li 0003 |
ACM Multimedia | 4 |
| 2023 | Salient Object Detection Using Reciprocal Learning
Changqun Xia, Tianshu Yu 0003, Jia Li 0003 |
PRCV (9) | 5 |
| 2023 | Semantic Contrastive Bootstrapping for Single-Positive Multi-label Recognition
Yifan Zhao 0002, Jia Li 0003 |
Int. J. Comput. Vis. | 3 |
| 2023 | Invariant and consistent: Unsupervised representation learning for few-shot visual recognition
Yifan Zhao 0002, Jia Li 0003 |
Neurocomputing | 3 |
| 2023 | From Pose to Part: Weakly-Supervised Pose Evolution for Human Part SegmentationabstractHuman part segmentation is a crucial but challenging task in computer vision. Recent works have achieved progress with the help of pixel-wise annotations. However, annotating pixel-wise masks especially at part-level is a tedious and labor-intensive procedure. To overcome this problem, we propose a part evolution framework to learn reliable predictions from weak pose annotations, which are much easier to collect. Our framework is composed of two essential modules: the first part adaptation module is designed to learn the deep prior knowledge from three related tasks, i.e., pose estimation, part-level and object-level segmentation; the second module is the part evolution module, which refines the part priors from deep predictions with the boundary-aware optimization algorithm. These two modules are conducted iteratively to evolve pose keypoint annotations into reliable part priors. Experimental evidence shows that our weakly-supervised approach generates comparable results with the state-of-the-art strongly-supervised methods on public benchmarks, and also validates the potential of notable improvements when combining weak labels with existing part segmentation masks. Yifan Zhao 0002, Jia Li 0003, Yu Zhang 0035, Yonghong Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Dual Adaptive Representation Alignment for Cross-Domain Few-Shot LearningabstractFew-shot learning aims to recognize novel queries with limited support samples by learning from base knowledge. Recent progress in this setting assumes that the base knowledge and novel query samples are distributed in the same domains, which are usually infeasible for realistic applications. Toward this issue, we propose to address the cross-domain few-shot learning problem where only extremely few samples are available in target domains. Under this realistic setting, we focus on the fast adaptation capability of meta-learners by proposing an effective dual adaptive representation alignment approach. In our approach, a prototypical feature alignment is first proposed to recalibrate support instances as prototypes and reproject these prototypes with a differentiable closed-form solution. Therefore feature spaces of learned knowledge can be adaptively transformed to query spaces by the cross-instance and cross-prototype relations. Besides the feature alignment, we further present a normalized distribution alignment module, which exploits prior statistics of query samples for solving the covariant shifts among the support and query samples. With these two modules, a progressive meta-learning framework is constructed to perform the fast adaptation with extremely few-shot samples while maintaining its generalization capabilities. Experimental evidence demonstrates our approach achieves new state-of-the-art results on 4 CDFSL benchmarks and 4 fine-grained cross-domain benchmarks. Yifan Zhao 0002, Jia Li 0003, Yonghong Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Rethinking Lightweight Salient Object Detection via Network Depth-Width TradeoffabstractExisting salient object detection methods often adopt deeper and wider networks for better performance, resulting in heavy computational burden and slow inference speed. This inspires us to rethink saliency detection to achieve a favorable balance between efficiency and accuracy. To this end, we design a lightweight framework while maintaining satisfying competitive accuracy. Specifically, we propose a novel trilateral decoder framework by decoupling the U-shape structure into three complementary branches, which are devised to confront the dilution of semantic context, loss of spatial structure and absence of boundary detail, respectively. Along with the fusion of three branches, the coarse segmentation results are gradually refined in structure details and boundary quality. Without adding additional learnable parameters, we further propose Scale-Adaptive Pooling Module to obtain multi-scale receptive field. In particular, on the premise of inheriting this framework, we rethink the relationship among accuracy, parameters and speed via network depth-width tradeoff. With these insightful considerations, we comprehensively design shallower and narrower models to explore the maximum potential of lightweight SOD. Our models are proposed for different application environments: 1) a tiny version CTD-S (1.7M, 125FPS) for resource constrained devices, 2) a fast version CTD-M (12.6M, 158FPS) for speed-demanding scenarios, 3) a standard version CTD-L (26.5M, 84FPS) for high-performance platforms. Extensive experiments validate the superiority of our method, which achieves better efficiency-accuracy balance across five benchmarks. Jia Li 0003, Shengye Qiao, Zhirui Zhao, Xiaowu Chen 0001, Changqun Xia |
IEEE Trans. Image Process. | 1 |
| 2023 | Boosting Broader Receptive Fields for Salient Object DetectionabstractSalient Object Detection has boomed in recent years and achieved impressive performance on regular-scale targets. However, existing methods encounter performance bottlenecks in processing objects with scale variation, especially extremely large- or small-scale objects with asymmetric segmentation requirements, since they are inefficient in obtaining more comprehensive receptive fields. With this issue in mind, this paper proposes a framework named BBRF for Boosting Broader Receptive Fields, which includes a Bilateral Extreme Stripping (BES) encoder, a Dynamic Complementary Attention Module (DCAM) and a Switch-Path Decoder (SPD) with a new boosting loss under the guidance of Loop Compensation Strategy (LCS). Specifically, we rethink the characteristics of the bilateral networks, and construct a BES encoder that separates semantics and details in an extreme way so as to get the broader receptive fields and obtain the ability to perceive extreme large- or small-scale objects. Then, the bilateral features generated by the proposed BES encoder can be dynamically filtered by the newly proposed DCAM. This module interactively provides spacial-wise and channel-wise dynamic attention weights for the semantic and detail branches of our BES encoder. Furthermore, we subsequently propose a Loop Compensation Strategy to boost the scale-specific features of multiple decision paths in SPD. These decision paths form a feature loop chain, which creates mutually compensating features under the supervision of boosting loss. Experiments on five benchmark datasets demonstrate that the proposed BBRF has a great advantage to cope with scale variation and can reduce the Mean Absolute Error over 20% compared with the state-of-the-art methods. Mingcan Ma, Changqun Xia, Xiaowu Chen 0001, Jia Li 0003 |
IEEE Trans. Image Process. | 5 |
| 2023 | MagConv: Mask-Guided Convolution for Image InpaintingabstractStandard convolution applied to image inpainting would lead to color discrepancy and blurriness for treating valid and invalid/hole regions without difference, which was partially amended by partial convolution (PConv). In PConv, a binary/hard mask was maintained as an indicator of valid and invalid pixels, where valid pixels and invalid pixels were treated differently. However, it can not describe validity degree of an impaired pixel. In addition, mask and image paths were separated, without sharing convolution kernel and exchanging information mutually, reducing data utilization efficiency. In this paper, a mask-guided convolution (MagConv) is proposed for image inpainting. In MagConv, mask and image paths share a convolution kernel to interact with each other and form a joint optimization scheme. In addition, a learnable piecewise activation function is raised to replace the reciprocal function of PConv, providing more flexible and adaptable compensation to convolution contaminated by invalid pixels. It also results in a soft mask of floating-point coefficients from 0 to 1 capable of indicating the validity degree of each pixel. Last but not least, MagConv splits the convolution kernel into positive and negative weights so that they can evaluate the validity of each pixel faithfully. Qualitative and quantitative experiments on the CelebA, Paris StreetView and Places2 datasets demonstrate that our method achieves favorable visual quality against state-of-the-art approaches. Xuexin Yu, Long Xu 0001, Jia Li 0003, Xiangyang Ji |
IEEE Trans. Image Process. | 3 |
| 2023 | View-Aware Salient Object Detection for $360^{\circ }$ Omnidirectional ImageabstractImage-based salient object detection (ISOD) in$360^{\circ }$scenarios is significant for understanding and applying panoramic information. However, research on$360^{\circ }$ISOD has not been widely explored due to the lack of large, complex, high-resolution, and well-labeled datasets. Towards this end, we construct a large scale$360^{\circ }$ISOD dataset with object-level pixel-wise annotation on equirectangular projection (ERP), which contains rich panoramic scenes with not less than 2K resolution and is the largest dataset for$360^{\circ }$ISOD by far to our best knowledge. By observing the data, we find current methods face three significant challenges in panoramic scenarios: diverse distortion degrees, discontinuous edge effects and changeable object scales. Inspired by humans' observing process, we propose a view-aware salient object detection method based on a Sample Adaptive View Transformer (SAVT) module with two sub-modules to mitigate these issues. Specifically, the sub-module View Transformer (VT) contains three transform branches based on different kinds of transformations to learn various features under different views and heighten the model's feature toleration of distortion, edge effects and object scales. Moreover, the sub-module Sample Adaptive Fusion (SAF) is to adjust the weights of different transform branches based on various sample features and make transformed enhanced features fuse more appropriately. The benchmark results of 20 state-of-the-art ISOD methods reveal the constructed dataset is very challenging. Moreover, exhaustive experiments verify the proposed approach is practical and outperforms the state-of-the-art methods. Changqun Xia, Tianshu Yu 0003, Jia Li 0003 |
IEEE Trans. Multim. | 4 |
| 2022 | Retinomorphic Object Detection in Asynchronous Visual StreamsabstractDue to high-speed motion blur and challenging illumination, conventional frame-based cameras have encountered an important challenge in object detection tasks. Neuromorphic cameras that output asynchronous visual streams instead of intensity frames, by taking the advantage of high temporal resolution and high dynamic range, have brought a new perspective to address the challenge. In this paper, we propose a novel problem setting, retinomorphic object detection, which is the first trial that integrates foveal-like and peripheral-like visual streams. Technically, we first build a large-scale multimodal neuromorphic object detection dataset (i.e., PKU-Vidar-DVS) over 215.5k spatio-temporal synchronized labels. Then, we design temporal aggregation representations to preserve the spatio-temporal information from asynchronous visual streams. Finally, we present a novel bio-inspired unifying framework to fuse two sensing modalities via a dynamic interaction mechanism. Our experimental evaluation shows that our approach has significant improvements over the state-of-the-art methods with the single-modality, especially in high-speed motion and low-light scenarios. We hope that our work will attract further research into this newly identified, yet crucial research direction. Our dataset can be available at https://www.pkuml.org/resources/pku-vidar-dvs.html. Jianing Li 0001, Xiao Wang 0014, Lin Zhu 0012, Jia Li 0003, Tiejun Huang 0001, Yonghong Tian 0001 |
AAAI | 4 |
| 2022 | Pyramid Grafting Network for One-Stage High Resolution Saliency DetectionabstractRecent salient object detection (SOD) methods based on deep neural network have achieved remarkable performance. However, most of existing SOD models designed for low-resolution input perform poorly on high-resolution images due to the contradiction between the sampling depth and the receptive field size. Aiming at resolving this con-tradiction, we propose a novel one-stage framework called Pyramid Grafting Network (PGNet), using transformer and CNN backbone to extract features from different resolution images independently and then graft the features from transformer branch to CNN branch. An attention-based Cross-Model Grafting Module (CMGM) is proposed to en-able CNN branch to combine broken detailed information more holistically, guided by different source feature during decoding process. Moreover, we design an Attention Guided Loss (AGL) to explicitly supervise the attention matrix generated by CMGM to help the network better interact with the attention from different models. We contribute a new Ultra-High-Resolution Saliency Detection dataset UHRSD, containing 5,920 images at 4K-SK resolutions. To our knowledge, it is the largest dataset in both quantity and resolution for high-resolution SOD task, which can be used for training and testing in future research. Sufficient exper-iments on UHRSD and widely-used SOD datasets demon-strate that our method achieves superior performance compared to the state-of-the-art methods. Changqun Xia, Mingcan Ma, Zhirui Zhao, Xiaowu Chen 0001, Jia Li 0003 |
CVPR | 6 |
| 2022 | ReAct: Temporal Action Detection with Relational Queries
Dingfeng Shi, Qiong Cao, Jing Zhang 0037, Lin Ma 0002, Jia Li 0003, Dacheng Tao |
ECCV (10) | 6 |
| 2022 | Revisiting Stochastic Learning for Generalizable Person Re-identificationabstractGeneralizable person re-identification aims to achieve a well generalization capability on target domains without accessing target data. Existing methods focus on suppressing domain-specific information or simulating unseen environments by meta-learning strategies, which could damage the capture ability on fine-grained visual patterns or lead to overfitting issues by the repetitive training of episodes. In this paper, we revisit the stochastic behaviors from two different perspectives: 1) Stochastic splitting-sliding sampler. It splits domain sources into approximately equal sample-size subsets and selects several subsets from various sources by a sliding window, forcing the model to step out of local minimums under stochastic sources. 2) Variance-varying gradient dropout. Gradients in parts of network are also selected by a sliding window and multiplied by binary masks generated from Bernoulli distribution, making gradients in varying variance and preventing the model from local minimums. By applying these two proposed stochastic behaviors, the model achieves a better generalization performance on unseen target domains without any additional computation costs or auxiliary modules. Extensive experiments demonstrate that our proposed model is effective and outperforms state-of-the-art methods on public domain generalizable person Re-ID benchmarks. Jiajian Zhao, Yifan Zhao 0002, Xiaowu Chen 0001, Jia Li 0003 |
ACM Multimedia | 4 |
| 2022 | Joint self-supervised and reference-guided learning for depth inpaintingabstractDepth information can benefit various computer vision tasks on both images and videos. However, depth maps may suffer from invalid values in many pixels, and also large holes. To improve such data, we propose a joint self-supervised and reference-guided learning approach for depth inpainting. For the self-supervised learning strategy, we introduce an improved spatial convolutional sparse coding module in which total variation regularization is employed to enhance the structural information while preserving edge information. This module alternately learns a convolutional dictionary and sparse coding from a corrupted depth map. Then, both the learned convolutional dictionary and sparse coding are convolved to yield an initial depth map, which is effectively smoothed using local contextual information. The reference-guided learning part is inspired by the fact that adjacent pixels with close colors in the RGB image tend to have similar depth values. We thus construct a hierarchical joint bilateral filter module using the corresponding color image to fill in large holes. In summary, our approach integrates a convolutional sparse coding module to preserve local contextual information and a hierarchical joint bilateral filter module for filling using specific adjacent information. Experimental results show that the proposed approach works well for both invalid value restoration and large hole inpainting. Kui Fu, Yifan Zhao 0002, Haokun Song, Jia Li 0003 |
Comput. Vis. Media | 5 |
| 2022 | Asynchronous Spatio-Temporal Memory Network for Continuous Event-Based Object DetectionabstractEvent cameras, offering extremely high temporal resolution and high dynamic range, have brought a new perspective to addressing common object detection challenges (e.g., motion blur and low light). However, how to learn a better spatio-temporal representation and exploit rich temporal cues from asynchronous events for object detection still remains an open issue. To address this problem, we propose a novel asynchronous spatio-temporal memory network (ASTMNet) that directly consumes asynchronous events instead of event images prior to processing, which can well detect objects in a continuous manner. Technically, ASTMNet learns an asynchronous attention embedding from the continuous event stream by adopting an adaptive temporal sampling strategy and a temporal attention convolutional module. Besides, a spatio-temporal memory module is designed to exploit rich temporal cues via a lightweight yet efficient inter-weaved recurrent-convolutional architecture. Empirically, it shows that our approach outperforms the state-of-the-art methods using the feed-forward frame-based detectors on three datasets by a large margin (i.e., 7.6% in the KITTI Simulated Dataset, 10.8% in the Gen1 Automotive Dataset, and 10.5% in the 1Mpx Detection Dataset). The results demonstrate that event cameras can perform robust object detection even in cases where conventional cameras fail, e.g., fast motion and challenging light conditions. Jianing Li 0001, Jia Li 0003, Lin Zhu 0012, Xijie Xiang, Tiejun Huang 0001, Yonghong Tian 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Pyramidal Feature Shrinking for Salient Object DetectionabstractRecently, we have witnessed the great progress of salient object detection (SOD), which benefits from the effectiveness of various feature aggregation strategies. However, existing methods usually aggregate the low-level features containing details and the high-level features containing semantics over a large span, which introduces noise into the aggregated features and generate inaccurate saliency map. To address this issue, we propose pyramidal feature shrinking network (PFSNet), which aims to aggregate adjacent feature nodes in pairs with layer-by-layer shrinkage, so that the aggregated features fuse effective details and semantics together and discard interference information. Specifically, pyramidal shrinking decoder (PSD) is proposed to aggregate adjacent features hierarchically in an asymptotic manner. Unlike other methods that aggregate features with significantly different information, this method only focuses on adjacent feature nodes in each layer and shrinks them to a final unique feature node. Besides, we propose adjacent fusion module (AFM) to perform mutual spatial enhancement between the adjacent features so as to dynamically weight the features and adaptively fuse the appropriate information. In addition, scale-aware enrichment module (SEM) based on the features extracted from backbone is utilized to obtain rich scale information and generate diverse initial features with dilated convolutions. Extensive quantitative and qualitative experiments demonstrate that the proposed intuitive framework outperforms 14 state-of-the-art approaches on 5 public datasets. Mingcan Ma, Changqun Xia, Jia Li 0003 |
AAAI | 3 |
| 2021 | Informative and Consistent Correspondence Mining for Cross-Domain Weakly Supervised Object DetectionabstractCross-domain weakly supervised object detection aims to adapt object-level knowledge from a fully labeled source domain dataset (i.e., with object bounding boxes) to train object detectors for target domains that are weakly labeled (i.e., with image-level tags). Instead of domain-level distribution matching, as popularly adopted in the literature, we propose to learn pixel-wise cross-domain correspondences for more precise knowledge transfer. It is realized through a novel cross-domain co-attention scheme trained as region competition. In this scheme, the cross-domain correspondence module seeks for informative features on the target domain image, which if warped to the source domain image, could best explain its annotations. Meanwhile, a collaborative mask generator competes to mask out the relevant target image region to make the remaining features uninformative. Such competitive learning strives to correlate the full foreground in cross-domain image pairs, revealing the accurate object extent in target domain. To alleviate the ambiguity of inter-domain correspondence learning, a domain-cycle consistency regularizer is further proposed to leverage the more reliable intra-domain correspondence. The proposed approach achieves consistent improvements over existing approaches by a considerable margin, demonstrated by the experiments on various datasets. Luwei Hou, Yu Zhang 0035, Kui Fu, Jia Li 0003 |
CVPR | 4 |
| 2021 | Graph-Based High-Order Relation Discovery for Fine-Grained RecognitionabstractFine-grained object recognition aims to learn effective features that can identify the subtle differences between visually similar objects. Most of the existing works tend to amplify discriminative part regions with attention mechanisms. Besides its unstable performance under complex backgrounds, the intrinsic interrelationship between different semantic features is less explored. Toward this end, we propose an effective graph-based relation discovery approach to build a contextual understanding of high-order relationships. In our approach, a high-dimensional feature bank is first formed and jointly regularized with semantic- and positional-aware high-order constraints, endowing rich attributes to feature representations. Second, to overcome the high-dimension curse, we propose a graph-based semantic grouping strategy to embed this high-order tensor bank into a low-dimensional space. Meanwhile, a group-wise learning strategy is proposed to regularize the features focusing on the cluster embedding center. With the collaborative learning of three modules, our module is able to grasp the stronger contextual details of fine-grained objects. Experimental evidence demonstrates our approach achieves new state-of-the-art on 4 widely-used fine-grained object recognition benchmarks. Yifan Zhao 0002, Ke Yan 0007, Feiyue Huang, Jia Li 0003 |
CVPR | 4 |
| 2021 | Amplitude-Phase Recombination: Rethinking Robustness of Convolutional Neural Networks in Frequency DomainabstractRecently, the generalization behavior of Convolutional Neural Networks (CNN) is gradually transparent through explanation techniques with the frequency components decomposition. However, the importance of the phase spectrum of the image for a robust vision system is still ignored. In this paper, we notice that the CNN tends to converge at the local optimum which is closely related to the high-frequency components of the training images, while the amplitude spectrum is easily disturbed such as noises or common corruptions. In contrast, more empirical studies found that humans rely on more phase components to achieve robust recognition. This observation leads to more explanations of the CNN’s generalization behaviors in both robustness to common perturbations and out-of-distribution detection, and motivates a new perspective on data augmentation designed by re-combing the phase spectrum of the current image and the amplitude spectrum of the distracter image. That is, the generated samples force the CNN to pay more attention to the structured information from phase components and keep robust to the variation of the amplitude. Experiments on several image datasets indicate that the proposed method achieves state-of-the-art performances on multiple generalizations and calibration tasks, including adaptability for common corruptions and surface variations, out-of-distribution detection, and adversarial attack. The code is released on github/iCGY96/APR. Peixi Peng, Li Ma 0009, Jia Li 0003, Lin Du 0010, Yonghong Tian 0001 |
ICCV | 4 |
| 2021 | Transformer-based Dual Relation Graph for Multi-label Image RecognitionabstractThe simultaneous recognition of multiple objects in one image remains a challenging task, spanning multiple events in the recognition field such as various object scales, inconsistent appearances, and confused inter-class relationships. Recent research efforts mainly resort to the statistic label co-occurrences and linguistic word embedding to enhance the unclear semantics. Different from these researches, in this paper, we propose a novel Transformer-based Dual Relation learning framework, constructing complementary relationships by exploring two aspects of correlation, i.e., structural relation graph and semantic relation graph. The structural relation graph aims to capture long-range correlations from object context, by developing a cross-scale transformer-based architecture. The semantic graph dynamically models the semantic meanings of image objects with explicit semantic-aware constraints. In addition, we also incorporate the learnt structural relationship into the semantic graph, constructing a joint relation graph for robust representations. With the collaborative learning of these two effective relation graphs, our approach achieves new state-of-the-art on two popular multi-label recognition benchmarks, i.e. MS-COCO and VOC 2007 dataset. Ke Yan 0007, Yifan Zhao 0002, Feiyue Huang, Jia Li 0003 |
ICCV | 6 |
| 2021 | Heterogeneous Relational Complement for Vehicle Re-identificationabstractThe crucial problem in vehicle re-identification is to find the same vehicle identity when reviewing this object from cross-view cameras, which sets a higher demand for learning viewpoint-invariant representations. In this paper, we propose to solve this problem from two aspects: constructing robust feature representations and proposing camera-sensitive evaluations. We first propose a novel Heterogeneous Relational Complement Network (HRCN) by incorporating region-specific features and cross-level features as complements for the original high-level output. Considering the distributional differences and semantic misalignment, we propose graph-based relation modules to embed these heterogeneous features into one unified high-dimensional space. On the other hand, considering the deficiencies of cross-camera evaluations in existing measures (i.e., CMC and AP), we then propose a Cross-camera Generalization Measure (CGM) to improve the evaluations by introducing position-sensitivity and cross-camera generalization penalties. We further construct a new benchmark of existing models with our proposed CGM and experimental results reveal that our proposed HRCN model achieves new state-of-the-art in VeRi-776, VehicleID, and VERI-Wild. Jiajian Zhao, Yifan Zhao 0002, Jia Li 0003, Ke Yan 0007, Yonghong Tian 0001 |
ICCV | 3 |
| 2021 | Exploring Driving-Aware Salient Object Detection via Knowledge TransferabstractRecently, general salient object detection (SOD) has made great progress with the rapid development of deep neural networks. However, task-aware SOD has hardly been studied due to the lack of task-specific datasets. In this paper, we construct a driving task-oriented dataset where pixel-level masks of salient objects have been annotated. Comparing with general SOD datasets, we find that the cross-domain knowledge difference and task-specific scene gap are two main challenges to focus the salient objects when driving. Inspired by these findings, we proposed a baseline model for the driving task-aware SOD via a knowledge transfer convolutional neural network. In this network, we construct an attention-based knowledge transfer module to make up the knowledge difference. In addition, an efficient boundary-aware feature decoding module is introduced to perform fine feature decoding for objects in the complex task-specific scenes. The whole network integrates the knowledge transfer and feature decoding modules in a progressive manner. Experiments show that the proposed dataset is very challenging, and the proposed method outperforms 12 state-of-the-art methods on the dataset, which facilitates the development of task-aware SOD. Jinming Su, Changqun Xia, Jia Li 0003 |
ICME | 3 |
| 2021 | Selective, Structural, Subtle: Trilinear Spatial-Awareness for Few-Shot Fine-Grained Visual RecognitionabstractFew-shot learning aims to recognize the novel categories from a few examples. However, most of the existing approaches usually focus on general image classification and fail to handle subtle differences between images. To alleviate this issue, we propose a trilinear spatial-awareness network for few-shot-grained visual recognition, called S3Net, which is composed of a spatial selection module, structural pyramid descriptor, and subtle difference mining module. Specifically, we first build the global relation to strengthen the features by spatial selection module. The structural pyramid descriptor then constructs a multi-scale representation for enhancing the rich contextual information by exploiting different receptive fields in the same feature layer. Furthermore, a similarity loss based on local descriptors and a global classification loss is design to help the network learn discrimination capability by handling subtle differences in confused or near-duplicated samples. Extensive experiments on 4 few-shot fine-grained benchmarks demonstrate that our proposed approach is effective and outperforms state-of-the-art models by large margins. Yifan Zhao 0002, Jia Li 0003 |
ICME | 3 |
| 2021 | How to Learn a Domain-Adaptive Event Simulator?abstractThe low-latency streams captured by event cameras have shown impressive potential in addressing vision tasks such as video reconstruction and optical flow estimation. However, these tasks often require massive training event streams, which are expensive to collect and largely bypassed by recently proposed event camera simulators. To align the statistics of synthetic events with that of target event cameras, existing simulators often need to be heuristically tuned with elaborative manual efforts and thus become incompetent to automatically adapt to various domains. To address this issue, this work proposes one of the first learning-based, domain-adaptive event simulator. Given a specific domain, the proposed simulator learns pixel-wise distributions of event contrast thresholds that, after stochastic sampling and paralleled rendering, can generate event representations well aligned with those from the data from realistic event cameras. To achieve such domain-specific alignment, we design a novel divide-and-conquer discrimination scheme that adaptively evaluates the synthetic-to-real consistency of event representations according to the local statistics of images and events. Trained with the data synthesized by the proposed simulator, the performances of state-of-the-art event-based video reconstruction and optical flow estimation approaches are boosted up to 22.9% and 2.8%, respectively. In addition, we show significantly improved domain adaptation capability over existing event simulators and tuning strategies, consistently on three real event datasets. Daxin Gu, Jia Li 0003, Yu Zhang 0035, Yonghong Tian 0001 |
ACM Multimedia | 2 |
| 2021 | DehazeFlow: Multi-scale Conditional Flow Network for Single Image DehazingabstractSingle image dehazing is a crucial and preliminary task for many computer vision applications, making progress with deep learning. The dehazing task is an ill-posed problem since the haze in the image leads to the loss of information. Thus, there are multiple feasible solutions for image restoration of a hazy image. Most existing methods learn a deterministic one-to-one mapping between a hazy image and its ground-truth, which ignores the ill-posedness of the dehazing task. To solve this problem, we propose DehazeFlow, a novel single image dehazing framework based on conditional normalizing flow. Our method learns the conditional distribution of haze-free images given a hazy image, enabling the model to sample multiple dehazed results. Furthermore, we propose an attention-based coupling layer to enhance the expression ability of a single flow step, which converts natural images into latent space and fuses features of paired data. These designs enable our model to achieve state-of-the-art performance while considering the ill-posedness of the task. We carry out sufficient experiments on both synthetic datasets and real-world hazy images to illustrate the effectiveness of our method. The extensive experiments indicate that DehazeFlow surpasses the state-of-the-art methods in terms of PSNR, SSIM, LPIPS, and subjective visual effects. Jia Li 0003, Dong Zhao 0008, Long Xu 0001 |
ACM Multimedia | 2 |
| 2021 | Pose-guided Inter- and Intra-part Relational Transformer for Occluded Person Re-IdentificationabstractPerson Re-Identification (Re-Id) in occlusion scenarios is a challenging problem because a pedestrian can be partially occluded. The use of local information for feature extraction and matching is still necessary. Therefore, we propose a Pose-guided inter- and intra-part relational transformer (Pirt) for occluded person Re-Id, which builds part-aware long-term correlations by introducing transformer. In our framework, we firstly develop a pose-guided feature extraction module with regional grouping and mask construction for robust feature representations. The positions of a pedestrian in the image under surveillance scenarios are relatively fixed, hence we propose intra-part and inter-part relational transformer. The intra-part module creates local relations with mask-guided features, while the inter-part relationship builds correlations with transformers, to develop cross relationships between part nodes. With the collaborative learning inter- and intra-part relationships, experiments reveal that our proposed Pirt model achieves a new state of the art on the public occluded dataset, and further extensions on standard non-occluded person Re-Id datasets also reveal our comparable performances. Zhongxing Ma, Yifan Zhao 0002, Jia Li 0003 |
ACM Multimedia | 3 |
| 2021 | Complementary Trilateral Decoder for Fast and Accurate Salient Object DetectionabstractSalient object detection (SOD) has made great progress, but most of existing SOD methods focus more on performance than efficiency. Besides, the U-shape structure exists some drawbacks and there is still a lot of room for improvement. Therefore, we propose a novel framework to treat semantic context, spatial detail and boundary information separately in the decoder part. Specifically, we propose an efficient and effective Complementary Trilateral Decoder (CTD) for saliency detection with three branches: Semantic Path, Spatial Path and Boundary Path. These three branches are designed to solve the dilution of semantic information, loss of spatial information and absence of boundary information, respectively. These three branches are complementary to each other and we design three distinctive fusion modules to gradually merge them according to "coarse-fine-finer'' strategy, which significantly improves the region accuracy and boundary quality. To facilitate the practical application in different environments, we provide two versions: CTDNet-18 (11.82M, 180FPS) and CTDNet-50 (24.63M, 110FPS). Experiments show that our model performs better than state-of-the-art approaches on five benchmarks, which achieves a favorable balance between speed and accuracy. Zhirui Zhao, Changqun Xia, Jia Li 0003 |
ACM Multimedia | 4 |
| 2021 | M3TR: Multi-modal Multi-label Recognition with TransformerabstractMulti-label image recognition aims to recognize multiple objects simultaneously in one image. Recent ideas to solve this problem have focused on learning dependencies of label co-occurrences to enhance the high-level semantic representations. However, these methods usually neglect the important relations of intrinsic visual structures and face difficulties in understanding contextual relationships. To build the global scope of visual context as well as interactions between visual modality and linguistic modality, we propose the Multi-Modal Multi-label recognition TRansformers (M3TR) with the ternary relationship learning for inter-and intra-modalities. For the intra-modal relationship, we make insightful conjunction of CNNs and Transformers, which embeds visual structures into the high-level features by learning the semantic cross-attention. For constructing the interactions between the visual and linguistic modalities, we propose a linguistic cross-attention to embed the class-wise linguistic information into the visual structure learning, and finally present a linguistic guided enhancement module to enhance the representation of high-level semantics. Experimental evidence reveals that with the collaborative learning of ternary relationship, our proposed M3TR achieves new state-of-the-art on two public multi-label recognition benchmarks. Yifan Zhao 0002, Jia Li 0003 |
ACM Multimedia | 3 |
| 2021 | Learning From Large-Scale Noisy Web Data With Ubiquitous Reweighting for Image ClassificationabstractMany important advances of deep learning techniques have originated from the efforts of addressing the image classification task on large-scale datasets. However, the construction of clean datasets is costly and time-consuming since the Internet is overwhelmed by noisy images with inadequate and inaccurate tags. In this paper, we propose a Ubiquitous Reweighting Network (URNet) that can learn an image classification model from noisy web data. By observing the web data, we find that there are five key challenges, i.e., imbalanced class sizes, high intra-classes diversity and inter-class similarity, imprecise instances, insufficient representative instances, and ambiguous class labels. With these challenges in mind, we assume every training instance has the potential to contribute positively by alleviating the data bias and noise via reweighting the influence of each instance according to different class sizes, large instance clusters, its confidence, small instance bags, and the labels. In this manner, the influence of bias and noise in the data can be gradually alleviated, leading to the steadily improving performance of URNet. Experimental results in the WebVision 2018 challenge with 16 million noisy training images from 5000 classes show that our approach outperforms state-of-the-art models and ranks first place in the image classification task. Jia Li 0003, Yafei Song 0002, Lele Cheng, Pengcheng Yuan, Shumin Han |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Ordinal Multi-Task Part Segmentation With Recurrent Prior GenerationabstractSemantic object part segmentation is a fundamental task in object understanding and geometric analysis. The clear understanding of part relationships can be of great use to the segmentation process. In this work, we propose a novel Ordinal Multi-task Part Segmentation (OMPS) approach which explicitly models the part ordinal relationship to guide the segmentation process in a recurrent manner. Quantitative and qualitative experiments are conducted first to explore the mutual impacts among object parts and then an ordinal part inference algorithm is formulated via experimental observations. Specifically, our framework is mainly composed of two modules, the forward module to segment multiple parts as individual subtasks with prior knowledge, and the recurrent module to generate appropriate part priors with the ordinal inference algorithm. These two modules work iteratively to optimize the segmentation performance and the network parameters. Experimental results show that our approach outperforms the state-of-the-art models on human and vehicle part parsing benchmarks. Comprehensive evaluations are conducted to demonstrate the effectiveness of our approach in object part segmentation. Yifan Zhao 0002, Jia Li 0003, Yu Zhang 0035, Yafei Song 0002, Yonghong Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Pyramid Global Context Network for Image DehazingabstractHaze caused by atmospheric scattering and absorption would severely affect scene visibility of an image. Thus, image dehazing for haze removal has been widely studied in the literature. Within a hazy image, haze is not confined in a small local patch/position, while widely diffusing in a whole image. Under this circumstance, global context is a crucial factor in the success of dehazing, which was seldom investigated in existing dehazing algorithms. In the literature, the global context (GC) block has been designed to learn point-wise long-range dependencies of an image for global context modeling; however, patch-wise long-range dependencies were ignored. To image dehazing, patch-wise long-range dependencies should be highlighted to cooperate with patch-wise operations of image dehazing. In this paper, we first extend the point-wise GC into a Pyramid Global Context (PGC), which is a multi-scale GC, after undergoing the pyramid pooling. Thus, patch-wise long-range dependencies can be explored by the PGC. Then, the proposed PGC is plugged into a U-Net, getting an attentive U-Net. Further, the attentive U-Net is optimized by importing ResNet's shortcut connection and dilated convolution. Thus, the finalized dehazing model can explore both long-range and patch-wise context dependencies for global context modeling, which is crucial for image dehazing. The extensive experiments on synthetic databases and real-world hazy images demonstrate the superiority of our model over other representative state-of-the-art models from both quantitative and qualitative comparisons. Dong Zhao 0016, Long Xu 0001, Lin Ma 0002, Jia Li 0003, Yihua Yan |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | DanceIt: Music-Inspired Dancing Video SynthesisabstractClose your eyes and listen to music, one can easily imagine an actor dancing rhythmically along with the music. These dance movements are usually made up of dance movements you have seen before. In this paper, we propose to reproduce such an inherent capability of the human-being within a computer vision system. The proposed system consists of three modules. To explore the relationship between music and dance movements, we propose a cross-modal alignment module that focuses on dancing video clips, accompanied by pre-designed music, to learn a system that can judge the consistency between the visual features of pose sequences and the acoustic features of music. The learned model is then used in the imagination module to select a pose sequence for the given music. Such pose sequence selected from the music, however, is usually discontinuous. To solve this problem, in the spatial-temporal alignment module we develop a spatial alignment algorithm based on the tendency and periodicity of dance movements to predict dance movements between discontinuous fragments. In addition, the selected pose sequence is often misaligned with the music beat. To solve this problem, we further develop a temporal alignment algorithm to align the rhythm of music and dance. Finally, the processed pose sequence is used to synthesize realistic dancing videos in the imagination module. The generated dancing videos match the content and rhythm of the music. Experimental results and subjective evaluations show that the proposed approach can perform the function of generating promising dancing videos by inputting music. Yifan Zhao 0002, Jia Li 0003 |
IEEE Trans. Image Process. | 3 |
| 2021 | Salient Object Detection With Purificatory Mechanism and Structural Similarity LossabstractImage-based salient object detection has made great progress over the past decades, especially after the revival of deep neural networks. By the aid of attention mechanisms to weight the image features adaptively, recent advanced deep learning-based models encourage the predicted results to approximate the ground-truth masks with as large predictable areas as possible, thus achieving the state-of-the-art performance. However, these methods do not pay enough attention to small areas prone to misprediction. In this way, it is still tough to accurately locate salient objects due to the existence of regions with indistinguishable foreground and background and regions with complex or fine structures. To address these problems, we propose a novel convolutional neural network with purificatory mechanism and structural similarity loss. Specifically, in order to better locate preliminary salient objects, we first introduce the promotion attention, which is based on spatial and channel attention mechanisms to promote attention to salient regions. Subsequently, for the purpose of restoring the indistinguishable regions that can be regarded as error-prone regions of one model, we propose the rectification attention, which is learned from the areas of wrong prediction and guide the network to focus on error-prone regions thus rectifying errors. Through these two attentions, we use the Purificatory Mechanism to impose strict weights with different regions of the whole salient objects and purify results from hard-to-distinguish regions, thus accurately predicting the locations and details of salient objects. In addition to paying different attention to these hard-to-distinguish regions, we also consider the structural constraints on complex regions and propose the Structural Similarity Loss. The proposed loss models the region-level pair-wise relationship between regions to assist these regions to calibrate their own saliency values. In experiments, the proposed purificatory mechanism and structural similarity loss can both effectively improve the performance, and the proposed approach outperforms 19 state-of-the-art methods on six datasets with a notable margin. Also, the proposed method is efficient and runs at over 27FPS on a single NVIDIA 1080Ti GPU. Jia Li 0003, Jinming Su, Changqun Xia, Mingcan Ma, Yonghong Tian 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Part-Guided Relational Transformers for Fine-Grained Visual RecognitionabstractFine-grained visual recognition is to classify objects with visually similar appearances into subcategories, which has made great progress with the development of deep CNNs. However, handling subtle differences between different subcategories still remains a challenge. In this paper, we propose to solve this issue in one unified framework from two aspects, i.e., constructing feature-level interrelationships, and capturing part-level discriminative features. This framework, namely PArt-guided Relational Transformers (PART), is proposed to learn the discriminative part features with an automatic part discovery module, and to explore the intrinsic correlations with a feature transformation module by adapting the Transformer models from the field of natural language processing. The part discovery module efficiently discovers the discriminative regions which are highly-corresponded to the gradient descent procedure. Then the second feature transformation module builds correlations within the global embedding and multiple part embedding, enhancing spatial interactions among semantic pixels. Moreover, our proposed approach does not rely on additional part branches in the inference time and reaches state-of-the-art performance on 3 widely-used fine-grained object recognition benchmarks. Experimental results and explainable visualizations demonstrate the effectiveness of our proposed approach. Yifan Zhao 0002, Jia Li 0003, Xiaowu Chen 0001, Yonghong Tian 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | RGB-D Salient Object Detection With Ubiquitous Target AwarenessabstractConventional RGB-D salient object detection methods aim to leverage depth as complementary information to find the salient regions in both modalities. However, the salient object detection results heavily rely on the quality of captured depth data which sometimes are unavailable. In this work, we make the first attempt to solve the RGB-D salient object detection problem with a novel depth-awareness framework. This framework only relies on RGB data in the testing phase, utilizing captured depth data as supervision for representation learning. To construct our framework as well as achieving accurate salient detection results, we propose a Ubiquitous Target Awareness (UTA) network to solve three important challenges in RGB-D SOD task: 1) a depth awareness module to excavate depth information and to mine ambiguous regions via adaptive depth-error weights, 2) a spatial-aware cross-modal interaction and a channel-aware cross-level interaction, exploiting the low-level boundary cues and amplifying high-level salient channels, and 3) a gated multi-scale predictor module to perceive the object saliency in different contextual scales. Besides its high performance, our proposed UTA network is depth-free for inference and runs in real-time with 43 FPS. Experimental evidence demonstrates that our proposed network not only surpasses the state-of-the-art methods on five public RGB-D SOD benchmarks by a large margin, but also verifies its extensibility on five public RGB SOD benchmarks. Yifan Zhao 0002, Jia Li 0003, Xiaowu Chen 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | Ultrafast Video Attention Prediction with Coupled Knowledge DistillationabstractLarge convolutional neural network models have recently demonstrated impressive performance on video attention prediction. Conventionally, these models are with intensive computation and large memory. To address these issues, we design an extremely light-weight network with ultrafast speed, named UVA-Net. The network is constructed based on depth-wise convolutions and takes low-resolution images as input. However, this straight-forward acceleration method will decrease performance dramatically. To this end, we propose a coupled knowledge distillation strategy to augment and train the network effectively. With this strategy, the model can further automatically discover and emphasize implicit useful cues contained in the data. Both spatial and temporal knowledge learned by the high-resolution complex teacher networks also can be distilled and transferred into the proposed low-resolution light-weight spatiotemporal network. Experimental results show that the performance of our model is comparable to 11 state-of-the-art models in video attention prediction, while it costs only 0.68 MB memory footprint, runs about 10,106 FPS on GPU and 404 FPS on CPU, which is 206 times faster than previous models. Kui Fu, Peipei Shi, Yafei Song 0002, Shiming Ge, Xiangju Lu, Jia Li 0003 |
AAAI | 6 |
| 2020 | Learning Open Set Network with Discriminative Reciprocal Points
Limeng Qiao, Yemin Shi 0001, Peixi Peng, Jia Li 0003, Tiejun Huang 0001, Shiliang Pu, Yonghong Tian 0001 |
ECCV (3) | 5 |
| 2020 | Reconstructing Part-Level 3D Models From a Single ImageabstractUnderstanding an image with 3D representations has been an increasingly attractive topic in computer vision. The state-of the-art 3D reconstruction methods usually focus on the reconstruction of the holistic object, while missing important part information, which is crucial in robotic interaction and virtual reality applications. To solve this issue, we make the first attempt to reconstruct the 3D models with part-level representations in a unified framework. With the input of the singleview images, we first develop a feature enhancement encoder to incorporate discriminative local features into the feature representation. The local features are selected adaptively by a learnable local awareness module. Then the enhanced local features are fused with the global branch to form the 3D representations. We then develop a 3D part generator to decode the image priors to 3D parts with a 3D focal loss, which enables the representations of small parts. Experimental results indicate that our model generates reliable part-level structures while achieving state-of-the-art performance in object-level recovering. Dingfeng Shi, Yifan Zhao 0002, Jia Li 0003 |
ICME | 3 |
| 2020 | Look Through Masks: Towards Masked Face Recognition with De-Occlusion DistillationabstractMany real-world applications today like video surveillance and urban governance need to address the recognition of masked faces, where content replacement by diverse masks often brings in incomplete appearance and ambiguous representation, leading to a sharp drop in accuracy. Inspired by recent progress on amodal perception, we propose to migrate the mechanism of amodal completion for the task of masked face recognition with an end-to-end de-occlusion distillation framework, which consists of two modules. The de-occlusion module applies a generative adversarial network to perform face completion, which recovers the content under the mask and eliminates appearance ambiguity. The distillation module takes a pre-trained general face recognition model as the teacher and transfers its knowledge to train a student for completed faces using massive online synthesized face pairs. Especially, the teacher knowledge is represented with structural relations among instances in multiple orders, which serves as a posterior regularization to enable the adaptation. In this way, the knowledge can be fully distilled and transferred to identify masked faces. Experiments on synthetic and realistic datasets show the efficacy of the proposed approach. Chenyu Li 0001, Shiming Ge, Daichi Zhang, Jia Li 0003 |
ACM Multimedia | 4 |
| 2020 | Cooperative Bi-path Metric for Few-shot LearningabstractGiven base classes with sufficient labeled samples, the target of few-shot classification is to recognize unlabeled samples of novel classes with only a few labeled samples. Most existing methods only pay attention to the relationship between labeled and unlabeled samples of novel classes, which do not make full use of information within base classes. In this paper, we make two contributions to investigate the few-shot classification problem. First, we report a simple and effective baseline trained on base classes in the way of traditional supervised learning, which can achieve comparable results to the state of the art. Second, based on the baseline, we propose a cooperative bi-path metric for classification, which leverages the correlations between base classes and novel classes to further improve the accuracy. Experiments on two widely used benchmarks show that our method is a simple and effective framework, and a new state of the art is established in the few-shot classification field. Yifan Zhao 0002, Jia Li 0003, Yonghong Tian 0001 |
ACM Multimedia | 3 |
| 2020 | Is Depth Really Necessary for Salient Object Detection?abstractSalient object detection (SOD) is a crucial and preliminary task for many computer vision applications, which have made progress with deep CNNs. Most of the existing methods mainly rely on the RGB information to distinguish the salient objects, which faces difficulties in some complex scenarios. To solve this, many recent RGBD-based networks are proposed by adopting the depth map as an independent input and fuse the features with RGB information. Taking the advantages of RGB and RGBD methods, we propose a novel depth-aware salient object detection framework, which has following superior designs: 1) It does not rely on depth data in the testing phase. 2) It comprehensively optimizes SOD features with multi-level depth-aware regularizations. 3) The depth information also serves as error-weighted map to correct the segmentation process. With these insightful designs combined, we make the first attempt in realizing an unified depth-aware framework with only RGB information as input for inference, which not only surpasses the state-of-the-art performance on five public RGB SOD benchmarks, but also surpasses the RGBD-based methods on five benchmarks by a large margin, while adopting less information and implementation light-weighted. Yifan Zhao 0002, Jia Li 0003, Xiaowu Chen 0001 |
ACM Multimedia | 3 |
| 2020 | Cartoon Face Recognition: A Benchmark DatasetabstractRecent years have witnessed increasing attention in cartoon media, powered by the strong demands of industrial applications. As the first step to understand this media, cartoon face recognition is a crucial but less-explored task with few datasets proposed. In this work, we first present a new challenging benchmark dataset, consisting of 389,678 images of 5,013 cartoon characters annotated with identity, bounding box, pose, and other auxiliary attributes. The dataset, named iCartoonFace, is currently the largest-scale, high-quality, rich-annotated, and spanning multiple occurrences in the field of image recognition, including near-duplications, occlusions, and appearance changes. In addition, we provide two types of annotations for cartoon media, i.e., face recognition, and face detection, with the help of a semi-automatic labeling algorithm. To further investigate this challenging dataset, we propose a multi-task domain adaptation approach that jointly utilizes the human and cartoon domain knowledge with three discriminative regularizations. We hence perform a benchmark analysis of the proposed dataset and verify the superiority of the proposed approach in the cartoon face recognition task. The dataset is available at https://iqiyi.cn/icartoonface. Yifan Zhao 0002, Mengyuan Ren, Xiangju Lu, Jia Li 0003 |
ACM Multimedia | 7 |
| 2020 | Model-Guided Multi-Path Knowledge Aggregation for Aerial Saliency PredictionabstractAs an emerging vision platform, a drone can look from many abnormal viewpoints which brings many new challenges into the classic vision task of video saliency prediction. To investigate these challenges, this paper proposes a large-scale video dataset for aerial saliency prediction, which consists of ground-truth salient object regions of 1,000 aerial videos, annotated by 24 subjects. To the best of our knowledge, it is the first large-scale video dataset that focuses on visual saliency prediction on drones. Based on this dataset, we propose a Model-guided Multi-path Network (MM-Net) that serves as a baseline model for aerial video saliency prediction. Inspired by the annotation process in eye-tracking experiments, MM-Net adopts multiple information paths, each of which is initialized under the guidance of a classic saliency model. After that, the visual saliency knowledge encoded in the most representative paths is selected and aggregated to improve the capability of MM-Net in predicting spatial saliency in aerial scenarios. Finally, these spatial predictions are adaptively combined with the temporal saliency predictions via a spatiotemporal optimization algorithm. Experimental results show that MM-Net outperforms ten state-of-the-art models in predicting aerial video saliency. Kui Fu, Jia Li 0003, Yu Zhang 0035, Hongze Shen, Yonghong Tian 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Efficient Low-Resolution Face Recognition via Bridge DistillationabstractFace recognition in the wild is now advancing towards light-weight models, fast inference speed and resolution-adapted capability. In this paper, we propose a bridge distillation approach to turn a complex face model pretrained on private high-resolution faces into a light-weight one for low-resolution face recognition. In our approach, such a cross-dataset resolution-adapted knowledge transfer problem is solved via two-step distillation. In the first step, we conduct cross-dataset distillation to transfer the prior knowledge from private high-resolution faces to public high-resolution faces and generate compact and discriminative features. In the second step, the resolution-adapted distillation is conducted to further transfer the prior knowledge to synthetic low-resolution faces via multi-task learning. By learning low-resolution face representations and mimicking the adapted high-resolution knowledge, a light-weight student model can be constructed with high efficiency and promising accuracy in recognizing low-resolution faces. Experimental results show that the student model performs impressively in recognizing low-resolution faces with only 0.21M parameters and 0.057MB memory. Meanwhile, its speed reaches up to 14,705, 934 and 763 faces per second on GPU, CPU and mobile phone, respectively. Shiming Ge, Shengwei Zhao, Chenyu Li 0001, Yu Zhang 0035, Jia Li 0003 |
IEEE Trans. Image Process. | 5 |
| 2020 | Spatiotemporal Knowledge Distillation for Efficient Estimation of Aerial Video SaliencyabstractThe performance of video saliency estimation techniques has achieved significant advances along with the rapid development of Convolutional Neural Networks (CNNs). However, devices like cameras and drones may have limited computational capability and storage space so that the direct deployment of complex deep saliency models becomes infeasible. To address this problem, this paper proposes a dynamic saliency estimation approach for aerial videos via spatiotemporal knowledge distillation. In this approach, five components are involved, including two teachers, two students and the desired spatiotemporal model. The knowledge of spatial and temporal saliency is first separately transferred from the two complex and redundant teachers to their simple and compact students, while the input scenes are also degraded from high-resolution to low-resolution to remove the probable data redundancy so as to greatly speed up the feature extraction process. After that, the desired spatiotemporal model is further trained by distilling and encoding the spatial and temporal saliency knowledge of two students into a unified network. In this manner, the inter-model redundancy can be removed for the effective estimation of dynamic saliency on aerial videos. Experimental results show that the proposed approach is comparable to 11 state-of-the-art models in estimating visual saliency on aerial videos, while its speed reaches up to 28,738 FPS and 1,490.5 FPS on the GPU and CPU platforms, respectively. Jia Li 0003, Kui Fu, Shengwei Zhao, Shiming Ge |
IEEE Trans. Image Process. | 1 |
| 2019 | Part-Regularized Near-Duplicate Vehicle Re-IdentificationabstractVehicle re-identification (Re-ID) has been attracting more interests in computer vision owing to its great contributions in urban surveillance and intelligent transportation. With the development of deep learning approaches, vehicle Re-ID still faces a near-duplicate challenge, which is to distinguish different instances with nearly identical appearances. Previous methods simply rely on the global visual features to handle this problem. In this paper, we proposed a simple but efficient part-regularized discriminative feature preserving method which enhances the perceptive ability of subtle discrepancies. We further develop a novel framework to integrate part constrains with the global Re-ID modules by introducing an detection branch. Our framework is trained end-to-end with combined local and global constrains. Specially, without the part-regularized local constrains in inference step, our Re-ID network outperforms the state-of-the-art method by a large margin on large benchmark datasets VehicleID and VeRi-776. Jia Li 0003, Yifan Zhao 0002, Yonghong Tian 0001 |
CVPR | 2 |
| 2019 | Transductive Episodic-Wise Adaptive Metric for Few-Shot LearningabstractFew-shot learning, which aims at extracting new concepts rapidly from extremely few examples of novel classes, has been featured into the meta-learning paradigm recently. Yet, the key challenge of how to learn a generalizable classifier with the capability of adapting to specific tasks with severely limited data still remains in this domain. To this end, we propose a Transductive Episodic-wise Adaptive Metric (TEAM) framework for few-shot learning, by integrating the meta-learning paradigm with both deep metric learning and transductive inference. With exploring the pairwise constraints and regularization prior within each task, we explicitly formulate the adaptation procedure into a standard semi-definite programming problem. By solving the problem with its closed-form solution on the fly with the setup of transduction, our approach efficiently tailors an episodic-wise metric for each task to adapt all features from a shared task-agnostic embedding space into a more discriminative task-specific metric space. Moreover, we further leverage an attention-based bi-directional similarity strategy for extracting the more robust relationship between queries and prototypes. Extensive experiments on three benchmark datasets show that our framework is superior to other existing approaches and achieves the state-of-the-art performance in the few-shot literature. Limeng Qiao, Yemin Shi 0001, Jia Li 0003, Yonghong Tian 0001, Tiejun Huang 0001, Yaowei Wang 0001 |
ICCV | 3 |
| 2019 | Selectivity or Invariance: Boundary-Aware Salient Object DetectionabstractTypically, a salient object detection (SOD) model faces opposite requirements in processing object interiors and boundaries. The features of interiors should be invariant to strong appearance change so as to pop-out the salient object as a whole, while the features of boundaries should be selective to slight appearance change to distinguish salient objects and background. To address this selectivity-invariance dilemma, we propose a novel boundary-aware network with successive dilation for image-based SOD. In this network, the feature selectivity at boundaries is enhanced by incorporating a boundary localization stream, while the feature invariance at interiors is guaranteed with a complex interior perception stream. Moreover, a transition compensation stream is adopted to amend the probable failures in transitional regions between interiors and boundaries. In particular, an integrated successive dilation module is proposed to enhance the feature invariance at interiors and transitional regions. Extensive experiments on six datasets show that the proposed approach outperforms 16 state-of-the-art methods. Jinming Su, Jia Li 0003, Yu Zhang 0035, Changqun Xia, Yonghong Tian 0001 |
ICCV | 2 |
| 2019 | Multi-Class Part Parsing With Joint Boundary-Semantic AwarenessabstractObject part parsing in the wild, which requires to simultaneously detect multiple object classes in the scene and accurately segments semantic parts within each class, is challenging for the joint presence of class-level and part-level ambiguities. Despite its importance, however, this problem is not sufficiently explored in existing works. In this paper, we propose a joint parsing framework with boundary and semantic awareness to address this challenging problem. To handle part-level ambiguity, a boundary awareness module is proposed to make mid-level features at multiple scales attend to part boundaries for accurate part localization, which are then fused with high-level features for effective part recognition. For class-level ambiguity, we further present a semantic awareness module that selects discriminative part features relevant to a category to prevent irrelevant features being merged together. The proposed modules are lightweight and implementation friendly, improving the performance substantially when plugged into various baseline architectures. Without bells and whistles, the full model sets new state-of-the-art results on the Pascal-Part dataset, in both multi-class and the conventional single-class setting, while running substantially faster than recent high-performance approaches. Yifan Zhao 0002, Jia Li 0003, Yu Zhang 0035, Yonghong Tian 0001 |
ICCV | 2 |
| 2019 | Learning Local Feature Descriptor with Motion Attribute For Vision-based LocalizationabstractIn recent years, camera-based localization has been widely used for robotic applications, and most proposed algorithms rely on local features extracted from recorded images. For better performance, the features used for open-loop localization are required to be short-term globally static, and the ones used for re-localization or loop closure detection need to be long-term static. Therefore, the motion attribute of a local feature point could be exploited to improve localization performance, e.g., the feature points extracted from moving persons or vehicles can be excluded from these systems due to their unsteadiness. In this paper, we design a fully convolutional network (FCN), named MD-Net, to perform motion attribute estimation and feature description simultaneously. MD-Net has a shared backbone network to extract features from the input image and two network branches to complete each sub-task. With MD-Net, we can obtain the motion attribute while avoiding increasing much more computation. Experimental results demonstrate that the proposed method can learn distinct local feature descriptor along with motion attribute only using an FCN, by outperforming competing methods by a wide margin. We also show that the proposed algorithm can be integrated into a vision-based localization algorithm to improve estimation accuracy significantly. Yafei Song 0002, Jia Li 0003, Yonghong Tian 0001, Mingyang Li 0001 |
IROS | 3 |
| 2019 | Fewer-Shots and Lower-Resolutions: Towards Ultrafast Face Recognition in the WildabstractIs it possible to train an effective face recognition model with fewer shots that works efficiently on low-resolution faces in the wild? To answer this question, this paper proposes a few-shot knowledge distillation approach to learn an ultrafast face recognizer via two steps. In the first step, we initialize a simple yet effective face recognition model on synthetic low-resolution faces by distilling knowledge from an existing complex model. By removing the redundancies in both face images and the model structure, the initial model can provide an ultrafast speed with impressive recognition accuracy. To further adapt this model into the wild scenarios with fewer faces per person, the second step refines the model via few-shot learning by incorporating a relation module that compares low-resolution query faces with faces in the support set. In this manner, the performance of the model can be further enhanced with only fewer low-resolution faces in the wild. Experimental results show that the proposed approach performs favorably against state-of-the-arts in recognizing low-resolution faces with an extremely low memory of 30KB and runs at an ultrafast speed of 1,460 faces per second on CPU or 21,598 faces per second on GPU. Shiming Ge, Shengwei Zhao, Xindi Gao, Jia Li 0003 |
ACM Multimedia | 4 |
| 2019 | Cross-Reference Stitching Quality Assessment for 360° Omnidirectional ImagesabstractAlong with the development of virtual reality (VR), omnidirectional images play an important role in producing multimedia content with an immersive experience. However, despite various existing approaches for omnidirectional image stitching, how to quantitatively assess the quality of stitched images is still insufficiently explored. To address this problem, we first establish a novel omnidirectional image dataset containing stitched images as well as dual-fisheye images captured from standard quarters of 0$^\circ$, 90$^\circ$, 180$^\circ$, and 270$^\circ$. In this manner, when evaluating the quality of an image stitched from a pair of fisheye images (\eg, 0$^\circ$ and 180$^\circ$), the other pair of fisheye images (\eg, 90$^\circ$ and 270$^\circ$) can be used as the cross-reference to provide ground-truth observations of the stitching regions. Based on this dataset, we propose a set of Omnidirectional Stitching Image Quality Assessment (OS-IQA) metrics. In these metrics, the stitching regions are assessed by exploring the local relationships between the stitched image and its cross-reference with histogram statistics, perceptual hash and sparse reconstruction, while the whole stitched images are assessed by the global indicators of color difference and fitness of blind zones.Qualitative and quantitative experiments show our method outperforms the classic IQA metrics and is highly consistent with human subjective evaluations. To the best of our knowledge, it is the first attempt that assesses the stitching quality of omnidirectional images by using cross-references. Jia Li 0003, Kaiwen Yu, Yifan Zhao 0002, Yu Zhang 0035, Long Xu 0001 |
ACM Multimedia | 1 |
| 2019 | Salient object detection: A surveyabstractDetecting and segmenting salient objects from natural scenes, often referred to as salient object detection, has attracted great interest in computer vision. While many models have been proposed and several applications have emerged, a deep understanding of achievements and issues remains lacking. We aim to provide a comprehensive review of recent progress in salient object detection and situate this field among other closely related areas such as generic scene segmentation, object proposal generation, and saliency for fixation prediction. Covering 228 publications, we survey i) roots, key concepts, and tasks, ii) core techniques and main modeling trends, and iii) datasets and evaluation metrics for salient object detection. We also discuss open problems such as evaluation metrics and dataset bias in model performance, and suggest future research directions. Ali Borji, Ming-Ming Cheng, Qibin Hou, Huaizu Jiang, Jia Li 0003 |
Comput. Vis. Media | 5 |
| 2019 | Deep3DSaliency: Deep Stereoscopic Video Saliency Detection Model by 3D Convolutional NetworksabstractStereoscopic saliency detection plays an important role in various stereoscopic video processing applications. However, conventional stereoscopic video saliency detection methods mainly use independent low-level features instead of extracting them automatically, and thus, they ignore the intrinsic relationship between the spatial and temporal information. In this paper, we propose a novel stereoscopic video saliency detection method based on 3D convolutional neural networks, namely Deep 3D Video Saliency (Deep3DSaliency). The proposed network consists of two sub-models: Spatiotemporal Saliency Model (STSM), and Stereoscopic Saliency Aware Model (SSAM). STSM directly takes three consecutive video frames as the input to extract visual spatiotemporal features, while SSAM attempts to further infer the depth and semantic features from the left and right video frames by shared parameters from STSM. The visual spatiotemporal features from STSM, and the depth and semantic features from SSAM are learned by an alternating optimization scheme. Finally, all these saliency-related features are combined together for the final stereoscopic saliency detection via 3D deconvolution. Experimental results show the superior performance of the proposed model over other existing ones in saliency estimation for 3D video sequences. Yuming Fang 0001, Guanqun Ding, Jia Li 0003, Zhijun Fang 0001 |
IEEE Trans. Image Process. | 3 |
| 2019 | Low-Resolution Face Recognition in the Wild via Selective Knowledge DistillationabstractTypically, the deployment of face recognition models in the wild needs to identify low-resolution faces with extremely low computational cost. To address this problem, a feasible solution is compressing a complex face model to achieve higher speed and lower memory at the cost of minimal performance drop. Inspired by that, this paper proposes a learning approach to recognize low-resolution faces via selective knowledge distillation. In this approach, a two-stream convolutional neural network (CNN) is first initialized to recognize high-resolution faces and resolution-degraded faces with a teacher stream and a student stream, respectively. The teacher stream is represented by a complex CNN for high-accuracy recognition, and the student stream is represented by a much simpler CNN for low-complexity recognition. To avoid significant performance drop at the student stream, we then selectively distil the most informative facial features from the teacher stream by solving a sparse graph optimization problem, which are then used to regularize the finetuning process of the student stream. In this way, the student stream is actually trained by simultaneously handling two tasks with limited computational resources: approximating the most informative facial cues via feature regression, and recovering the missing facial cues via low-resolution face classification. Experimental results show that the student stream performs impressively in recognizing low-resolution faces and costs only 0.15MB memory and runs at 418 faces per second on CPU and 9; 433 faces per second on GPU. Shiming Ge, Shengwei Zhao, Chenyu Li 0001, Jia Li 0003 |
IEEE Trans. Image Process. | 4 |
| 2019 | Rearranging Online Tubes for Streaming Video Synopsis: A Dynamic Graph Coloring ApproachabstractTo efficiently browse long surveillance videos, the video synopsis technique is often used to rearrange tubes (i.e., tracks of moving objects) along the temporal axis to form a much shorter video. In this process, two key issues need to be addressed, i.e., the minimization of spatial tube collision and the maximization of temporal video condensation. In addition, when a surveillance video comes as a stream, an online algorithm with the capability of dynamically rearranging tubes is also required. Toward this end, this paper proposes a novel graph-based tube rearrangement approach for online video synopsis. The relationships among tubes are modeled with a dynamic graph, whose nodes (i.e., object masks of tubes) and edges (i.e., relationships) can be progressively inserted and updated. Based on this graph, we propose a dynamic graph coloring algorithm to efficiently rearrange all tubes by determining when they should appear. Extensive experimental results show that our approach can condense online surveillance video streams in real time with less tube collision and high compact ratio. Shikui Wei, Jia Li 0003, Yao Zhao 0001 |
IEEE Trans. Image Process. | 3 |
| 2019 | Saliency Inside: Learning Attentive CNNs for Content-Based Image RetrievalabstractIn content-based image retrieval (CBIR), one of the most challenging and ambiguous tasks are to correctly understand the human query intention and measure its semantic relevance with images in the database. Due to the impressive capability of visual saliency in predicting human visual attention that is closely related to the query intention, this paper attempts to explicitly discover the essential effect of visual saliency in CBIR via qualitative and quantitative experiments. Toward this end, we first generate the fixation density maps of images from a widely used CBIR dataset by using an eye-tracking apparatus. These ground-truth saliency maps are then used to measure the influence of visual saliency to the task of CBIR by exploring several probable ways of incorporating such saliency cues into the retrieval process. We find that visual saliency is indeed beneficial to the CBIR task, and the best saliency involving scheme is possibly different for different image retrieval models. Inspired by the findings, this paper presents two-stream attentive CNNs with saliency embedded inside for CBIR. The proposed network has two streams that simultaneously handle two tasks. The main stream focuses on extracting discriminative visual features that are tightly related to semantic attributes. Meanwhile, the auxiliary stream aims to facilitate the main stream by redirecting the feature extraction to the salient image content that human may pay attention to. By fusing these two streams into the Main and Auxiliary CNNs (MAC), image similarity can be computed as the human being does by reserving conspicuous content and suppressing irrelevant regions. Extensive experiments show that the proposed model achieves impressive performance in image retrieval on four public datasets. Shikui Wei, Lixin Liao, Jia Li 0003, Qinjie Zheng, Yao Zhao 0001 |
IEEE Trans. Image Process. | 3 |
| 2018 | Collaborative Annotation of Semantic Objects in Images with Multi-granularity SupervisionsabstractPer-pixel masks of semantic objects are very useful in many applications, which, however, are tedious to be annotated. In this paper, we propose a collaborative annotation approach to efficiently generate per-pixel masks of semantic objects in tagged images with multi-granularity supervisions. Given a set of tagged images, a computer agent is dynamically generated to roughly localize the semantic objects described by the tag. The agent first extracts massive object proposals and then infer the tag-related ones under the weak and strong supervisions from linguistically and visually similar images as well as previously annotated objects. By representing such supervisions by over-complete dictionaries, tag-related proposals can pop-out according to their sparse coding length, which are then converted to superpixels with binary labels. After that, human annotators participate in the annotation by flipping labels and dividing superpixels with clicks, which are used as click supervisions that teaches the agent to recover false positives/negatives in processing images with the same tags. Experimental results show that our approach can facilitate the annotation and generate object masks that are consistent with those generated by the LabelMe toolbox. Lishi Zhang, Chenghan Fu, Jia Li 0003 |
ACM Multimedia | 3 |
| 2018 | Semantic Object Segmentation in Tagged Videos via DetectionabstractSemantic object segmentation (SOS) is a challenging task in computer vision that aims to detect and segment all pixels of the objects within predefined semantic categories. In image-based SOS, many supervised models have been proposed and achieved impressive performances due to the rapid advances of well-annotated training images and machine learning theories. However, in video-based SOS it is often difficult to directly train a supervised model since most videos are weakly annotated by tags. To handle such tagged videos, this paper proposes a novel approach that adopts a segmentation-by-detection framework. In this framework, object detection and segment proposals are first generated using the models pre-trained on still images, which provide useful cues to roughly localize the semantic objects. Based on these proposals, we propose an efficient algorithm to initialize object tracks by solving a joint assignment problem. As such tracks provide rough spatiotemporal configurations of the semantic objects, a voting-based refinement algorithm is further proposed to improve their spatiotemporal consistency. Extensive experiments demonstrate that the proposed framework can robustly and effectively segment semantic objects in tagged videos, even when the image-based object detectors provide inaccurate proposals. On various public benchmarks, the proposed approach obtains substantial improvements over the state-of-the-arts. Yu Zhang 0035, Xiaowu Chen 0001, Jia Li 0003, Changqun Xia |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | A Benchmark Dataset and Saliency-Guided Stacked Autoencoders for Video-Based Salient Object DetectionabstractImage-based salient object detection (SOD) has been extensively studied in past decades. However, video-based SOD is much less explored due to the lack of large-scale video datasets within which salient objects are unambiguously defined and annotated. Toward this end, this paper proposes a video-based SOD dataset that consists of 200 videos. In constructing the dataset, we manually annotate all objects and regions over 7650 uniformly sampled keyframes and collect the eye-tracking data of 23 subjects who free-view all videos. From the user data, we find that salient objects in a video can be defined as objects that consistently pop-out throughout the video, and objects with such attributes can be unambiguously annotated by combining manually annotated object/region masks with eye-tracking data of multiple subjects. To the best of our knowledge, it is currently the largest dataset for video-based salient object detection. Based on this dataset, this paper proposes an unsupervised baseline approach for video-based SOD by using saliency-guided stacked autoencoders. In the proposed approach, multiple spatiotemporal saliency cues are first extracted at the pixel, superpixel, and object levels. With these saliency cues, stacked autoencoders are constructed in an unsupervised manner that automatically infers a saliency score for each pixel by progressively encoding the high-dimensional saliency cues gathered from the pixel and its spatiotemporal neighbors. In experiments, the proposed unsupervised approach is compared with 31 state-of-the-art models on the proposed dataset and outperforms 30 of them, including 19 image-based classic (unsupervised or non-deep learning) models, six image-based deep learning models, and five video-based unsupervised models. Moreover, benchmarking results show that the proposed dataset is very challenging and has the potential to boost the development of video-based SOD. Jia Li 0003, Changqun Xia, Xiaowu Chen 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Exploring Weakly Labeled Images for Video Object Segmentation With Submodular Proposal SelectionabstractVideo object segmentation (VOS) is important for various computer vision problems, and handling it with minimal human supervision is highly desired for the large-scale applications. To bring down the supervision, existing approaches largely follow a data mining perspective by assuming the availability of multiple videos sharing the same object categories. It, however, would be problematic for the tasks that consume a single video. To address this problem, this paper proposes a novel approach that explores weakly labeled images to solve video object segmentation. Given a video labeled with a target category, images labeled with the same category are collected, from which noisy object exemplars are automatically discovered. After that the proposed approach extracts a set of region proposals on various frames and efficiently matches them with massive noisy exemplars in terms of appearance and spatial context. We then jointly select the best proposals across the video by solving a novel submodular problem that combines region voting and global region matching. Finally, the localization results are leveraged as strong supervision to guide pixel-level segmentation. Extensive experiments are conducted on two challenging public databases: Youtube-Objects and DAVIS. The results suggest that the proposed approach improves over previous weakly supervised/unsupervised approaches significantly, showing a performance even comparable with the several approaches supervised by the costly manual segmentations. Yu Zhang 0035, Xiaowu Chen 0001, Jia Li 0003, Wei Teng, Haokun Song |
IEEE Trans. Image Process. | 3 |
| 2018 | Single Image Dehazing Using Ranking Convolutional Neural NetworkabstractSingle image dehazing, which aims to recover the clear image solely from an input hazy or foggy image, is a challenging ill-posed problem. Analyzing existing approaches, the common key step is to estimate the haze density of each pixel. To this end, various approaches oftenheuristically designedhaze-relevant features. Several recent works also automatically learn the features via directly exploiting convolutional neural networks (CNN). However, it may be insufficient to fully capture the intrinsic attributes of hazy images. To obtain effective features for single image dehazing, this paper presents a novel ranking convolutional neural network (Ranking-CNN). In Ranking-CNN, a novel ranking layer is proposed to extend the structure of CNN so that the statistical and structural attributes of hazy images can be simultaneously captured. By training Ranking-CNN in a well-designed manner, powerful haze-relevant features can beautomatically learnedfrom massive hazy image patches. Based on these features, haze can be effectively removed by using a haze density prediction model trained through the random forest regression. Experimental results show that our approach outperforms several previous dehazing approaches on synthetic and real-world benchmark images. Comprehensive analyses are also conducted to interpret the proposed Ranking-CNN from both the theoretical and experimental aspects. Yafei Song 0002, Jia Li 0003, Xiaogang Wang 0005, Xiaowu Chen 0001 |
IEEE Trans. Multim. | 2 |
| 2017 | Detecting Masked Faces in the Wild with LLE-CNNsabstractDetecting faces with occlusions is a challenging task due to two main reasons: 1) the absence of large datasets of masked faces, and 2) the absence of facial cues from the masked regions. To address these two issues, this paper first introduces a dataset, denoted as MAFA, with 30, 811 Internet images and 35, 806 masked faces. Faces in the dataset have various orientations and occlusion degrees, while at least one part of each face is occluded by mask. Based on this dataset, we further propose LLE-CNNs for masked face detection, which consist of three major modules. The Proposal module first combines two pre-trained CNNs to extract candidate facial regions from the input image and represent them with high dimensional descriptors. After that, the Embedding module is incorporated to turn such descriptors into a similarity-based descriptor by using locally linear embedding (LLE) algorithm and the dictionaries trained on a large pool of synthesized normal faces, masked faces and non-faces. In this manner, many missing facial cues can be largely recovered and the influences of noisy cues introduced by diversified masks can be greatly alleviated. Finally, the Verification module is incorporated to identify candidate facial regions and refine their positions by jointly performing the classification and regression tasks within a unified CNN. Experimental results on the MAFA dataset show that the proposed approach remarkably outperforms 6 state-of-the-arts by at least 15.6%. Shiming Ge, Jia Li 0003, Qiting Ye, Zhao Luo |
CVPR | 2 |
| 2017 | What is and What is Not a Salient Object? Learning Salient Object Detector by Ensembling Linear Exemplar RegressorsabstractFinding what is and what is not a salient object can be helpful in developing better features and models in salient object detection (SOD). In this paper, we investigate the images that are selected and discarded in constructing a new SOD dataset and find that many similar candidates, complex shape and low objectness are three main attributes of many non-salient objects. Moreover, objects may have diversified attributes that make them salient. As a result, we propose a novel salient object detector by ensembling linear exemplar regressors. We first select reliable foreground and background seeds using the boundary prior and then adopt locally linear embedding (LLE) to conduct manifold-preserving foregroundness propagation. In this manner, a foregroundness map can be generated to roughly pop-out salient objects and suppress non-salient ones with many similar candidates. Moreover, we extract the shape, foregroundness and attention descriptors to characterize the extracted object proposals, and a linear exemplar regressor is trained to encode how to detect salient proposals in a specific image. Finally, various linear exemplar regressors are ensembled to form a single detector that adapts to various scenarios. Extensive experimental results on 5 dataset and the new SOD dataset show that our approach outperforms 9 state-of-art methods. Changqun Xia, Jia Li 0003, Xiaowu Chen 0001, Anlin Zheng, Yu Zhang 0035 |
CVPR | 2 |
| 2017 | Primary Video Object Segmentation via Complementary CNNs and Neighborhood Reversible FlowabstractThis paper proposes a novel approach for segmenting primary video objects by using Complementary Convolutional Neural Networks (CCNN) and neighborhood reversible flow. The proposed approach first pre-trains CCNN on massive images with manually annotated salient objects in an end-to-end manner, and the trained CCNN has two separate branches that simultaneously handle two complementary tasks, i.e., foregroundness and backgroundness estimation. By applying CCNN on each video frame, the spatial foregroundness and backgroundness maps can be initialized, which are then propagated between various frames so as to segment primary video objects and suppress distractors. To enforce efficient temporal propagation, we divide each frame into superpixels and construct neighborhood reversible flow that reflects the most reliable temporal correspondences between superpixels in far-away frames. Within such flow, the initialized foregroundness and backgroundness can be efficiently and accurately propagated along the temporal axis so that primary video objects gradually pop-out and distractors are well suppressed. Extensive experimental results on three video datasets show that the proposed approach achieves impressive performance in comparisons with 18 state-of-the-art models. Jia Li 0003, Anlin Zheng, Xiaowu Chen 0001 |
ICCV | 1 |
| 2017 | Look, Perceive and Segment: Finding the Salient Objects in Images via Two-stream Fixation-Semantic CNNsabstractRecently, CNN-based models have achieved remarkable success in image-based salient object detection (SOD). In these models, a key issue is to find a proper network architecture that best fits for the task of SOD. Toward this end, this paper proposes two-stream fixation-semantic CNNs, whose architecture is inspired by the fact that salient objects in complex images can be unambiguously annotated by selecting the pre-segmented semantic objects that receive the highest fixation density in eye-tracking experiments. In the two-stream CNNs, a fixation stream is pre-trained on eye-tracking data whose architecture well fits for the task of fixation prediction, and a semantic stream is pre-trained on images with semantic tags that has a proper architecture for semantic perception. By fusing these two streams into an inception-segmentation module and jointly fine-tuning them on images with manually annotated salient objects, the proposed networks show impressive performance in segmenting salient objects. Experimental results show that our approach outperforms 10 state-of-the-art models (5 deep, 5 non-deep) on 4 datasets. Xiaowu Chen 0001, Anlin Zheng, Jia Li 0003, Feng Lu 0005 |
ICCV | 3 |
| 2017 | Embedding 3D Geometric Features for Rigid Object Part SegmentationabstractObject part segmentation is a challenging and fundamental problem in computer vision. Its difficulties may be caused by the varying viewpoints, poses, and topological structures, which can be attributed to an essential reason, i.e., a specific object is a 3D model rather than a 2D figure. Therefore, we conjecture that not only 2D appearance features but also 3D geometric features could be helpful. With this in mind, we propose a 2-stream FCN. One stream, named AppNet, is to extract 2D appearance features from the input image. The other stream, named GeoNet, is to extract 3D geometric features. However, the problem is that the input is just an image. To this end, we design a 2D convolution based CNN structure to extract 3D geometric features from 3D volume, which is named VolNet. Then a teacher-student strategy is adopted and VolNet teaches GeoNet how to extract 3D geometric features from an image. To perform this teaching process, we synthesize training data using 3D models. Each training sample consists of an image and its corresponding volume. A perspective voxelization algorithm is further proposed to align them. Experimental results verify our conjecture and the effectiveness of both the proposed 2-stream CNN and VolNet. Yafei Song 0002, Xiaowu Chen 0001, Jia Li 0003, Qinping Zhao |
ICCV | 3 |
| 2017 | Two-stream Attentive CNNs for Image RetrievalabstractIn content-based image retrieval, the most challenging (and ambiguous) part is to define the similarity between images. For the human-being, such similarity can be defined with respect to where they pay attention to and what semantic attributes they understand. Inspired by this fact, this paper presents two-stream attentive CNNs for image retrieval. As the human-being does, the proposed network has two streams that simultaneously handle two tasks. The Main stream focuses on extracting discriminative visual features that are tightly correlated with semantic attributes. Meanwhile, the Auxiliary stream aims to facilitate the main stream by redirecting the feature extraction operation mainly to the image content that human may pay attention to. By fusing these two streams into the Main and Auxiliary CNNs (MAC), image similarity can be computed as the human-being does by reserving the conspicuous content and suppressing the irrelevant regions. Extensive experiments show that the proposed model achieves impressive performance in image retrieval on four public datasets. Jia Li 0003, Shikui Wei, Qinjie Zheng, Ting Liu 0012, Yao Zhao 0001 |
ACM Multimedia | 2 |
| 2017 | Learning visual saliency from human fixations for stereoscopic images
Yuming Fang 0001, Jianjun Lei 0001, Jia Li 0003, Long Xu 0001, Weisi Lin, Patrick Le Callet |
Neurocomputing | 3 |
| 2017 | Multi-Task Rank Learning for Image Quality AssessmentabstractIn practice, images are distorted by more than one distortion. For image quality assessment (IQA), existing machine learning (ML)-based methods generally establish a unified model for all the distortion types, or each model is trained independently for each distortion type, which is therefore distortion aware. In distortion-aware methods, the common features among different distortions are not exploited. In addition, there are fewer training samples for each model training task, which may result in overfitting. To address these problems, we propose a multi-task learning framework to train multiple IQA models together, where each model is for each distortion type; however, all the training samples are associated with each model training task. Thus, the common features among different distortion types and the said underlying relatedness among all the learning tasks are exploited, which would benefit the generalization ability of trained models and prevent overfitting possibly. In addition, pairwise image quality ranking instead of image quality rating is optimized in our learning task, which is fundamentally departed from traditional ML-based IQA methods toward better performance. The experimental results confirm that the proposed multi-task rank-learning-based IQA metric is prominent against all state-of-the-art nonreference IQA approaches. Long Xu 0001, Jia Li 0003, Weisi Lin, Yongbing Zhang 0002, Lin Ma 0002, Yuming Fang 0001, Yihua Yan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Learning Discriminative Subspaces on Random Contrasts for Image Saliency AnalysisabstractIn visual saliency estimation, one of the most challenging tasks is to distinguish targets and distractors that share certain visual attributes. With the observation that such targets and distractors can sometimes be easily separated when projected to specific subspaces, we propose to estimate image saliency by learning a set of discriminative subspaces that perform the best in popping out targets and suppressing distractors. Toward this end, we first conduct principal component analysis on massive randomly selected image patches. The principal components, which correspond to the largest eigenvalues, are selected to construct candidate subspaces since they often demonstrate impressive abilities to separate targets and distractors. By projecting images onto various subspaces, we further characterize each image patch by its contrasts against randomly selected neighboring and peripheral regions. In this manner, the probable targets often have the highest responses, while the responses at background regions become very low. Based on such random contrasts, an optimization framework with pairwise binary terms is adopted to learn the saliency model that best separates salient targets and distractors by optimally integrating the cues from various subspaces. Experimental results on two public benchmarks show that the proposed approach outperforms 16 state-of-the-art methods in human fixation prediction. Shu Fang, Jia Li 0003, Yonghong Tian 0001, Tiejun Huang 0001, Xiaowu Chen 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2016 | Local Shape Transfer for Image Co-segmentation
Wei Teng, Yu Zhang 0035, Xiaowu Chen 0001, Jia Li 0003, Zhiqiang He 0002 |
BMVC | 4 |
| 2016 | Structure-adaptive Shape Editing for Man-made ObjectsabstractAbstract One of the challenging problems for shape editing is to adapt shapes with diversified structures for various editing needs. In this paper we introduce a shape editing approach that automatically adapts the structure of a shape being edited with respect to user inputs. Given a category of shapes, our approach first classifies them into groups based on the constituent parts. The group‐sensitive priors, including both inter‐group and intra‐group priors, are then learned through statistical structure analysis and multivariate regression. By using these priors, the inherent characteristics and typical variations of shape structures can be well captured. Based on such group‐sensitive priors, we propose a framework for real‐time shape editing, which adapts the structure of shape to continuous user editing operations. Experimental results show that the proposed approach is capable of both structure‐preserving and structure‐varying shape editing. Qiang Fu 0004, Xiaowu Chen 0001, Xiaoyu Su, Jia Li 0003, Hongbo Fu 0001 |
Comput. Graph. Forum | 4 |
| 2016 | Measuring Visual Surprise Jointly from Intrinsic and Extrinsic Contexts for Image Saliency Estimation
Jia Li 0003, Yonghong Tian 0001, Xiaowu Chen 0001, Tiejun Huang 0001 |
Int. J. Comput. Vis. | 1 |
| 2016 | Light-weight binary code embedding of local feature distribution in image search
Shikui Wei, Yao Zhao 0001, Jia Li 0003 |
Neurocomputing | 3 |
| 2016 | Fuzzy community detection via modularity guided membership-degree propagation
Xiaowu Chen 0001, Jia Li 0003 |
Pattern Recognit. Lett. | 3 |
| 2016 | 6-DOF Image Localization From Massive Geo-Tagged Reference ImagesabstractThe 6-degrees of freedom (DOF) image localization, which aims to calculate the spatial position and rotation of a camera, is a challenging problem for most location-based services. In existing approaches, this problem is often tackled by finding the matches between 2D image points and 3D structure points so as to derive the location information via direct linear transformation algorithm. However, as these 2D-to-3D-based approaches need to reconstruct the 3D structure points of the scene, they may not be flexible enough to employ massive and increasing geo-tagged data. To this end, this paper presents a novel approach for 6-DOF image localization by fusing candidate poses relative to reference images. In this approach, we propose to localize an input image according to the position and rotation information of multiple geo-tagged images retrieved from a reference dataset. From the reference images, an efficient relative pose estimation algorithm is proposed to derive a set of candidate poses for the input image. Each candidate pose encodes the relative rotation and direction of the input image with respect to a specific reference image. Finally, these candidate poses can be fused together by minimizing a well-defined geometry error so that the 6-DOF location of the input image is effectively derived. Experimental results show that our method can obtain satisfactory localization accuracy. In addition, the proposed relative pose estimation algorithm is much faster than existing work. Yafei Song 0002, Xiaowu Chen 0001, Xiaogang Wang 0005, Yu Zhang 0035, Jia Li 0003 |
IEEE Trans. Multim. | 5 |
| 2015 | Semantic object segmentation via detection in weakly labeled videoabstractSemantic object segmentation in video is an important step for large-scale multimedia analysis. In many cases, however, semantic objects are only tagged at video-level, making them difficult to be located and segmented. To address this problem, this paper proposes an approach to segment semantic objects in weakly labeled video via object detection. In our approach, a novel video segmentation-by-detection framework is proposed, which first incorporates object and region detectors pre-trained on still images to generate a set of detection and segmentation proposals. Based on the noisy proposals, several object tracks are then initialized by solving a joint binary optimization problem with min-cost flow. As such tracks actually provide rough configurations of semantic objects, we thus refine the object segmentation while preserving the spatiotemporal consistency by inferring the shape likelihoods of pixels from the statistical information of tracks. Experimental results on Youtube-Objects dataset and SegTrack v2 dataset demonstrate that our method outperforms state-of-the-arts and shows impressive results. Yu Zhang 0035, Xiaowu Chen 0001, Jia Li 0003, Changqun Xia |
CVPR | 3 |
| 2015 | Multi-task rank learning for image quality assessmentabstractIn practice, multiple types of distortions are associated with an image quality degradation process. The existing machine learning (ML) based image quality assessment (IQA) approaches generally established a unified model for all distortion types, or each model is trained independently for each distortion type by using single-task learning, which lead to the poor generalization ability of the models as applied to practical image processing. There are often the underlying cross relatedness amongst these single-task learnings in IQA, which is ignored by the previous approaches. To solve this problem, we propose a multi-task learning framework to train IQA models simultaneously across individual tasks each of which concerns one distortion type. These relatedness can be therefore exploited to improve the generalization ability of IQA models from single-task learning. In addition, pairwise image quality rank instead of image quality rating is optimized in learning task. By mapping image quality rank to image quality rating, a novel no-reference (NR) IQA approach can be derived. The experimental results confirm that the proposed Multi-task Rank Learning based IQA (MRLIQ) approach is prominent among all state-of-the-art NR-IQA approaches. Long Xu 0001, Jia Li 0003, Weisi Lin, Yongbing Zhang 0002, Lin Ma 0002, Yuming Fang 0001, Yun Zhang 0002, Yihua Yan |
ICASSP | 2 |
| 2015 | A Data-Driven Metric for Comprehensive Evaluation of Saliency ModelsabstractIn the past decades, hundreds of saliency models have been proposed for fixation prediction, along with dozens of evaluation metrics. However, existing metrics, which are often heuristically designed, may draw conflict conclusions in comparing saliency models. As a consequence, it becomes somehow confusing on the selection of metrics in comparing new models with state-of-the-arts. To address this problem, we propose a data-driven metric for comprehensive evaluation of saliency models. Instead of heuristically designing such a metric, we first conduct extensive subjective tests to find how saliency maps are assessed by the human-being. Based on the user data collected in the tests, nine representative evaluation metrics are directly compared by quantizing their performances in assessing saliency maps. Moreover, we propose to learn a data-driven metric by using Convolutional Neural Network. Compared with existing metrics, experimental results show that the data-driven metric performs the most consistently with the human-being in evaluating saliency maps as well as saliency models. Jia Li 0003, Changqun Xia, Yafei Song 0002, Shu Fang, Xiaowu Chen 0001 |
ICCV | 1 |
| 2015 | Cuboids detection in RGB-D images via Maximum Weighted CliqueabstractCuboid detection is an essential step for understanding 3D structure of scenes. As most of indoor scene cuboids are actually objects, we propose in this paper an object-based approach to detect 3D cuboids in indoor RGB-D images. The proposed approach is learning-free and can handle general object classes rather than a limited pre-defined category set. In our approach, we first apply an extended version of the CPMC framework to generate a set of segment hypotheses, and fit a set of cuboid candidates. Given the candidate set, we select several cuboids that can provide plausible interpretations of the images by solving a Maximum Weighted Clique (MWC) problem. With this formulation, a set of ranked mid-level representations of the input image is obtained, and are further re-ranked by Maximal Marginal Relevance (MMR) measure to improve their diversity. Experimental results on NYU-V2 dataset shows that our method significantly outperforms the state-of-the-art, and shows impressive results. Xiaowu Chen 0001, Yu Zhang 0035, Jia Li 0003, Xiaogang Wang 0005 |
ICME | 4 |
| 2015 | Learning Complementary Saliency Priors for Foreground Object Segmentation in Complex Scenes
Yonghong Tian 0001, Jia Li 0003, Shui Yu 0001, Tiejun Huang 0001 |
Int. J. Comput. Vis. | 2 |
| 2015 | Finding the Secret of Image Saliency in the Frequency DomainabstractThere are two sides to every story of visual saliency modeling in the frequency domain. On the one hand, image saliency can be effectively estimated by applying simple operations to the frequency spectrum. On the other hand, it is still unclear which part of the frequency spectrum contributes the most to popping-out targets and suppressing distractors. Toward this end, this paper tentatively explores the secret of image saliency in the frequency domain. From the results obtained in several qualitative and quantitative experiments, we find that the secret of visual saliency may mainly hide in the phases of intermediate frequencies. To explain this finding, we reinterpret the concept of discrete Fourier transform from the perspective of template-based contrast computation and thus develop several principles for designing the saliency detector in the frequency domain. Following these principles, we propose a novel approach to design the saliency detector under the assistance of prior knowledge obtained through both unsupervised and supervised learning processes. Experimental results on a public image benchmark show that the learned saliency detector outperforms 18 state-of-the-art approaches in predicting human fixations. Jia Li 0003, Ling-Yu Duan, Xiaowu Chen 0001, Tiejun Huang 0001, Yonghong Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Robust multiple cameras pedestrian detection with multi-view Bayesian network
Peixi Peng, Yonghong Tian 0001, Yaowei Wang 0001, Jia Li 0003, Tiejun Huang 0001 |
Pattern Recognit. | 4 |
| 2015 | Image saliency estimation via random walk guided by informativeness and latent signal correlations
Jia Li 0003, Shu Fang, Yonghong Tian 0001, Tiejun Huang 0001, Xiaowu Chen 0001 |
Signal Process. Image Commun. | 1 |
| 2015 | Salient Object Detection: A BenchmarkabstractWe extensively compare, qualitatively and quantitatively, 41 state-of-the-art models (29 salient object detection, 10 fixation prediction, 1 objectness, and 1 baseline) over seven challenging data sets for the purpose of benchmarking salient object detection and segmentation methods. From the results obtained so far, our evaluation shows a consistent rapid progress over the last few years in terms of both accuracy and running time. The top contenders in this benchmark significantly outperform the models identified as the best in the previous benchmark conducted three years ago. We find that the models designed specifically for salient object detection generally work better than models in closely related areas, which in turn provides a precise definition and suggests an appropriate treatment of this problem that distinguishes it from other problems. In particular, we analyze the influences of center bias and scene complexity in model performance, which, along with the hard cases for the state-of-the-art models, provide useful hints toward constructing more challenging large-scale data sets and better saliency models. Finally, we propose probable solutions for tackling several open problems, such as evaluation scores and data set bias, which also suggest future research directions in the rapidly growing field of salient object detection. Ali Borji, Ming-Ming Cheng, Huaizu Jiang, Jia Li 0003 |
IEEE Trans. Image Process. | 4 |
| 2014 | Rank learning on training set selection and image quality assessmentabstractMachine learning (ML) techniques are widely used in recent no-reference visual quality assessment (NR-VQA) metrics by training on subjective image quality databases. In these metrics, the optimization function is constructed based on L2norm of the distance between subjective image quality and predicted image quality. There are two problems in these L2norm based methods: (1) human's opinion on subjective image quality rating is not reliable at fine-scale level. A small difference between subjective image qualities represented by mean opinion scores (MOSs) of two images may not truly reflect the real quality difference between these two images, but acts as noise. The optimization process should avoid such noise. (2) Generally, human's opinion on pairwise comparison (PC) for image quality is more reliable and believable than MOS. The importance of PC is ignored during the optimization process of existing ML-based studies, which are designed based on the numerical rating system. In this paper, we introduce image quality ranking concept to establish a new optimization objective instead of L2norm optimization, and then a novel NR-VQA is constructed based on ranking learning. The proposed metric firstly suggests a reasonable training set for ML, which is ignored by existing ML-based NR-VQA. The ranking theory is adopted to build optimization function, which reflects the properties of PC over the numerical ranting system used by traditional NR-VQA. By ignoring the small difference between MOSs from two images during the optimization process, the proposed ranking-based NR-VQA can also well address the first problem from the existing related metrics. Experimental results show that the proposed ranking-based NR-VQA can obtain better performance over the state-of-the-art NR-VQA approaches. Long Xu 0001, Weisi Lin, Jia Li 0003, Xu Wang 0006, Yihua Yan, Yuming Fang 0001 |
ICME | 3 |
| 2014 | Visual Saliency with Statistical Priors
Jia Li 0003, Yonghong Tian 0001, Tiejun Huang 0001 |
Int. J. Comput. Vis. | 1 |
| 2013 | Estimating Visual Saliency Through Single Image OptimizationabstractThis letter presents a novel approach for visual saliency estimation through single image optimization. Instead of directly mapping visual features to saliency values with a unified model, we treat regional saliency values as the optimization objective on each single image. By using a quadratic programming framework, our approach can adaptively optimize the regional saliency values on each specific image to simultaneously meet multiple saliency hypotheses on visual rarity, center-bias and mutual correlation. Experimental results show that our approach can outperform 14 state-of-the-art approaches on a public image benchmark. Jia Li 0003, Yonghong Tian 0001, Ling-Yu Duan, Tiejun Huang 0001 |
IEEE Signal Process. Lett. | 1 |
| 2012 | Removing Label Ambiguity in Learning-Based Visual Saliency EstimationabstractVisual saliency is a useful clue to depict visually important image/video contents in many multimedia applications. In visual saliency estimation, a feasible solution is to learn a "feature-saliency" mapping model from the user data obtained by manually labeling activities or eye-tracking devices. However, label ambiguities may also arise due to the inaccurate and inadequate user data. To process the noisy training data, we propose a multi-instance learning to rank approach for visual saliency estimation. In our approach, the correlations between various image patches are incorporated into an ordinal regression framework. By iteratively refining a ranking model and relabeling the image patches with respect to their mutual correlations, the label ambiguities can be effectively removed from the training data. Consequently, visual saliency can be effectively estimated by the ranking model, which can pop out real targets and suppress real distractors. Extensive experiments on two public image data sets show that our approach outperforms 11 state-of-the-art methods remarkably in visual saliency estimation. Jia Li 0003, Dong Xu 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2011 | Salient region detection and segmentation for general object recognition and image understanding
Tiejun Huang 0001, Yonghong Tian 0001, Jia Li 0003, Haonan Yu |
Sci. China Inf. Sci. | 3 |
| 2011 | Multi-Task Rank Learning for Visual Saliency EstimationabstractVisual saliency plays an important role in various video applications such as video retargeting and intelligent video advertising. However, existing visual saliency estimation approaches often construct a unified model for all scenes, thus leading to poor performance for the scenes with diversified contents. To solve this problem, we propose a multi-task rank learning approach which can be used to infer multiple saliency models that apply to different scene clusters. In our approach, the problem of visual saliency estimation is formulated in a pair-wise rank learning framework, in which the visual features can be effectively integrated to distinguish salient targets from distractors. A multi-task learning algorithm is then presented to infer multiple visual saliency models simultaneously. By an appropriate sharing of information across models, the generalization ability of each model can be greatly improved. Extensive experiments on a public eye-fixation dataset show that our multi-task rank learning approach outperforms 12 state-of-the-art methods remarkably in visual saliency estimation. Jia Li 0003, Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2010 | Saliency detection based on 2D log-gabor wavelets and center biasabstractVisual saliency can be a useful tool for image content analysis such as automatic image cropping and image compression. In existing methods on visual saliency detection, most of them are related to the model of receptive field. In this paper, we propose a bottom-up model which introduces 2D Log-Gabor wavelets for saliency detection. Compared with the traditional model of receptive field, the 2D Log-Gabor wavelets can better simulate the biological characteristics of the simple cortical cell in the receptive filed. Moreover, we also incorporate the influence of center bias into our model, which is a common phenomenon that directs visual attention to the center of images in natural scenes. Experimental results show that our approach outperforms three state-of-the-art approaches remarkably. Jia Li 0003, Tiejun Huang 0001, Yonghong Tian 0001, Ling-Yu Duan, Guochen Jia |
ACM Multimedia | 2 |
| 2010 | Automatic interesting object extraction from images using complementary saliency mapsabstractAutomatic interesting object extraction is widely used in many image applications. Among various extraction approaches, saliency-based ones usually have a better performance since they well accord with human visual perception. However, nearly all existing saliency-based approaches suffer the integrity problem, namely, the extracted result is either a small part of the object (referred to as sketch-like) or a large region that contains some redundant part of the background (referred to as envelope-like). In this paper, we propose a novel object extraction approach by integrating two kinds of "complementary" saliency maps (i.e., sketch-like and envelope-like maps). In our approach, the extraction process is decomposed into two sub-processes, one used to extract a high-precision result based on the sketch-like map, and the other used to extract a high-recall result based on the envelope-like map. Then a classification step is used to extract an exact object based on the two results. By transferring the complex extraction task to an easier classification problem, our approach can effectively break down the integrity problem. Experimental results show that the proposed approach outperforms six state-of-art saliency-based methods remarkably in automatic object extraction, and is even comparable to some interactive approaches. Haonan Yu, Jia Li 0003, Yonghong Tian 0001, Tiejun Huang 0001 |
ACM Multimedia | 2 |
| 2010 | Probabilistic Multi-Task Learning for Visual Saliency Estimation in Video
Jia Li 0003, Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001 |
Int. J. Comput. Vis. | 1 |
| 2010 | Salient object extraction for user-targeted video content associationabstractThe increasing amount of videos on the Internet and digital libraries highlights the necessity and importance of interactive video services such as automatically associating additional materials (e.g., advertising logos and relevant selling information) with the video content so as to enrich the viewing experience. Toward this end, this paper presents a novel approach for user-targeted video content association (VCA). In this approach, the salient objects are extracted automatically from the video stream using complementary saliency maps. According to these salient objects, the VCA system can push the related logo images to the users. Since the salient objects often correspond to important video content, the associated images can be considered as content-related. Our VCA system also allows users to associate images to the preferred video content through simple interactions by the mouse and an infrared pen. Moreover, by learning the preference of each user through collecting feedbacks on the pulled or pushed images, the VCA system can provide user-targeted services. Experimental results show that our approach can effectively and efficiently extract the salient objects. Moreover, subjective evaluations show that our system can provide content-related and user-targeted VCA services in a less intrusive way. Jia Li 0003, Han-nan Yu, Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001 |
J. Zhejiang Univ. Sci. C | 1 |
| 2010 | Cost-Sensitive Rank Learning From Positive and Unlabeled Data for Visual Saliency EstimationabstractThis paper presents a cost-sensitive rank learning approach for visual saliency estimation. This approach avoids the explicit selection of positive and negative samples, which is often used by existing learning-based visual saliency estimation approaches. Instead, both the positive and unlabeled data are directly integrated into a rank learning framework in a cost-sensitive manner. Compared with existing approaches, the rank learning framework can take the influences of both the local visual attributes and the pair-wise contexts into account simultaneously. Experimental results show that our algorithm outperforms several state-of-the-art approaches remarkably in visual saliency estimation. Jia Li 0003, Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001 |
IEEE Signal Process. Lett. | 1 |
| 2009 | A dataset and evaluation methodology for visual saliency in videoabstractRecently, visual saliency has drawn great research interest in the field of computer vision and multimedia. Various approaches aiming at calculating visual saliency have been proposed. To evaluate these approaches, several datasets have been presented for visual saliency in images. However, there are few datasets to capture spatiotemporal visual saliency in video. Intuitively, visual saliency in video is strongly affected by temporal context and might vary significantly even in visually similar frames. In this paper, we present an extensive dataset with 7.5-hour videos to capture spatiotemporal visual saliency. The salient regions in frames sequentially sampled from these videos are manually labeled by 23 subjects and then averaged to generate the ground-truth saliency maps. We also present three metrics to evaluate competing approaches. Several typical algorithms were evaluated on the dataset. The experimental results show that this dataset is very suitable for evaluating visual saliency. We also discover some interesting findings that would be addressed in future research. Currently, the dataset is freely available online together with the source code for evaluation. Jia Li 0003, Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001 |
ICME | 1 |
| 2008 | Multi-polarity text segmentation using graph theoryabstractText segmentation, or named text binarization, is usually an essential step for text information extraction from images and videos. However, most existing text segmentation methods have difficulties in extracting multi-polarity texts, where multi-polarity texts mean those texts with multiple colors or intensities in the same line. In this paper, we propose a novel algorithm for multi-polarity text segmentation based on graph theory. By representing a text image with an undirected weighted graph and partitioning it iteratively, multi-polarity text image can be effectively split into several single-polarity text images. As a result, these text images are then segmented by single-polarity text segmentation algorithms. Experiments on thousands of multi-polarity text images show that our algorithm can effectively segment multi-polarity texts. Jia Li 0003, Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001 |
ICIP | 1 |