Xu Jia 0012

dblp:95/3616-12 · DBLP profile ↗
← Back
76ranked-venue papers
6as first author
57since 2021 · last 2026
0000-0003-3168-3505ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 62 · 5 first-author · 47 since 2021Artificial intelligence and machine learning · 50 · 4 first-author · 34 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Temporal Inconsistency Guidance for Super-resolution Video Quality Assessment
abstract
As super-resolution (SR) techniques introduce unique distortions that fundamentally differ from those caused by traditional degradation processes (e.g., compression), there is an increasing demand for specialized video quality assessment (VQA) methods tailored to SR-generated content. One critical factor affecting perceived quality is temporal inconsistency, which refers to irregularities between consecutive frames. However, existing VQA approaches rarely quantify this phenomenon or explicitly investigate its relationship with human perception. Moreover, SR videos exhibit amplified inconsistency levels as a result of enhancement processes. In this paper, we propose Temporal Inconsistency Guidance for Super-resolution Video Quality Assessment (TIG-SVQA) that underscores the critical role of temporal inconsistency in guiding the quality assessment of SR videos. We first design a perception-oriented approach to quantify frame-wise temporal inconsistency. Based on this, we introduce the Inconsistency Highlighted Spatial Module, which localizes inconsistent regions at both coarse and fine scales. Inspired by the human visual system, we further develop an Inconsistency Guided Temporal Module that performs progressive temporal feature aggregation: (1) a consistency-aware fusion stage in which a visual memory capacity block adaptively determines the information load of each temporal segment based on inconsistency levels, and (2) an informative filtering stage for emphasizing quality-related features. Extensive experiments on both single-frame and multi-frame SR video scenarios demonstrate that our method significantly outperforms state-of-the-art VQA approaches.
Xiaoyuan Yang 0003, Weide Liu, Xin Jin 0014, Xu Jia 0012, Yukun Lai, Paul L. Rosin, Hantao Liu, Wei Zhou 0021
AAAI5
2026 The Avengers: A Routing Recipe for Collective Intelligence in Language Models
abstract
Proprietary models are increasingly dominating the race for ever-larger language models. Can open-source, smaller models remain competitive across a broad range of tasks? In this paper, we present the Avengers---a lightweight framework that leverages the collective intelligence of these smaller models. The Avengers builds upon four lightweight operations: (i) embedding: encode queries using a text embedding model; (ii) clustering: group queries based on their semantic similarity; (iii) scoring: scores each model's performance within each cluster; and (iv) voting: improve outputs via repeated sampling and voting. At inference time, each query is embedded and assigned to its nearest cluster. The top-performing model(s) within that cluster are selected to generate the response with repeated sampling. Remarkably, with 10 open-source models (~7B parameters each), the Avengers surpasses GPT-4o, 4.1, and 4.5 in average performance across 15 diverse datasets spanning mathematics, coding, logical reasoning, general knowledge, and affective tasks. In particular, it surpasses GPT-4.1 on mathematics tasks by 18.21% and on code tasks by 7.46%. Furthermore, the Avengers delivers superior out-of-distribution generalization, and remains robust across various embedding models, clustering algorithms, ensemble strategies, data efficiency, and values of its sole parameter---the number of clusters.
Hao Li 0069, Linyao Chen, Qiaosheng Zhang 0002, Peng Ye 0006, Shi Feng 0001, Xinrun Wang, Xu Jia 0012, Lei Bai 0001, Shuyue Hu
AAAI9
2026 Self-distilled learning of adaptive interval 3D lookup tables on real-time image enhancement
Ruikai Zhou, Canqian Yang, Meiguang Jin, Xu Jia 0012, Ying Chen 0011, Yi Xu 0001
Pattern Recognit.5
2026 Video Demoiréing With Spatial-Temporal Filtering in Frequency Domain
abstract
When acquiring images or videos of electronic displays, moiré patterns often arise due to aliasing between overlapping pixel grids, substantially compromising the perceptual quality of the captured content. Although frequency domain techniques have demonstrated high efficacy in image demoiréing, existing video approaches often overlook inter-frame frequency contextual relationships. This limitation restricts their capacity to achieve consistent temporal coherence and reconstruction fidelity. To overcome these challenges, we introduce a novel network (STFNet) with spatial-temporal filtering in frequency domain for video demoiréing. The proposed architecture comprises two dedicated stages: (1) Temporal-Guided Filtering (TGF), aims to adaptively incorporate temporal cues into learnable bandpass filters; and (2) Joint Filtering with Partially Shared Passbands (JFPS), which enhances representation learning of low-frequency moiré textures through strategic parameter sharing. Comprehensive evaluations on multiple public benchmarks confirm the superiority of our method. STFNet consistently outperforms state-of-the-art alternatives across both quantitative metrics and perceptual quality assessments, demonstrating robust performance in dynamic moiré suppression and detail preservation.
Zhongqi Liu, Bolun Zheng, Qianyu Zhang 0002, Heng Jin, Qiankun Li 0005, Xu Jia 0012, Jiyong Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.8
2026 Exploiting Cross-Task Synergy via Frequency-Driven Hierarchical Learning for Multi-Task Dense Prediction
abstract
Multi-task dense prediction improves pixel-level performance by leveraging shared representations and inter-task collaboration. However, existing approaches either rely on implicit task relationships or neglect frequency-domain cues that are essential for preserving fine-grained details and enhancing cross-task feature learning at multiple scales. As a result, they face persistent challenges in multi-scale feature fusion, effective task interaction, and accurate decoding. To address these issues, we propose a hierarchical frequency-driven framework, termed Hierarchical Frequency-Adaptive Network (HiFAN), that facilitates cross-task collaborative optimization via frequency-domain analysis. Specifically, we first design a task-adaptive fusion module that exploits multi-scale frequency-domain information to enhance spatial details. This module generates dynamic convolutional kernels with task-specific parameters and positional biases to adaptively accommodate diverse task requirements. Next, we introduce an efficient cross-task interaction module that leverages compact low-frequency representations to enable global context exchange across tasks. Finally, we present a high-frequency-aware decoder that mitigates feature smoothing and detail loss commonly introduced by Transformer-based decoders. We demonstrate the effectiveness of HiFAN on two standard multi-task learning benchmarks, PASCAL-Context and NYUD-v2, achieving strong and competitive performance across multiple tasks. The code and model weights are available in HiFAN.
Yunzhi Zhuge, Xinzhuo Yu, Lu Zhang 0053, Xu Jia 0012, Jin Zhan, Huchuan Lu
IEEE Trans. Image Process.4
2025 ReNeg: Learning Negative Embedding with Reward Guidance
abstract
In text-to-image (T2I) generation applications, negative embeddings have proven to be a simple yet effective approach for enhancing generation quality. Typically, these negative embeddings are derived from user-defined negative prompts, which, while being functional, are not necessarily optimal. In this paper, we introduce ReNeg, an end-to-end method designed to learn improved Negative embeddings guided by a Reward model. We employ a reward feedback learning framework and integrate classifier-free guidance (CFG) into the training process, which was previously utilized only during inference, thus enabling the effec tive learning of negative embeddings. We also propose two strategies for learning both global and per-sample negative embeddings. Extensive experiments show that the learned negative embedding significantly outperforms null-text and handcrafted counterparts, achieving substantial improvements in human preference alignment. Additionally, the negative embedding learned within the same text embedding space exhibits strong generalization capabilities. For example, using the same CLIP text encoder, the negative embedding learned on SD1.5 can be seamlessly transferred to text-to-image or even text-to-video models such as ControlNet, ZeroScope, and VideoCrafter2, resulting in consistent performance improvements across the board. Code is available at https://github.com/AMD-AIG-AIMA/ReNeg.
Xiaomin Li 0001, Yixuan Liu 0004, Takashi Isobe, Xu Jia 0012, Qinpeng Cui, Dong Zhou 0003, Dong Li 0025, You He 0002, Huchuan Lu, Zhongdao Wang, Emad Barsoum
CVPR4
2025 CCL-LGS: Contrastive Codebook Learning for 3D Language Gaussian Splatting
abstract
Recent advances in 3D reconstruction techniques and vision-language models have fueled significant progress in 3D semantic understanding, a capability critical to robotics, autonomous driving, and virtual/augmented reality. However, methods that rely on 2D priors are prone to a critical challenge: cross-view semantic inconsistencies induced by occlusion, image blur, and view-dependent variations. These inconsistencies, when propagated via projection supervision, deteriorate the quality of 3D Gaussian semantic fields and introduce artifacts in the rendered outputs. To mitigate this limitation, we propose CCL-LGS, a novel framework that enforces view-consistent semantic supervision by integrating multi-view semantic cues. Specifically, our approach first employs a zero-shot tracker to align a set of SAM-generated 2D masks and reliably identify their corresponding categories. Next, we utilize CLIP to extract robust semantic encodings across views. Finally, our Contrastive Codebook Learning (CCL) module distills discriminative semantic features by enforcing intra-class compactness and inter-class distinctiveness. In contrast to previous methods that directly apply CLIP to imperfect masks, our framework explicitly resolves semantic conflicts while preserving category discriminability. Extensive experiments demonstrate that CCL-LGS outperforms previous state-of-the-art methods. Our project page is available at https://epsilontl.github.io/CCL-LGS/.
Xiaomin Li 0001, Liqian Ma, Zirui Zheng, Hefei Huang, Taiqing Li, Huchuan Lu, Xu Jia 0012
ICCV9
2025 VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical Prior
abstract
Video diffusion models (VDMs) have advanced significantly in recent years, enabling the generation of highly realistic videos and drawing the attention of the community in their potential as world simulators. However, despite their capabilities, VDMs often fail to produce physically plausible videos due to an inherent lack of understanding of physics, resulting in incorrect dynamics and event sequences. To address this limitation, we propose a novel two-stage image-to-video generation framework that explicitly incorporates physics with vision and language informed physical prior. In the first stage, we employ a Vision Language Model (VLM) as a coarse-grained motion planner, integrating chain-of-thought and physics-aware reasoning to predict a rough motion trajectories/changes that approximate real-world physical dynamics while ensuring the inter-frame consistency. In the second stage, we use the predicted motion trajectories/changes to guide the video generation of a VDM. As the predicted motion trajectories/changes are rough, noise is added during inference to provide freedom to the VDM in generating motion with more fine details. Extensive experimental results demonstrate that our framework can produce physically plausible motion, and comparative evaluations highlight the notable superiority of our approach over existing methods. More video results are available on our Project Page: https://madaoer.github.io/projects/physically_plausible_video_generation.
Xindi Yang, Baolu Li 0001, Zhenfei Yin, Lei Bai 0001, Liqian Ma, Zhiyong Wang 0001, Jianfei Cai 0001, Tien-Tsin Wong, Huchuan Lu, Xu Jia 0012
ICCV11
2025 Evagaussians: Event Stream Assisted Gaussian Splatting from Blurry Images
Wangbo Yu, Chaoran Feng 0001, Jianing Li 0001, Jiye Tang, Jiashu Yang, Zhenyu Tang 0004, Meng Cao 0002, Xu Jia 0012, Li Yuan 0007, Yonghong Tian 0001
ICCV8
2025 Towards Survivability in Complex Motion Scenarios: RGB-Event Object Tracking via Historical Trajectory Prompting
abstract
Event data has recently emerged as a valuable complement to object tracking, offering dense temporal resolution and a high dynamic range. However, existing RGB-Event trackers struggle with targets exhibiting complex motion trajectories, where RGB features alone fail to provide sufficient discrimination. To address this, we propose EventTPT, an innovative RGB-Event tracking framework that leverages pivotal prompts embedded in historical trajectories for enhanced tracking. Specifically, EventTPT integrates the trajectories of multiple adjacent frames into a single event image using a time-weighted aggregation and subsequently inputs this as a visual prompt into the tracker for current frame locating. A cross-modal adaptive fusion module is further designed for object perception in scenarios with photometric inconsistency. Additionally, we introduce EventUAV, a novel and challenging RGB-Event tracking benchmark featuring objects with intricate motion dynamics and poor visibility in RGB-only modalities. Extensive experiments demonstrate that EventTPT surpasses state-of-the-art trackers on EventUAV and achieves competitive performance on other benchmarks (e.g., COESOT and VisEvent), underscoring its strong generalizability and robustness for resilient robotic vision systems. The code can be found at https://github.com/xiawenhao2022/EventTPT.
Wenhao Xia, Jiawen Zhu 0003, Jinqing Qi, You He 0002, Xu Jia 0012
ICRA6
2025 Regularizing Subspace Redundancy of Low-Rank Adaptation
abstract
Low-Rank Adaptation (LoRA) and its variants have delivered strong capability in Parameter-Efficient Transfer Learning (PETL) by minimizing trainable parameters and benefiting from reparameterization. However, their projection matrices remain unrestricted during training, causing high representation redundancy and diminishing the effectiveness of feature adaptation in the resulting subspaces. While existing methods mitigate this by manually adjusting the rank or implicitly applying channel-wise masks, they lack flexibility and generalize poorly across various datasets and architectures. Hence, we propose ReSoRA, a method that explicitly models redundancy between mapping subspaces and adaptively Regularizes Subspace redundancy of Low-Rank Adaptation. Specifically, it theoretically decomposes the low-rank submatrices into multiple equivalent subspaces and systematically applies de-redundancy constraints to the feature distributions across different projections. Extensive experiments validate that our proposed method consistently facilitates existing state-of-the-art PETL methods across various backbones and datasets in vision-language retrieval and standard visual classification benchmarks. Besides, as a training supervision, ReSoRA can be seamlessly integrated into existing approaches in a plug-and-play manner, with no additional inference costs. Code is publicly available at: https://github.com/Lucenova/ReSoRA.
Yue Zhu 0012, Haiwen Diao, Shang Gao 0012, Jiazuo Yu 0001, Jiawen Zhu 0003, Yunzhi Zhuge, Shuai Hao 0007, Xu Jia 0012, Lu Zhang 0053, Ying Zhang 0021, Huchuan Lu
ACM Multimedia8
2025 Automated Evaluation of Large Vision-Language Models on Self-Driving Corner Cases
abstract
Large Vision-Language Models (LVLMs) have received widespread attentions for advancing the interpretable self-driving. Existing evaluations of LVLMs primarily focus on multi-faceted capabilities in natural circumstances, lacking automated and quantifiable assessment for self-driving, let alone the severe road corner cases. In this work, we propose CODA-LM, the very first benchmark for the automatic evaluation of LVLMs for self-driving corner cases. We adopt a hierarchical data structure and prompt powerful LVLMs to analyze complex driving scenes and generate high-quality pre-annotations for the human annotators, while for LVLM evaluation, we show that using the text-only large language models (LLMs) as judges reveals even better alignment with human preferences than the LVLM judges. Moreover, with our CODA-LM, we build CODA-VLM, a new driving LVLM surpassing all open-sourced counterparts on CODA-LM. Our CODA-VLM performs comparably with GPT-4V, even surpassing GPT-4V by +21.42% on the regional perception task. We hope CODA-LM can become the catalyst to promote interpretable self-driving empowered by LVLMs.
Kai Chen 0023, Yanxin Liu, Ruiyuan Gao 0001, Lanqing Hong, Xinhai Zhao, Zhenguo Li, Dit-Yan Yeung, Huchuan Lu, Xu Jia 0012
WACV13
2025 TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion Models
abstract
Despite remarkable achievements in video synthesis, achieving granular control over complex dynamics, such as nuanced movement among multiple interacting objects, still presents a significant hurdle for dynamic world modeling, compounded by the necessity to manage appearance and disappearance, drastic scale changes, and ensure consistency for instances across frames. These challenges hinder the development of video generation that can faithfully mimic real-world complexity, limiting utility for applications requiring high-level realism and controllability, including advanced scene simulation and training of perception systems. To address that, we propose TrackDiffusion, a novel video generation framework affording fine-grained trajectory-conditioned motion control via diffusion models, which facilitates the precise manipulation of the object trajectories and interactions, overcoming the prevalent limitation of scale and continuity disruptions. A pivotal component of TrackDiffusion is the instance enhancer, which explicitly ensures inter-frame consistency of multiple objects, a critical factor overlooked in the current literature. More-over, we demonstrate that generated video sequences by our TrackDiffusion can be used as training data for visual per-ception models. To the best of our knowledge, this is the first work to apply video diffusion models with tracklet conditions and demonstrate that generated frames can be beneficial for improving the performance of object trackers. 1
Kai Chen 0023, Zhili Liu, Ruiyuan Gao 0001, Lanqing Hong, Dit-Yan Yeung, Huchuan Lu, Xu Jia 0012
WACV8
2025 MoBox: Enhancing Video Object Segmentation With Motion-Augmented Box Supervision
abstract
We propose MoBox, a low-cost solution for semi-supervised video object segmentation that requires only bounding boxes as manual annotations for training. Built upon a mature semi-supervised video object segmentation network, we redesign the training losses and employ a more stringent training strategy. Specifically, we introduce a well-designed constraint term that enhances traditional spatial projection by simultaneously leveraging the projections of both the ground-truth box and the predicted mask across two axes, rather than evaluating discrepancies along the x-axis and y-axis independently. To harness the intrinsic properties of videos, considering the underlying correspondence between motion represented by optical flow and the original image, we incorporate motion coherence information into the color consistency loss as supplementary information and propose a motion discrepancy loss to obtain accurate boundaries. Additionally, to mitigate the ambiguity of weak supervision, we further introduce the pseudo strict constraint during training, which significantly improves model performance. Our approach yields competitive scores on popular benchmarks, achieving a$\mathcal {J}\& \mathcal {F}$score of 78.6 on the DAVIS 2017 validation set and an Overall score of 78.0 on the YouTube-VOS 2018 validation set. These results highlight the efficacy of MoBox, demonstrating that the semi-supervised video object segmentation model can be effectively trained using only motion-augmented box supervision and intrinsic information of videos.
Xiaomin Li 0001, Dezhuang Li, Mengmeng Ge 0002, Xu Jia 0012, You He 0002, Huchuan Lu
IEEE Trans. Circuits Syst. Video Technol.5
2025 Pyramid Learnable Bandpass Filters for Ultra-High-Definition Image Demoiréing
abstract
Moiré patterns usually depend on the style of display grids and the position of shooting camera, appearing in the form of stripes, meshes or ripples, with various and irregular colors. Compared with low-resolution moiré images, high-definition (HD) and ultra-high-definition (UHD) moiré images exhibit more complex moiré patterns, e.g., wider distribution of moiré frequencies and higher coupling degree of moirés of different scales, which poses a greater challenge to the modeling capabilities of the model. To address these challenges, we propose a novel Pyramid Learnable Bandpass Filtering Network (PBNet) for demoiréing UHD images. Specifically, we propose a pyramid learnable bandpass filter (P-LBF) to perform multi-scale filtering in the same semantic context to obtain richer frequency domain information. The P-LBF contains three stages: aligning, filtering and fusing. First, we introduce a pyramid alignment (DA) to align neighbor pixels for eliminating the deviations raised by different styles of display grids and relative position of the shooting camera. Then, a pyramid filtering (PF) is conducted to model the complex and variable moiré patterns with aligned neighbor pixels. Finally, the frequency domain responses of these different scales are fused with a multi-dimensional feature fusion (MFF). The PBNet is constructed based on the P-LBF, incorporating a cross-layer feature fusion (CLF) module to facilitate more effective information interaction between features at different depths. Extensive experiments on four public datasets show that our model achieves state-of-the-art performance for both high- and low-resolution moiré images. The code is publicly available at:https://github.com/liuzhongqi1/PBNet.
Zhongqi Liu, Bolun Zheng, Qianyu Zhang 0002, Xu Jia 0012, Jiyong Zhang 0001, Chenggang Yan 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 Hierarchical Frequency-Based Upsampling and Refining for HEVC Compressed Video Enhancement
abstract
Video compression artifacts arise from quantization applied in the frequency domain. Video quality enhancement aims to reduce such compression artifacts and reconstruct a visually pleasant result. While existing methods effectively reduce artifacts in the spatial domain, they often overlook the rich frequency domain information, especially in addressing multi-scale compression artifacts. This work introduces a frequency-domain upsampling strategy within a multi-scale framework, specifically designed to focus on high-frequency details rather than simply blending neighboring pixels during the upsampling process. Our proposed hierarchical frequency-based upsampling and refinement neural network (HFUR) consists of two modules: implicit frequency upsampling (ImpFreqUp) and hierarchical and iterative refinement (HIR). ImpFreqUp exploits the DCT-domain prior derived through an implicit DCT transform, and accurately reconstructs the DCT-domain signal via a coarse-to-fine transfer. Additionally, HIR is introduced to facilitate cross-collaboration and information compensation between the scales, further refining the feature maps and promoting the visual quality of the final output. We demonstrate the effectiveness of the proposed modules via ablation experiments and visualized results. Experimental results demonstrate that HFUR outperforms the state-of-the-art methods up to 0.13dB/0.17dB on both constant bit rate and constant QP modes. The code is available athttps://github.com/zqqqyu/HFUR.
Qianyu Zhang 0002, Bolun Zheng, Xingying Chen, Zunjie Zhu, Canjin Wang, Zongpeng Li, Xu Jia 0012, Chengang Yan
IEEE Trans. Circuits Syst. Video Technol.8
2025 Perception-Oriented Bidirectional Attention Network for Image Super-Resolution Quality Assessment
abstract
Many super-resolution (SR) algorithms have been proposed to increase image resolution. However, full-reference (FR) image quality assessment (IQA) metrics for comparing and evaluating different SR algorithms are limited. In this work, we propose the Perception-oriented Bidirectional Attention Network (PBAN) for image SR FR-IQA, which is composed of three modules: an image encoder module, a perception-oriented bidirectional attention (PBA) module, and a quality prediction module. First, we encode the input images for feature representations. Inspired by the characteristics of the human visual system, we then construct the perception-oriented PBA module. Specifically, different from existing attention-based SR IQA methods, we conceive a Bidirectional Attention to bidirectionally construct visual attention to distortion, which is consistent with the generation and evaluation processes of SR images. To further guide the quality assessment towards the perception of distorted information, we propose Grouped Multi-scale Deformable Convolution, enabling the proposed method to adaptively perceive distortion. Moreover, we design Sub-information Excitation Convolution to direct visual perception to both sub-pixel and sub-channel attention. Finally, the quality prediction module is exploited to integrate quality-aware features and regress quality scores. Extensive experiments demonstrate that our proposed PBAN outperforms state-of-the-art quality assessment methods.
Xiaoyuan Yang 0003, Guanghui Yue 0001, Jun Fu 0007, Qiuping Jiang, Xu Jia 0012, Paul L. Rosin, Hantao Liu, Wei Zhou 0021
IEEE Trans. Image Process.6
2025 CharacterFactory: Sampling Consistent Characters With GANs for Diffusion Models
abstract
Recent advances in text-to-image models have opened new frontiers in human-centric generation. However, these models cannot be directly employed to generate images with consistent newly coined identities. In this work, we propose CharacterFactory, a framework that allows sampling new characters with consistent identities in the latent space of GANs for diffusion models. More specifically, we consider the word embeddings of celeb names as ground truths for the identity-consistent generation task and train a GAN model to learn the mapping from a latent space to the celeb embedding space. In addition, we design a context-consistent loss to ensure that the generated identity embeddings can produce identity-consistent images in various contexts. Remarkably, the whole model only takes 10 minutes for training, and can sample infinite characters end-to-end during inference. Extensive experiments demonstrate excellent performance of the proposed CharacterFactory on character creation in terms of identity consistency and editability. Furthermore, the generated characters can be seamlessly combined with the off-the-shelf image/video/3D diffusion models. We believe that the proposed CharacterFactory is an important step for identity-consistent character generation. Code and Gradio demo are available at: https://qinghew.github.io/CharacterFactory/.
Baolu Li 0001, Xiaomin Li 0001, Bing Cao 0002, Liqian Ma, Huchuan Lu, Xu Jia 0012
IEEE Trans. Image Process.7
2025 StableIdentity: Inserting Anybody Into Anywhere at First Sight
abstract
Recent advances in large pretrained text-to-image generation models have shown unprecedented capabilities for high-quality human-centric generation, however, customizing face identity is still an intractable problem. Existing methods cannot ensure stable identity preservation and flexible editability, even with several images for each subject during training. In this work, we propose StableIdentity, which allows identity-consistent recontextualization with just one face image from a person seen for the first time. More specifically, we employ a face encoder with the identity prior to encode the input face, and then calibrate the face representation to align the distribution of a space with the editability prior, which is constructed from celeb names. By incorporating identity prior and editability prior, the learned identity can be injected anywhere with various contexts. In addition, we design a masked two-phase diffusion loss to boost the pixel-level perception of the input face and maintain the diversity of generation. Extensive experiments demonstrate our method outperforms previous customization methods. In addition, the learned identity can be flexibly combined with the off-theshelf modules such as ControlNet. Notably, to the best of our knowledge, we are the first to directly inject the identity learned from a single image into video/3D generation without finetuning. We believe that the proposed StableIdentity is an important step to unify image, video, and 3D customized generation models. The code is available: https://github.com/qinghew/StableIdentity.
Xu Jia 0012, Xiaomin Li 0001, Taiqing Li, Liqian Ma, Yunzhi Zhuge, Huchuan Lu
IEEE Trans. Multim.2
2025 Event-Assisted Recurrent Network for Arbitrary-Temporal-Scale Blurry Image Unfolding
abstract
Recovering a sequence of latent sharp frames from a motion-blurred image is a challenging task. The bio-inspired event camera, which produces an event stream with high temporal resolution, has been exploited to promote the recovery performance. However, recovering sharp sequences with arbitrary temporal scales has been ignored for a long time. Existing works can only recover a fixed number of latent frames from a blurry image once they are trained. In this work, we propose an event-assisted blurry image unfolding framework that can work across arbitrary temporal scales. A bi-directional recurrent network is employed to encode events corresponding to each latent frame, which gathers information over all events in the exposure time. Features of both the blurry image and events are fused together and fed to a bi-directional latent sequence decoder (BiLSD) to produce a sequence of latent sharp frames. Extensive experiments show that the proposed method not only performs favorably against state-of-the-art methods in recovering a fixed number of frames from a blurry image but can be well generalized to arbitrary-temporal-scale blurry image unfolding.
Hao Ju 0004, Weihua He, Yaoyuan Wang, Shengming Li, Dong Wang 0004, Huchuan Lu, Xu Jia 0012
IEEE Trans. Neural Networks Learn. Syst.9
2024 UniPT: Universal Parallel Tuning for Transfer Learning with Efficient Parameter and Memory
abstract
Parameter-efficient transfer learning (PETL), i.e., finetuning a small portion of parameters, is an effective strategy for adapting pre-trained models to downstream domains. To further reduce the memory demand, recent PETL works focus on the more valuable memory-efficient characteristic. In this paper, we argue that the scalability, adaptability, and generalizability of state-of-the-art methods are hindered by structural dependency and pertinency on specific pretrained backbones. To this end, we propose a new memoryefficient PETL strategy, Universal Parallel Tuning (UniPT), to mitigate these weaknesses. Specifically, we facilitate the transfer process via a lightweight and learnable parallel network, which consists of: 1) A parallel interaction module that decouples the sequential connections and processes the intermediate activations detachedly from the pre-trained network. 2) A confidence aggregation module that learns optimal strategies adaptively for integrating cross-layer features. We evaluate UniPT with different backbones (e.g., T5 [69], VSE∞[12], CLIP4Clip [58], Clip-ViL [73], and MDETR [42]) on various vision-and-language and pure NLP tasks. Extensive ablations on 18 datasets have validated that UniPT can not only dramatically reduce memory consumption and outperform the best competitor, but also achieve competitive performance over other plain PETL methods with lower training memory overhead. Our code is publicly available at: https://github.com/Paranioar/UniPT.
Haiwen Diao, Ying Zhang 0021, Xu Jia 0012, Huchuan Lu, Long Chen 0016
CVPR4
2024 SHERL: Synthesizing High Accuracy and Efficient Memory for Resource-Limited Transfer Learning
Haiwen Diao, Xu Jia 0012, Yunzhi Zhuge, Ying Zhang 0021, Huchuan Lu, Long Chen 0016
ECCV (44)3
2024 EvSign: Sign Language Recognition and Translation with Streaming Events
Zeren Wang, Wenyue Chen, Shengming Li, Dong Wang 0004, Huchuan Lu, Xu Jia 0012
ECCV (5)8
2024 Multi-Stage Fusion for Event-based Multimodal Tracker
abstract
Event cameras are bio-inspired sensors with high dynamic range and time resolution, which are favorable properties for visual object tracking. There are already some methods that fuse the event modality and RGB modality with cross-domain feature integrator to achieve improved tracking performance. Researchers have developed some architectures for event modality processing or fusion, successfully boosting the tracking performance. In this work, we design a RGB-E tracker with multi-stage fusion. In the early stage, frames are enhanced with aid of events to mitigate blur or under/over-exposure degradation. During the middle stage, we utilize a fusion module for feature-level integration. At the late stage, we carry out decision-level fusion by predicting tracking boxes based on frame features, event features, and fused features, and the one with highest score is taken as the final estimation. Our design thoroughly integrate information from various levels, allowing each modality to contribute to the tracking process as much as possible. Extensive experiments demonstrate that the proposed method performs favorably against state-of-the-art RGB-E trackers in both accuracy and efficiency.
Xinyu Zhang 0017, Hefei Huang, Xu Jia 0012, Wenyue Chen, Dong Wang 0004, Shengming Li, Huchuan Lu
ICME3
2024 Customizing Text-to-Image Generation with Inverted Interaction
abstract
Subject-driven image generation, aimed at customizing user-specified subjects, has experienced rapid progress. However, most of them focus on transferring the customized appearance of subjects. In this work, we consider a novel concept customization task, that is, capturing the interaction between subjects in exemplar images and transferring the learned concept of interaction to achieve customized text-to-image generation. Intrinsically, the interaction between subjects is diverse and is difficult to describe in only a few words. In addition, typical exemplar images are about the interaction between humans, which further intensifies the challenge of interaction-driven image generation with various categories of subjects. To address this task, we adopt a divide-and-conquer strategy and propose a two-stage interaction inversion framework. The framework begins by learning a pseudo-word for a single pose of each subject in the interaction. This is then employed to promote the learning of the concept for the interaction. In addition, language prior and cross-attention loss are incorporated into the optimization process to encourage the modeling of interaction. Extensive experiments demonstrate that the proposed methods are able to effectively invert the interactive pose from exemplar images and apply it to the customized generation with user-specified interaction.
Mengmeng Ge 0002, Xu Jia 0012, Takashi Isobe, Xiaomin Li 0001, Dong Zhou 0003, Li Wang 0125, Huchuan Lu, Ashish Sirasao, Emad Barsoum
ACM Multimedia2
2024 Event-Guided Rolling Shutter Correction with Time-Aware Cross-Attentions
abstract
Many consumer cameras with rolling shutter (RS) CMOS would suffer undesired distortion and artifacts, particularly when objects experiences fast motion. The neuromorphic event camera, with high temporal resolution events, could bring much benefit to the RS correction process. In this work, we explore the characteristics of RS images and event data for the design of the rolling shutter correction (RSC) model. Specifically, the relationship between RS images and event data is modeled by incorporating time encoding to the computation of cross-attention in transformer encoder to achieve time-aware multi-modal information fusion. Features from RS images enhanced by event data are adopted as keys and values in transformer decoder, providing source for appearance, while features from event data enhanced by RS images are adopted as queries, providing spatial transition information. By embedding the time information of the desired global shutter (GS) image into the query, the transformer with deformable attention is capable of producing the target GS image.To enhance the model's generalization ability, we propose to further self-supervise the model by cycling between time coordinate systems corresponding to RS images and GS images. Extensive evaluations over both synthetic and real datasets demonstrate that the proposed method performs favorably against state-of-the-art approaches.
Hefei Huang, Xu Jia 0012, Xinyu Zhang 0017, Shengming Li, Huchuan Lu
ACM Multimedia2
2024 MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models
abstract
Existing pretrained text-to-video (T2V) models have demonstrated impressive abilities in generating realistic videos with basic motion or camera movement. However, these models exhibit significant limitations when generating intricate, human-centric motions. Current efforts primarily focus on fine-tuning models on a small set of videos containing a specific motion. They often fail to effectively decouple motion and the appearance in the limited reference videos, thereby weakening the modeling capability of motion patterns. To this end, we propose MoTrans, a customized motion transfer method enabling video generation of similar motion in new context. Specifically, we introduce a multimodal large language model (MLLM)-based recaptioner to expand the initial prompt to focus more on appearance and an appearance injection module to adapt appearance prior from video frames to the motion modeling process. These complementary multimodal representations from recaptioned prompt and video frames promote the modeling of appearance and facilitate the decoupling of appearance and motion. In addition, we devise a motion-specific embedding for further enhancing the modeling of the specific motion. Experimental results demonstrate that our method effectively learns specific motion pattern from singular or multiple reference videos, performing favorably against existing methods in customized video generation.
Xiaomin Li 0001, Xu Jia 0012, Haiwen Diao, Mengmeng Ge 0002, You He 0002, Huchuan Lu
ACM Multimedia2
2024 Video Frame Interpolation for Large Motion with Generative Prior
Xu Jia 0012, Lu Zhang 0053, Xiaomin Li 0001, Huchuan Lu
PRCV (10)2
2024 Wavelet-based network for high dynamic range imaging
abstract
High dynamic range (HDR) imaging from multiple low dynamic range (LDR) images has been suffering from ghosting artifacts caused by scene and objects motion. Existing methods, such as optical flow based and end-to-end deep learning based solutions, are error-prone either in detail restoration or ghosting artifacts removal. Comprehensive empirical evidence shows that ghosting artifacts caused by large foreground motion are mainly low-frequency signals and the details are mainly high-frequency signals. In this work, we propose a novel frequency-guided end-to-end deep neural network (FHDRNet) to conduct HDR fusion in the frequency domain, and Discrete Wavelet Transform (DWT) is used to decompose inputs into different frequency bands. The low-frequency signals are used to avoid specific ghosting artifacts, while the high-frequency signals are used for preserving details. Using a U-Net as the backbone, we propose two novel modules: merging module and frequency-guided upsampling module. The merging module applies the attention mechanism to the low-frequency components to deal with the ghost caused by large foreground motion. The frequency-guided upsampling module reconstructs details from multiple frequency-specific components with rich details. In addition, a new RAW dataset is created for training and evaluating multi-frame HDR imaging algorithms in the RAW domain. Extensive experiments are conducted on public datasets and our RAW dataset, showing that the proposed FHDRNet achieves state-of-the-art performance.
Tianhong Dai, Wei Li 0002, Xilei Cao, Jianzhuang Liu, Xu Jia 0012, Ales Leonardis, Youliang Yan, Shanxin Yuan
Comput. Vis. Image Underst.5
2024 Neural image re-exposure
Xinyu Zhang 0017, Hefei Huang, Xu Jia 0012, Dong Wang 0004, Lihe Zhang, Bolun Zheng, Wei Zhou 0021, Huchuan Lu
Comput. Vis. Image Underst.3
2024 Efficient Adaptive Feature Fusion Network for Remote-Sensing Image Super-Resolution
abstract
Image super-resolution is a fundamental low-level vision task aimed at recovering high-resolution images with fine details. Deep learning has significantly enhanced the performance of super-resolution techniques for remote sensing imagery. However, increasing the depth of networks and the size of their parameters has resulted in substantial computational and storage burdens. To address this challenge, we propose an adaptive approach that learns both local and global information for each region. We introduce a lightweight hybrid model named the Efficient Adaptive Feature Fusion Network, which combines CNNs and Transformers to fully exploit the texture information in remote sensing images. This model leverages local details and long-range dependencies within images in an adaptive manner to achieve superior super-resolution. Specifically, a set of Transformers is employed to model the self-similarity between pixels and perform dense texture pattern predictions at each pixel, while a set of CNNs captures local details within the images. The computed global and local features serve as inputs to the proposed Adaptive Contextual Fusion Block, which learns to fuse local and global information across different regions to generate robust image super-resolution features. We conduct extensive experimental evaluations of the proposed method on the UCMerced and AID datasets, demonstrating its outstanding performance in terms of PSNR and SSIM metrics. Comprehensive experiments validate the effectiveness of our approach, showing that the proposed method achieves an excellent balance between performance and complexity.
Shuai Hao 0007, Shuai Liu 0009, Xu Jia 0012, Huchuan Lu, You He 0002
IEEE Signal Process. Lett.3
2024 Event-Based Shutter Unrolling and Motion Deblurring in Dynamic Scenes
abstract
The Rolling Shutter (RS) effect and motion blur are common challenges in images captured by CMOS cameras during dynamic scenes. Inspired by biological vision principles, event cameras capture intensity changes asynchronously with low latency, providing valuable insights into image degradation during exposure. This study addresses the dual challenges of rolling shutter correction and deblurring using event data, merging them into a unified one-stage network. This streamlined approach reduces cumulative errors and inference time compared to traditional two-stage methods. To achieve this, we introduce an Event Representation for Rolling Shutter Deblurring, which explicitly models the conversion relationship between the input RS blurry frame and the latent image using events. To enhance the fusion of image and event information, we present a Time-guided Cross-Modal Attention module. Furthermore, we improve performance by incorporating a Multi-Scale Context-Aware Transformer Block, effectively addressing varying degrees of distortion and blurriness using a multi-scale attention mechanism. Extensive experiments validate that our method outperforms existing state-of-the-art approaches.
Yangguang Wang, Chenxu Jiang, Xu Jia 0012, Yufei Guo 0001, Lei Yu 0006
IEEE Signal Process. Lett.3
2024 Event-Assisted Blurriness Representation Learning for Blurry Image Unfolding
abstract
The goal of blurry image deblurring and unfolding task is to recover a single sharp frame or a sequence from a blurry one. Recently, its performance is greatly improved with introduction of a bio-inspired visual sensor, event camera. Most existing event-assisted deblurring methods focus on the design of powerful network architectures and effective training strategy, while ignoring the role of blur modeling in removing various blur in dynamic scenes. In this work, we propose to implicitly model blur in an image by computing blurriness representation with an event-assisted blurriness encoder. The learning of blurriness representation is formulated as a ranking problem based on specially synthesized pairs. Blurriness-aware image unfolding is achieved by integrating blur relevant information contained in the representation into a base unfolding network. The integration is mainly realized by the proposed blurriness-guided modulation and multi-scale aggregation modules. Experiments on GOPRO and HQF datasets show favorable performance of the proposed method against state-of-the-art approaches. More results on real-world data validate its effectiveness in recovering a sequence of latent sharp frames from a blurry image.
Hao Ju 0004, Lei Yu 0006, Weihua He, Yaoyuan Wang, Qi Xu 0008, Shengming Li, Dong Wang 0004, Huchuan Lu, Xu Jia 0012
IEEE Trans. Image Process.11
2024 Deformable Dynamic Sampling and Dynamic Predictable Mask Mining for Image Inpainting
abstract
Existing image inpainting methods often produce artifacts that are caused by using vanilla convolution layers as building blocks that treat all image regions equally and generate holes at random locations with equal probability. This design does not differentiate the missing regions and valid regions in inference and does not consider the predictability of missing regions in training. To address these issues, we propose a deformable dynamic sampling (DDS) mechanism which is built on deformable convolutions (DCs), and a constraint is proposed to avoid the deformably sampled elements falling into the corrupted regions. Furthermore, to select both valid sample locations and suitable kernels dynamically, we equip DCs with content-aware dynamic kernel selection (DKS). In addition, to further encourage the DDS mechanism to find meaningful sampling locations, we propose to train the inpainting model with mined predictable regions as holes. During training, we jointly train a mask generator with the inpainting network to generate hole masks dynamically for each training sample. Thus, the mask generator can find large yet predictable missing regions as a better alternative to random masks. Extensive experiments demonstrate the advantages of our method over state-of-the-art methods qualitatively and quantitatively.
Cai Cai, Yu Zeng 0001, Shu Yang 0004, Xu Jia 0012, Huchuan Lu, You He 0002
IEEE Trans. Neural Networks Learn. Syst.4
2023 Dual Memory Aggregation Network for Event-Based Object Detection with Learnable Representation
abstract
Event-based cameras are bio-inspired sensors that capture brightness change of every pixel in an asynchronous manner. Compared with frame-based sensors, event cameras have microsecond-level latency and high dynamic range, hence showing great potential for object detection under high-speed motion and poor illumination conditions. Due to sparsity and asynchronism nature with event streams, most of existing approaches resort to hand-crafted methods to convert event data into 2D grid representation. However, they are sub-optimal in aggregating information from event stream for object detection. In this work, we propose to learn an event representation optimized for event-based object detection. Specifically, event streams are divided into grids in the x-y-t coordinates for both positive and negative polarity, producing a set of pillars as 3D tensor representation. To fully exploit information with event streams to detect objects, a dual-memory aggregation network (DMANet) is proposed to leverage both long and short memory along event streams to aggregate effective information for object detection. Long memory is encoded in the hidden state of adaptive convLSTMs while short memory is modeled by computing spatial-temporal correlation between event pillars at neighboring time intervals. Extensive experiments on the recently released event-based automotive detection dataset demonstrate the effectiveness of the proposed method.
Xu Jia 0012, Xinyu Zhang 0017, Yaoyuan Wang, Dong Wang 0004, Huchuan Lu
AAAI2
2023 GM-NeRF: Learning Generalizable Model-Based Neural Radiance Fields from Multi-View Images
abstract
In this work, we focus on synthesizing high-fidelity novel view images for arbitrary human performers, given a set of sparse multi-view images. It is a challenging task due to the large variation among articulated body poses and heavy self-occlusions. To alleviate this, we introduce an effective generalizable framework Generalizable Model-based Neural Radiance Fields (GM-NeRF) to synthesize free-viewpoint images. Specifically, we propose a geometry-guided attention mechanism to register the appearance code from multi-view 2D images to a geometry proxy which can alleviate the misalignment between inaccurate geometry prior and pixel space. On top of that, we further conduct neural rendering and partial gradient back propagation for efficient perceptual supervision and improvement of the perceptual quality of synthesis. To evaluate our method, we conduct experiments on synthesized datasets THuman2.0 and Multi-garment, and real-world datasets Genebody and ZJUMocap. The results demonstrate that our approach outperforms state-of-the-art methods in terms of novel view synthesis and geometric reconstruction.
Jianchuan Chen, Wentao Yi, Liqian Ma, Xu Jia 0012, Huchuan Lu
CVPR4
2023 Compression-Aware Video Super-Resolution
abstract
Videos stored on mobile devices or delivered on the Internet are usually in compressed format and are of various unknown compression parameters, but most video super-resolution (VSR) methods often assume ideal inputs resulting in large performance gap between experimental settings and real-world applications. In spite of a few pioneering works being proposed recently to super-resolve the compressed videos, they are not specially designed to deal with videos of various levels of compression. In this paper, we propose a novel and practical compression-aware video super-resolution model, which could adapt its video enhancement process to the estimated compression level. A compression encoder is designed to model compression levels of input frames, and a base VSR model is then conditioned on the implicitly computed representation by inserting compression-aware modules. In addition, we propose to further strengthen the VSR model by taking full advantage of meta data that is embedded naturally in compressed video streams in the procedure of information fusion. Extensive experiments are conducted to demonstrate the effectiveness and efficiency of the proposed method on compressed VSR benchmarks. The codes will be available at https://github.com/aprBlue/CAVSR
Takashi Isobe, Xu Jia 0012, Xin Tao 0001, Huchuan Lu, Yu-Wing Tai
CVPR3
2023 Image Super-Resolution with Implicit Texture Pattern Modulation
abstract
Image super-resolution is one of the classical low-level vision tasks with the purpose of restoring a high-resolution image with fine details. Being aware of texture patterns with an image would benefit super-resolution performance a lot. However, it would be difficult to predict texture patterns for each region because of lack of annotations with different kinds of categories. In this work, we propose to implicitly model texture information with each region and take that as prior to promote super-resolution performance. In order to fully explore texture patterns, a hybrid model of convolutional neural networks and transformers is proposed. It is able to take advantage of both local and long range dependencies within an image to super-resolve an image. Specifically, a set of transformers are employed to model self-similarity among pixels within an image and to make dense texture patterns prediction at each pixel. The computed texture pattern representation then works as condition to modulate convolution-based residual blocks. In this way texture patterns could be integrated into the CNNs to obtain powerful features for image super-resolution. Extensive experiments on several benchmark datasets demonstrate its favorable performance against state-of-the-art methods and show its potential as a generic design for hybrid of transformers and CNNs.
Shuai Hao 0007, Xu Jia 0012, You He 0002, Huchuan Lu
ICME3
2023 A uniform transformer-based structure for feature fusion and enhancement for RGB-D saliency detection
Yue Wang 0038, Xu Jia 0012, Lu Zhang 0053, James H. Elder, Huchuan Lu
Pattern Recognit.2
2023 Event-Based Semantic Segmentation With Posterior Attention
abstract
In the past years, attention-based Transformers have swept across the field of computer vision, starting a new stage of backbones in semantic segmentation. Nevertheless, semantic segmentation under poor light conditions remains an open problem. Moreover, most papers about semantic segmentation work on images produced by commodity frame-based cameras with a limited framerate, hindering their deployment to auto-driving systems that require instant perception and response at milliseconds. An event camera is a new sensor that generates event data at microseconds and can work in poor light conditions with a high dynamic range. It looks promising to leverage event cameras to enable perception where commodity cameras are incompetent, but algorithms for event data are far from mature. Pioneering researchers stack event data as frames so that event-based segmentation is converted to frame-based segmentation, but characteristics of event data are not explored. Noticing that event data naturally highlight moving objects, we propose a posterior attention module that adjusts the standard attention by the prior knowledge provided by event data. The posterior attention module can be readily plugged into many segmentation backbones. Plugging the posterior attention module into a recently proposed SegFormer network, we get EvSegFormer (the event-based version of SegFormer) with state-of-the-art performance in two datasets (MVSEC and DDD-17) collected for event-based segmentation. Code is available at https://github.com/zexiJia/EvSegFormer to facilitate research on event-based vision.
Zexi Jia, Kaichao You, Weihua He, Yang Tian 0002, Yongxiang Feng, Yaoyuan Wang, Xu Jia 0012, Yihang Lou, Guoqi Li 0002
IEEE Trans. Image Process.7
2022 Blind Image Super-Resolution with Degradation-Aware Adaptation
Yue Wang 0038, Jiawen Ming, Xu Jia 0012, James H. Elder, Huchuan Lu
ACCV (3)3
2022 Multi-granularity Transformer for Image Super-Resolution
Yunzhi Zhuge, Xu Jia 0012
ACCV (3)2
2022 TimeReplayer: Unlocking the Potential of Event Cameras for Video Interpolation
abstract
Recording fast motion in a high FPS (frame-per-second) requires expensive high-speed cameras. As an alternative, interpolating low-FPS videos from commodity cameras has attracted significant attention. If only low-FPS videos are available, motion assumptions (linear or quadratic) are necessary to infer intermediate frames, which fail to model complex motions. Event camera, a new camera with pixels producing events of brightness change at the temporal resolution of μs (10–6second), is a game-changing device to enable video interpolation at the presence of arbitrarily complex motion. Since event camera is a novel sensor, its potential has not been fulfilled due to the lack of processing algorithms. The pioneering work Time Lens introduced event cameras to video interpolation by designing optical devices to collect a large amount of paired training data of high-speed frames and events, which is too costly to scale. To fully unlock the potential of event cameras, this paper proposes a novel TimeReplayer algorithm to interpolate videos captured by commodity cameras with events. It is trained in an unsupervised cycleconsistent style, canceling the necessity of high-speed training data and bringing the additional ability of video extrapolation. Its state-of-the-art results and demo videos in supplementary reveal the promising future of event-based vision.
Weihua He, Kaichao You, Zhendong Qiao, Xu Jia 0012, Wenhui Wang 0001, Huchuan Lu, Yaoyuan Wang, Jianxing Liao
CVPR4
2022 Look Back and Forth: Video Super-Resolution with Explicit Temporal Difference Modeling
abstract
Temporal modeling is crucial for video super-resolution. Most of the video super-resolution methods adopt the optical flow or deformable convolution for explicitly motion compensation. However, such temporal modeling techniques increase the model complexity and might fail in case of occlusion or complex motion, resulting in serious distortion and artifacts. In this paper, we propose to explore the role of explicit temporal difference modeling in both LR and HR space. Instead of directly feeding consecutive frames into a VSR model, we propose to compute the temporal difference between frames and divide those pixels into two subsets according to the level of difference. They are separately processed with two branches of different receptive fields in order to better extract complementary information. To further enhance the super-resolution result, not only spatial residual features are extracted, but the difference between consecutive frames in high-frequency domain is also computed. It allows the model to exploit intermediate SR results in both future and past to refine the current SR output. The difference at different time steps could be cached such that information from further distance in time could be propagated to the current frame for refinement. Experiments on several video super-resolution benchmark datasets demonstrate the effectiveness of the proposed method and its favorable performance against state-of-the-art methods.
Takashi Isobe, Xu Jia 0012, Xin Tao 0001, Ruihuang Li, Yongjie Shi, Huchuan Lu, Yu-Wing Tai
CVPR2
2022 Class-Balanced Pixel-Level Self-Labeling for Domain Adaptive Semantic Segmentation
abstract
Domain adaptive semantic segmentation aims to learn a model with the supervision of source domain data, and produce satisfactory dense predictions on unlabeled target domain. One popular solution to this challenging task is self-training, which selects high-scoring predictions on target samples as pseudo labels for training. However, the produced pseudo labels often contain much noise because the model is biased to source domain as well as majority categories. To address the above issues, we propose to di-rectly explore the intrinsic pixel distributions of target do-main data, instead of heavily relying on the source domain. Specifically, we simultaneously cluster pixels and rectify pseudo labels with the obtained cluster assignments. This process is done in an online fashion so that pseudo labels could co-evolve with the segmentation model without extra training rounds. To overcome the class imbalance problem on long-tailed categories, we employ a distribution align-ment technique to enforce the marginal class distribution of cluster assignments to be close to that of pseudo labels. The proposed method, namely Class-balanced Pixel-level Self-Labeling (CPSL), improves the segmentation performance on target domain over state-of-the-arts by a large margin, especially on long-tailed categories. The source code is available at ht tps: / / gi thub. com/lslrh/CPSL.
Ruihuang Li, Shuai Li 0014, Chenhang He, Yabin Zhang 0001, Xu Jia 0012, Lei Zhang 0006
CVPR5
2022 AdaInt: Learning Adaptive Intervals for 3D Lookup Tables on Real-time Image Enhancement
abstract
The 3D Lookup Table (3D LUT) is a highly-efficient tool for real-time image enhancement tasks, which models a non-linear 3D color transform by sparsely sampling it into a discretized 3D lattice. Previous works have made efforts to learn image-adaptive output color values of LUTs for flexible enhancement but neglect the importance of sampling strategy. They adopt a sub-optimal uniform sampling point allocation, limiting the expressiveness of the learned LUTs since the (tri-)linear interpolation between uniform sampling points in the LUT transform might fail to model local non-linearities of the color transform. Focusing on this problem, we present AdaInt (Adaptive Intervals Learning), a novel mechanism to achieve a more flexible sampling point allocation by adaptively learning the non-uniform sampling intervals in the 3D color space. In this way, a 3D LUT can increase its capability by conducting dense sampling in color ranges requiring highly non-linear transforms and sparse sampling for near-linear transforms. The proposed AdaInt could be implemented as a compact and efficient plug-and-play module for a 3D LUT-based method. To enable the end-to-end learning of AdaInt, we design a novel differentiable operator called AiLUT-Transform (Adaptive Interval LUT Transform) to locate input colors in the non-uniform 3D LUT and provide gradients to the sampling intervals. Experiments demonstrate that methods equipped with AdaInt can achieve state-of-the-art performance on two public benchmark datasets with a negligible overhead increase. Our source code is available at https://github.com/ImCharlesY/AdaInt.
Canqian Yang, Meiguang Jin, Xu Jia 0012, Yi Xu 0001, Ying Chen 0011
CVPR3
2022 Refactoring ISP for High-Level Vision Tasks
abstract
The image signal processing (ISP) pipeline, which transforms raw sensor measurement to a color image, is composed of a sequence of processing modules. Traditionally, the ISP pipeline is manually tuned by experts for human perception. The resulting handcrafted ISP configuration does not necessarily benefit the downstream high-level vision tasks. To mitigate these problems, this paper presents a simple yet effective framework based on Evolutionary Algorithm to search for a set of compact ISP configurations for high-level vision tasks. In particular, we encode ISP structure into a binary string and ISP parameters into a set of float numbers. Then we jointly optimize them with task-specific loss and ISP computation budgets (e.g., running time) through solving a nonlinear multi-objective optimization problem. By mutating the configurations of the ISP pipeline, we are able to remove redundant modules and design an ISP with both low cost and high accuracy. We validate the proposed method on extreme noisy and low-light raw images, and experimental results show that our framework can help find effective and efficient ISP configurations for both object detection and semantic segmentation tasks. We further provide a detailed analysis on the importance of different modules in the ISP configurations, which benefits the design of ISP for downstream tasks in the future.
Yongjie Shi, Songjiang Li, Xu Jia 0012, Jianzhuang Liu
ICRA3
2022 A Continual Learning Survey: Defying Forgetting in Classification Tasks
abstract
Artificial neural networks thrive in solving the classification problem for a particular rigid task, acquiring knowledge through generalized learning behaviour from a distinct training phase. The resulting network resembles a static entity of knowledge, with endeavours to extend this knowledge without targeting the original task resulting in a catastrophic forgetting. Continual learning shifts this paradigm towards networks that can continually accumulate knowledge over different tasks without the need to retrain from scratch. We focus on task incremental classification, where tasks arrive sequentially and are delineated by clear boundaries. Our main contributions concern: (1) a taxonomy and extensive overview of the state-of-the-art; (2) a novel framework to continually determine the stability-plasticity trade-off of the continual learner; (3) a comprehensive experimental comparison of 11 state-of-the-art continual learning methods; and (4) baselines. We empirically scrutinize method strengths and weaknesses on three benchmarks, considering Tiny Imagenet and large-scale unbalanced iNaturalist and a sequence of recognition datasets. We study the influence of model capacity, weight decay and dropout regularization, and the order in which the tasks are presented, and qualitatively compare methods in terms of required memory, computation time, and storage.
Matthias De Lange, Rahaf Aljundi, Marc Masana, Sarah Parisot, Xu Jia 0012, Ales Leonardis, Gregory Slabaugh, Tinne Tuytelaars
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Neighbor2Neighbor: A Self-Supervised Framework for Deep Image Denoising
abstract
In recent years, image denoising has benefited a lot from deep neural networks. However, these models need large amounts of noisy-clean image pairs for supervision. Although there have been attempts in training denoising networks with only noisy images, existing self-supervised algorithms suffer from inefficient network training, heavy computational burden, or dependence on noise modeling. In this paper, we proposed a self-supervised framework named Neighbor2Neighbor for deep image denoising. We develop a theoretical motivation and prove that by designing specific samplers for training image pairs generation from only noisy images, we can train a self-supervised denoising network similar to the network trained with clean images supervision. Besides, we propose a regularizer in the perspective of optimization to narrow the optimization gap between the self-supervised denoiser and the supervised denoiser. We present a very simple yet effective self-supervised training scheme based on the theoretical understandings: training image pairs are generated by random neighbor sub-samplers, and denoising networks are trained with a regularized loss. Moreover, we propose a training strategy named BayerEnsemble to adapt the Neighbor2Neighbor framework in raw image denoising. The proposed Neighbor2Neighbor framework can enjoy the progress of state-of-the-art supervised denoising networks in network architecture design. It also avoids heavy dependence on the assumption of the noise distribution. We evaluate the Neighbor2Neighbor framework through extensive experiments, including synthetic experiments with different noise distributions and real-world experiments under various scenarios. The code is available online: https://github.com/TaoHuang2018/Neighbor2Neighbor.
Tao Huang 0022, Songjiang Li, Xu Jia 0012, Huchuan Lu, Jianzhuang Liu
IEEE Trans. Image Process.3
2021 Semi-Supervised Domain Adaptation Based on Dual-Level Domain Mixing for Semantic Segmentation
abstract
Data-driven based approaches, in spite of great success in many tasks, have poor generalization when applied to unseen image domains, and require expensive cost of annotation especially for dense pixel prediction tasks such as semantic segmentation. Recently, both unsupervised domain adaptation (UDA) from large amounts of synthetic data and semi-supervised learning (SSL) with small set of labeled data have been studied to alleviate this issue. However, there is still a large gap on performance compared to their supervised counterparts. We focus on a more practical setting of semi-supervised domain adaptation (SSDA) where both a small set of labeled target data and large amounts of labeled source data are available. To address the task of SSDA, a novel framework based on dual-level domain mixing is proposed. The proposed framework consists of three stages. First, two kinds of data mixing methods are proposed to reduce domain gap in both region-level and sample-level respectively. We can obtain two complementary domain-mixed teachers based on dual-level mixed data from holistic and partial views respectively. Then, a student model is learned by distilling knowledge from these two teachers. Finally, pseudo labels of unlabeled data are generated in a self-training manner for another few rounds of teachers training. Extensive experimental results have demonstrated the effectiveness of our proposed framework on synthetic-to-real semantic segmentation benchmarks.
Shuaijun Chen, Xu Jia 0012, Yongjie Shi, Jianzhuang Liu
CVPR2
2021 Multi-Source Domain Adaptation With Collaborative Learning for Semantic Segmentation
abstract
Multi-source unsupervised domain adaptation (MSDA) aims at adapting models trained on multiple labeled source domains to an unlabeled target domain. In this paper, we propose a novel multi-source domain adaptation framework based on collaborative learning for semantic segmentation. Firstly, a simple image translation method is introduced to align the pixel value distribution to reduce the gap between source domains and target domain to some extent. Then, to fully exploit the essential semantic information across source domains, we propose a collaborative learning method for domain adaptation without seeing any data from target domain. In addition, similar to the setting of unsupervised domain adaptation, unlabeled target domain data is leveraged to further improve the performance of domain adaptation. This is achieved by additionally constraining the outputs of multiple adaptation models with pseudo labels online generated by an ensembled model. Extensive experiments and ablation studies are conducted on the widely-used domain adaptation benchmark datasets in semantic segmentation. Our proposed method achieves 59.0% mIoU on the validation set of Cityscapes by training on the labeled Synscapes and GTA5 datasets and unlabeled training set of Cityscapes. It significantly outperforms all previous state-of-the-arts single-source and multi-source unsupervised domain adaptation methods.
Xu Jia 0012, Shuaijun Chen, Jianzhuang Liu
CVPR2
2021 Neighbor2Neighbor: Self-Supervised Denoising From Single Noisy Images
abstract
In the last few years, image denoising has benefited a lot from the fast development of neural networks. However, the requirement of large amounts of noisy-clean image pairs for supervision limits the wide use of these models. Although there have been a few attempts in training an image denoising model with only single noisy images, existing self-supervised denoising approaches suffer from inefficient network training, loss of useful information, or dependence on noise modeling. In this paper, we present a very simple yet effective method named Neighbor2Neighbor to train an effective image denoising model with only noisy images. Firstly, a random neighbor sub-sampler is proposed for the generation of training image pairs. In detail, input and target used to train a network are images sub-sampled from the same noisy image, satisfying the requirement that paired pixels of paired images are neighbors and have very similar appearance with each other. Secondly, a denoising network is trained on sub-sampled training pairs generated in the first stage, with a proposed regularizer as additional loss for better performance. The proposed Neighbor2Neighbor framework is able to enjoy the progress of state-of-the-art supervised denoising networks in network architecture design. Moreover, it avoids heavy dependence on the assumption of the noise distribution. We explain our approach from a theoretical perspective and further validate it through extensive experiments, including synthetic experiments with different noise distributions in sRGB space and real-world experiments on a denoising benchmark dataset in raw-RGB space.
Tao Huang 0022, Songjiang Li, Xu Jia 0012, Huchuan Lu, Jianzhuang Liu
CVPR3
2021 Multi-Target Domain Adaptation With Collaborative Consistency Learning
abstract
Recently unsupervised domain adaptation for the semantic segmentation task has become more and more popular due to high-cost of pixel-level annotation on real-world images. However, most domain adaptation methods are only restricted to single-source-single-target pair, and can not be directly extended to multiple target domains. In this work, we propose a collaborative learning framework to achieve unsupervised multi-target domain adaptation. An unsupervised domain adaptation expert model is first trained for each source-target pair and is further encouraged to collaborate with each other through a bridge built between different target domains. These expert models are further improved by adding the regularization of making the consistent pixel-wise prediction for each sample with the same structured context. To obtain a single model that works across multiple target domains, we propose to simultaneously learn a student model which is trained to not only imitate the output of each expert on the corresponding target domain, but also to pull different expert close to each other with regularization on their weights. Extensive experiments demonstrate that the proposed method can effectively exploit rich structured information contained in both labeled source domain and multiple unlabeled target domains. Not only does it perform well across multiple target domains but also performs favorably against state-of-the-art unsupervised domain adaptation methods specially trained on a single source-target pair. Code is available at https://github.com/junpan19/MTDA.
Takashi Isobe, Xu Jia 0012, Shuaijun Chen, Yongjie Shi, Jianzhuang Liu, Huchuan Lu, Shengjin Wang
CVPR2
2021 T-SVDNet: Exploring High-Order Prototypical Correlations for Multi-Source Domain Adaptation
abstract
Most existing domain adaptation methods focus on adaptation from only one source domain, however, in practice there are a number of relevant sources that could be leveraged to help improve performance on target domain. We propose a novel approach named T-SVDNet to address the task of Multi-source Domain Adaptation (MDA), which is featured by incorporating Tensor Singular Value Decomposition (T-SVD) into a neural network’s training pipeline. Overall, high-order correlations among multiple domains and categories are fully explored so as to better bridge the domain gap. Specifically, we impose Tensor-Low-Rank (TLR) constraint on a tensor obtained by stacking up a group of prototypical similarity matrices, aiming at capturing consistent data structure across different domains. Furthermore, to avoid negative transfer brought by noisy source data, we propose a novel uncertainty-aware weighting strategy to adaptively assign weights to different source domains and samples based on the result of uncertainty estimation. Extensive experiments conducted on public benchmarks demonstrate the superiority of our model in addressing the task of MDA compared to state-of-the-art methods. Code is available at https://github.com/lslrh/T-SVDNet.
Ruihuang Li, Xu Jia 0012, Shuaijun Chen, Qinghua Hu
ICCV2
2021 Motion Deblurring with Real Events
abstract
In this paper, we propose an end-to-end learning framework for event-based motion deblurring in a self-supervised manner, where real-world events are exploited to alleviate the performance degradation caused by data inconsistency. To achieve this end, optical flows are predicted from events, with which the blurry consistency and photometric consistency are exploited to enable self-supervision on the deblurring network with real-world data. Furthermore, a piecewise linear motion model is proposed to take into account motion non-linearities and thus leads to an accurate model for the physical formation of motion blurs in the real-world scenario. Extensive evaluation on both synthetic and real motion blur datasets demonstrates that the proposed algorithm bridges the gap between simulated and real-world motion blurs and shows remarkable performance for eventbased motion deblurring in real-world scenarios.
Lei Yu 0006, Bishan Wang, Wen Yang 0001, Gui-Song Xia, Xu Jia 0012, Zhendong Qiao, Jianzhuang Liu
ICCV6
2021 Weakly-Supervised Temporal Action Localization via Cross-Stream Collaborative Learning
abstract
Weakly supervised temporal action localization (WTAL) is a challenging task as only video-level category labels are available during training stage. Without precise temporal annotations, most approaches rely on complementary RGB and optical flow features to predict the start and end frame of each action category in a video. However, existing approaches simply resort to either concatenation or weighted sum to learn how to take advantages of these two modalities for accurate action localization, which ignore the substantial variance between such two modalities. In this paper, we present Cross-Stream Collaborative Learning (CSCL) to address these issues. The proposed CSCL introduce a cross-stream weighting module to identify which modality is more robust during training and take advantage of the robust modality to guide the weaker one. Furthermore, we suppress the snippets which has high action-ness scores in both modalities to further exploiting the complementary property between two modalities. In addition, we bring the concept of co-training for WTAL and take both modalities into account for pseudo label generation to help training a stronger model. Extensive experiments conducted on THUMOS14 and ActivityNet dataset demonstrate that CSCL achieves a favorable performance against state-of-the-arts methods.
Xu Jia 0012, Huchuan Lu, Xiang Ruan
ACM Multimedia2
2021 Towards effective learning for face super-resolution with shape and pose perturbations
Xiyuan Hu, Zhenfeng Fan, Xu Jia 0012, Xuyun Zhang, Lianyong Qi, Zuxing Xuan
Knowl. Based Syst.3
2020 Efficient Residual Dense Block Search for Image Super-Resolution
abstract
Although remarkable progress has been made on single image super-resolution due to the revival of deep convolutional neural networks, deep learning methods are confronted with the challenges of computation and memory consumption in practice, especially for mobile devices. Focusing on this issue, we propose an efficient residual dense block search algorithm with multiple objectives to hunt for fast, lightweight and accurate networks for image super-resolution. Firstly, to accelerate super-resolution network, we exploit the variation of feature scale adequately with the proposed efficient residual dense blocks. In the proposed evolutionary algorithm, the locations of pooling and upsampling operator are searched automatically. Secondly, network architecture is evolved with the guidance of block credits to acquire accurate super-resolution network. The block credit reflects the effect of current block and is earned during model evaluation process. It guides the evolution by weighing the sampling probability of mutation to favor admirable blocks. Extensive experimental results demonstrate the effectiveness of the proposed searching method and the found efficient super-resolution models achieve better performance than the state-of-the-art methods with limited number of parameters and FLOPs.
Dehua Song, Chang Xu 0002, Xu Jia 0012, Yiyi Chen 0003, Chunjing Xu, Yunhe Wang 0001
AAAI3
2020 Video Super-Resolution With Temporal Group Attention
abstract
Video super-resolution, which aims at producing a high-resolution video from its corresponding low-resolution version, has recently drawn increasing attention. In this work, we propose a novel method that can effectively incorporate temporal information in a hierarchical way. The input sequence is divided into several groups, with each one corresponding to a kind of frame rate. These groups provide complementary information to recover missing details in the reference frame, which is further integrated with an attention module and a deep intra-group fusion module. In addition, a fast spatial alignment is proposed to handle videos with large motion. Extensive results demonstrate the capability of the proposed model in handling videos with various motion. It achieves favorable performance against state-of-the-art methods on several benchmark datasets.
Takashi Isobe, Songjiang Li, Xu Jia 0012, Shanxin Yuan, Gregory Slabaugh, Chunjing Xu, Yali Li 0001, Shengjin Wang, Qi Tian 0001
CVPR3
2020 Unsupervised Model Personalization While Preserving Privacy and Scalability: An Open Problem
abstract
This work investigates the task of unsupervised model personalization, adapted to continually evolving, unlabeled local user images. We consider the practical scenario where a high capacity server interacts with a myriad of resource-limited edge devices, imposing strong requirements on scalability and local data privacy. We aim to address this challenge within the continual learning paradigm and provide a novel Dual User-Adaptation framework (DUA) to explore the problem. This framework flexibly disentangles user-adaptation into model personalization on the server and local data regularization on the user device, with desirable properties regarding scalability and privacy constraints. First, on the server, we introduce incremental learning of task-specific expert models, subsequently aggregated using a concealed unsupervised user prior. Aggregation avoids retraining, whereas the user prior conceals sensitive raw user data, and grants unsupervised adaptation. Second, local user-adaptation incorporates a domain adaptation point of view, adapting regularizing batch normalization parameters to the user data. We explore various empirical user configurations with different priors in categories and a tenfold of transforms for MIT Indoor Scene recognition, and classify numbers in a combined MNIST and SVHN setup. Extensive experiments yield promising results for data-driven local adaptation and elicit user priors for server adaptation to depend on the model rather than user data. Hence, although user-adaptation remains a challenging open problem, the DUA framework formalizes a principled foundation for personalizing both on server and user device, while maintaining privacy and scalability.
Matthias De Lange, Xu Jia 0012, Sarah Parisot, Ales Leonardis, Gregory Slabaugh, Tinne Tuytelaars
CVPR2
2020 Joint Demosaicing and Denoising With Self Guidance
abstract
Usually located at the very early stages of the computational photography pipeline, demosaicing and denoising play important parts in the modern camera image processing. Recently, some neural networks have shown the effectiveness in joint demosaicing and denoising (JDD). Most of them first decompose a Bayer raw image into a four-channel RGGB image and then feed it into a neural network. This practice ignores the fact that the green channels are sampled at a double rate compared to the red and the blue channels. In this paper, we propose a self-guidance network (SGNet), where the green channels are initially estimated and then works as a guidance to recover all missing values in the input image. In addition, as regions of different frequencies suffer different levels of degradation in image restoration. We propose a density-map guidance to help the model deal with a wide range of frequencies. Our model outperforms state-of-the-art joint demosaicing and denoising methods on four public datasets, including two real and two synthetic data sets. Finally, we also verify that our method obtains best results in joint demosaicing , denoising and super-resolution.
Lin Liu 0016, Xu Jia 0012, Jianzhuang Liu, Qi Tian 0001
CVPR2
2020 Learning to Select Base Classes for Few-Shot Classification
abstract
Few-shot learning has attracted intensive research attention in recent years. Many methods have been proposed to generalize a model learned from provided base classes to novel classes, but no previous work studies how to select base classes, or even whether different base classes will result in different generalization performance of the learned model. In this paper, we utilize a simple yet effective measure, the Similarity Ratio, as an indicator for the generalization performance of a few-shot model. We then formulate the base class selection problem as a submodular optimization problem over Similarity Ratio. We further provide theoretical analysis on the optimization lower bound of different optimization methods, which could be used to identify the most appropriate algorithm for different experimental settings. The extensive experiments on ImageNet, Caltech256 and CUB-200-2011 demonstrate that our proposed method is effective in selecting a better base dataset.
Linjun Zhou, Peng Cui 0001, Xu Jia 0012, Shiqiang Yang, Qi Tian 0001
CVPR3
2020 More Classifiers, Less Forgetting: A Generic Multi-classifier Paradigm for Incremental Learning
Yu Liu 0012, Sarah Parisot, Gregory Slabaugh, Xu Jia 0012, Ales Leonardis, Tinne Tuytelaars
ECCV (26)4
2020 Video Super-Resolution with Recurrent Structure-Detail Network
Takashi Isobe, Xu Jia 0012, Shuhang Gu, Songjiang Li, Shengjin Wang, Qi Tian 0001
ECCV (12)2
2019 Video Generation From Single Semantic Label Map
abstract
This paper proposes the novel task of video generation conditioned on a SINGLE semantic label map, which provides a good balance between flexibility and quality in the generation process. Different from typical end-to-end approaches, which model both scene content and dynamics in a single step, we propose to decompose this difficult task into two sub-problems. As current image generation methods do better than video generation in terms of detail, we synthesize high quality content by only generating the first frame. Then we animate the scene based on its semantic meaning to obtain temporally coherent video, giving us excellent results overall. We employ a cVAE for predicting optical flow as a beneficial intermediate step to generate a video sequence conditioned on the initial single frame. A semantic label map is integrated into the flow prediction module to achieve major improvements in the image-to-video generation process. Extensive experiments on the Cityscapes dataset show that our method outperforms all competing methods.
Junting Pan, Chengyu Wang 0003, Xu Jia 0012, Lu Sheng, Xiaogang Wang 0001
CVPR3
2019 Co-Evolutionary Compression for Unpaired Image Translation
abstract
Generative adversarial networks (GANs) have been successfully used for considerable computer vision tasks, especially the image-to-image translation. However, generators in these networks are of complicated architectures with large number of parameters and huge computational complexities. Existing methods are mainly designed for compressing and speeding-up deep neural networks in the classification task, and cannot be directly applied on GANs for image translation, due to their different objectives and training procedures. To this end, we develop a novel co-evolutionary approach for reducing their memory usage and FLOPs simultaneously. In practice, generators for two image domains are encoded as two populations and synergistically optimized for investigating the most important convolution filters iteratively. Fitness of each individual is calculated using the number of parameters, a discriminator-aware regularization, and the cycle consistency. Extensive experiments conducted on benchmark datasets demonstrate the effectiveness of the proposed method for obtaining compact and effective generators.
Han Shu, Yunhe Wang 0001, Xu Jia 0012, Kai Han 0002, Hanting Chen, Chunjing Xu, Qi Tian 0001, Chang Xu 0002
ICCV3
2019 Exemplar Guided Unsupervised Image-to-Image Translation with Semantic Consistency
Liqian Ma, Xu Jia 0012, Stamatios Georgoulis, Tinne Tuytelaars, Luc Van Gool
ICLR (Poster)2
2017 Pose Guided Person Image Generation
abstract
This paper proposes the novel Pose Guided Person Generation Network (PG$^2$) that allows to synthesize person images in arbitrary poses, based on an image of that person and a novel pose. Our generation framework PG$^2$ utilizes the pose information explicitly and consists of two key stages: pose integration and image refinement. In the first stage the condition image and the target pose are fed into a U-Net-like network to generate an initial but coarse image of the person with the target pose. The second stage then refines the initial and blurry result by training a U-Net-like generator in an adversarial way. Extensive experimental results on both 128$\times$64 re-identification images and 256$\times$256 fashion photos show that our model generates high-quality person images with convincing details.
Liqian Ma, Xu Jia 0012, Qianru Sun, Bernt Schiele, Tinne Tuytelaars, Luc Van Gool
NIPS2
2017 Visual tracking with structured patch-based model
Fu Li 0003, Xu Jia 0012, Cheng Xiang 0001, Huchuan Lu
Image Vis. Comput.2
2016 Towards Automatic Image Editing: Learning to See another You
Xu Jia 0012, Amir Ghodrati, Marco Pedersoli, Tinne Tuytelaars
BMVC1
2016 Dynamic Filter Networks
abstract
In a traditional convolutional layer, the learned filters stay fixed after training. In contrast, we introduce a new framework, the Dynamic Filter Network, where filters are generated dynamically conditioned on an input. We show that this architecture is a powerful one, with increased flexibility thanks to its adaptive nature, yet without an excessive increase in the number of model parameters. A wide variety of filtering operation can be learned this way, including local spatial transformations, but also others like selective (de)blurring or adaptive feature extraction. Moreover, multiple such layers can be combined, e.g. in a recurrent architecture. We demonstrate the effectiveness of the dynamic filter network on the tasks of video and stereo prediction, and reach state-of-the-art performance on the moving MNIST dataset with a much smaller model. By visualizing the learned filters, we illustrate that the network has picked up flow information by only looking at unlabelled training data. This suggests that the network can be used to pretrain networks for various supervised tasks in an unsupervised way, like optical flow and depth estimation.
Xu Jia 0012, Bert De Brabandere, Tinne Tuytelaars, Luc Van Gool
NIPS1
2016 Visual Tracking via Coarse and Fine Structural Local Sparse Appearance Models
abstract
Sparse representation has been successfully applied to visual tracking by finding the best candidate with a minimal reconstruction error using target templates. However, most sparse representation-based tracking methods only consider holistic rather than local appearance to discriminate between target and background regions, and hence may not perform well when target objects are heavily occluded. In this paper, we develop a simple yet robust tracking algorithm based on a coarse and fine structural local sparse appearance model. The proposed method exploits both partial and structural information of a target object based on sparse coding using the dictionary composed of patches from multiple target templates. The likelihood obtained by averaging and pooling operations exploits consistent appearance of object parts, thereby helping not only locate targets accurately but also handle partial occlusion. To update templates more accurately without introducing occluding regions, we introduce an occlusion detection scheme to account for pixels belonging to the target objects. The proposed method is evaluated on a large benchmark data set with three evaluation metrics. Experimental results demonstrate that the proposed tracking algorithm performs favorably against several state-of-the-art methods.
Xu Jia 0012, Huchuan Lu, Ming-Hsuan Yang 0001
IEEE Trans. Image Process.1
2015 Guiding the Long-Short Term Memory Model for Image Caption Generation
abstract
In this work we focus on the problem of image caption generation. We propose an extension of the long short term memory (LSTM) model, which we coin gLSTM for short. In particular, we add semantic information extracted from the image as extra input to each unit of the LSTM block, with the aim of guiding the model towards solutions that are more tightly coupled to the image content. Additionally, we explore different length normalization strategies for beam search to avoid bias towards short sentences. On various benchmark datasets such as Flickr8K, Flickr30K and MS COCO, we obtain results that are on par with or better than the current state-of-the-art.
Xu Jia 0012, Efstratios Gavves, Basura Fernando, Tinne Tuytelaars
ICCV1
2015 Swap Retrieval: Retrieving Images of Cats When the Query Shows a Dog
abstract
Query-by-example remains popular in image retrieval because it can exploit contextual information encoded in the image, that is difficult to express in a traditional textual query. Textual queries, on the other hand, give more flexibility in that it's easy to reformulate and refine a text query based on initial results.
Amir Ghodrati, Xu Jia 0012, Marco Pedersoli, Tinne Tuytelaars
ICMR2
2012 Visual tracking via adaptive structural local sparse appearance model
abstract
Sparse representation has been applied to visual tracking by finding the best candidate with minimal reconstruction error using target templates. However most sparse representation based trackers only consider the holistic representation and do not make full use of the sparse coefficients to discriminate between the target and the background, and hence may fail with more possibility when there is similar object or occlusion in the scene. In this paper we develop a simple yet robust tracking method based on the structural local sparse appearance model. This representation exploits both partial information and spatial information of the target based on a novel alignment-pooling method. The similarity obtained by pooling across the local patches helps not only locate the target more accurately but also handle occlusion. In addition, we employ a template update strategy which combines incremental subspace learning and sparse representation. This strategy adapts the template to the appearance change of the target with less possibility of drifting and reduces the influence of the occluded target template as well. Both qualitative and quantitative evaluations on challenging benchmark image sequences demonstrate that the proposed tracking algorithm performs favorably against several state-of-the-art methods.
Xu Jia 0012, Huchuan Lu, Ming-Hsuan Yang 0001
CVPR1
2012 Fragment-based tracking using online multiple kernel learning
abstract
Fragment-based tracking methods have shown its robustness in handling partial occlusion and pose change. In this paper, we propose a novel fragment-based tracking approach using on online multiple kernel learning (MKL) method. An online MKL method for object tracking is implemented by considering temporal continuity explicitly. Instead of directly using multiple features of objects, we employ MKL to make full use of multiple fragments of the object. This can automatically assign different weights to the fragments according to their discriminative power. In addition, for better robustness two kinds of independent features are computed to enrich the representation of patches. We build a classifier for each type of feature and assign them different weights according to their performance on classification. Both qualitative and quantitative evaluations on challenging image sequences demonstrate that the proposed tracking approach performs favorably against several state-of-the-art methods.
Xu Jia 0012, Dong Wang 0004, Huchuan Lu
ICIP1