Lingqiao Liu

dblp:45/7776 · DBLP profile ↗
← Back
131ranked-venue papers
11as first author
70since 2021 · last 2026
0000-0003-3584-795XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 99 · 10 first-author · 45 since 2021Graphics, computer vision, multimedia, augmented reality and games · 74 · 6 first-author · 37 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 6 since 2021Systems, architecture and hardware · 5 · 2 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ReFINE: A Reward-Based Framework for Interpretable and Nuanced Evaluation of Radiology Report Generation
abstract
Automated radiology report generation (R2Gen) has advanced significantly, yet evaluation remains challenging due to the complexity of assessing report quality. Traditional metrics often misalign with human judgments, failing to identify specific deficiencies. To address this, we introduce ReFINE, a framework for training an Evaluation Model using a novel margin-based reward enforcement loss. This approach decomposes report quality into fine-grained sub-scores across user-defined criteria, improving interpretability. Leveraging GPT-4, we generate diverse training data with paired accepted and rejected reports to train our model under a reward-based system. The trained ReFINE Score provides both granular sub-scores and an aggregated quality assessment, enabling criterion-specific evaluation. Experimental results demonstrate ReFINE's superior alignment with human judgments, outperforming traditional metrics in model selection. Its robustness is validated across three expert-annotated datasets—including chest X-rays and multimodal reports covering 9 imaging modalities—and under two distinct scoring systems.
Yunyi Liu, Yingshu Li 0001, Zhanyu Wang, Lingqiao Liu, Lei Wang 0001, Luping Zhou
AAAI5
2026 HiTVideo: Hierarchical Tokenizers for Enhancing Text-to-Video Generation with Autoregressive Large Language Models
abstract
Text-to-video generation poses significant challenges due to the inherent complexity of video data, which spans both temporal and spatial dimensions. It introduces additional redundancy, abrupt variations, and a domain gap between language and vision tokens while generation. Addressing these challenges requires an effective video tokenizer that can efficiently encode video data while preserving essential semantic and spatiotemporal information, serving as a critical bridge between text and vision. Inspired by the observation in VQ-VAE-2, we propose HiTVideo, a novel approach for text-to-video generation with hierarchical tokenizers. It utilizes a 3D causal VAE with a multi-layer discrete token framework, encoding video content into hierarchically structured codebooks. Higher layers capture semantic information with higher compression, while lower layers focus on fine-grained spatiotemporal details, striking a balance between compression efficiency and reconstruction quality. Our approach efficiently encodes longer video sequences (e.g., 8 seconds, 64 frames), reducing bits per pixel (bpp) by approximately 70% compared to previous tokenizers, while maintaining competitive reconstruction quality. We explore the trade-offs between compression and reconstruction, while emphasizing the advantages of high-compressed semantic tokens in text-to-video tasks. HiTVideo aims to address the potential limitations of existing video tokenizers in text-to-video generation tasks, striving for higher compression ratios, improved token quality, and simplify LLMs modeling under language guidance, offering a scalable and promising framework for advancing text to video generation.
Ziqin Zhou, Yifan Yang 0004, Yuqing Yang 0001, Tianyu He, Houwen Peng, Qi Dai 0001, Lili Qiu, Chong Luo 0001, Lingqiao Liu
AAAI10
2026 Revisiting Vision-Language Foundations for No-Reference Image Quality Assessment
abstract
Large-scale vision–language pre-training has recently shown promise for no-reference image-quality assessment (NR-IQA), yet the relative merits of modern Vision Transformer foundations remain poorly understood. In this work, we present the first systematic evaluation of six prominent pretrained backbones, CLIP, SigLIP2, DINOv2, DINOv3, Perception, and ResNet, for the task of No-Reference Image Quality Assessment (NR-IQA), each fine-tuned using an identical lightweight MLP head. Our study uncovers two previously overlooked factors: (1) SigLIP2 consistently achieves strong performance; and (2) the choice of activation function plays a surprisingly crucial role, particularly for enhancing the generalization ability of image quality assessment models. Notably, we find that simple sigmoid activations outperform commonly used ReLU and GELU on several benchmarks. Motivated by this finding, we introduce a learnable activation selection mechanism that adaptively determines the nonlinearity for each channel, eliminating the need for manual activation design, and achieving new state-of-the-art SRCC on CLIVE, KADID10K, and AGIQA3K. Extensive ablations confirm the benefits across architectures and regimes, establishing strong, resource-efficient NR-IQA baselines. Code is available at https://github.com/drkkgy/NR_IQA_AGM
Ta Duc Huy, Lingqiao Liu
WACV3
2026 MoSCo: Real-time and Efficient Text-to-Motion Synthesis via Delta Training
abstract
Generating expressive, fine-grained human motion from text remains a formidable challenge, particularly when aiming for high fidelity without incurring excessive computational cost. Existing methods often rely on complex, multi-stage pipelines with slow inference and large memory footprints, hindering real-time deployment. To address these limitations, we introduce MoSCo, a simple autoregressive text-to-motion framework that discretizes motion into part-level token sequences and models temporal dynamics via delta-based training strategy —i.e., predicting the motion difference from the previous time step—before fusing these tokens with textual embeddings through our Part-Aware Coordinator (PAO) and generating the full sequence with a single, lightweight transformer decoder. MoSCo sets a new milestone in text-to-motion inference speed—achieving an AITS of just 0.002s (vs 0.03s), over an order of magnitude faster than all prior methods—while maintaining a compact model footprint and delivering highly realistic motions (FID 0.085), making real-time, high-quality generation practical.
Lingqiao Liu
WACV2
2026 LearnMat: Semantic-Aware Self-Supervision Fine-Grained Visual Recognition
abstract
Self-supervised learning has shown potential in fine-grained visual recognition (FGVR). However, existing self-supervised learning methods are often susceptible to irrelevant patterns during training and lack the ability to capture the critical subtle differences in FGVR, leading to suboptimal performance. Moreover, existing approaches focus primarily on uni-modal visual concepts. Despite the emergence of powerful vision-language models (VLMs) in various high-level vision tasks, their potential in self-supervised FGVR remains largely unexplored. To this end, we propose a novel self-supervised learning (LearnMat) framework, that effectively filters out irrelevant feature interference and extracts more important and subtle discriminative features during training. Specifically, LearnMat consists of two key modules: the semantic awareness module (SAM) and the insight extraction module (IEM). In the SAM, we introduce a novel vision-language-grounded semantic distillation strategy using a corpus of generic, category-agnostic textual attributes, that injects explicit semantic constraints into self-supervised training and improves robustness to background interference. Complementarily, the IEM exploits gradient-based signals from the input image to highlight subtle differences and localize key discriminative regions, mitigating inter-class similarity and intra-class variation, and enhancing fine-grained discrimination. Extensive experiments across multiple popular FGVR datasets show that LearnMat significantly outperforms recent state-of-the-art methods, highlighting its marked effectiveness. Our code is avaliable at https://github.com/Heng-CHY/LearnMat.
ShuaiHeng Li, Fan Zhang 0070, Yangyang Shu, Guanbin Li, Junyu Dong, Lingqiao Liu, David Zhang 0001
IEEE Trans. Image Process.7
2026 Learning Feature Encoder With Synthetic Anomalies for Weakly Supervised Graph Anomaly Detection
abstract
Weakly supervised graph anomaly detection aims to unveil unusual graph instances, e.g., nodes, whose behaviors significantly differ from normal ones, given only a limited number of annotated anomalies and abundant unlabeled samples. A major challenge is to learn a meaningful latent feature representation that reduces intra-class variance among normal data while remaining highly sensitive to anomalies. Although recent works have applied self-supervised feature learning for graph anomaly detection, their strategies are not specifically tailored to its unique requirements, motivating our exploration of a more domain-specific approach. In this paper, we introduce a weakly supervised graph anomaly detection method that leverages a feature learning strategy tailored for graph anomalies. Our approach is built upon a multi-task learning scheme that extracts robust feature representations through synthesized anomalies. We generate synthetic anomalies by perturbing the normal graph in various ways and assign a dedicated detection head to each anomaly type, ensuring that learned features are sensitive to potential deviations from normal patterns. Although synthetic anomalies may not perfectly replicate real-world patterns, they provide valuable auxiliary data for effective feature learnin, much like features learned from ImageNet classification transfer to downstream vision tasks. Additionally, we adopt a two-phase learning strategy: an initial warm-up phase using only synthetic samples, followed by a full-training phase integrating both tasks, to balance the influence of synthetic and real data. Extensive experiments on public datasets demonstrate the superior performance of our method over its competitors. Code is available at https://github.com/yj-zhou/SAWGAD.
Yingjie Zhou 0001, Yuqin Xie, Fanxing Liu, Dongjin Song, Ce Zhu, Lingqiao Liu
IEEE Trans. Knowl. Data Eng.6
2026 Towards High-Quality Novel View Synthesis From Nonuniformly Distributed Input Views
abstract
Novel view synthesis is a crucial task in the field of visualization and computer graphics. Although current methods, such as neural radiance fields (NeRF), have proven versatile for novel view synthesis, they often assume that input views are uniformly distributed in space. This means that each view captures a portion of the real-world scene, with each area being equally represented in the input views. In contrast, real-world scenarios involve randomly captured images with an uneven spatial distribution of input views, where some views of the scene are observed more frequently than others. This discrepancy poses significant challenges for current methods, underscoring the need to address the issue of nonuniformity in practical applications. To tackle this challenge, we propose a new dataset and an effective approach. Our extensive dataset comprises diverse scenes, capturing images in various real-world environments. Unlike previous novel view synthesis datasets, which have uniform or near-uniform input view distributions, our dataset effectively reflects the nonuniform distribution of images in real-world settings. Additionally, we introduce a simple yet effective method that significantly enhances the performance of current techniques under nonuniform conditions without modifying their pipelines. We anticipate that our data and method will serve as a valuable source of inspiration for future research in this area.
Libo Sun 0002, Jiatong Xia, Lingqiao Liu
IEEE Trans. Vis. Comput. Graph.3
2025 Enhancing Close-up Novel View Synthesis via Pseudo-labeling
abstract
Recent methods, such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS), have demonstrated remarkable capabilities in novel view synthesis. However, despite their success in producing high-quality images for viewpoints similar to those seen during training, they struggle when generating detailed images from viewpoints that significantly deviate from the training set, particularly in close-up views. The primary challenge stems from the lack of specific training data for close-up views, leading to the inability of current methods to render these views accurately. To address this issue, we introduce a novel pseudo-label-based learning strategy. This approach leverages pseudo-labels derived from existing training data to provide targeted supervision across a wide range of close-up viewpoints. Recognizing the absence of benchmarks for this specific challenge, we also present a new dataset designed to assess the effectiveness of both current and future methods in this area. Our extensive experiments demonstrate the efficacy of our approach.
Jiatong Xia, Libo Sun 0002, Lingqiao Liu
AAAI3
2025 Attention-Driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models Without Fine-Tuning
abstract
Recent advancements in Multimodal Large Language Models (MLLMs) have generated significant interest in their ability to autonomously interact with and interpret Graphical User Interfaces (GUIs). A major challenge in these systems is grounding—accurately identifying critical GUI components such as text or icons based on a GUI image and a corresponding text query. Traditionally, this task has relied on fine-tuning MLLMs with specialized training data to predict component locations directly. However, in this paper, we propose a novel Tuning-free Attention-driven Grounding (TAG) method that leverages the inherent attention patterns in pretrained MLLMs to accomplish this task without the need for additional fine-tuning. Our method involves identifying and aggregating attention maps from specific tokens within a carefully constructed query prompt. Applied to MiniCPM-Llama3-V 2.5, a state-of-the-art MLLM, our tuning-free approach achieves performance comparable to tuning-based methods, with notable success in text localization. Additionally, we demonstrate that our attention map-based grounding technique significantly outperforms direct localization predictions from MiniCPM-Llama3-V 2.5, highlighting the potential of using attention maps from pretrained MLLMs and paving the way for future innovations in this domain.
Qi Chen 0014, Lei Wang 0001, Lingqiao Liu
AAAI4
2025 One Last Attention for Your Vision-Language Model
abstract
Pretrained vision-language models (VLMs), such as CLIP, achieve remarkable zero-shot performance, yet their downstream potential hinges on effective fine-tuning. Most adaptation methods typically focus on refining representation from separate modalities (text or vision) but neglect the critical role of their fused representations in the decision-making process, \emph{\ie} rational matrix that drives the final prediction. To bridge the gap, we propose a simple yet effective \textbf{R}ational \textbf{Ada}ptaion ({RAda}) to explicitly exploit the final fused representation during fine-tuning. RAda employs a learned mask, obtained from a lightweight attention layer attached at the end of a VLM, to dynamically calibrate the contribution of each element in the rational matrix, enabling targeted adjustments to the final cross-modal interactions without incurring costly modifications to intermediate features. Experiments in different settings (i.e., updating, or freezing pretrained encoders in adaptation, and test-time training that can only access the unlabeled test data) show that RAda serves as a versatile fine-tuning technique, improving the baseline with minimal code and performing comparably against current arts in most settings. Code is available at \href{https://github.com/khufia/RAda/tree/main}{github.com/khufia/RAda}.
Liang Chen 0030, Ghazi Shazan Ahmad, Tianjun Yao, Lingqiao Liu
ICCV4
2025 Effective Tuning Strategies for Generalist Robot Manipulation Policies
abstract
Generalist robot manipulation policies (GMPs) have the potential to generalize across a wide range of tasks, devices, and environments. However, existing policies continue to struggle with out-of-distribution scenarios due to the inherent difficulty of collecting sufficient action data to cover extensively diverse domains. While fine-tuning offers a practical way to quickly adapt a GMPs to novel domains and tasks with limited samples, we observe that the performance of the resulting GMPs differs significantly with respect to the design choices of fine-tuning strategies. In this work, we first conduct an indepth empirical study to investigate the effect of key factors in GMPs fine-tuning strategies, covering the action space, policy head, supervision signal and the choice of tunable parameters, where 2,500 rollouts are evaluated for a single configuration. We systematically discuss and summarize our findings and identify the key design choices, which we believe give a practical guideline for GMPs fine-tuning. We observe that in a lowdata regime, with carefully chosen fine-tuning strategies, a GMPs significantly outperforms the state-of-the-art imitation learning algorithms. The results presented in this work establish a new baseline for future studies on fine-tuned GMPs.
Wenbo Zhang 0009, Yang Li 0184, Yanyuan Qiao, Siyuan Huang 0004, Jiajun Liu 0004, Feras Dayoub, Lingqiao Liu
ICRA8
2025 Embodied Domain Adaptation for Object Detection
abstract
Mobile robots rely on object detectors for perception and object localization in indoor environments. However, standard closed-set methods struggle to handle the diverse objects and dynamic conditions encountered in real homes and labs. Open-vocabulary object detection (OVOD), driven by Vision Language Models (VLMs), extends beyond fixed labels but still struggles with domain shifts in indoor environments. We introduce a Source-Free Domain Adaptation (SFDA) approach that adapts a pre-trained model without accessing source data. We refine pseudo labels via temporal clustering, employ multi-scale threshold fusion, and apply a Mean Teacher framework with contrastive learning. Our Embodied Domain Adaptation for Object Detection (EDAOD) benchmark evaluates adaptation under sequential changes in lighting, layout, and object diversity. Our experiments show significant gains in zero-shot detection performance and flexible adaptation to dynamic indoor conditions.
Yanyuan Qiao, Lingqiao Liu, Feras Dayoub
IROS3
2025 PedCLIP: A Vision-Language Model for Pediatric X-Rays with Mixture of Body Part Experts
Ta Duc Huy, Abin Shoby, Sen Kim Tran, Yutong Xie 0001, Qi Chen 0014, Phi-Le Nguyen, Akshay Gole, Lingqiao Liu, Antonios Perperidis, Mark Friswell, Rebecca Linke, Andrea Glynn, Minh-Son To, Anton van den Hengel, Johan Verjans, Zhibin Liao, Minh Hieu Phan
MICCAI (5)8
2025 Let Your Video Listen to Your Music! - Beat-Aligned, Content-Preserving Video Editing with Arbitrary Music
abstract
Aligning the rhythm of visual motion in a video with a given music track is a practical need in multimedia production, yet remains an underexplored task in autonomous video editing. Effective alignment between motion and musical beats enhances viewer engagement and visual appeal, particularly in music videos, promotional content, and cinematic editing. Existing methods typically depend on labor-intensive manual cutting, speed adjustments, or heuristic-based editing techniques to achieve synchronization. While some generative models handle joint video and music generation, they often entangle the two modalities, limiting flexibility in aligning video to music beats while preserving the full visual content. In this paper, we propose a novel and efficient framework-termed MVAA (Music-Video Auto-Alignment)-that automatically edits video to align with the rhythm of a given music track while preserving the original visual content. To enhance flexibility, we modularize the task into a two-step process in our MVAA: aligning motion keyframes with audio beats, followed by rhythm-aware video inpainting. Specifically, we first insert keyframes at timestamps aligned with musical beats, then use a frame-conditioned diffusion model to generate coherent intermediate frames, preserving the original video's semantic content. Since comprehensive test-time training can be time-consuming, we adopt a two-stage strategy: pretraining the inpainting module on a small video set to learn general motion priors, followed by rapid inference-time fine-tuning for video-specific adaptation. This hybrid approach enables adaptation within ~10 minutes with one epoch on a single NVIDIA 4090 GPU using CogVideoX-5b-I2V [77] as the backbone. Extensive experiments show that our approach can achieve high-quality beat alignment and visual smoothness. User studies further validate the natural rhythmic quality of the results, confirming their effectiveness for practical music-video editing. The code is available at: zhangxinyu-xyz.github.io/MVAA
Xinyu Zhang 0017, Dong Gong, Zicheng Duan, Anton van den Hengel, Lingqiao Liu
ACM Multimedia5
2025 You Can Generate It Again: Data-to-Text Generation with Verification and Correction Prompting
abstract
Small language models like T5 excel in generating high-quality text for data-to-text tasks, offering adaptability and cost-efficiency compared to Large Language Models (LLMs). However, they frequently miss keywords, which is considered one of the most severe and common errors in this task. In this work, we explore the potential of using feedback systems to enhance semantic fidelity in smaller language models for data-to-text generation tasks, through our Verification and Correction Prompting (VCP) approach. In the inference stage, our approach involves a multi-step process, including generation, verification, and regeneration stages. During the verification stage, we implement a simple rule to check for the presence of every keyword in the prediction. Recognizing that this rule can be inaccurate, we have developed a carefully designed training procedure, which enabling the model to incorporate feedback from the error-correcting prompt effectively, despite its potential inaccuracies. The VCP approach effectively reduces the Semantic Error Rate (SER) while maintaining the text’s quality.
Xuan Ren, Zeyu Zhang 0006, Lingqiao Liu
MMAsia3
2025 Chain-of-Action: Trajectory Autoregressive Modeling for Robotic Manipulation
abstract
We present Chain-of-Action (CoA), a novel visuomotor policy paradigm built upon Trajectory Autoregressive Modeling. Unlike conventional approaches that predict next step action(s) forward, CoA generates an entire trajectory by explicit backward reasoning with task-specific goals through an action-level Chain-of-Thought (CoT) process. This process is unified within a single autoregressive structure: (1) the first token corresponds to a stable keyframe action that encodes the task-specific goals; and (2) subsequent action tokens are generated autoregressively, conditioned on the initial keyframe and previously predicted actions. This backward action reasoning enforces a global-to-local structure, allowing each local action to be tightly constrained by the final goal. To further realize the action reasoning structure, CoA incorporates four complementary designs: continuous action token representation; dynamic stopping for variable-length trajectory generation; reverse temporal ensemble; and multi-token prediction to balance action chunk modeling with global structure. As a result, CoA gives strong spatial generalization capabilities while preserving the flexibility and simplicity of a visuomotor policy. Empirically, we observe that CoA outperforms representative imitation learning algorithms such as ACT and Diffusion Policy across 60 RLBench tasks and 8 real-world tasks.
Wenbo Zhang 0009, Tianrun Hu, Hanbo Zhang, Yanyuan Qiao, Yuchu Qin, Yang Li 0184, Jiajun Liu 0004, Tao Kong, Lingqiao Liu
NeurIPS9
2025 Training-Free Instance-Aware 3D Scene Reconstruction and Diffusion-Based View Synthesis from Sparse Images
abstract
We introduce a novel, training-free system for reconstructing, understanding, and rendering 3D indoor scenes from a sparse set of unposed RGB images. Unlike traditional radiance field approaches that require dense views and per-scene optimization, our pipeline achieves high-fidelity results without any training or pose preprocessing. The system integrates three key innovations: (1) A robust point cloud reconstruction module that filters unreliable geometry using a warping-based anomaly removal strategy; (2) A warping-guided 2D-to-3D instance lifting mechanism that propagates 2D segmentation masks into a consistent, instance-aware 3D representation; and (3) A novel rendering approach that projects the point cloud into new views and refines the renderings with a 3D-aware diffusion model. Our method leverages the generative power of diffusion to compensate for missing geometry and enhances realism, especially under sparse input conditions. We further demonstrate that object-level scene editing—such as instance removal—can be naturally supported in our pipeline by modifying only the point cloud, enabling the synthesis of consistent, edited views without retraining. Our results establish a new direction for efficient, editable 3D content generation without relying on scene-specific optimization.
Jiatong Xia, Lingqiao Liu
SIGGRAPH Asia2
2025 A Simple-but-Effective Baseline for Training-Free Class-Agnostic Counting
abstract
Class-Agnostic Counting (CAC) seeks to accurately count objects in a given image with only a few reference examples. While previous methods achieving this relied on additional training, recent efforts have shown that it's possible to accomplish this without training by utilizing pre-existing foundation models, particularly the Segment Any-thing Model (SAM), for counting via instance-level segmentation. Although promising, current training-free methods still lag behind their training-based counterparts in terms of performance. In this research, we present a straightforward training-free solution that effectively bridges this performance gap, serving as a strong baseline. The primary contribution of our work lies in the discovery of four key technologies that can enhance performance. Specifically, we suggest employing a superpixel algorithm to generate more precise initial point prompts, utilizing an image encoder with richer semantic knowledge to replace the SAM encoder for representing candidate objects, and adopting a multiscale mechanism and a transductive prototype scheme to update the representation of reference examples. By combining these four technologies, our approach achieves significant improvements over existing training-free methods and delivers performance on par with training-based ones.
Lingqiao Liu, Qinfeng Shi
WACV3
2025 Learning Dual-Stream Conditional Concepts in Compositional Zero-Shot Learning
abstract
Compositional Zero-Shot Learning (CZSL) aims to recognize unseen compositional concepts composed of seen single concepts. One of the problems of CZSL is to model attributes interacting with objects and objects interacting with attributes. In this work, we focus on this problem and propose Dual-Stream Conditional Network (DSCNet) that learns dual-stream conditional concepts as a solution, where the conditional visual and semantic embeddings of attributes and objects are learned. First, we argue that the condition of the attribute or object is supposed to contain the recognized object and input image, or the recognized attribute and input image. Next, for each concept which can either be an attribute or object, in the semantic stream, we propose to encode the recognized object or attribute semantic features and the input image visual features as the encoded condition, which is then injected into all concept semantic embeddings by a semantic cross encoder to acquire conditional semantic embeddings. In the visual stream, the conditional attribute or object visual embeddings are acquired by injecting the semantic features of the recognized object or attribute into the mapped attribute or object visual features. Experimental results on CZSL benchmarks demonstrate the superiority of our proposed method.
Qingsheng Wang, Lingqiao Liu, Chenchen Jing, Peng Wang 0015, Yanning Zhang 0001, Chunhua Shen
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Generalizable Person Re-Identification From a 3D Perspective: Addressing Unpredictable Viewpoint Changes
abstract
Most existing Domain Generalizable Person Re-identification (DG-ReID) methods focus on addressing style disparities between domains but often overlook the impact of unpredictable camera view changes, which we have identified as a significant factor responsible for poor generalization performance. To address this issue, we propose a novel approach from a 3D perspective, utilizing a customized 2D-to-3D reconstruction model to convert images captured from arbitrary camera views into canonical view images. However, merely applying a 3D reconstruction model in isolation may not result in improved DG-ReID performance, as reconstruction quality can be influenced by multiple factors, such as insufficient image resolution, extreme viewpoint, and environmental variations. These factors may lead to error accumulation and the loss of critical discriminative clues in the reconstructed results. To address this difficulty, we propose fusing the canonical view image with the original image using a transformer-based module. The transformer’s cross-attention mechanism is ideal for aligning and fusing the key semantic clues of the original image with the canonical view image, compensating for reconstruction errors. We demonstrate the effectiveness of our method through extensive experiments in various evaluation settings, achieving superior DG-ReID performance compared to existing approaches. Our approach addresses the impact of unpredictable camera view changes and provides a new perspective for designing DG-ReID methods.
Bingliang Jiao, Lingqiao Liu, Liying Gao, Dapeng Oliver Wu, Guosheng Lin, Peng Wang 0015, Yanning Zhang 0001
IEEE Trans. Inf. Forensics Secur.2
2025 CoLeCLIP: Open-Domain Continual Learning via Joint Task Prompt and Vocabulary Learning
abstract
This article investigates the problem of continual learning (CL) of vision-language models (VLMs) in open domains, where models are required to perform continual updating and inference on a stream of datasets from diverse seen and unseen domains with novel classes. Such a capability is crucial for various applications in open environments, e.g., AI assistants, autonomous driving systems, and robotics. Current CL studies mostly focus on closed-set scenarios in a single domain with known classes. Large pretrained VLMs such as CLIP have showcased exceptional zero-shot recognition capabilities, and several recent studies have leveraged the unique characteristics of VLMs to mitigate catastrophic forgetting in CL. However, they primarily focus on closed-set CL in a single-domain dataset. Open-domain CL of large VLMs is significantly more challenging due to 1) large class correlations and domain gaps across the datasets and 2) the forgetting of zero-shot knowledge in the pretrained VLMs and the knowledge learned from the newly adapted datasets. In this work, we introduce a novel approach, termed CoLeCLIP, which learns an open-domain CL model based on CLIP. It addresses these challenges through joint learning of a set of task prompts and a cross-domain class vocabulary. Extensive experiments on 11 domain datasets show that CoLeCLIP achieves new state-of-the-art performance for open-domain CL under both task- and class-incremental learning (CIL) settings.
Guansong Pang, Wei Suo, Chenchen Jing, Yuling Xi, Lingqiao Liu, Hao Chen 0041, Guoqiang Liang 0001, Peng Wang 0015
IEEE Trans. Neural Networks Learn. Syst.6
2024 Semantic Role Labeling Guided Out-of-distribution Detection
abstract
Identifying unexpected domain-shifted instances in natural language processing is crucial in real-world applications. Previous works identify the out-of-distribution (OOD) instance by leveraging a single global feature embedding to represent the sentence, which cannot characterize subtle OOD patterns well. Another major challenge current OOD methods face is learning effective low-dimensional sentence representations to identify the hard OOD instances that are semantically similar to the in-distribution (ID) data. In this paper, we propose a new unsupervised OOD detection method, namely Semantic Role Labeling Guided Out-of-distribution Detection (SRLOOD), that separates, extracts, and learns the semantic role labeling (SRL) guided fine-grained local feature representations from different arguments of a sentence and the global feature representations of the full sentence using a margin-based contrastive loss. A novel self-supervised approach is also introduced to enhance such global-local feature learning by predicting the SRL extracted role. The resulting model achieves SOTA performance on four OOD benchmarks, indicating the effectiveness of our approach. The code is publicly accessible via https://github.com/cytai/SRLOOD.
Jinan Zou, Maihao Guo, Yu Tian 0001, Haiyao Cao, Lingqiao Liu, Ehsan Abbasnejad, Qinfeng Shi
LREC/COLING6
2024 Decomposing Disease Descriptions for Enhanced Pathology Detection: A Multi-Aspect Vision-Language Pre-Training Framework
abstract
Medical vision language pre-training (VLP) has emerged as a frontier of research, enabling zero-shot pathological recognition by comparing the query image with the textual descriptions for each disease. Due to the complex semantics of biomedical texts, current methods struggle to align medical images with key pathological findings in un-structured reports. This leads to the misalignment with the target disease's textual representation. In this paper, we introduce a novel VLP framework designed to dissect disease descriptions into their fundamental aspects, leveraging prior knowledge about the visual manifestations of pathologies. This is achieved by consulting a large language model and medical experts. Integrating a Transformer module, our approach aligns an input image with the diverse elements of a disease, generating aspect-centric image representations. By consolidating the matches from each aspect, we improve the compatibility between an image and its associated disease. Additionally, capitalizing on the aspect-oriented representations, we present a dual-head Transformer tailored to process known and unknown diseases, optimizing the comprehensive detection efficacy. Conducting experiments on seven downstream datasets, ours improves the accuracy of recent methods by up to 8.56% and 17.26% for seen and unseen categories, respectively. Our code is released at https://github.com/HieuPhan33/MAVL.
Vu Minh Hieu Phan, Yutong Xie 0001, Yuankai Qi, Lingqiao Liu, Liyang Liu, Bowen Zhang 0009, Zhibin Liao, Qi Wu 0001, Minh-Son To, Johan Verjans
CVPR4
2024 Unlocking the Potential of Pre-Trained Vision Transformers for Few-Shot Semantic Segmentation through Relationship Descriptors
abstract
The recent advent of pre-trained vision transformers has unveiled a promising property: their inherent capability to group semantically related visual concepts. In this paper, we explore to harnesses this emergent feature to tackle few-shot semantic segmentation, a task focused on classifying pixels in a test image with a few example data. A critical hurdle in this endeavor is preventing overfitting to the limited classes seen during training the few-shot segmentation model. As our main discovery, we find that the concept of “relationship descriptors”, initially conceived for enhancing the CLIP model for zero-shot semantic segmentation, offers a potential solution. We adapt and refine this concept to craft a relationship descriptor construction tailored for few-shot semantic segmentation, extending its application across multiple layers to enhance performance. Building upon this adaptation, we proposed a few-shot semantic segmentation framework that is not only easy to implement and train but also effectively scales with the number of support examples and categories. Through rigorous experimentation across various datasets, including PASCAL-5iand COCO-20i, we demonstrate a clear advantage of our method in diverse few-shot semantic segmentation scenarios, and a range of pre-trained vision transformer models. The findings clearly show that our method significantly outperforms current state-of-the-art techniques, highlighting the effectiveness of harnessing the emerging capabilities of vision transformers for few-shot semantic segmentation. We release the code at https://github.com/ZiqinZhou66/FewSegwithRD.git.
Ziqin Zhou, Yangyang Shu, Lingqiao Liu
CVPR4
2024 On Learning Discriminative Features from Synthesized Data for Self-supervised Fine-Grained Visual Recognition
Zihu Wang, Lingqiao Liu, Scott Ricardo Figueroa Weston, Samuel Tian, Peng Li 0001
ECCV (89)2
2024 I Learn Better If You Speak My Language: Understanding the Superior Performance of Fine-Tuning Large Language Models with LLM-Generated Responses
abstract
This paper explores an intriguing observation: fine-tuning a large language model (LLM) with responses generated by a LLM often yields better results than using responses generated by humans, particularly in reasoning tasks.We conduct an in-depth investigation to understand why this occurs.Contrary to the common belief that these instances is due to the more detailed nature of LLM-generated content, our study identifies another contributing factor: an LLM is inherently more "familiar" with LLM generated responses.This familiarity is evidenced by lower perplexity before fine-tuning.We design a series of experiments to understand the impact of the "familiarity" and our conclusion reveals that this "familiarity" significantly impacts learning performance.Training with LLM-generated responses not only enhances performance but also helps maintain the model's capabilities in other reasoning tasks after fine-tuning on a specific task.Our code and data are public at https://github.com/XuanRen4470/I-Learn- Better-If-You-Speak-My-Language.
Xuan Ren, Biao Wu 0006, Lingqiao Liu
EMNLP3
2024 KARGEN: Knowledge-Enhanced Automated Radiology Report Generation Using Large Language Models
Yingshu Li 0002, Zhanyu Wang, Yunyi Liu, Lei Wang 0001, Lingqiao Liu, Luping Zhou
MICCAI (5)5
2024 MRScore: Evaluating Medical Report with LLM-Based Reward System
Yunyi Liu, Zhanyu Wang, Yingshu Li 0002, Lingqiao Liu, Lei Wang 0001, Luping Zhou
MICCAI (3)5
2024 LFME: A Simple Framework for Learning from Multiple Experts in Domain Generalization
abstract
Domain generalization (DG) methods aim to maintain good performance in an unseen target domain by using training data from multiple source domains. While success on certain occasions are observed, enhancing the baseline across most scenarios remains challenging. This work introduces a simple yet effective framework, dubbed learning from multiple experts (LFME), that aims to make the target model an expert in all source domains to improve DG. Specifically, besides learning the target model used in inference, LFME will also train multiple experts specialized in different domains, whose output probabilities provide professional guidance by simply regularizing the logit of the target model. Delving deep into the framework, we reveal that the introduced logit regularization term implicitly provides effects of enabling the target model to harness more information, and mining hard samples from the experts during training. Extensive experiments on benchmarks from different DG tasks demonstrate that LFME is consistently beneficial to the baseline and can achieve comparable performance to existing arts. Code is available at https://github.com/liangchen527/LFME.
Liang Chen 0030, Yong Zhang 0034, Yibing Song, Lingqiao Liu
NeurIPS5
2024 A Causal Inspired Early-Branching Structure for Domain Generalization
Liang Chen 0030, Yong Zhang 0034, Yibing Song, Zhen Zhang 0008, Lingqiao Liu
Int. J. Comput. Vis.5
2024 CAT: A Simple yet Effective Cross-Attention Transformer for One-Shot Object Detection
Yuyan Deng, Yang Gao 0001, Ning Wang 0020, Lingqiao Liu, Lei Zhang 0054, Peng Wang 0015
J. Comput. Sci. Technol.5
2024 ReFs: A hybrid pre-training paradigm for 3D medical image segmentation
Yutong Xie 0001, Lingqiao Liu, Hu Wang 0005, Yiwen Ye, Johan Verjans, Yong Xia 0001
Medical Image Anal.3
2024 Semi-Supervised Adversarial Learning for Attribute-Aware Photo Aesthetic Assessment
abstract
Aesthetic attributes are crucial for aesthetics because they explicitly present some photo quality cues that a human expert might use to evaluate a photo’s aesthetic quality. However, annotating aesthetic attributes is a time-consuming, costly, and error-prone task, which leads to the issue that photos available are partially annotated with attributes. To alleviate this issue, we propose a novel semi-supervised adversarial learning method for photo aesthetic assessment from partially attribute-annotated photos, which can greatly reduce the reliance on manual attribute annotation. Specifically, the proposed method consists of a score-attributes generator$R$, a photo generator$G$, and a discriminator$D$. The score-attributes generator learns the aesthetic score and attributes simultaneously to capture their dependencies and construct better feature representations. The photo generator reconstructs the photo by feeding aesthetic attributes, score, and informative feature representation. A discriminator is used to force the convergence of the features-attributes-score tuples generated from the score-attributes generator, the photo generator, and the ground-truth distribution in labeled data for training data. The proposed method significantly outperforms the state of the art, increasing the Spearman rank-order correlation coefficient (SRCC) from the existing best reported of 0.726 to 0.761 onAesthetic and attributes databaseand 0.756 to 0.774 onAesthetic visual analysis database, respectively.
Yangyang Shu, Qian Li 0003, Lingqiao Liu, Guandong Xu
IEEE Trans. Multim.3
2023 Zero-Shot Slot Filling with Slot-Prefix Prompting and Attention Relationship Descriptor
abstract
This paper addresses zero-shot slot filling, which tries to build a system that can generalize to unseen slot types without any training data. The key to zero-shot slot-filling is to match the tokens from the utterance with the semantic definition of the slot without training data in the target domain. This paper tackles this problem by devising a scheme to fully leverage pre-trained language models (PLMs). To this end, we propose a new prompting scheme that utilizes both learnable tokens and slot names to guide the model to focus on the relevant text spans for a given slot. Furthermore, we use attention values between tokens to form a feature descriptor for each token, which is motivated by the fact that the attention value in a PLM naturally characterizes various relationships, e.g., syntactic or semantic, between tokens. By further consolidating those features with an additional transformer-based aggregation module, we create a simple-but-effective zero-shot slot filling system that can achieve significantly better performance than the previous methods, as demonstrated by our experimental studies.
Qiaoyang Luo, Lingqiao Liu
AAAI2
2023 Improved Test-Time Adaptation for Domain Generalization
abstract
The main challenge in domain generalization (DG) is to handle the distribution shift problem that lies between the training and test data. Recent studies suggest that test-time training (TTT), which adapts the learned model with test data, might be a promising solution to the problem. Generally, a TTT strategy hinges its performance on two main factors: selecting an appropriate auxiliary TTT task for updating and identifying reliable parameters to update during the test phase. Both previous arts and our experiments indicate that TTT may not improve but be detrimental to the learned model if those two factors are not properly considered. This work addresses those two factors by proposing an Improved Test-Time Adaptation (ITTA) method. First, instead of heuristically defining an auxiliary objective, we propose a learnable consistency loss for the TTT task, which contains learnable parameters that can be adjusted toward better alignment between our TTT task and the main prediction task. Second, we introduce additional adaptive parameters for the trained model, and we suggest only updating the adaptive parameters during the test phase. Through extensive experiments, we show that the proposed two strategies are beneficial for the learned model (see Figure 1), and ITTA could achieve superior performance to the current state-of-the-art methods on several DG benchmarks. Code is available at https://github.com/liangchen527/ITTA.
Liang Chen 0030, Yong Zhang 0034, Yibing Song, Ying Shan, Lingqiao Liu
CVPR5
2023 Learning Common Rationale to Improve Self-Supervised Representation for Fine-Grained Visual Recognition Problems
abstract
Self-supervised learning (SSL) strategies have demonstrated remarkable performance in various recognition tasks. However, both our preliminary investigation and recent studies suggest that they may be less effective in learning representations for fine-grained visual recognition (FGVR) since many features helpful for optimizing SSL objectives are not suitable for characterizing the subtle differences in FGVR. To overcome this issue, we propose learning an additional screening mechanism to identify discriminative clues commonly seen across instances and classes, dubbed as common rationales in this paper. Intuitively, common rationales tend to correspond to the discriminative patterns from the key parts of foreground objects. We show that a common rationale detector can be learned by simply exploiting the GradCAM induced from the SSL objective without using any pre-trained object parts or saliency detectors, making it seamlessly to be integrated with the existing SSL process. Specifically, we fit the GradCAM with a branch with limited fitting capacity, which allows the branch to capture the common rationales and discard the less common discriminative patterns. At the test stage, the branch generates a set of spatial weights to selectively aggregate features representing an instance. Extensive experimental results on four visual tasks demonstrate that the proposed method can lead to a significant improvement in different evaluation settings.11The source code will be publicly available at:https://github.com/GANPerf/LCR
Yangyang Shu, Anton van den Hengel, Lingqiao Liu
CVPR3
2023 Learning Conditional Attributes for Compositional Zero-Shot Learning
abstract
Compositional Zero-Shot Learning (CZSL) aims to train models to recognize novel compositional concepts based on learned concepts such as attribute-object combinations. One of the challenges is to model attributes interacted with different objects, e.g., the attribute “wet” in “wet apple” and “wet cat” is different. As a solution, we provide analysis and argue that attributes are conditioned on the recognized object and input image and explore learning conditional attribute embeddings by a proposed attribute learning framework containing an attribute hyper learner and an attribute base learner. By encoding conditional attributes, our model enables to generate flexible attribute embeddings for generalization from seen to unseen compositions. Experiments on CZSL benchmarks, including the more challenging C-GQA dataset, demonstrate better performances compared with other state-of-the-art approaches and validate the importance of learning conditional attributes. Code‡1Gllee:https://gitee.com/wqshmzh/canet-czsl is available at https://github.com/wqshmzh/CANet-CZSL.
Qingsheng Wang, Lingqiao Liu, Chenchen Jing, Hao Chen 0041, Guoqiang Liang 0001, Peng Wang 0015, Chunhua Shen
CVPR2
2023 METransformer: Radiology Report Generation by Transformer with Multiple Learnable Expert Tokens
abstract
In clinical scenarios, multi-specialist consultation could significantly benefit the diagnosis, especially for intricate cases. This inspires us to explore a “multi-expert joint diagnosis” mechanism to upgrade the existing “single expert” framework commonly seen in the current literature. To this end, we propose METransformer, a method to realize this idea with a transformer-based backbone. The key design of our method is the introduction of multiple learnable “expert” tokens into both the transformer encoder and decoder. In the encoder, each expert token interacts with both vision tokens and other expert tokens to learn to attend different image regions for image representation. These expert tokens are encouraged to capture complementary information by an orthogonal loss that minimizes their over-lap. In the decoder, each attended expert token guides the cross-attention between input words and visual tokens, thus influencing the generated report. A metrics-based expert voting strategy is further developed to generate the final report. By the multi-experts concept, our model enjoys the merits of an ensemble-based approach but through a manner that is computationally more efficient and supports more sophisticated interactions among experts. Experimental results demonstrate the promising performance of our proposed model on two widely used benchmarks. Last but not least, the framework-level innovation makes our work ready to incorporate advances on existing “single-expert” models to further improve its performance.
Zhanyu Wang, Lingqiao Liu, Lei Wang 0001, Luping Zhou
CVPR2
2023 ZegCLIP: Towards Adapting CLIP for Zero-shot Semantic Segmentation
abstract
Recently, CLIP has been applied to pixel-level zero-shot learning tasks via a two-stage scheme. The general idea is to first generate class-agnostic region proposals and then feed the cropped proposal regions to CLIP to utilize its image-level zero-shot classification capability. While effective, such a scheme requires two image encoders, one for proposal generation and one for CLIP, leading to a complicated pipeline and high computational cost. In this work, we pursue a simpler-and-efficient one-stage solution that directly extends CLIP's zero-shot prediction capability from image to pixel level. Our investigation starts with a straightforward extension as our baseline that generates semantic masks by comparing the similarity between text and patch embeddings extracted from CLIP. However, such a paradigm could heavily overfit the seen classes and fail to generalize to unseen classes. To handle this issue, we propose three simple-but-effective designs and figure out that they can significantly retain the inherent zero-shot capacity of CLIP and improve pixel-level generalization ability. Incorporating those modifications leads to an efficient zero-shot semantic segmentation system called ZegCLIP. Through extensive experiments on three public benchmarks, ZegCLIP demonstrates superior performance, outperforming the state-of-the-art methods by a large margin under both “inductive” and “transductive” zero-shot settings. In addition, compared with the two-stage method, our one-stage ZegCLIP achieves a speedup of about 5 times faster during inference. We release the code at https://github.com/ZiqinZhou66/ZegCLIP.git.
Ziqin Zhou, Yinjie Lei, Bowen Zhang 0009, Lingqiao Liu, Yifan Liu 0001
CVPR4
2023 Out-of-Distribution Generalization in Natural Language Processing: Past, Present, and Future
abstract
Linyi Yang, Yaoxian Song, Xuan Ren, Chenyang Lyu, Yidong Wang, Jingming Zhuo, Lingqiao Liu, Jindong Wang, Jennifer Foster, Yue Zhang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Linyi Yang, Yaoxian Song, Xuan Ren, Chenyang Lyu, Jingming Zhuo, Lingqiao Liu, Jindong Wang 0001, Jennifer Foster, Yue Zhang 0004
EMNLP7
2023 Domain Generalization via Rationale Invariance
abstract
This paper offers a new perspective to ease the challenge of domain generalization, which involves maintaining robust results even in unseen environments. Our design focuses on the decision-making process in the final classifier layer. Specifically, we propose treating the element-wise contributions to the final results as the rationale for making a decision and representing the rationale for each sample as a matrix. For a well-generalized model, we suggest the rationale matrices for samples belonging to the same category should be similar, indicating the model relies on domain-invariant clues to make decisions, thereby ensuring robust results. To implement this idea, we introduce a rationale invariance loss as a simple regularization technique, requiring only a few lines of code. Our experiments demonstrate that the proposed approach achieves competitive results across various datasets, despite its simplicity. Code is available at https://github.com/liangchen527/RIDG.
Liang Chen 0030, Yong Zhang 0034, Yibing Song, Anton van den Hengel, Lingqiao Liu
ICCV5
2023 Transformer-based Open-world Instance Segmentation with Cross-task Consistency Regularization
abstract
Open-World Instance Segmentation (OWIS) is an emerging research topic that aims to segment class-agnostic object instances from images. The mainstream approaches use a two-stage segmentation framework, which first locates the candidate object bounding boxes and then performs instance segmentation. In this work, we instead promote a single-stage transformer-based framework for OWIS. We argue that the end-to-end training process in the single-stage framework can be more convenient for directly regularizing the localization of class-agnostic object pixels. Based on the transformer-based instance segmentation framework, we propose a regularization model to predict foreground pixels and use its relation to instance segmentation to construct a cross-task consistency loss. We show that such a consistency loss could alleviate the problem of incomplete instance annotation - a common problem in the existing OWIS datasets. We also show that the proposed loss lends itself to an effective solution to semi-supervised OWIS that could be considered an extreme case that all object annotations are absent for some images. Our extensive experiments demonstrate that the proposed method achieves impressive results in both fully-supervised and semi-supervised settings. Compared to SOTA methods, the proposed method significantly improves the AP_100 score by 4.75% in UVO dataset →UVO dataset setting and 4.05% in COCO dataset →UVO dataset setting.
Xizhe Xue, Dongdong Yu, Lingqiao Liu, Yu Liu 0015, Satoshi Tsutsui, Ying Li 0017, Zehuan Yuan, Zheng Shou 0001
ACM Multimedia3
2023 The CLIP Model is Secretly an Image-to-Prompt Converter
abstract
The Stable Diffusion model is a prominent text-to-image generation model that relies on a text prompt as its input, which is encoded using the Contrastive Language-Image Pre-Training (CLIP). However, text prompts have limitations when it comes to incorporating implicit information from reference images. Existing methods have attempted to address this limitation by employing expensive training procedures involving millions of training samples for image-to-image generation. In contrast, this paper demonstrates that the CLIP model, as utilized in Stable Diffusion, inherently possesses the ability to instantaneously convert images into text prompts. Such an image-to-prompt conversion can be achieved by utilizing a linear projection matrix that is calculated in a closed form. Moreover, the paper showcases that this capability can be further enhanced by either utilizing a small amount of similar-domain training data (approximately 100 images) or incorporating several online training steps (around 30 iterations) on the reference images. By leveraging these approaches, the proposed method offers a simple and flexible solution to bridge the gap between images and text prompts. This methodology can be applied to various tasks such as image variation and image editing, facilitating more effective and seamless interaction between images and textual prompts.
Chunna Tian, Haoxuan Ding, Lingqiao Liu
NeurIPS4
2023 Toward Re-Identifying Any Animal
abstract
The current state of re-identification (ReID) models poses limitations to their applicability in the open world, as they are primarily designed and trained for specific categories like person or vehicle. In light of the importance of ReID technology for tracking wildlife populations and migration patterns, we propose a new task called ``Re-identify Any Animal in the Wild'' (ReID-AW). This task aims to develop a ReID model capable of handling any unseen wildlife category it encounters. To address this challenge, we have created a comprehensive dataset called Wildlife-71, which includes ReID data from 71 different wildlife categories. This dataset is the first of its kind to encompass multiple object categories in the realm of ReID. Furthermore, we have developed a universal re-identification model named UniReID specifically for the ReID-AW task. To enhance the model's adaptability to the target category, we employ a dynamic prompting mechanism using category-specific visual prompts. These prompts are generated based on knowledge gained from a set of pre-selected images within the target category. Additionally, we leverage explicit semantic knowledge derived from the large-scale pre-trained language model, GPT-4. This allows UniReID to focus on regions that are particularly useful for distinguishing individuals within the target category. Extensive experiments have demonstrated the remarkable generalization capability of our UniReID model. It showcases promising performance in handling arbitrary wildlife categories, offering significant advancements in the field of ReID for wildlife conservation and research purposes.
Bingliang Jiao, Lingqiao Liu, Liying Gao, Ruiqi Wu 0001, Guosheng Lin, Peng Wang 0015, Yanning Zhang 0001
NeurIPS2
2023 Balanced image captioning with task-aware decoupled learning and fusion
Lingqiao Liu, Chunna Tian, Xiangnan Zhang, Xilan Tian
Neurocomputing2
2023 Learning Resolution-Adaptive Representations for Cross-Resolution Person Re-Identification
abstract
Cross-resolution person re-identification (CRReID) is a challenging and practical problem that involves matching low-resolution (LR) query identity images against high-resolution (HR) gallery images. Query images often suffer from resolution degradation due to the different capturing conditions from real-world cameras. State-of-the-art solutions for CRReID either learn a resolution-invariant representation or adopt a super-resolution (SR) module to recover the missing information from the LR query. In this paper, we propose an alternative SR-free paradigm to directly compare HR and LR images via a dynamic metric that is adaptive to the resolution of a query image. We realize this idea by learning resolution-adaptive representations for cross-resolution comparison. We propose two resolution-adaptive mechanisms to achieve this. The first mechanism encodes the resolution specifics into different subvectors in the penultimate layer of the deep neural network, creating a varying-length representation. To better extract resolution-dependent information, we further propose to learn resolution-adaptive masks for intermediate residual feature blocks. A novel progressive learning strategy is proposed to train those masks properly. These two mechanisms are combined to boost the performance of CRReID. Experimental results show that the proposed method outperforms existing approaches and achieves state-of-the-art performance on multiple CRReID benchmarks.
Lin Wu 0001, Lingqiao Liu, Yang Wang 0023, Zheng Zhang 0006, Farid Boussaïd, Mohammed Bennamoun, Xianghua Xie
IEEE Trans. Image Process.2
2023 Verbal-Person Nets: Pose-Guided Multi-Granularity Language-to-Person Generation
abstract
Person image generation conditioned on natural language allows us to personalize image editing in a user-friendly manner. This fashion, however, involves different granularities of semantic relevance between texts and visual content. Given a sentence describing an unknown person, we propose a novel pose-guided multi-granularity attention architecture to synthesize the person image in an end-to-end manner. To determine what content to draw at a global outline, the sentence-level description and pose feature maps are incorporated into a U-Net architecture to generate a coarse person image. To further enhance the fine-grained details, we propose to draw the human body parts with highly correlated textual nouns and determine the spatial positions with respect to target pose points. Our model is premised on a conditional generative adversarial network (GAN) that translates language description into a realistic person image. The proposed model is coupled with two-stream discriminators: 1) text-relevant local discriminators to improve the fine-grained appearance by identifying the region-text correspondences at the finer manipulation and 2) a global full-body discriminator to regulate the generation via a pose-weighting feature selection. Extensive experiments conducted on benchmarks validate the superiority of our method for person image generation.
Deyin Liu, Lin Wu 0001, Feng Zheng 0001, Lingqiao Liu, Meng Wang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2022 Dual Decision Improves Open-Set Panoptic Segmentation
Hao Chen 0041, Lingqiao Liu, Yufei Yin
BMVC3
2022 Self-supervised Learning of Adversarial Example: Towards Good Generalizations for Deepfake Detection
abstract
Recent studies in deepfake detection have yielded promising results when the training and testing face forgeries are from the same dataset. However, the problem remains challenging when one tries to generalize the detector to forgeries created by unseen methods in the training dataset. This work addresses the generalizable deepfake detection from a simple principle: a generalizable representation should be sensitive to diverse types of forgeries. Following this principle, we propose to enrich the “diversity” of forgeries by synthesizing augmented forgeries with a pool of forgery configurations and strengthen the “sensitivity” to the forgeries by enforcing the model to predict the forgery configurations. To effectively explore the large forgery augmentation space, we further propose to use the adversarial training strategy to dynamically synthesize the most challenging forgeries to the current model. Through extensive experiments, we show that the proposed strategies are surprisingly effective (see Figure 1), and they could achieve superior performance than the current state-of-the-art methods. Code is available at https://github.com/liangchen527/SLADD.
Liang Chen 0030, Yong Zhang 0034, Yibing Song, Lingqiao Liu, Jue Wang 0001
CVPR4
2022 Dynamically Transformed Instance Normalization Network for Generalizable Person Re-Identification
Bingliang Jiao, Lingqiao Liu, Liying Gao, Guosheng Lin, Lu Yang 0016, Shizhou Zhang, Peng Wang 0015, Yanning Zhang 0001
ECCV (14)2
2022 Improving Fine-Grained Visual Recognition in Low Data Regimes via Self-boosting Attention Mechanism
Yangyang Shu, Baosheng Yu, Lingqiao Liu
ECCV (25)4
2022 Pluggable Weakly-Supervised Cross-View Learning for Accurate Vehicle Re-Identification
abstract
Learning cross-view consistent feature representation is the key for accurate vehicle Re-identification (ReID), since the visual appearance of vehicles changes significantly under different viewpoints. To this end, many existing approaches resort to the supervised cross-view learning using extensive extra viewpoints annotations, which however, is difficult to deploy in real applications due to the expensive labelling cost and the continous viewpoint variation that makes it hard to define discrete viewpoint labels. In this study, we present a pluggable Weakly-supervised Cross-View Learning (WCVL) module for vehicle ReID. Through hallucinating the cross-view samples as the hardest positive counterparts with small luminance difference and large local feature variance, we can learn the consistent feature representation via minimizing the cross-view feature distance based on vehicle IDs only without using any viewpoint annotation. More importantly, the proposed method can be seamlessly plugged into most existing vehicle ReID baselines for cross-view learning without re-training the baselines. To demonstrate its efficacy, we plug the proposed method into a bunch of off-the-shelf baselines and obtain significant performance improvement on four public benchmark datasets, i.e., VeRi-776, VehicleID, VRIC and VRAI.
Lu Yang 0016, Hongbang Liu, Lingqiao Liu, Jinghao Zhou, Lei Zhang 0054, Peng Wang 0015, Yanning Zhang 0001
ICMR3
2022 Progressive Class Semantic Matching for Semi-supervised Text Classification
abstract
Semi-supervised learning is a promising way to reduce the annotation cost for textclassification.Combining with pre-trained language models (PLMs), e.g., BERT, recent semi-supervised learning methods achieved impressive performance.In this work, we further investigate the marriage between semisupervised learning and a pre-trained language model.Unlike existing approaches that utilize PLMs only for model parameter initialization, we explore the inherent topic matching capability inside PLMs for building a more powerful semi-supervised learning approach.Specifically, we propose a joint semi-supervised learning process that can progressively build a standard K-way classifier and a matching network for the input text and the Class Semantic Representation (CSR).The CSR will be initialized from the given labeled sentences and progressively updated through the training process.By means of extensive experiments, we show that our method can not only bring remarkable improvement to baselines, but also overall be more stable, and achieves state-ofthe-art performance in semi-supervised text classification.
Lingqiao Liu, Ehsan Abbasnejad
NAACL-HLT2
2022 OST: Improving Generalization of DeepFake Detection via One-Shot Test-Time Training
abstract
State-of-the-art deepfake detectors perform well in identifying forgeries when they are evaluated on a test set similar to the training set, but struggle to maintain good performance when the test forgeries exhibit different characteristics from the training images e.g., forgeries are created by unseen deepfake methods. Such a weak generalization capability hinders the applicability of deepfake detectors. In this paper, we introduce a new learning paradigm specially designed for the generalizable deepfake detection task. Our key idea is to construct a test-sample-specific auxiliary task to update the model before applying it to the sample. Specifically, we synthesize pseudo-training samples from each test image and create a test-time training objective to update the model. Moreover, we proposed to leverage meta-learning to ensure that a fast single-step test-time gradient descent, dubbed one-shot test-time training (OST), can be sufficient for good deepfake detection performance. Extensive results across several benchmark datasets demonstrate that our approach performs favorably against existing arts in terms of generalization to unseen data and robustness to different post-processing steps.
Liang Chen 0030, Yong Zhang 0034, Yibing Song, Jue Wang 0001, Lingqiao Liu
NeurIPS5
2022 Semi-supervised Semantic Segmentation with Prototype-based Consistency Regularization
abstract
Semi-supervised semantic segmentation requires the model to effectively propagate the label information from limited annotated images to unlabeled ones. A challenge for such a per-pixel prediction task is the large intra-class variation, i.e., regions belonging to the same class may exhibit a very different appearance even in the same picture. This diversity will make the label propagation hard from pixels to pixels. To address this problem, we propose a novel approach to regularize the distribution of within-class features to ease label propagation difficulty. Specifically, our approach encourages the consistency between the prediction from a linear predictor and the output from a prototype-based predictor, which implicitly encourages features from the same pseudo-class to be close to at least one within-class prototype while staying far from the other between-class prototypes. By further incorporating CutMix operations and a carefully-designed prototype maintenance strategy, we create a semi-supervised semantic segmentation algorithm that demonstrates superior performance over the state-of-the-art methods from extensive experimental evaluation on both Pascal VOC and Cityscapes benchmarks.
Lingqiao Liu, Qiuchen Bian
NeurIPS2
2022 RPC: a large-scale and fine-grained retail product checkout dataset
Xiu-Shen Wei, Quan Cui, Lei Yang 0048, Peng Wang 0023, Lingqiao Liu, Jian Yang 0003
Sci. China Inf. Sci.5
2022 Structured Binary Neural Networks for Image Recognition
Bohan Zhuang, Chunhua Shen, Mingkui Tan, Peng Chen 0037, Lingqiao Liu, Ian D. Reid 0001
Int. J. Comput. Vis.5
2022 Effective Training of Convolutional Neural Networks With Low-Bitwidth Weights and Activations
abstract
This paper tackles the problem of training a deep convolutional neural network of both low-bitwidth weights and activations. Optimizing a low-precision network is very challenging due to the non-differentiability of the quantizer, which may result in substantial accuracy loss. To address this, we propose three practical approaches, including (i) progressive quantization; (ii) stochastic precision; and (iii) joint knowledge distillation to improve the network training. First, for progressive quantization, we propose two schemes to progressively find good local minima. Specifically, we propose to first optimize a network with quantized weights and subsequently quantize activations. This is in contrast to the traditional methods which optimize them simultaneously. Furthermore, we propose a second progressive quantization scheme which gradually decreases the bitwidth from high-precision to low-precision during training. Second, to alleviate the excessive training burden due to the multi-round training stages, we further propose a one-stage stochastic precision strategy to randomly sample and quantize sub-networks while keeping other parts in full-precision. Finally, we adopt a novel learning scheme to jointly train a full-precision model alongside the low-precision one. By doing so, the full-precision model provides hints to guide the low-precision model training and significantly improves the performance of the low-precision network. Extensive experiments on various datasets (e.g., CIFAR-100, ImageNet) show the effectiveness of the proposed methods.
Bohan Zhuang, Mingkui Tan, Jing Liu 0048, Lingqiao Liu, Ian D. Reid 0001, Chunhua Shen
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Privileged multi-task learning for attribute-aware aesthetic assessment
Yangyang Shu, Qian Li 0003, Lingqiao Liu, Guandong Xu
Pattern Recognit.3
2022 Center Prediction Loss for Re-identification
Lu Yang 0016, Yunlong Wang 0008, Lingqiao Liu, Peng Wang 0015, Yanning Zhang 0001
Pattern Recognit.3
2022 Computationally Efficient Dilated Convolutional Model for Melody Extraction
abstract
In this paper we propose a dilated convolutional model for music melody extraction. Taking variable-q transforms (VQTs) as inputs, it first uses consecutive layers of convolution to capture local temporal-frequency patterns, and then a single layer of dilated convolution to capture global frequency patterns contributed by the pitches and harmonics of active notes. Compared with the contrast model without dilation, the proposed model can remarkably cut down the computational cost, and at the same time does not compromise the performance. Its advantages over existing models are two fold. First, it performs best on most datasets, for both general and vocal melody extraction. Second, it can achieve the best performance with least training data.
Lingqiao Liu, Qinfeng Shi
IEEE Signal Process. Lett.2
2022 Feature Encoding With Autoencoders for Weakly Supervised Anomaly Detection
abstract
Weakly supervised anomaly detection aims at learning an anomaly detector from a limited amount of labeled data and abundant unlabeled data. Recent works build deep neural networks for anomaly detection by discriminatively mapping the normal samples and abnormal samples to different regions in the feature space or fitting different distributions. However, due to the limited number of annotated anomaly samples, directly training networks with the discriminative loss may not be sufficient. To overcome this issue, this article proposes a novel strategy to transform the input data into a more meaningful representation that could be used for anomaly detection. Specifically, we leverage an autoencoder to encode the input data and utilize three factors, hidden representation, reconstruction residual vector, and reconstruction error, as the new representation for the input data. This representation amounts to encode a test sample with its projection on the training data manifold, its direction to its projection, and its distance to its projection. In addition to this encoding, we also propose a novel network architecture to seamlessly incorporate those three factors. From our extensive experiments, the benefits of the proposed strategy are clearly demonstrated by its superior performance over the competitive methods. Code is available at: https://github.com/yj-zhou/Feature_Encoding_with_AutoEncoders_for_Weakly-supervised_Anomaly_Detection.
Yingjie Zhou 0001, Xucheng Song, Yanru Zhang, Fanxing Liu, Ce Zhu, Lingqiao Liu
IEEE Trans. Neural Networks Learn. Syst.6
2021 Label Definitions Augmented Interaction Model for Legal Charge Prediction
Liangyi Kang, Jie Liu 0008, Lingqiao Liu, Dan Ye 0004
ECIR (1)3
2021 Contextualize Knowledge Bases with Transformer for End-to-end Task-Oriented Dialogue Systems
abstract
Incorporating knowledge bases (KB) into endto-end task-oriented dialogue systems is challenging, since it requires to properly represent the entity of KB, which is associated with its KB context and dialogue context.The existing works represent the entity with only perceiving a part of its KB context, which can lead to the less effective representation due to the information loss, and adversely favor KB reasoning and response generation.To tackle this issue, we explore to fully contextualize the entity representation by dynamically perceiving all the relevant entities and dialogue history.To achieve this, we propose a COntextaware Memory Enhanced Transformer framework (COMET), which treats the KB as a sequence and leverages a novel Memory Mask to enforce the entity to only focus on its relevant entities and dialogue history, while avoiding the distraction from the irrelevant entities.Through extensive experiments, we show that our COMET framework can achieve superior performance over the state of the arts.
Yanjie Gou, Yinjie Lei, Lingqiao Liu, Yong Dai 0001, Chunxu Shen
EMNLP (1)3
2021 Exploring the auxiliary learning for long-tailed visual recognition
Junjie Zhang 0002, Lingqiao Liu, Peng Wang 0023, Jian Zhang 0002
Neurocomputing2
2021 Semi-supervised emotion recognition in textual conversation via a context-augmented auxiliary training task
Liangyi Kang, Jie Liu 0008, Lingqiao Liu, Dan Ye 0004
Inf. Process. Manag.3
2021 Towards using count-level weak supervision for crowd counting
Yinjie Lei, Yan Liu 0015, Lingqiao Liu
Pattern Recognit.4
2021 Boundarymix: Generating pseudo-training images for improving segmentation with scribble annotations
Wanxuan Lu, Dong Gong, Kun Fu 0001, Xian Sun 0001, Wenhui Diao, Lingqiao Liu
Pattern Recognit.6
2021 Where to Look and How to Describe: Fashion Image Retrieval With an Attentional Heterogeneous Bilinear Network
abstract
Fashion products typically feature in compositions of a variety of styles at different clothing parts. In order to distinguish images of different fashion products, we need to extract both appearance (i.e., “how to describe”) and localization (i.e., “where to look”) information, and their interactions. To this end, we propose a biologically inspired framework for image-based fashion product retrieval, which mimics the hypothesized two-stream visual processing system of human brain. The proposed attentional heterogeneous bilinear network (AHBN) consists of two branches: a deep CNN branch to extract fine-grained appearance attributes and a fully convolutional branch to extract landmark localization information. A joint channel-wise attention mechanism is further applied to the extracted heterogeneous features to focus on important channels, followed by a compact bilinear pooling layer to model the interaction of the two streams. Our proposed framework achieves satisfactory performance on three image-based fashion product retrieval benchmarks.
Haibo Su, Peng Wang 0015, Lingqiao Liu, Hui Li 0031, Zhen Li 0068, Yanning Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2021 Global and Local Texture Randomization for Synthetic-to-Real Semantic Segmentation
abstract
Semantic segmentation is a crucial image understanding task, where each pixel of image is categorized into a corresponding label. Since the pixel-wise labeling for ground-truth is tedious and labor intensive, in practical applications, many works exploit the synthetic images to train the model for real-word image semantic segmentation, i.e., Synthetic-to-Real Semantic Segmentation (SRSS). However, Deep Convolutional Neural Networks (CNNs) trained on the source synthetic data may not generalize well to the target real-world data. To address this problem, there has been rapidly growing interest in Domain Adaption technique to mitigate the domain mismatch between the synthetic and real-world images. Besides, Domain Generalization technique is another solution to handle SRSS. In contrast to Domain Adaption, Domain Generalization seeks to address SRSS without accessing any data of the target domain during training. In this work, we propose two simple yet effective texture randomization mechanisms, Global Texture Randomization (GTR) and Local Texture Randomization (LTR), for Domain Generalization based SRSS. GTR is proposed to randomize the texture of source images into diverse unreal texture styles. It aims to alleviate the reliance of the network on texture while promoting the learning of the domain-invariant cues. In addition, we find the texture difference is not always occurred in entire image and may only appear in some local areas. Therefore, we further propose a LTR mechanism to generate diverse local regions for partially stylizing the source images. Finally, we implement a regularization of Consistency between GTR and LTR (CGL) aiming to harmonize the two proposed mechanisms during training. Extensive experiments on five publicly available datasets (i.e., GTA5, SYNTHIA, Cityscapes, BDDS and Mapillary) with various SRSS settings (i.e., GTA5/SYNTHIA to Cityscapes/BDDS/Mapillary) demonstrate that the proposed method is superior to the state-of-the-art methods for domain generalization based SRSS.
Duo Peng, Yinjie Lei, Lingqiao Liu, Jun Liu 0036
IEEE Trans. Image Process.3
2020 V-PROM: A Benchmark for Visual Reasoning Using Visual Progressive Matrices
abstract
Advances in machine learning have generated increasing enthusiasm for tasks that require high-level reasoning on top of perceptual capabilities, particularly over visual data. Such tasks include, for example, image captioning, visual question answering, and visual navigation. Their evaluation is however hindered by task-specific confounding factors and dataset biases. In parallel, the existing benchmarks for abstract reasoning are limited to synthetic stimuli (e.g. images of simple shapes) and do not capture the challenges of real-world data. We propose a new large-scale benchmark to evaluates abstract reasoning over real visual data. The test involves visual questions that require operations fundamental to many high-level vision tasks, such as comparisons of counts and logical operations on complex visual properties. The benchmark measures a method's ability to infer high-level relationships and to generalise them over image-based concepts. We provide multiple training/test splits that require controlled levels of generalization. We evaluate a range of deep learning architectures, and find that existing models, including those popular for vision-and-language tasks, are unable to solve seemingly-simple instances. Models using relational networks fare better but leave substantial room for improvement.
Damien Teney, Peng Wang 0023, Jiewei Cao, Lingqiao Liu, Chunhua Shen, Anton van den Hengel
AAAI4
2020 Training Quantized Neural Networks With a Full-Precision Auxiliary Module
abstract
In this paper, we seek to tackle a challenge in training low-precision networks: the notorious difficulty in propagating gradient through a low-precision network due to the non-differentiable quantization function. We propose a solution by training the low-precision network with a full-precision auxiliary module. Specifically, during training, we construct a mix-precision network by augmenting the original low-precision network with the full precision auxiliary module. Then the augmented mix-precision network and the low-precision network are jointly optimized. This strategy creates additional full-precision routes to update the parameters of the low-precision model, thus making the gradient back-propagates more easily. At the inference time, we discard the auxiliary module without introducing any computational complexity to the low-precision network. We evaluate the proposed method on image classification and object detection over various quantization approaches and show consistent performance increase. In particular, we achieve near lossless performance to the full-precision model by using a 4-bit detector, which is of great practical value.
Bohan Zhuang, Lingqiao Liu, Mingkui Tan, Chunhua Shen, Ian D. Reid 0001
CVPR2
2020 Semi-supervised Crowd Counting via Self-training on Surrogate Tasks
Yan Liu 0015, Lingqiao Liu, Peng Wang 0023, Yinjie Lei
ECCV (15)2
2020 Harmonic Structure-Based Neural Network Model for Music Pitch Detection
abstract
Pitched music notes are characterized by a harmonic structure. That is, the waveform of each note is composed of a number of sinusoids with different frequencies. Among these frequencies, the lowest is called the pitch or the fundamental frequency of the note, and all the others are integer multiples of the pitch and referred to as the harmonics of the note. In this paper we propose a neural network model for pitch detection that detects the active pitches present in a music recording. The model first uses conventional convolution whose receptive field is continuous to capture local frequency patterns in the spectrogram of the recording, and then uses sparse convolution with discontinuous receptive field to capture global harmonic frequency patterns. When the MAPS dataset is used both for training and for test, the model outperforms all existing models by large margins. When the MAESTRO dataset is used for training and the MAPS dataset is used for test, the model achieves the state-of-the-art performance as well.
Lingqiao Liu, Qinfeng Shi
ICMLA2
2020 Enhancing Piano Transcription by Dilated Convolution
abstract
When detecting music pitch with deep learning, the most challenging issue is how to aggregate the multi-scale frequency information scattered in the spectrogram of a music signal. Traditionally, this aggregation was achieved via strided pooling and fully connected operations, or domain-specific sparse convolutions. In this paper we propose to use more general-purpose dilated convolution to this end. In particular, first we design a dilated convolutional acoustic model for piano pitch detection. This model leads existing acoustic models by large margins. Based on this model, we then design an automatic music transcription (AMT) system. When trained and tested on the MAPS dataset, this system outperforms existing AMT systems again by large margins. When trained and tested on the MAESTRO dataset, this system performs equally well as a state- of-the-art system, but has much better generalization capability.
Lingqiao Liu, Qinfeng Shi
ICMLA2
2020 Meta Learning with Differentiable Closed-form Solver for Fast Video Object Segmentation
abstract
Video object segmentation plays a vital role to many robotic tasks, beyond the satisfied accuracy, quickly adapt to the new scenario with very limited annotations and conduct a quick inference are also important. In this paper, we are specifically concerned with the task of fast segmenting all pixels of a target object in all frames, given the annotation mask in the first frame. Even when such annotation is available, this remains a challenging problem because of the changing appearance and shape of the object over time. In this paper, we tackle this task by formulating it as a meta-learning problem, where the base learner grasping the semantic scene understanding for a general type of objects, and the meta learner quickly adapting the appearance of the target object with a few examples. Our proposed meta-learning method uses a closed form optimizer, the so-called "ridge regression", which has been shown to be conducive for fast and better training convergence. Moreover, we propose a mechanism, named "block splitting", to further speed up the training process as well as to reduce the number of learning parameters. In comparison with the state-of-the art methods, our proposed framework achieves significant boost up in processing speed, while having highly comparable performance compared to the best performing methods on the widely used datasets. Video demo can be found here1.
Yu Liu 0029, Lingqiao Liu, Haokui Zhang, Seyed Hamid Rezatofighi, Qingsen Yan, Ian D. Reid 0001
IROS2
2020 GADE: A Generative Adversarial Approach to Density Estimation and its Applications
Ehsan Abbasnejad, Qinfeng Shi, Anton van den Hengel, Lingqiao Liu
Int. J. Comput. Vis.4
2020 Adaptive Importance Learning for Improving Lightweight Image Super-Resolution Network
Lei Zhang 0054, Peng Wang 0023, Chunhua Shen, Lingqiao Liu, Wei Wei 0008, Yanning Zhang 0001, Anton van den Hengel
Int. J. Comput. Vis.4
2020 A Dynamic Parameter Enhanced Network for distant supervised relation extraction
Yanjie Gou, Yinjie Lei, Lingqiao Liu, Xi Peng 0001
Knowl. Based Syst.3
2020 Adversarial Learning of Structure-Aware Fully Convolutional Networks for Landmark Localization
abstract
Landmark/pose estimation in single monocular images has received much effort in computer vision due to its important applications. It remains a challenging task when input images come with severe occlusions caused by, e.g., adverse camera views. Under such circumstances, biologically implausible pose predictions may be produced. In contrast, human vision is able to predict poses by exploiting geometric constraints of landmark point inter-connectivity. To address the problem, by incorporating priors about the structure of pose components, we propose a novel structure-aware fully convolutional network to implicitly take such priors into account during training of the deep network. Explicit learning of such constraints is typically challenging. Instead, inspired by how human identifies implausible poses, we design discriminators to distinguish the real poses from the fake ones (such as biologically implausible ones). If the pose generator G generates results that the discriminator fails to distinguish from real ones, the network successfully learns the priors. Training of the network follows the strategy of conditional Generative Adversarial Networks (GANs). The effectiveness of the proposed network is evaluated on three pose-related tasks: 2D human pose estimation, 2D facial landmark estimation and 3D human pose estimation. The proposed approach significantly outperforms several state-of-the-art methods and almost always generates plausible pose predictions, demonstrating the usefulness of implicit learning of structures using GANs.
Yu Chen 0037, Chunhua Shen, Hao Chen 0041, Xiu-Shen Wei, Lingqiao Liu, Jian Yang 0003
IEEE Trans. Pattern Anal. Mach. Intell.5
2020 Mask-Aware Networks for Crowd Counting
abstract
Crowd counting problem aims to count the number of objects within an image or a frame in the videos and is usually solved by estimating the density map generated from the object location annotations. The values in the density map, by nature, take two possible states: zero indicating no object around, a non-zero value indicating the existence of objects and the value denoting the local object density. In contrast to traditional methods which do not differentiate the density prediction of these two states, we propose to use a dedicated network branch to predict the object/non-object mask and then combine its prediction with the input image to produce the density map. Our rationale is that the mask prediction could be better modeled as a binary segmentation problem and the difficulty of estimating the density could be reduced if the mask is known. A key to the proposed scheme is the strategy of incorporating the mask prediction into the density map estimator. To this end, we study five possible solutions, and via analysis and experimental validation we identify the most effective one. Through extensive experiments on three public datasets, we demonstrate the superior performance of the proposed approach over the baselines and show that our network could achieve the state-of-the-art performance.
Shengqin Jiang, Xiaobo Lu, Yinjie Lei, Lingqiao Liu
IEEE Trans. Circuits Syst. Video Technol.4
2020 Towards Effective Deep Embedding for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) can be formulated as a cross-domain matching problem: after being projected into a joint embedding space, a visual sample will match against all candidate class-level semantic descriptions and be assigned to the nearest class. In this process, the embedding space underpins the success of such matching and is crucial for ZSL. In this paper, we conduct an in-depth study on the construction of embedding space for ZSL and posit that an ideal embedding space should satisfy two criteria: intra-class compactness and inter-class separability. While the former encourages the embeddings of visual samples of one class to distribute tightly close to the semantic description embedding of this class, the latter requires embeddings from different classes to be well separated from each other. Towards this goal, we present a simple but effective two-branch network to simultaneously map semantic descriptions and visual samples into a joint space, on which visual embeddings are forced to regress to their class-level semantic embeddings and the embeddings crossing classes are required to be distinguishable by a trainable classifier. Furthermore, we extend our method to a transductive setting to better handle the model bias problem in ZSL (i.e., samples from unseen classes tend to be categorized into seen classes) with minimal extra supervision. Specifically, we propose a pseudo labeling strategy to progressively incorporate the testing samples into the training process and thus balance the model between seen and unseen classes. Experimental results on five standard ZSL datasets show the superior performance of the proposed method and its transductive extension.
Lei Zhang 0054, Peng Wang 0023, Lingqiao Liu, Chunhua Shen, Wei Wei 0008, Yanning Zhang 0001, Anton van den Hengel
IEEE Trans. Circuits Syst. Video Technol.3
2019 A Generative Adversarial Density Estimator
abstract
Density estimation is a challenging unsupervised learning problem. Current maximum likelihood approaches for density estimation are either restrictive or incapable of producing high-quality samples. On the other hand, likelihood-free models such as generative adversarial networks, produce sharp samples without a density model. The lack of a density estimate limits the applications to which the sampled data can be put, however. We propose a Generative Adversarial Density Estimator, a density estimation approach that bridges the gap between the two. Allowing for a prior on the parameters of the model, we extend our density estimator to a Bayesian model where we can leverage the predictive variance to measure our confidence in the likelihood. Our experiments on challenging applications such as visual dialog where the density and the confidence in predictions are crucial shows the effectiveness of our approach.
Ehsan Abbasnejad, Qinfeng Shi, Anton van den Hengel, Lingqiao Liu
CVPR4
2019 Structured Binary Neural Networks for Accurate Image Classification and Semantic Segmentation
abstract
In this paper, we propose to train convolutional neural networks (CNNs) with both binarized weights and activations, leading to quantized models specifically for mobile devices with limited power capacity and computation resources. By assuming the same architecture to full-precision networks, previous works on quantizing CNNs seek to preserve the floating-point information using a set of discrete values, which we call value approximation. However, we take a novel ``structure approximation'' view for quantization--- it is very likely that a different architecture may be better for best performance. In particular, we propose a ``network decomposition'' strategy, named Group-Net, in which we divide the network into groups. In this way, each full-precision group can be effectively reconstructed by aggregating a set of homogeneous binary branches. In addition, we learn effect connections among groups to improve the representational capability. Moreover, the proposed Group-Net shows strong generalization to other tasks. For instance, we extend Group-Net for highly accurate semantic segmentation by embedding rich context into the binary structure. Experiments on both classification and semantic segmentation tasks demonstrate the superior performance of the proposed methods over various popular architectures. In particular, we outperform the previous best binary neural networks in terms of accuracy and huge computation saving.
Bohan Zhuang, Chunhua Shen, Mingkui Tan, Lingqiao Liu, Ian D. Reid 0001
CVPR4
2019 Memorizing Normality to Detect Anomaly: Memory-Augmented Deep Autoencoder for Unsupervised Anomaly Detection
abstract
Deep autoencoder has been extensively used for anomaly detection. Training on the normal data, the autoencoder is expected to produce higher reconstruction error for the abnormal inputs than the normal ones, which is adopted as a criterion for identifying anomalies. However, this assumption does not always hold in practice. It has been observed that sometimes the autoencoder "generalizes" so well that it can also reconstruct anomalies well, leading to the miss detection of anomalies. To mitigate this drawback for autoencoder based anomaly detector, we propose to augment the autoencoder with a memory module and develop an improved autoencoder called memory-augmented autoencoder, i.e. MemAE. Given an input, MemAE firstly obtains the encoding from the encoder and then uses it as a query to retrieve the most relevant memory items for reconstruction. At the training stage, the memory contents are updated and are encouraged to represent the prototypical elements of the normal data. At the test stage, the learned memory will be fixed, and the reconstruction is obtained from a few selected memory records of the normal data. The reconstruction will thus tend to be close to a normal sample. Thus the reconstructed errors on anomalies will be strengthened for anomaly detection. MemAE is free of assumptions on the data type and thus general to be applied to different tasks. Experiments on various datasets prove the excellent generalization and high effectiveness of the proposed MemAE.
Dong Gong, Lingqiao Liu, Vuong Le, Budhaditya Saha, Moussa Reda Mansour, Svetha Venkatesh, Anton van den Hengel
ICCV2
2019 Exploiting Stereo Sound Channels to Boost Performance of Neural Network-Based Music Transcription
abstract
In recent years deep learning begins to show great potential for automatic music transcription that reproduces MIDI-like music composition information, such as note pitches and onset and offset times, from music recordings. In the literature without exception the two stereo sound channels coming with music recordings were averaged into a single channel to alleviate the computation overhead, which, from an entropy standpoint, definitely sacrifices information. In this paper we propose a method to properly combine the two sound channels for deep learning-based pitch detection. In particular, through modifying the loss function the network is forced to focus on the worse performing sound channel. This method achieves start-of-the-art frame-wise pitch detection performance on the MAPS dataset.
Lingqiao Liu, Qinfeng Shi
ICMLA2
2019 Deep point-to-subspace metric learning for sketch-based 3D shape retrieval
Yinjie Lei, Ziqin Zhou, Yulan Guo, Zijun Ma, Lingqiao Liu
Pattern Recognit.6
2019 Order-aware convolutional pooling for video based action recognition
Peng Wang 0023, Lingqiao Liu, Chunhua Shen, Heng Tao Shen
Pattern Recognit.2
2019 Piecewise Classifier Mappings: Learning Fine-Grained Learners for Novel Categories With Few Examples
abstract
Humans are capable of learning a new fine-grained concept with very little supervision, e.g., few exemplary images for a species of bird, yet our best deep learning systems need hundreds or thousands of labeled examples. In this paper, we try to reduce this gap by studying the fine-grained image recognition problem in a challenging few-shot learning setting, termed few-shot fine-grained recognition (FSFG). The task of FSFG requires the learning systems to build classifiers for the novel fine-grained categories from few examples (only one or less than five). To solve this problem, we propose an end-to-end trainable deep network, which is inspired by the state-of-the-art fine-grained recognition model and is tailored for the FSFG task. Specifically, our network consists of a bilinear feature learning module and a classifier mapping module: while the former encodes the discriminative information of an exemplar image into a feature vector, the latter maps the intermediate feature into the decision boundary of the novel category. The key novelty of our model is a "piecewise mappings" function in the classifier mapping module, which generates the decision boundary via learning a set of more attainable sub-classifiers in a more parameter-economic way. We learn the exemplar-to-classifier mapping based on an auxiliary dataset in a meta-learning fashion, which is expected to be able to generalize to novel categories. By conducting comprehensive experiments on three fine-grained datasets, we demonstrate that the proposed method achieves superior performance over the competing baselines.
Xiu-Shen Wei, Peng Wang 0023, Lingqiao Liu, Chunhua Shen, Jianxin Wu 0001
IEEE Trans. Image Process.3
2018 Coarse-to-Fine: A RNN-Based Hierarchical Attention Model for Vehicle Re-identification
Xiu-Shen Wei, Chen-Lin Zhang, Lingqiao Liu, Chunhua Shen, Jianxin Wu 0001
ACCV (2)3
2018 Towards Effective Low-Bitwidth Convolutional Neural Networks
abstract
This paper tackles the problem of training a deep convolutional neural network with both low-precision weights and low-bitwidth activations. Optimizing a low-precision network is very challenging since the training process can easily get trapped in a poor local minima, which results in substantial accuracy loss. To mitigate this problem, we propose three simple-yet-effective approaches to improve the network training. First, we propose to use a two-stage optimization strategy to progressively find good local minima. Specifically, we propose to first optimize a net with quantized weights and then quantized activations. This is in contrast to the traditional methods which optimize them simultaneously. Second, following a similar spirit of the first method, we propose another progressive optimization approach which progressively decreases the bit-width from high-precision to low-precision during the course of training. Third, we adopt a novel learning scheme to jointly train a full-precision model alongside the low-precision one. By doing so, the full-precision model provides hints to guide the low-precision model training. Extensive experiments on various datasets (i.e., CIFAR-100 and ImageNet) show the effectiveness of the proposed methods. To highlight, using our methods to train a 4-bit precision network leads to no performance decrease in comparison with its full-precision counterpart with standard network architectures (i.e., AlexNet and ResNet-50).
Bohan Zhuang, Chunhua Shen, Mingkui Tan, Lingqiao Liu, Ian D. Reid 0001
CVPR4
2018 Seeing Deeply and Bidirectionally: A Deep Learning Approach for Single Image Reflection Removal
Jie Yang 0002, Dong Gong, Lingqiao Liu, Qinfeng Shi
ECCV (3)3
2018 Semantic Segmentation from Limited Training Data
abstract
We present our approach for robotic perception in cluttered scenes that led to winning the recent Amazon Robotics Challenge (ARC) 2017. Next to small objects with shiny and transparent surfaces, the biggest challenge of the 2017 competition was the introduction of unseen categories. In contrast to traditional approaches which require large collections of annotated data and many hours of training, the task here was to obtain a robust perception pipeline with only few minutes of data acquisition and training time. To that end, we present two strategies that we explored. One is a deep metric learning approach that works in three separate steps: semantic-agnostic boundary detection, patch classification and pixel-wise voting. The other is a fully-supervised semantic segmentation approach with efficient dataset collection. We conduct an extensive analysis of the two methods on our ARC 2017 dataset. Interestingly, only few examples of each class are sufficient to fine-tune even very deep convolutional neural networks for this specific task.
Anton Milan, Trung Pham, Kumar Vijay, Douglas Morrison, Adam W. Tow, Lingqiao Liu, Jordan Erskine, Riccardo Grinover, Alec Gurman, Thomas Hunn, Norton Kelly-Boxall, Darryl Qijun Lee, Matthew McTaggart, Gerald Rallos, Andrew Razjigaev, Thomas James Rowntree, Rohan Smith, Sean Wade-McCue, Zheyu Zhuang, Chris Lehnert, Guosheng Lin, Ian D. Reid 0001, Peter I. Corke, Jürgen Leitner
ICRA6
2017 Weakly Supervised Semantic Segmentation Based on Co-segmentation
Guosheng Lin, Lingqiao Liu, Chunhua Shen, Ian D. Reid 0001
BMVC3
2017 From Motion Blur to Motion Flow: A Deep Learning Solution for Removing Heterogeneous Motion Blur
abstract
Removing pixel-wise heterogeneous motion blur is challenging due to the ill-posed nature of the problem. The predominant solution is to estimate the blur kernel by adding a prior, but extensive literature on the subject indicates the difficulty in identifying a prior which is suitably informative, and general. Rather than imposing a prior based on theory, we propose instead to learn one from the data. Learning a prior over the latent image would require modeling all possible image content. The critical observation underpinning our approach, however, is that learning the motion flow instead allows the model to focus on the cause of the blur, irrespective of the image content. This is a much easier learning task, but it also avoids the iterative process through which latent image priors are typically applied. Our approach directly estimates the motion flow from the blurred image through a fully-convolutional deep neural network (FCN) and recovers the unblurred image from the estimated motion flow. Our FCN is the first universal end-to-end mapping from the blurred image to the dense motion flow. To train the FCN, we simulate motion flows to generate synthetic blurred-image-motion-flow pairs thus avoiding the need for human labeling. Extensive experiments on challenging realistic blurred images demonstrate that the proposed method outperforms the state-of-the-art.
Dong Gong, Jie Yang 0002, Lingqiao Liu, Yanning Zhang 0001, Ian D. Reid 0001, Chunhua Shen, Anton van den Hengel, Qinfeng Shi
CVPR3
2017 Sequential Person Recognition in Photo Albums with a Recurrent Network
abstract
Recognizing the identities of people in everyday photos is still a very challenging problem for machine vision, due to issues such as non-frontal faces, changes in clothing, location, lighting. Recent studies have shown that rich relational information between people in the same photo can help in recognizing their identities. In this work, we propose to model the relational information between people as a sequence prediction task. At the core of our work is a novel recurrent network architecture, in which relational information between instances labels and appearance are modeled jointly. In addition to relational cues, scene context is incorporated in our sequence prediction model with no additional cost. In this sense, our approach is a unified framework for modeling both contextual cues and visual appearance of person instances. Our model is trained end-to-end with a sequence of annotated instances in a photo as inputs, and a sequence of corresponding labels as targets. We demonstrate that this simple but elegant formulation achieves state-of-the-art performance on the newly released People In Photo Albums (PIPA) dataset.
Yao Li 0003, Guosheng Lin, Bohan Zhuang, Lingqiao Liu, Chunhua Shen, Anton van den Hengel
CVPR4
2017 Graph-Structured Representations for Visual Question Answering
abstract
This paper proposes to improve visual question answering (VQA) with structured representations of both scene contents and questions. A key challenge in VQA is to require joint reasoning over the visual and text domains. The predominant CNN/LSTM-based approach to VQA is limited by monolithic vector representations that largely ignore structure in the scene and in the question. CNN feature vectors cannot effectively capture situations as simple as multiple object instances, and LSTMs process questions as series of words, which do not reflect the true complexity of language structure. We instead propose to build graphs over the scene objects and over the question words, and we describe a deep neural network that exploits the structure in these representations. We show that this approach achieves significant improvements over the state-of-the-art, increasing accuracy from 71.2% to 74.4% in accuracy on the abstract scenes multiple-choice benchmark, and from 34.7% to 39.1% in accuracy over pairs of balanced scenes, i.e. images with fine-grained differences and opposite yes/no answers to a same question.
Damien Teney, Lingqiao Liu, Anton van den Hengel
CVPR2
2017 Multi-attention Network for One Shot Learning
abstract
One-shot learning is a challenging problem where the aim is to recognize a class identified by a single training image. Given the practical importance of one-shot learning, it seems surprising that the rich information present in the class tag itself has largely been ignored. Most existing approaches restrict the use of the class tag to finding similar classes and transferring classifiers or metrics learned thereon. We demonstrate here, in contrast, that the class tag can inform one-shot learning as a guide to visual attention on the training image for creating the image representation. This is motivated by the fact that human beings can better interpret a training image if the class tag of the image is understood. Specifically, we design a neural network architecture which takes the semantic embedding of the class tag to generate attention maps and uses those attention maps to create the image features for one-shot learning. Note that unlike other applications, our task requires that the learned attention generator can be generalized to novel classes. We show that this can be realized by representing class tags with distributed word embeddings and learning the attention map generator from an auxiliary training set. Also, we design a multiple-attention scheme to extract richer information from the exemplar image and this leads to substantial performance improvement. Through comprehensive experiments, we show that the proposed approach leads to superior performance over the baseline methods.
Peng Wang 0023, Lingqiao Liu, Chunhua Shen, Zi Huang, Anton van den Hengel, Heng Tao Shen
CVPR2
2017 Attend in Groups: A Weakly-Supervised Deep Learning Framework for Learning from Web Data
abstract
Large-scale datasets have driven the rapid development of deep neural networks for visual recognition. However, annotating a massive dataset is expensive and time-consuming. Web images and their labels are, in comparison, much easier to obtain, but direct training on such automatically harvested images can lead to unsatisfactory performance, because the noisy labels of Web images adversely affect the learned recognition models. To address this drawback we propose an end-to-end weakly-supervised deep learning framework which is robust to the label noise in Web images. The proposed framework relies on two unified strategies - random grouping and attention - to effectively reduce the negative impact of noisy web image annotations. Specifically, random grouping stacks multiple images into a single training instance and thus increases the labeling accuracy at the instance level. Attention, on the other hand, suppresses the noisy signals from both incorrectly labeled images and less discriminative image regions. By conducting intensive experiments on two challenging datasets, including a newly collected fine-grained dataset with Web images of different car models,1, the superior performance of the proposed methods over competitive baselines is clearly demonstrated.
Bohan Zhuang, Lingqiao Liu, Yao Li 0003, Chunhua Shen, Ian D. Reid 0001
CVPR2
2017 Adversarial PoseNet: A Structure-Aware Convolutional Network for Human Pose Estimation
abstract
For human pose estimation in monocular images, joint occlusions and overlapping upon human bodies often result in deviated pose predictions. Under these circumstances, biologically implausible pose predictions may be produced. In contrast, human vision is able to predict poses by exploiting geometric constraints of joint inter-connectivity. To address the problem by incorporating priors about the structure of human bodies, we propose a novel structure-aware convolutional network to implicitly take such priors into account during training of the deep network. Explicit learning of such constraints is typically challenging. Instead, we design discriminators to distinguish the real poses from the fake ones (such as biologically implausible ones). If the pose generator (G) generates results that the discriminator fails to distinguish from real ones, the network successfully learns the priors.,,To better capture the structure dependency of human body joints, the generator G is designed in a stacked multi-task manner to predict poses as well as occlusion heatmaps. Then, the pose and occlusion heatmaps are sent to the discriminators to predict the likelihood of the pose being real. Training of the network follows the strategy of conditional Generative Adversarial Networks (GANs). The effectiveness of the proposed network is evaluated on two widely used human pose estimation benchmark datasets. Our approach significantly outperforms the state-of-the-art methods and almost always generates plausible human pose predictions.
Yu Chen 0037, Chunhua Shen, Xiu-Shen Wei, Lingqiao Liu, Jian Yang 0003
ICCV4
2017 Towards Context-Aware Interaction Recognition for Visual Relationship Detection
abstract
Recognizing how objects interact with each other is a crucial task in visual recognition. If we define the context of the interaction to be the objects involved, then most current methods can be categorized as either: (i) training a single classifier on the combination of the interaction and its context; or (ii) aiming to recognize the interaction independently of its explicit context. Both methods suffer limitations: the former scales poorly with the number of combinations and fails to generalize to unseen combinations, while the latter often leads to poor interaction recognition performance due to the difficulty of designing a contextindependent interaction classifier.,,To mitigate those drawbacks, this paper proposes an alternative, context-aware interaction recognition framework. The key to our method is to explicitly construct an interaction classifier which combines the context, and the interaction. The context is encoded via word2vec into a semantic space, and is used to derive a classification result for the interaction. The proposed method still builds one classifier for one interaction (as per type (ii) above), but the classifier built is adaptive to context via weights which are context dependent. The benefit of using the semantic space is that it naturally leads to zero-shot generalizations in which semantically similar contexts (subject-object pairs) can be recognized as suitable contexts for an interaction, even if they were not observed in the training set. Our method also scales with the number of interaction-context pairs since our model parameters do not increase with the number of interactions. Thus our method avoids the limitation of both approaches. We demonstrate experimentally that the proposed framework leads to improved performance for all investigated interaction representations and datasets.
Bohan Zhuang, Lingqiao Liu, Chunhua Shen, Ian D. Reid 0001
ICCV2
2017 Deep learning features at scale for visual place recognition
abstract
The success of deep learning techniques in the computer vision domain has triggered a range of initial investigations into their utility for visual place recognition, all using generic features from networks that were trained for other types of recognition tasks. In this paper, we train, at large scale, two CNN architectures for the specific place recognition task and employ a multi-scale feature encoding method to generate condition- and viewpoint-invariant features. To enable this training to occur, we have developed a massive Specific PlacEs Dataset (SPED) with hundreds of examples of place appearance change at thousands of different places, as opposed to the semantic place type datasets currently available. This new dataset enables us to set up a training regime that interprets place recognition as a classification problem. We comprehensively evaluate our trained networks on several challenging benchmark place recognition datasets and demonstrate that they achieve an average 10% increase in performance over other place recognition algorithms and pre-trained CNNs. By analyzing the network responses and their differences from pre-trained networks, we provide insights into what a network learns when training for place recognition, and what these results signify for future research in this area.
Zetao Chen, Adam Jacobson, Niko Sünderhauf, Ben Upcroft, Lingqiao Liu, Chunhua Shen, Ian D. Reid 0001, Michael Milford
ICRA5
2017 Mining Mid-level Visual Patterns with Deep CNN Activations
Yao Li 0003, Lingqiao Liu, Chunhua Shen, Anton van den Hengel
Int. J. Comput. Vis.2
2017 Cross-Convolutional-Layer Pooling for Image Recognition
abstract
Recent studies have shown that a Deep Convolutional Neural Network (DCNN) trained on a large image dataset can be used as a universal image descriptor and that doing so leads to impressive performance for a variety of image recognition tasks. Most of these studies adopt activations from a single DCNN layer, usually a fully-connected layer, as the image representation. In this paper, we proposed a novel way to extract image representations from two consecutive convolutional layers: one layer is used for local feature extraction and the other serves as guidance to pool the extracted features. By taking different viewpoints of convolutional layers, we further develop two schemes to realize this idea. The first directly uses convolutional layers from a DCNN. The second applies the pre-trained CNN on densely sampled image regions and treats the fully-connected activations of each image region as a convolutional layer's feature activations. We then train another convolutional layer on top of that as the pooling-guidance convolutional layer. By applying our method to three popular visual classification tasks, we find that our first scheme tends to perform better on applications which demand strong discrimination on lower-level visual patterns while the latter excels in cases that require discrimination on category-level patterns. Overall, the proposed method achieves superior performance over existing approaches for extracting image representations from a DCNN. In addition, we apply cross-layer pooling to the problem of image retrieval and propose schemes to reduce the computational cost. Experimental results suggest that the proposed method achieves promising results for the image retrieval task.
Lingqiao Liu, Chunhua Shen, Anton van den Hengel
IEEE Trans. Pattern Anal. Mach. Intell.1
2017 Compositional Model Based Fisher Vector Coding for Image Classification
abstract
Deriving from the gradient vector of a generative model of local features, Fisher vector coding (FVC) has been identified as an effective coding method for image classification. Most, if not all, FVC implementations employ the Gaussian mixture model (GMM) as the generative model for local features. However, the representative power of a GMM can be limited because it essentially assumes that local features can be characterized by a fixed number of feature prototypes, and the number of prototypes is usually small in FVC. To alleviate this limitation, in this work, we break the convention which assumes that a local feature is drawn from one of a few Gaussian distributions. Instead, we adopt a compositional mechanism which assumes that a local feature is drawn from a Gaussian distribution whose mean vector is composed as a linear combination of multiple key components, and the combination weight is a latent random variable. In doing so we greatly enhance the representative power of the generative model underlying FVC. To implement our idea, we design two particular generative models following this compositional approach. In our first model, the mean vector is sampled from the subspace spanned by a set of bases and the combination weight is drawn from a Laplace distribution. In our second model, we further assume that a local feature is composed of a discriminative part and a residual part. As a result, a local feature is generated by the linear combination of discriminative part bases and residual part bases. The decomposition of the discriminative and residual parts is achieved via the guidance of a pre-trained supervised coding method. By calculating the gradient vector of the proposed models, we derive two new Fisher vector coding strategies. The first is termed Sparse Coding-based Fisher Vector Coding (SCFVC) and can be used as the substitute of traditional GMM based FVC. The second is termed Hybrid Sparse Coding-based Fisher vector coding (HSCFVC) since it combines the merits of both pre-trained supervised coding methods and FVC. Using pre-trained Convolutional Neural Network (CNN) activations as local features, we experimentally demonstrate that the proposed methods are superior to traditional GMM based FVC and achieve state-of-the-art performance in various image classification tasks.
Lingqiao Liu, Peng Wang 0023, Chunhua Shen, Lei Wang 0001, Anton van den Hengel, Chao Wang 0063, Heng Tao Shen
IEEE Trans. Pattern Anal. Mach. Intell.1
2017 Learning discriminative trajectorylet detector sets for accurate skeleton-based action recognition
Ruizhi Qiao, Lingqiao Liu, Chunhua Shen, Anton van den Hengel
Pattern Recognit.2
2017 Temporal Pyramid Pooling-Based Convolutional Neural Network for Action Recognition
abstract
Encouraged by the success of convolutional neural networks (CNNs) in image classification, recently much effort is spent on applying the CNNs to the video-based action recognition problems. One challenge is that a video contains a varying number of frames, which is incompatible to the standard input format of the CNNs. Existing methods handle this issue either by directly sampling a fixed number of frames or bypassing this issue by introducing a 3D convolutional layer, which conducts convolution in spatial-temporal domain. In this paper, we propose a novel network structure, which allows an arbitrary number of frames as the network input. The key to our solution is to introduce a module consisting of an encoding layer and a temporal pyramid pooling layer. The encoding layer maps the activation from the previous layers to a feature vector suitable for pooling, whereas the temporal pyramid pooling layer converts multiple frame-level activations into a fixed-length video-level representation. In addition, we adopt a feature concatenation layer that combines the appearance and motion information. Compared with the frame sampling strategy, our method avoids the risk of missing any important frames. Compared with the 3D convolutional method, which requires a huge video data set for network training, our model can be learned on a small target data set because we can leverage the off-the-shelf image-level CNN for model parameter initialization. Experiments on three challenging data sets, Hollywood2, HMDB51, and UCF101 demonstrate the effectiveness of the proposed network.
Peng Wang 0023, Yuanzhouhan Cao, Chunhua Shen, Lingqiao Liu, Heng Tao Shen
IEEE Trans. Circuits Syst. Video Technol.4
2017 A Graph-Embedding Approach to Hierarchical Visual Word Mergence
abstract
Appropriately merging visual words are an effective dimension reduction method for the bag-of-visual-words model in image classification. The approach of hierarchically merging visual words has been extensively employed, because it gives a fully determined merging hierarchy. Existing supervised hierarchical merging methods take different approaches and realize the merging process with various formulations. In this paper, we propose a unified hierarchical merging approach built upon the graph-embedding framework. Our approach is able to merge visual words for any scenario, where a preferred structure and an undesired structure are defined, and, therefore, can effectively attend to all kinds of requirements for the word-merging process. In terms of computational efficiency, we show that our algorithm can seamlessly integrate a fast search strategy developed in our previous work and, thus, well maintain the state-of-the-art merging speed. To the best of our survey, the proposed approach is the first one that addresses the hierarchical visual word mergence in such a flexible and unified manner. As demonstrated, it can maintain excellent image classification performance even after a significant dimension reduction, and outperform all the existing comparable visual word-merging methods. In a broad sense, our work provides an open platform for applying, evaluating, and developing new criteria for hierarchical word-merging tasks.
Lei Wang 0001, Lingqiao Liu, Luping Zhou
IEEE Trans. Neural Networks Learn. Syst.2
2016 Less is More: Zero-Shot Learning from Online Textual Documents with Noise Suppression
abstract
Classifying a visual concept merely from its associated online textual source, such as a Wikipedia article, is an attractive research topic in zero-shot learning because it alleviates the burden of manually collecting semantic attributes. Recent work has pursued this approach by exploring various ways of connecting the visual and text domains. In this paper, we revisit this idea by going further to consider one important factor: the textual representation is usually too noisy for the zero-shot learning application. This observation motivates us to design a simple yet effective zero-shot learning method that is capable of suppressing noise in the text. Specifically, we propose an l2,1-norm based objective function which can simultaneously suppress the noisy signal in the text and learn a function to match the text document and visual features. We also develop an optimization algorithm to efficiently solve the resulting problem. By conducting experiments on two large datasets, we demonstrate that the proposed method significantly outperforms those competing methods which rely on online information sources but with no explicit noise suppression. Furthermore, we make an in-depth analysis of the proposed method and provide insight as to what kind of information in documents is useful for zero-shot learning.
Ruizhi Qiao, Lingqiao Liu, Chunhua Shen, Anton van den Hengel
CVPR2
2016 What's Wrong with That Object? Identifying Images of Unusual Objects by Modelling the Detection Score Distribution
abstract
This paper studies the challenging problem of identifying unusual instances of known objects in images within an "open world" setting. That is, we aim to find objects that are members of a known class, but which are not typical of that class. Thus the "unusual object" should be distinguished from both the "regular object" and the "other objects". Such unusual objects may be of interest in many applications such as surveillance or quality control. We propose to identify unusual objects by inspecting the distribution of object detection scores at multiple image regions. The key observation motivating our approach is that "regular object" images, "unusual object" images and "other objects" images exhibit different region-level scores in terms of both the score values and the spatial distributions. To model these distributions we propose to use Gaussian Processes (GP) to construct two separate generative models, one for the "regular object" and the other for the "other objects". More specifically, we design a new covariance function to simultaneously model the detection score at a single location and the score dependencies between multiple regions. We demonstrate that the proposed approach outperforms comparable methods on a new large dataset constructed for the purpose.
Peng Wang 0023, Lingqiao Liu, Chunhua Shen, Zi Huang, Anton van den Hengel, Heng Tao Shen
CVPR2
2016 What Value Do Explicit High Level Concepts Have in Vision to Language Problems?
abstract
Much recent progress in Vision-to-Language (V2L) problems has been achieved through a combination of Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs). This approach does not explicitly represent high-level semantic concepts, but rather seeks to progress directly from image features to text. In this paper we investigate whether this direct approach succeeds due to, or despite, the fact that it avoids the explicit representation of high-level information. We propose a method of incorporating high-level concepts into the successful CNN-RNN approach, and show that it achieves a significant improvement on the state-of-the-art in both image captioning and visual question answering. We also show that the same mechanism can be used to introduce external semantic information and that doing so further improves performance. We achieve the best reported results on both image captioning and VQA on several benchmark datasets, and provide an analysis of the value of explicit high-level concepts in V2L problems.
Qi Wu 0001, Chunhua Shen, Lingqiao Liu, Anthony R. Dick, Anton van den Hengel
CVPR3
2016 Image Co-localization by Mimicking a Good Detector's Confidence Score Distribution
Yao Li 0003, Lingqiao Liu, Chunhua Shen, Anton van den Hengel
ECCV (2)2
2016 A Generalized Probabilistic Framework for Compact Codebook Creation
abstract
Compact and discriminative visual codebooks are preferred in many visual recognition tasks. In the literature, a number of works have taken the approach of hierarchically merging visual words of an initial large-sized codebook, but implemented this approach with different merging criteria. In this work, we propose a single probabilistic framework to unify these merging criteria, by identifying two key factors: the function used to model the class-conditional distribution and the method used to estimate the distribution parameters. More importantly, by adopting new distribution functions and/or parameter estimation methods, our framework can readily produce a spectrum of novel merging criteria. Three of them are specifically discussed in this paper. For the first criterion, we adopt the multinomial distribution with the Bayesian method; For the second criterion, we integrate the Gaussian distribution with maximum likelihood parameter estimation. For the third criterion, which shows the best merging performance, we propose a max-margin-based parameter estimation method and apply it with the multinomial distribution. Extensive experimental study is conducted to systematically analyze the performance of the above three criteria and compare them with existing ones. As demonstrated, the best criterion within our framework achieves the overall best merging performance among the compared merging criteria developed in the literature.
Lingqiao Liu, Lei Wang 0001, Chunhua Shen
IEEE Trans. Pattern Anal. Mach. Intell.1
2016 Learning Discriminative Bayesian Networks from High-Dimensional Continuous Neuroimaging Data
abstract
Due to its causal semantics, Bayesian networks (BN) have been widely employed to discover the underlying data relationship in exploratory studies, such as brain research. Despite its success in modeling the probability distribution of variables, BN is naturally a generative model, which is not necessarily discriminative. This may cause the ignorance of subtle but critical network changes that are of investigation values across populations. In this paper, we propose to improve the discriminative power of BN models for continuous variables from two different perspectives. This brings two general discriminative learning frameworks for Gaussian Bayesian networks (GBN). In the first framework, we employ Fisher kernel to bridge the generative models of GBN and the discriminative classifiers of SVMs, and convert the GBN parameter learning to Fisher kernel learning via minimizing a generalization error bound of SVMs. In the second framework, we employ the max-margin criterion and build it directly upon GBN models to explicitly optimize the classification performance of the GBNs. The advantages and disadvantages of the two frameworks are discussed and experimentally compared. Both of them demonstrate strong power in learning discriminative parameters of GBNs for neuroimaging based brain network analysis, as well as maintaining reasonable representation capacity. The contributions of this paper also include a new Directed Acyclic Graph (DAG) constraint with theoretical guarantee to ensure the graph validity of GBN.
Luping Zhou, Lei Wang 0001, Lingqiao Liu, Philip Ogunbona, Dinggang Shen
IEEE Trans. Pattern Anal. Mach. Intell.3
2015 Mid-level deep pattern mining
abstract
Mid-level visual element discovery aims to find clusters of image patches that are both representative and discriminative. In this work, we study this problem from the prospective of pattern mining while relying on the recently popularized Convolutional Neural Networks (CNNs). Specifically, we find that for an image patch, activation extracted from the first fully-connected layer of a CNN have two appealing properties which enable its seamless integration with pattern mining. Patterns are then discovered from a large number of CNN activations of image patches through the well-known association rule mining. When we retrieve and visualize image patches with the same pattern (See Fig. 1), surprisingly, they are not only visually similar but also semantically consistent. We apply our approach to scene and object classification tasks, and demonstrate that our approach outperforms all previous works on mid-level visual element discovery by a sizeable margin with far fewer elements being used. Our approach also outperforms or matches recent works using CNN for these tasks. Source code of the complete system is available online.
Yao Li 0003, Lingqiao Liu, Chunhua Shen, Anton van den Hengel
CVPR2
2015 The treasure beneath convolutional layers: Cross-convolutional-layer pooling for image classification
abstract
A number of recent studies have shown that a Deep Convolutional Neural Network (DCNN) pretrained on a large dataset can be adopted as a universal image descriptor, and that doing so leads to impressive performance at a range of image classification tasks. Most of these studies, if not all, adopt activations of the fully-connected layer of a DCNN as the image or region representation and it is believed that convolutional layer activations are less discriminative. This paper, however, advocates that if used appropriately, convolutional layer activations constitute a powerful image representation. This is achieved by adopting a new technique proposed in this paper called cross-convolutional-layer pooling. More specifically, it extracts subarrays of feature maps of one convolutional layer as local features, and pools the extracted features with the guidance of the feature maps of the successive convolutional layer. Compared with existing methods that apply DCNNs in the similar local feature setting, the proposed method avoids the input image style mismatching issue which is usually encountered when applying fully connected layer activations to describe local regions. Also, the proposed method is easier to implement since it is codebook free and does not have any tuning parameters. By applying our method to four popular visual classification tasks, it is demonstrated that the proposed method can achieve comparable or in some cases significantly better performance than existing fully-connected layer based image representations.
Lingqiao Liu, Chunhua Shen, Anton van den Hengel
CVPR1
2015 Density Maximization for Improving Graph Matching With Its Applications
abstract
Graph matching has been widely used in both image processing and computer vision domain due to its powerful performance for structural pattern representation. However, it poses three challenges to image sparse feature matching: 1) the combinatorial nature limits the size of the possible matches; 2) it is sensitive to outliers because its objective function prefers more matches; and 3) it works poorly when handling many-to-many object correspondences, due to its assumption of one single cluster of true matches. In this paper, we address these challenges with a unified framework called density maximization (DM), which maximizes the values of a proposed graph density estimator both locally and globally. DM leads to the integration of feature matching, outlier elimination, and cluster detection. Experimental evaluation demonstrates that it significantly boosts the true matches and enables graph matching to handle both outliers and many-to-many object correspondences. We also extend it to dense correspondence estimation and obtain large improvement over the state-of-the-art methods. We further demonstrate the usefulness of our methods using three applications: 1) instance-level image retrieval; 2) mask transfer; and 3) image enhancement.
Chao Wang 0063, Lei Wang 0001, Lingqiao Liu
IEEE Trans. Image Process.3
2014 Progressive Mode-Seeking on Graphs for Sparse Feature Matching
Chao Wang 0063, Lei Wang 0001, Lingqiao Liu
ECCV (2)3
2014 Max-Margin Based Learning for Discriminative Bayesian Network from Neuroimaging Data
Luping Zhou, Lei Wang 0001, Lingqiao Liu, Philip Ogunbona, Dinggang Shen
MICCAI (3)3
2014 Encoding High Dimensional Local Features by Sparse Coding Based Fisher Vectors
Lingqiao Liu, Chunhua Shen, Lei Wang 0001, Anton van den Hengel, Chao Wang 0063
NIPS1
2014 A Hierarchical Word-Merging Algorithm with Class Separability Measure
abstract
In image recognition with the bag-of-features model, a small-sized visual codebook is usually preferred to obtain a low-dimensional histogram representation and high computational efficiency. Such a visual codebook has to be discriminative enough to achieve excellent recognition performance. To create a compact and discriminative codebook, in this paper we propose to merge the visual words in a large-sized initial codebook by maximally preserving class separability. We first show that this results in a difficult optimization problem. To deal with this situation, we devise a suboptimal but very efficient hierarchical word-merging algorithm, which optimally merges two words at each level of the hierarchy. By exploiting the characteristics of the class separability measure and designing a novel indexing structure, the proposed algorithm can hierarchically merge 10,000 visual words down to two words in merely 90 seconds. Also, to show the properties of the proposed algorithm and reveal its advantages, we conduct detailed theoretical analysis to compare it with another hierarchical word-merging algorithm that maximally preserves mutual information, obtaining interesting findings. Experimental studies are conducted to verify the effectiveness of the proposed algorithm on multiple benchmark data sets. As shown, it can efficiently produce more compact and discriminative codebooks than the state-of-the-art hierarchical word-merging algorithms, especially when the size of the codebook is significantly reduced.
Lei Wang 0001, Luping Zhou, Chunhua Shen, Lingqiao Liu, Huan Liu 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2014 HEp-2 cell image classification with multiple linear descriptors
Lingqiao Liu, Lei Wang 0001
Pattern Recognit.1
2013 Discriminative Brain Effective Connectivity Analysis for Alzheimer's Disease: A Kernel Learning Approach upon Sparse Gaussian Bayesian Network
abstract
Analyzing brain networks from neuroimages is becoming a promising approach in identifying novel connectivity-based biomarkers for the Alzheimer's disease (AD). In this regard, brain ``effective connectivity" analysis, which studies the causal relationship among brain regions, is highly challenging and of many research opportunities. Most of the existing works in this field use generative methods. Despite their success in data representation and other important merits, generative methods are not necessarily discriminative, which may cause the ignorance of subtle but critical disease-induced changes. In this paper, we propose a learning-based approach that integrates the benefits of generative and discriminative methods to recover effective connectivity. In particular, we employ Fisher kernel to bridge the generative models of sparse Bayesian networks (SBN) and the discriminative classifiers of SVMs, and convert the SBN parameter learning to Fisher kernel learning via minimizing a generalization error bound of SVMs. Our method is able to simultaneously boost the discriminative power of both the generative SBN models and the SBN-induced SVM classifiers via Fisher kernel. The proposed method is tested on analyzing brain effective connectivity for AD from ADNI data, and demonstrates significant improvements over the state-of-the-art work.
Luping Zhou, Lei Wang 0001, Lingqiao Liu, Philip Ogunbona, Dinggang Shen
CVPR3
2013 A Scalable Unsupervised Feature Merging Approach to Efficient Dimensionality Reduction of High-Dimensional Visual Data
abstract
To achieve a good trade-off between recognition accuracy and computational efficiency, it is often needed to reduce high-dimensional visual data to medium-dimensional ones. For this task, even applying a simple full-matrix-based linear projection causes significant computation and memory use. When the number of visual data is large, how to efficiently learn such a projection could even become a problem. The recent feature merging approach offers an efficient way to reduce the dimensionality, which only requires a single scan of features to perform reduction. However, existing merging algorithms do not scale well with high-dimensional data, especially in the unsupervised case. To address this problem, we formulate unsupervised feature merging as a PCA problem imposed with a special structure constraint. By exploiting its connection with k-means, we transform this constrained PCA problem into a feature clustering problem. Moreover, we employ the hashing technique to improve its scalability. These produce a scalable feature merging algorithm for our dimensionality reduction task. In addition, we develop an extension of this method by leveraging the neighborhood structure in the data to further improve dimensionality reduction performance. In further, we explore the incorporation of bipolar merging - a variant of merging function which allows the subtraction operation - into our algorithms. Through three applications in visual recognition, we demonstrate that our methods can not only achieve good dimensionality reduction performance with little computational cost but also help to create more powerful representation at both image level and local feature level.
Lingqiao Liu, Lei Wang 0001
ICCV1
2013 Improving Graph Matching via Density Maximization
abstract
Graph matching has been widely used in various applications in computer vision due to its powerful performance. However, it poses three challenges to image sparse feature matching: (1) The combinatorial nature limits the size of the possible matches, (2) It is sensitive to outliers because the objective function prefers more matches, (3) It works poorly when handling many-to-many object correspondences, due to its assumption of one single cluster for each graph. In this paper, we address these problems with a unified framework-Density Maximization. We propose a graph density local estimator (DLE) to measure the quality of matches. Density Maximization aims to maximize the DLE values both locally and globally. The local maximization of DLE finds the clusters of nodes as well as eliminates the outliers. The global maximization of DLE efficiently refines the matches by exploring a much larger matching space. Our Density Maximization is orthogonal to specific graph matching algorithms. Experimental evaluation demonstrates that it significantly boosts the true matches and enables graph matching to handle both outliers and many-to-many object correspondences.
Chao Wang 0063, Lei Wang 0001, Lingqiao Liu
ICCV3
2013 An Adaptive Approach to Learning Optimal Neighborhood Kernels
abstract
Learning an optimal kernel plays a pivotal role in kernel-based methods. Recently, an approach called optimal neighborhood kernel learning (ONKL) has been proposed, showing promising classification performance. It assumes that the optimal kernel will reside in the neighborhood of a "pre-specified" kernel. Nevertheless, how to specify such a kernel in a principled way remains unclear. To solve this issue, this paper treats the pre-specified kernel as an extra variable and jointly learns it with the optimal neighborhood kernel and the structure parameters of support vector machines. To avoid trivial solutions, we constrain the pre-specified kernel with a parameterized model. We first discuss the characteristics of our approach and in particular highlight its adaptivity. After that, two instantiations are demonstrated by modeling the pre-specified kernel as a common Gaussian radial basis function kernel and a linear combination of a set of base kernels in the way of multiple kernel learning (MKL), respectively. We show that the optimization in our approach is a min-max problem and can be efficiently solved by employing the extended level method and Nesterov's method. Also, we give the probabilistic interpretation for our approach and apply it to explain the existing kernel learning methods, providing another perspective for their commonness and differences. Comprehensive experimental results on 13 UCI data sets and another two real-world data sets show that via the joint learning process, our approach not only adaptively identifies the pre-specified kernel, but also achieves superior classification performance to the original ONKL and the related MKL algorithms.
Xinwang Liu 0002, Jianping Yin, Lei Wang 0001, Lingqiao Liu, Jun Liu 0003, Chenping Hou, Jian Zhang 0002
IEEE Trans. Cybern.4
2012 What has my classifier learned? Visualizing the classification rules of bag-of-feature model by support region detection
abstract
In the past decade, the bag-of-feature model has established itself as the state-of-the-art method in various visual classification tasks. Despite its simplicity and high performance, it normally works as a black box and the classification rule is not transparent to users. However, to better understand the classification process, it is favorable to look into the black box to see how an image is recognized. To fill this gap, we developed a tool called Restricted Support Region Set (RSRS) Detection which can be utilized to visualize the image regions that are critical to the classification decision. More specifically, we define the Restricted Support Region Set for a given image as such a set of size-restricted and non-overlapped regions that if any one of them is removed the image will be wrongly classified. Focusing on the state-of-the-art bag-of-feature classification system, we developed an efficient RSRS detection algorithm and discussed its applications. We showed that it can be used to identify the limitation of a classifier, predict its failure mode, discover the classification rules and reveal the database bias. Moreover, as experimentally demonstrated, this tool also enables common users to efficiently tune the classifier by removing the inappropriate support regions, which can lead to a better generalization performance.
Lingqiao Liu, Lei Wang 0001
CVPR1
2012 Incorporation of radius-info can be simple with SimpleMKL
Xinwang Liu 0002, Lei Wang 0001, Jianping Yin, Lingqiao Liu
Neurocomputing4
2011 A generalized probabilistic framework for compact codebook creation
abstract
Compact and discriminative visual codebooks are preferred in many visual recognition tasks. In the literature, a few researchers have taken the approach of hierarchically merging visual words of a initial large-size code-book, but implemented this idea with different merging criteria. In this work, we show that by defining different class-conditional distribution function and parameter estimation method, these merging criteria can be unified under a single probabilistic framework. More importantly, by adopting new distribution functions and/or parameter estimation methods, we can generalize this framework to produce a spectrum of novel merging criteria. Two of them are particularly focused in this work. For one criterion, we adopt the multinomial distribution to model each object class, and for the other criterion we propose a max-margin-based parameter estimation method. Both theoretical analysis and experimental study demonstrate the superior performance of the two new merging criteria and the general applicability of our probabilistic framework.
Lingqiao Liu, Lei Wang 0001, Chunhua Shen
CVPR1
2011 In defense of soft-assignment coding
abstract
In object recognition, soft-assignment coding enjoys computational efficiency and conceptual simplicity. However, its classification performance is inferior to the newly developed sparse or local coding schemes. It would be highly desirable if its classification performance could become comparable to the state-of-the-art, leading to a coding scheme which perfectly combines computational efficiency and classification performance. To achieve this, we revisit soft-assignment coding from two key aspects: classification performance and probabilistic interpretation. For the first aspect, we argue that the inferiority of soft-assignment coding is due to its neglect of the underlying manifold structure of local features. To remedy this, we propose a simple modification to localize the soft-assignment coding, which surprisingly achieves comparable or even better performance than existing sparse or local coding schemes while maintaining its computational advantage. For the second aspect, based on our probabilistic interpretation of the soft-assignment coding, we give a probabilistic explanation to the magic max-pooling operation, which has successfully been used by sparse or local coding schemes but still poorly understood. This probability explanation motivates us to develop a new mix-order max-pooling operation which further improves the classification performance of the proposed coding scheme. As experimentally demonstrated, the localized soft-assignment coding achieves the state-of-the-art classification performance with the highest computational efficiency among the existing coding schemes.
Lingqiao Liu, Lei Wang 0001, Xinwang Liu 0002
ICCV1
2011 Exploring latent class information for image retrieval using the bag-of-feature model
abstract
Recently, the Bag-of-Feature (BoF) model has shown promising performance in object and generic image retrieval. The similarity between two images is typically measured by the distance between the two histograms. Due to the imperfection of local descriptor and quantization error, visually similar image patches can be wrongly quantized into different visual words, making this distance-based measure less accurate. To address this issue, this paper explores the information of latent class, which is formed by all the database images that share the same visual concept with the one being compared to a given query. We then cast image similarity as the probability of the query and a database image belonging to a same latent class. Considering that a class of images together can better depict a visual concept, the shift from image-to-image to image-to-class comparison is expected to bring a more robust similarity measure. Because the ground truth of the latent class is not accessible in image retrieval, we define a latent class prior in our probabilistic model and derive its marginal distribution. This gives rise to a novel and efficient image similarity measure. It can significantly improve retrieval performance without prolonging retrieval process. Experimental study on multiple benchmark data sets demonstrates its advantages.
Lingqiao Liu, Lei Wang 0001
ACM Multimedia1