EDBT 2026 Demo / reviewers in the wild / expert
Qinfeng Shi
dblp:51/5804 · also Javen Qinfeng Shi, Javen Shi, Qinfeng (Javen) Shi
· DBLP profile ↗
124ranked-venue papers
9as first author
44since 2021 · last 2026
0000-0002-9126-2107ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 93 · 8 first-author · 34 since 2021Graphics, computer vision, multimedia, augmented reality and games · 69 · 4 first-author · 17 since 2021Databases, data management, data science and information retrieval · 5Human-computer interaction and ubiquitous computing · 3Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Truth as a Trajectory: What Internal Representations Reveal About Large Language Model ReasoningabstractHamed Damirchi, Imezadelajara, Ehsan Abbasnejad, Afshar Shamsi, Zhen Zhang, Javen Qinfeng Shi. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hamed Damirchi, Imezadelajara, Ehsan Abbasnejad, Afshar Shamsi, Zhen Zhang 0008, Qinfeng Shi |
ACL (1) | 6 |
| 2026 | Identifying Weight-Variant Latent Causal ModelsabstractThe task of causal representation learning aims to uncover latent higher-level causal variables that affect lower-level observations. Identifying the true latent causal variables from observed data, while allowing instantaneous causal relations among latent variables, remains a challenge, however. To this end, we start with the analysis of three intrinsic indeterminacies in identifying latent variables from observations: transitivity, permutation indeterminacy, and scaling indeterminacy. We find that transitivity acts as a key role in impeding the identifiability of latent causal variables. To address the unidentifiable issue due to transitivity, we introduce a novel identifiability condition where the underlying latent causal model satisfies a linear-Gaussian model, in which the causal coefficients and the distribution of Gaussian noise are modulated by an additional observed variable. Under certain assumptions, including the existence of a reference condition under which latent causal influences vanish, we can show that the latent causal variables can be identified up to trivial permutation and scaling, and that partial identifiability results can still be obtained when this reference condition is violated for a subset of latent variables. Furthermore, based on these theoretical results, we propose a novel method, termed Structural caUsAl Variational autoEncoder (SuaVE), which directly learns causal representations and causal relationships among them, together with the mapping from the latent causal variables to the observed ones. Experimental results on synthetic and real data demonstrate the identifiability and consistency results and the efficacy of SuaVE in learning causal representations. Yuhang Liu 0002, Zhen Zhang 0008, Dong Gong, Mingming Gong, Biwei Huang, Anton van den Hengel, Kun Zhang 0001, Qinfeng Shi |
J. Mach. Learn. Res. | 8 |
| 2026 | Parameter-efficient action planning with large language models for vision-and-language navigationabstract• In REVERIE, the agent needs fine-grained instructions to act more efficiently. • Large Language Models (LLMs) show great potential for the navigation task. • Generating goal plans and single-step instructions improves the agent’s performance. • Parameter-efficient fine-tuning the LLMs results in more accurate navigation plans. • Designing precise prompts alongside fine-tuning mitigates hallucination generation. The remote embodied referring expression (REVERIE) task requires an agent to navigate through complex indoor environments and localize a remote object specified by high-level instructions, such as “ bring me a spoon ”, without pre-exploration. Hence, an efficient navigation plan is essential for the final success. This paper proposes a novel parameter-efficient action planner using large language models (PEAP-LLM) to generate a single-step instruction at each location. The proposed model consists of two modules, LLM goal planner (LGP) and LoRA action planner (LAP). Initially, LGP extracts the goal-oriented plan from REVERIE instructions, including the target object and room. Then, LAP generates a single-step instruction with the goal-oriented plan, high-level instruction, and current visual observation as input. PEAP-LLM enables the embodied agent to interact with LAP as the path planner on the fly. A simple direct application of LLMs hardly achieves good performance. Moreover, existing hard-prompt-based methods are prone to errors in complex scenarios and require human intervention. To address these issues and prevent the LLM from generating hallucinations and biased information, we propose a novel two-stage method for fine-tuning the LLM, consisting of supervised fine-tuning (SFT) and direct preference optimization (DPO). SFT improves the quality of the generated instructions, while DPO incorporates environmental feedback into fine-tuning. Experimental results demonstrate a noticeable enhancement over the baseline model on the REVERIE benchmark. Bahram Mohammadi, Ehsan Abbasnejad, Yuankai Qi, Qi Wu 0001, Anton van den Hengel, Qinfeng Shi |
Pattern Recognit. | 6 |
| 2026 | Link prediction on multi-relational graphs from an influence propagation perspective
Zidu Yin, Yuankai Qi, Dong Gong, Ehsan Abbasnejad, Kun Yue, Qinfeng Shi |
Pattern Recognit. | 6 |
| 2025 | Causal Disentanglement and Cross-Modal Alignment for Enhanced Few-Shot Learning
Tianjiao Jiang, Zhen Zhang 0008, Yuhang Liu 0002, Qinfeng Shi |
ICCV | 4 |
| 2025 | Analytic DAG Constraints for Differentiable DAG LearningabstractRecovering the underlying Directed Acyclic Graph (DAG)
structures from observational data presents a formidable challenge, partly due
to the combinatorial nature of the DAG-constrained optimization
problem. Recently, researchers have identified gradient vanishing as
one of the primary obstacles in differentiable DAG learning and have
proposed several DAG constraints to mitigate this issue. By developing
the necessary theory to establish a connection between analytic
functions and DAG constraints, we demonstrate that analytic functions
from the set $\\{f(x) = c_0 + \\sum_{i=1}^{\infty}c_ix^i | \\forall i > 0, c_i > 0; r = \\lim_{i\\rightarrow \\infty}c_{i}/c_{i+1} > 0\\}$ can be employed to
formulate effective DAG constraints. Furthermore, we establish that
this set of functions is closed under several functional operators,
including differentiation, summation, and
multiplication. Consequently, these operators can be leveraged to
create novel DAG constraints based on existing ones. Using these
properties, we design a series of DAG constraints and develop an
efficient algorithm to evaluate them. Experiments
in various settings demonstrate that our DAG constraints
outperform previous state-of-the-art comparators. Our implementation is available at https://github.com/zzhang1987/AnalyticDAGLearning. Zhen Zhang 0008, Ignavier Ng, Dong Gong, Yuhang Liu 0002, Mingming Gong, Biwei Huang, Kun Zhang 0001, Anton van den Hengel, Qinfeng Shi |
ICLR | 9 |
| 2025 | Coreset Selection via Reducible Loss in Continual LearningabstractRehearsal-based continual learning (CL) aims to mitigate catastrophic forgetting by maintaining a subset of samples from previous tasks and replaying them. The rehearsal memory can be naturally constructed as a coreset, designed to form a compact subset that enables training with performance comparable to using the full dataset. The coreset selection task can be formulated as bilevel optimization that solves for the subset to minimize the outer objective of the learning task. Existing methods primarily rely on inefficient probabilistic sampling or local gradient-based scoring to approximate sample importance through an iterative process that can be susceptible to ambiguity or noise. Specifically, non-representative samples like ambiguous or noisy samples are difficult to learn and incur high loss values even when training on the full dataset. However, existing methods relying on local gradient tend to highlight these samples in an attempt to minimize the outer loss, leading to a suboptimal coreset. To enhance coreset selection, especially in CL where high-quality samples are essential, we propose a coreset selection method that measures sample importance using reducible loss (ReL) that quantifies the impact of adding a sample to model performance. By leveraging ReL and a process derived from bilevel optimization, we identify and retain samples that yield the highest performance gain. They are shown to be informative and representative. Furthermore, ReL requires only forward computation, making it significantly more efficient than previous methods. To better apply coreset selection in CL, we extend our method to address key challenges such as task interference, streaming data, and knowledge distillation. Experiments on data summarization and continual learning demonstrate the effectiveness and efficiency of our approach. Ruilin Tong, Yuhang Liu 0002, Qinfeng Shi, Dong Gong |
ICLR | 3 |
| 2025 | On the Value of Cross-Modal Misalignment in Multimodal Representation LearningabstractMultimodal representation learning, exemplified by multimodal contrastive learning (MMCL) using image-text pairs, aims to learn powerful representations by aligning cues across modalities. This approach relies on the core assumption that the exemplar image-text pairs constitute two representations of an identical concept. However, recent research has revealed that real-world datasets often exhibit cross-modal misalignment. There are two distinct viewpoints on how to address this issue: one suggests mitigating the misalignment, and the other leveraging it. We seek here to reconcile these seemingly opposing perspectives, and to provide a practical guide for practitioners. Using latent variable models we thus formalize cross-modal misalignment by introducing two specific mechanisms: Selection bias, where some semantic variables are absent in the text, and perturbation bias, where semantic variables are altered—both leading to misalignment in data pairs. Our theoretical analysis demonstrates that, under mild assumptions, the representations learned by MMCL capture exactly the information related to the subset of the semantic variables invariant to selection and perturbation biases. This provides a unified perspective for understanding misalignment. Based on this, we further offer actionable insights into how misalignment should inform the design of real-world ML systems. We validate our theoretical findings via extensive empirical studies on both synthetic data and real image-text datasets, shedding light on the nuanced impact of cross-modal misalignment on multimodal representation learning. Yichao Cai 0001, Yuhang Liu 0002, Erdun Gao, Tianjiao Jiang, Zhen Zhang 0008, Anton van den Hengel, Qinfeng Shi |
NeurIPS | 7 |
| 2025 | Solving the Asymmetric Traveling Salesman Problem via Trace-Guided Cost AugmentationabstractThe Asymmetric Traveling Salesman Problem (ATSP) ranks among the most fundamental and notoriously difficult problems in combinatorial optimization. We propose a novel continuous relaxation framework for the Asymmetric Traveling Salesman Problem (ATSP) by leveraging differentiable constraints that encourage acyclic structures and valid permutations. Our approach integrates a differentiable trace-based Directed Acyclic Graph (DAG) constraint with a doubly stochastic matrix relaxation of the assignment problem, enabling gradient-based optimization over soft permutations. We develop a projected exponentiated gradient method with adaptive step size to minimize tour cost while satisfying the relaxed constraints. To recover high-quality discrete tours, we introduce a greedy post-processing procedure that iteratively corrects subtours using cost-aware cycle merging. Our method achieves state-of-the-art performance on standard asymmetric TSP benchmarks and demonstrates competitive scalability and accuracy, particularly on large or asymmetric instances where heuristic solvers such as LKH-3 struggle. Zhen Zhang 0008, Qinfeng Shi, Wee Sun Lee |
NeurIPS | 2 |
| 2025 | A Simple-but-Effective Baseline for Training-Free Class-Agnostic CountingabstractClass-Agnostic Counting (CAC) seeks to accurately count objects in a given image with only a few reference examples. While previous methods achieving this relied on additional training, recent efforts have shown that it's possible to accomplish this without training by utilizing pre-existing foundation models, particularly the Segment Any-thing Model (SAM), for counting via instance-level segmentation. Although promising, current training-free methods still lag behind their training-based counterparts in terms of performance. In this research, we present a straightforward training-free solution that effectively bridges this performance gap, serving as a strong baseline. The primary contribution of our work lies in the discovery of four key technologies that can enhance performance. Specifically, we suggest employing a superpixel algorithm to generate more precise initial point prompts, utilizing an image encoder with richer semantic knowledge to replace the SAM encoder for representing candidate objects, and adopting a multiscale mechanism and a transductive prototype scheme to update the representation of reference examples. By combining these four technologies, our approach achieves significant improvements over existing training-free methods and delivers performance on par with training-based ones. Lingqiao Liu, Qinfeng Shi |
WACV | 4 |
| 2025 | One-stage keypoint detection network for end-to-end cow body measurement
Guangyuan Yang, Yongliang Qiao, Hongxing Deng, Qinfeng Shi, Huaibo Song |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | Augmented Commonsense Knowledge for Remote Object GroundingabstractThe vision-and-language navigation (VLN) task necessitates an agent to perceive the surroundings, follow natural language instructions, and act in photo-realistic unseen environments. Most of the existing methods employ the entire image or object features to represent navigable viewpoints. However, these representations are insufficient for proper action prediction, especially for the REVERIE task, which uses concise high-level instructions, such as “Bring me the blue cushion in the master bedroom”. To address enhancing representation, we propose an augmented commonsense knowledge model (ACK) to leverage commonsense information as a spatio-temporal knowledge graph for improving agent navigation. Specifically, the proposed approach involves constructing a knowledge base by retrieving commonsense information from ConceptNet, followed by a refinement module to remove noisy and irrelevant knowledge. We further present ACK which consists of knowledge graph-aware cross-modal and concept aggregation modules to enhance visual representation and visual-textual data alignment by integrating visible objects, commonsense knowledge, and concept history, which includes object and knowledge temporal information. Moreover, we add a new pipeline for the commonsense-based decision-making process which leads to more accurate local action prediction. Experimental results demonstrate our proposed model noticeably outperforms the baseline and archives the state-of-the-art on the REVERIE benchmark. The source code is available at https://github.com/Bahram-Mohammadi/ACK. Bahram Mohammadi, Yicong Hong, Yuankai Qi, Qi Wu 0001, Shirui Pan, Qinfeng Shi |
AAAI | 6 |
| 2024 | Semantic Role Labeling Guided Out-of-distribution DetectionabstractIdentifying unexpected domain-shifted instances in natural language processing is crucial in real-world applications. Previous works identify the out-of-distribution (OOD) instance by leveraging a single global feature embedding to represent the sentence, which cannot characterize subtle OOD patterns well. Another major challenge current OOD methods face is learning effective low-dimensional sentence representations to identify the hard OOD instances that are semantically similar to the in-distribution (ID) data. In this paper, we propose a new unsupervised OOD detection method, namely Semantic Role Labeling Guided Out-of-distribution Detection (SRLOOD), that separates, extracts, and learns the semantic role labeling (SRL) guided fine-grained local feature representations from different arguments of a sentence and the global feature representations of the full sentence using a margin-based contrastive loss. A novel self-supervised approach is also introduced to enhance such global-local feature learning by predicting the SRL extracted role. The resulting model achieves SOTA performance on four OOD benchmarks, indicating the effectiveness of our approach. The code is publicly accessible via https://github.com/cytai/SRLOOD. Jinan Zou, Maihao Guo, Yu Tian 0001, Haiyao Cao, Lingqiao Liu, Ehsan Abbasnejad, Qinfeng Shi |
LREC/COLING | 8 |
| 2024 | CLAP: Isolating Content from Style Through Contrastive Learning with Augmented Prompts
Yichao Cai 0001, Yuhang Liu 0002, Zhen Zhang 0008, Qinfeng Shi |
ECCV (21) | 4 |
| 2024 | Invariant Representation Learning for Generalizable ImitationabstractWe address the problem of learning imitation policies that generalize across environments sharing the same underlying causal structure between the system dynamics and the task.We introduce a novel loss for learning invariant state representations that draws inspiration from adversarial robustness.Our approach is algorithm-agnostic and does not require knowledge of domain labels.Yet, evaluation in visual and non-visual environments reveals improved zero-shot generalization in the presence of spurious features compared to previous works. Mohamed Khalil Jabri, Panagiotis Papadakis, Ehsan Abbasnejad, Gilles Coppin, Qinfeng Shi |
ESANN | 5 |
| 2024 | Identifiable Latent Polynomial Causal Models through the Lens of ChangeabstractCausal representation learning aims to unveil latent high-level causal representations from observed low-level data. One of its primary tasks is to provide reliable assurance of identifying these latent causal models, known as \textit{identifiability}. A recent breakthrough explores identifiability by leveraging the change of causal influences among latent causal variables across multiple environments \citep{liu2022identifying}. However, this progress rests on the assumption that the causal relationships among latent causal variables adhere strictly to linear Gaussian models. In this paper, we extend the scope of latent causal models to involve nonlinear causal relationships, represented by polynomial models, and general noise distributions conforming to the exponential family. Additionally, we investigate the necessity of imposing changes on all causal parameters and present partial identifiability results when part of them remains unchanged. Further, we propose a novel empirical estimation method, grounded in our theoretical finding, that enables learning consistent latent causal representations. Our experimental results, obtained from both synthetic and real-world data, validate our theoretical contributions concerning identifiability and consistency. Yuhang Liu 0002, Zhen Zhang 0008, Dong Gong, Mingming Gong, Biwei Huang, Anton van den Hengel, Kun Zhang 0001, Qinfeng Shi |
ICLR | 8 |
| 2024 | Influence Function Based Second-Order Channel Pruning: Evaluating True Loss Changes for Pruning is Possible Without RetrainingabstractA challenge of channel pruning is designing efficient and effective criteria to select channels to prune. A widely used criterion is minimal performance degeneration, e.g., loss changes before and after pruning being the smallest. To accurately evaluate the truth performance degeneration requires retraining the survived weights to convergence, which is prohibitively slow. Hence existing pruning methods settle to use previous weights (without retraining) to evaluate the performance degeneration. However, we observe that the loss changes differ significantly with and without retraining. It motivates us to develop a technique to evaluate true loss changes without retraining, using which to select channels to prune with more reliability and confidence. We first derive a closed-form estimator of the true loss change per mask change, using influence functions without retraining. Influence function is a classic technique from robust statistics that reveals the impacts of a training sample on the model's prediction and is repurposed by us to assess impacts on true loss changes. We then show how to assess the importance of all channels simultaneously and develop a novel global channel pruning algorithm accordingly. We conduct extensive experiments to verify the effectiveness of the proposed algorithm, which significantly outperforms the competing channel pruning methods on both image classification and object detection tasks. To the best of our knowledge, we are the first that shows evaluating true loss changes or pruning without retraining is possible. This finding will open up opportunities for a series of new paradigms to emerge that differ from existing pruning methods. Hongrong Cheng, Miao Zhang 0022, Qinfeng Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | A Survey on Deep Neural Network Pruning: Taxonomy, Comparison, Analysis, and RecommendationsabstractModern deep neural networks, particularly recent large language models, come with massive model sizes that require significant computational and storage resources. To enable the deployment of modern models on resource-constrained environments and to accelerate inference time, researchers have increasingly explored pruning techniques as a popular research direction in neural network compression. More than three thousand pruning papers have been published from 2020 to 2024. However, there is a dearth of up-to-date comprehensive review papers on pruning. To address this issue, in this survey, we provide a comprehensive review of existing research works on deep neural network pruning in a taxonomy of 1) universal/specific speedup, 2) when to prune, 3) how to prune, and 4) fusion of pruning and other compression techniques. We then provide a thorough comparative analysis of eight pairs of contrast settings for pruning (e.g., unstructured/structured, one-shot/iterative, data-free/data-driven, initialized/pre-trained weights, etc.) and explore several emerging topics, including pruning for large language models, vision transformers, diffusion models, and large multimodal models, post-training pruning, and different levels of supervision for pruning to shed light on the commonalities and differences of existing methods and lay the foundation for further method development. Finally, we provide some valuable recommendations on selecting pruning methods and prospect several promising research directions for neural network pruning. To facilitate future research on deep neural network pruning, we summarize broad pruning applications (e.g., adversarial robustness, natural language understanding, etc.) and build a curated collection of datasets, networks, and evaluations on different applications. We maintain a repository on https://github.com/hrcheng1066/awesome-pruning that serves as a comprehensive resource for neural network pruning papers and corresponding open-source codes. We will keep updating this repository to include the latest advancements in the field. Hongrong Cheng, Miao Zhang 0022, Qinfeng Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Improving Reward Estimation in Goal-Conditioned Imitation Learning with Counterfactual Data and Structural Causal ModelsabstractInternational audience Mohamed Khalil Jabri, Panagiotis Papadakis, Ehsan Abbasnejad, Gilles Coppin, Qinfeng Shi |
ICINCO (2) | 5 |
| 2023 | Factor Graph Neural NetworksabstractIn recent years, we have witnessed a surge of Graph Neural Networks (GNNs), most of which can learn powerful representations in an end-to-end fashion with great success in many real-world applications. They have resemblance to Probabilistic Graphical Models (PGMs), but break free from some limitations of PGMs. By aiming to provide expressive methods for representation learning instead of computing marginals or most likely configurations, GNNs provide flexibility in the choice of information flowing rules while maintaining good performance. Despite their success and inspirations, they lack efficient ways to represent and learn higher-order relations among variables/nodes. More expressive higher-order GNNs which operate on k-tuples of nodes need increased computational resources in order to process higher-order tensors. We propose Factor Graph Neural Networks (FGNNs) to effectively capture higher-order relations for inference and learning. To do so, we first derive an efficient approximate Sum-Product loopy belief propagation inference algorithm for discrete higher-order PGMs. We then neuralize the novel message passing scheme into a Factor Graph Neural Network (FGNN) module by allowing richer representations of the message update rules; this facilitates both efficient inference and powerful end-to-end learning. We further show that with a suitable choice of message aggregation operators, our FGNN is also able to represent Max-Product belief propagation, providing a single family of architecture that can represent both Max and Sum-Product loopy belief propagation. Our extensive experimental evaluation on synthetic as well as real datasets demonstrates the potential of the proposed model. Zhen Zhang 0008, Mohammed Haroon Dupty, Fan Wu 0011, Qinfeng Shi, Wee Sun Lee |
J. Mach. Learn. Res. | 4 |
| 2023 | Human Interaction Understanding With Consistency-Aware LearningabstractCompared with the progress made on human activity classification, much less success has been achieved on human interaction understanding (HIU). Apart from the latter task is much more challenging, the main causation is that recent approaches learn human interactive relations via shallow graphical representations, which are inadequate to model complicated human interactive-relations. This paper proposes a deep consistency-aware framework aiming at tackling the grouping and labelling inconsistencies in HIU. This framework consists of three components, including a backbone CNN to extract image features, a factor graph network to implicitly learn higher-order consistencies among labelling and grouping variables, and a consistency-aware reasoning module to explicitly enforcing consistencies. The last module is inspired by our key observation that the consistency-aware reasoning bias can be embedded into an energy function or a particular loss function, minimizing which delivers consistent predictions. An efficient mean-field inference algorithm is proposed, such that all modules of our network could be trained in an end-to-end fashion. Experimental results demonstrate that the two proposed consistency-learning modules complement each other, and both make considerable contributions in achieving leading performance on three benchmarks of HIU. The effectiveness of the proposed approach is further validated by experiments on detecting human-object interactions. Jiajun Meng, Zhenhua Wang 0003, Kaining Ying, Jianhua Zhang 0002, Dongyan Guo, Zhen Zhang 0008, Qinfeng Shi, Shengyong Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | 3D Medical image segmentation using parallel transformers
Qingsen Yan, Shengqiang Liu, Songhua Xu, Caixia Dong, Zongfang Li, Qinfeng Shi, Yanning Zhang 0001, Duwei Dai |
Pattern Recognit. | 6 |
| 2023 | SharpFormer: Learning Local Feature Preserving Global Representations for Image DeblurringabstractThe goal of dynamic scene deblurring is to remove the motion blur presented in a given image. To recover the details from the severe blurs, conventional convolutional neural networks (CNNs) based methods typically increase the number of convolution layers, kernel-size, or different scale images to enlarge the receptive field. However, these methods neglect the non-uniform nature of blurs, and cannot extract varied local and global information. Unlike the CNNs-based methods, we propose a Transformer-based model for image deblurring, named SharpFormer, that directly learns long-range dependencies via a novel Transformer module to overcome large blur variations. Transformer is good at learning global information but is poor at capturing local information. To overcome this issue, we design a novel Locality preserving Transformer (LTransformer) block to integrate sufficient local information into global features. In addition, to effectively apply LTransformer to the medium-resolution features, a hybrid block is introduced to capture intermediate mixed features. Furthermore, we use a dynamic convolution (DyConv) block, which aggregates multiple parallel convolution kernels to handle the non-uniform blur of inputs. We leverage a powerful two-stage attentive framework composed of the above blocks to learn the global, hybrid, and local features effectively. Extensive experiments on the GoPro and REDS datasets show that the proposed SharpFormer performs favourably against the state-of-the-art methods in blurred image restoration. Qingsen Yan, Dong Gong, Zhen Zhang 0008, Yanning Zhang 0001, Qinfeng Shi |
IEEE Trans. Image Process. | 6 |
| 2022 | Active Learning by Feature MixingabstractThe promise of active learning (AL) is to reduce labelling costs by selecting the most valuable examples to annotate from a pool of unlabelled data. Identifying these examples is especially challenging with high-dimensional data (e.g. images, videos) and in low-data regimes. In this paper, we propose a novel method for batch AL called ALFA-Mix. We identify unlabelled instances with sufficiently-distinct features by seeking inconsistencies in predictions resulting from interventions on their representations. We construct interpolations between representations of labelled and unlabelled instances then examine the predicted labels. We show that inconsistencies in these predictions help discovering features that the model is unable to recognise in the unlabelled instances. We derive an efficient implementation based on a closed-form solution to the optimal interpolation causing changes in predictions. Our method outperforms all recent AL approaches in 30 different settings on 12 benchmarks of images, videos, and non-visual data. The improvements are especially significant in low-data regimes and on self-trained vision transformers, where ALFA-Mix outperforms the state-of-the-art in 59% and 43% of the experiments respectively11The code is available at https://github.com/aminparvaneh/alpha_mix_active_learning. Amin Parvaneh, Ehsan Abbasnejad, Damien Teney, Gholamreza Haffari, Anton van den Hengel, Qinfeng Shi |
CVPR | 6 |
| 2022 | Learning Bayesian Sparse Networks with Full Experience Replay for Continual LearningabstractContinual Learning (CL) methods aim to enable machine learning models to learn new tasks without catastrophic forgetting of those that have been previously mastered. Existing CL approaches often keep a buffer of previously-seen samples, perform knowledge distillation, or use regularization techniques towards this goal. Despite their performance, they still suffer from interference across tasks which leads to catastrophic forgetting. To ameliorate this problem, we propose to only activate and select sparse neurons for learning current and past tasks at any stage. More parameters space and model capacity can thus be reserved for the future tasks. This minimizes the interference between parameters for different tasks. To do so, we propose a Sparse neural Network for Continual Learning (SNCL), which employs variational Bayesian sparsity priors on the activations of the neurons in all layers. Full Experience Replay (FER) provides effective supervision in learning the sparse activations of the neurons in different layers. A loss-aware reservoir-sampling strategy is developed to maintain the memory buffer. The proposed method is agnostic as to the network structures and the task boundaries. Experiments on different datasets show that SNCL achieves state-of-the-art result for mitigating forgetting. Qingsen Yan, Dong Gong, Yuhang Liu 0002, Anton van den Hengel, Qinfeng Shi |
CVPR | 5 |
| 2022 | Implicit Sample Extension for Unsupervised Person Re-IdentificationabstractMost existing unsupervised person re-identification (Re-ID) methods use clustering to generate pseudo labels for model training. Unfortunately, clustering sometimes mixes different true identities together or splits the same identity into two or more sub clusters. Training on these noisy clusters substantially hampers the Re-ID accuracy. Due to the limited samples in each identity, we suppose there may lack some underlying information to well reveal the accurate clusters. To discover these information, we propose an Implicit Sample Extension (ISE) method to generate what we call support samples around the cluster boundaries. Specifically, we generate support samples from actual samples and their neighbouring clusters in the embedding space through a progressive linear interpolation (PLI) strategy. PLI controls the generation with two critical factors, i.e., 1) the direction from the actual sample towards its K-nearest clusters and 2) the degree for mixing up the context information from the K-nearest clusters. Meanwhile, given the support samples, ISE further uses a label-preserving loss to pull them towards their corresponding actual samples, so as to compact each cluster. Consequently, ISE reduces the “sub and mixed” clustering errors, thus improving the Re-ID performance. Extensive experiments demonstrate that the proposed method is effective and achieves state-of-the-art performance for unsupervised person Re-ID. Code is available at: https://github.com/PaddlePaddle/PaddleClas. Xinyu Zhang 0015, Zhigang Wang 0002, Jian Wang 0066, Errui Ding, Qinfeng Shi, Zhaoxiang Zhang 0001, Jingdong Wang 0001 |
CVPR | 6 |
| 2022 | Bayesian Learning with Information Gain Provably Bounds Risk for a Robust Adversarial DefenseabstractWe present a new algorithm to learn a deep neural network model robust against adversarial attacks. Previous algorithms demonstrate an adversarially trained Bayesian Neural Network (BNN) provides improved robustness. We recognize the learning approach for approximating the multi-modal posterior distribution of an adversarially trained Bayesian model can lead to mode collapse; consequently, the model’s achievements in robustness and performance are sub-optimal. Instead, we first propose preventing mode collapse to better approximate the multi-modal posterior distribution. Second, based on the intuition that a robust model should ignore perturbations and only consider the informative content of the input, we conceptualize and formulate an information gain objective to measure and force the information learned from both benign and adversarial training instances to be similar. Importantly. we prove and demonstrate that minimizing the information gain objective allows the adversarial risk to approach the conventional empirical risk. We believe our efforts provide a step towards a basis for a principled method of adversarially training BNNs. Our extensive experimental results demonstrate significantly improved robustness up to 20% compared with adversarial training and Adv-BNN under PGD attacks with 0.035 distortion on both CIFAR-10 and STL-10 dataset. Bao Gia Doan, Ehsan Abbasnejad, Qinfeng Shi, Damith Chinthana Ranasinghe |
ICML | 3 |
| 2022 | Truncated Matrix Power Iteration for Differentiable DAG LearningabstractRecovering underlying Directed Acyclic Graph (DAG) structures from observational data is highly challenging due to the combinatorial nature of the DAG-constrained optimization problem. Recently, DAG learning has been cast as a continuous optimization problem by characterizing the DAG constraint as a smooth equality one, generally based on polynomials over adjacency matrices. Existing methods place very small coefficients on high-order polynomial terms for stabilization, since they argue that large coefficients on the higher-order terms are harmful due to numeric exploding. On the contrary, we discover that large coefficients on higher-order terms are beneficial for DAG learning, when the spectral radiuses of the adjacency matrices are small, and that larger coefficients for higher-order terms can approximate the DAG constraints much better than the small counterparts. Based on this, we propose a novel DAG learning method with efficient truncated matrix power iteration to approximate geometric series based DAG constraints. Empirically, our DAG learning method outperforms the previous state-of-the-arts in various settings, often by a factor of $3$ or more in terms of structural Hamming distance. Zhen Zhang 0008, Ignavier Ng, Dong Gong, Yuhang Liu 0002, Ehsan Abbasnejad, Mingming Gong, Kun Zhang 0001, Qinfeng Shi |
NeurIPS | 8 |
| 2022 | ForeSI: Success-Aware Visual Navigation AgentabstractIn this work, we present a method to improve the efficiency and robustness of the previous model-free Reinforcement Learning (RL) algorithms for the task of object-goal visual navigation. Despite achieving state-of-the-art results, one of the major drawbacks of those approaches is the lack of a forward model that informs the agent about the potential consequences of its actions, i.e., being model-free. In this work, we augment the model-free RL with such a forward model that can predict a representation of a future state, from the beginning of a navigation episode, if the episode were to be successful. Furthermore, in order for efficient training, we develop an algorithm to integrate a replay buffer into the model-free RL that alternates between training the policy and the forward model. We call our agent ForeSI; ForeSI is trained to imagine a future latent state that leads to success. By explicitly imagining such a state, during the navigation, our agent is able to take better actions leading to two main advantages: first, in the absence of an object detector, ForeSI presents a more robust policy, i.e., it leads to about 5% absolute improvement on the Success Rate (SR); second, when combined with an off-the-shelf object detector to help better distinguish the target object, our method leads to about 3% absolute improvement on the SR and about 2% absolute improvement on Success weighted by inverse Path Length (SPL), i.e., presents higher efficiency. Mohammad Mahdi Kazemi Moghaddam, Ehsan Abbasnejad, Qi Wu 0001, Qinfeng Shi, Anton van den Hengel |
WACV | 4 |
| 2022 | Dual-Attention-Guided Network for Ghost-Free High Dynamic Range Imaging
Qingsen Yan, Dong Gong, Qinfeng Shi, Anton van den Hengel, Chunhua Shen, Ian D. Reid 0001, Yanning Zhang 0001 |
Int. J. Comput. Vis. | 3 |
| 2022 | Learn to Predict Sets Using Feed-Forward Neural NetworksabstractThis paper addresses the task of set prediction using deep feed-forward neural networks. A set is a collection of elements which is invariant under permutation and the size of a set is not fixed in advance. Many real-world problems, such as image tagging and object detection, have outputs that are naturally expressed as sets of entities. This creates a challenge for traditional deep neural networks which naturally deal with structured outputs such as vectors, matrices or tensors. We present a novel approach for learning to predict sets with unknown permutation and cardinality using deep neural networks. In our formulation we define a likelihood for a set distribution represented by a) two discrete distributions defining the set cardinally and permutation variables, and b) a joint distribution over set elements with a fixed cardinality. Depending on the problem under consideration, we define different training models for set prediction using deep neural networks. We demonstrate the validity of our set formulations on relevant vision problems such as: 1) multi-label image classification where we outperform the other competing methods on the PASCAL VOC and MS COCO datasets, 2) object detection, for which our formulation outperforms popular state-of-the-art detectors, and 3) a complex CAPTCHA test, where we observe that, surprisingly, our set-based network acquired the ability of mimicking arithmetics without any rules being coded. Seyed Hamid Rezatofighi, Tianyu Zhu 0001, Roman Kaskman, Farbod T. Motlagh, Qinfeng Shi, Anton Milan, Daniel Cremers, Laura Leal-Taixé, Ian D. Reid 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Video super-resolution via mixed spatial-temporal convolution and selective fusion
Wei Sun 0036, Dong Gong, Qinfeng Shi, Anton van den Hengel, Yanning Zhang 0001 |
Pattern Recognit. | 3 |
| 2022 | High dynamic range imaging via gradient-aware context aggregation network
Qingsen Yan, Dong Gong, Qinfeng Shi, Anton van den Hengel, Jinqiu Sun, Yu Zhu 0004, Yanning Zhang 0001 |
Pattern Recognit. | 3 |
| 2022 | Computationally Efficient Dilated Convolutional Model for Melody ExtractionabstractIn this paper we propose a dilated convolutional model for music melody extraction. Taking variable-q transforms (VQTs) as inputs, it first uses consecutive layers of convolution to capture local temporal-frequency patterns, and then a single layer of dilated convolution to capture global frequency patterns contributed by the pitches and harmonics of active notes. Compared with the contrast model without dilation, the proposed model can remarkably cut down the computational cost, and at the same time does not compromise the performance. Its advantages over existing models are two fold. First, it performs best on most datasets, for both general and vocal melody extraction. Second, it can achieve the best performance with least training data. Lingqiao Liu, Qinfeng Shi |
IEEE Signal Process. Lett. | 3 |
| 2022 | Show, Price and Negotiate: A Negotiator With Online Value Look-AheadabstractNegotiation, as an essential and complicated aspect of online shopping, is still challenging for an intelligent agent. To that end, we propose thePrice Negotiator, a modular deep neural network that addresses the unsolved problems in recent studies by (1) considering images of the items as a crucial, though neglected, source of information in a negotiation, (2) heuristically finding the most similar items from an external online source to predict the potential value and an acceptable agreement price, (3) predicting a general price-based “action” at each turn which is fed into the language generator to output the supporting natural language, and (4) adjusting the prices based on the predicted actions. Empirically, we show that our model, that is trained in both supervised and reinforcement learning setting, significantly improves negotiation on the CraigslistBargain dataset, in terms of the agreement price, price consistency, and dialogue quality. Amin Parvaneh, Ehsan Abbasnejad, Qi Wu 0001, Qinfeng Shi, Anton van den Hengel |
IEEE Trans. Multim. | 4 |
| 2021 | A Comprehensive CT Dataset for Liver Computer Assisted Diagnosis
Qingsen Yan, Bo Wang 0011, Dong Gong, Dingwen Zhang, Yang Yang 0009, Zheng You, Yanning Zhang 0001, Qinfeng Shi |
BMVC | 8 |
| 2021 | Consistency-Aware Graph Network for Human Interaction UnderstandingabstractCompared with the progress made on human activity classification, much less success has been achieved on human interaction understanding (HIU). Apart from the latter task is much more challenging, the main cause is that recent approaches learn human interactive relations via shallow graphical models, which is inadequate to model complicated human interactions. In this paper, we propose a consistency-aware graph network, which combines the representative ability of graph network and the consistency-aware reasoning to facilitate HIU. Our network consists of three components, a backbone CNN to extract image features, a factor graph network to learn third-order interactive relations among participants, and a consistency-aware reasoning module to enforce labeling and grouping consistencies. Our key observation is that the consistency-aware-reasoning bias for HIU can be embedded into an energy, minimizing which delivers consistent predictions. An efficient mean-field inference algorithm is proposed, such that all modules of our network could be trained jointly in an end-to-end manner. Experimental results show that our approach achieves leading performance on three benchmarks. Code is available at https://git.io/CAGNet. Zhenhua Wang 0003, Jiajun Meng, Dongyan Guo, Jianhua Zhang 0002, Qinfeng Shi, Shengyong Chen |
ICCV | 5 |
| 2021 | Memory-augmented Dynamic Neural Relational InferenceabstractDynamic interacting systems are prevalent in vision tasks. These interactions are usually difficult to observe and measure directly, and yet understanding latent interactions is essential for performing inference tasks on dynamic systems like forecasting. Neural relational inference (NRI) techniques are thus introduced to explicitly estimate interpretable relations between the entities in the system for trajectory prediction. However, NRI assumes static relations; thus, dynamic neural relational inference (DNRI) was proposed to handle dynamic relations using LSTM. Unfortunately, the older information will be washed away when the LSTM updates the latent variable as a whole, which is why DNRI struggles with modeling long-term dependences and forecasting long sequences. This motivates us to propose a memory-augmented dynamic neural relational inference method, which maintains two associative memory pools: one for the interactive relations and the other for the individual entities. The two memory pools help retain useful relation features and node features for the estimation in the future steps. Our model dynamically estimates the relations by learning better embeddings and utilizing the long-range information stored in the memory. With the novel memory modules and customized structures, our memory-augmented DNRI can update and access the memory adaptively as required. The memory pools also serve as global latent variables across time to maintain detailed long-term temporal relations readily available for other components to use. Experiments on synthetic and real-world datasets show the effectiveness of the proposed method on modeling dynamic relations and forecasting complex trajectories. Dong Gong, Zhen Zhang 0008, Qinfeng Shi, Anton van den Hengel |
ICCV | 3 |
| 2021 | Optimistic Agent: Accurate Graph-Based Value Estimation for More Successful Visual NavigationabstractWe humans can impeccably search for a target object, given its name only, even in an unseen environment. We argue that this ability is largely due to three main reasons: the incorporation of prior knowledge (or experience), the adaptation of it to the new environment using the observed visual cues and most importantly optimistically searching without giving up early. This is currently missing in the state-of-the-art visual navigation methods based on Reinforcement Learning (RL). In this paper, we propose to use externally learned prior knowledge of the relative object locations and integrate it into our model by constructing a neural graph. In order to efficiently incorporate the graph without increasing the state-space complexity, we propose Graph-based Value Estimation (GVE) module. GVE provides a more accurate baseline for estimating the Advantage function in actor-critic RL algorithm. This results in reduced value estimation error and, consequently, convergence to a more optimal policy. Through empirical studies, we show that our agent, dubbed as the optimistic agent, has a more realistic estimate of the state value during a navigation episode which leads to a higher success rate. Our extensive ablation studies show the efficacy of our simple method which achieves the state-of-the-art results measured by the conventional visual navigation metrics, e.g. Success Rate (SR) and Success weighted by Path Length (SPL), in AI2THOR environment. Mohammad Mahdi Kazemi Moghaddam, Qi Wu 0001, Ehsan Abbasnejad, Qinfeng Shi |
WACV | 4 |
| 2021 | Towards accurate HDR imaging with learning generator constraints
Qingsen Yan, Bo Wang 0011, Lei Zhang 0054, Zheng You, Qinfeng Shi, Yanning Zhang 0001 |
Neurocomputing | 6 |
| 2021 | COVID-19 Chest CT Image Segmentation Network by Multi-Scale Fusion and Enhancement OperationsabstractA novel coronavirus disease 2019 (COVID-19) was detected and has spread rapidly across various countries around the world since the end of the year 2019. Computed Tomography (CT) images have been used as a crucial alternative to the time-consuming RT-PCR test. However, pure manual segmentation of CT images faces a serious challenge with the increase of suspected cases, resulting in urgent requirements for accurate and automatic segmentation of COVID-19 infections. Unfortunately, since the imaging characteristics of the COVID-19 infection are diverse and similar to the backgrounds, existing medical image segmentation methods cannot achieve satisfactory performance. In this article, we try to establish a new deep convolutional neural network tailored for segmenting the chest CT images with COVID-19 infections. We first maintain a large and new chest CT image dataset consisting of 165,667 annotated chest CT images from 861 patients with confirmed COVID-19. Inspired by the observation that the boundary of the infected lung can be enhanced by adjusting the global intensity, in the proposed deep CNN, we introduce a feature variation block which adaptively adjusts the global properties of the features for segmenting COVID-19 infection. The proposed FV block can enhance the capability of feature representation effectively and adaptively for diverse cases. We fuse features at different scales by proposing Progressive Atrous Spatial Pyramid Pooling to handle the sophisticated infection areas with diverse appearance and shapes. The proposed method achieves state-of-the-art performance. Dice similarity coefficients are 0.987 and 0.726 for lung and COVID-19 segmentation, respectively. We conducted experiments on the data collected in China and Germany and show that the proposed deep CNN can produce impressive performance effectively. The proposed network enhances the segmentation ability of the COVID-19 infection, makes the connection with other techniques and contributes to the development of remedying COVID-19 infection. Qingsen Yan, Bo Wang 0011, Dong Gong, Chuan Luo 0003, Jianhu Shen, Jingyang Ai, Qinfeng Shi, Yanning Zhang 0001, Liang Zhang 0010, Zheng You |
IEEE Trans. Big Data | 8 |
| 2021 | Learning to Zoom-In via Learning to Zoom-Out: Real-World Super-Resolution by Generating and Adapting DegradationabstractMost learning-based super-resolution (SR) methods aim to recover high-resolution (HR) image from a given low-resolution (LR) image via learning on LR-HR image pairs. The SR methods learned on synthetic data do not perform well in real-world, due to the domain gap between the artificially synthesized and real LR images. Some efforts are thus taken to capture real-world image pairs. However, the captured LR-HR image pairs usually suffer from unavoidable misalignment, which hampers the performance of end- to-end learning. Here, focusing on the real-world SR, we ask a different question: since misalignment is unavoidable, can we propose a method that does not need LR-HR image pairing and alignment at all and utilizes real images as they are? Hence we propose a framework to learn SR from an arbitrary set of unpaired LR and HR images and see how far a step can go in such a realistic and "unsupervised" setting. To do so, we firstly train a degradation generation network to generate realistic LR images and, more importantly, to capture their distribution (i.e., learning to zoom out). Instead of assuming the domain gap has been eliminated, we minimize the discrepancy between the generated data and real data while learning a degradation adaptive SR network (i.e., learning to zoom in). The proposed unpaired method achieves state-of- the-art SR results on real-world images, even in the datasets that favour the paired-learning methods more. Wei Sun 0036, Dong Gong, Qinfeng Shi, Anton van den Hengel, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Attention-Guided Deep Neural Network With Multi-Scale Feature Fusion for Liver Vessel SegmentationabstractLiver vessel segmentation is fast becoming a key instrument in the diagnosis and surgical planning of liver diseases. In clinical practice, liver vessels are normally manual annotated by clinicians on each slice of CT images, which is extremely laborious. Several deep learning methods exist for liver vessel segmentation, however, promoting the performance of segmentation remains a major challenge due to the large variations and complex structure of liver vessels. Previous methods mainly using existing UNet architecture, but not all features of the encoder are useful for segmentation and some even cause interferences. To overcome this problem, we propose a novel deep neural network for liver vessel segmentation, called LVSNet, which employs special designs to obtain the accurate structure of the liver vessel. Specifically, we design Attention-Guided Concatenation (AGC) module to adaptively select the useful context features from low-level features guided by high-level features. The proposed AGC module focuses on capturing rich complemented information to obtain more details. In addition, we introduce an innovative multi-scale fusion block by constructing hierarchical residual-like connections within one single residual block, which is of great importance for effectively linking the local blood vessel fragments together. Furthermore, we construct a new dataset containing 40 thin thickness cases (0.625 mm) which consist of CT volumes and annotated vessels. To evaluate the effectiveness of the method with minor vessels, we also propose an automatic stratification method to split major and minor liver vessels. Extensive experimental results demonstrate that the proposed LVSNet outperforms previous methods on liver vessel segmentation datasets. Additionally, we conduct a series of ablation studies that comprehensively support the superiority of the underlying concepts. Qingsen Yan, Bo Wang 0011, Wei Zhang 0098, Chuan Luo 0003, Wei Xu 0005, Zhengqing Xu, Yanning Zhang 0001, Qinfeng Shi, Liang Zhang 0010, Zheng You |
IEEE J. Biomed. Health Informatics | 8 |
| 2021 | Deep Single Image Deraining via Modeling Haze-Like EffectabstractRemoving rain from images is of a great importance to various applications such as autonomous driving, drone piloting, and photo editing. Conventional methods rely on some heuristics to handcraft various priors to remove or separate rain from images. Recently, deep learning models are proposed to learn various end-to-end methods to complete this task. However, these methods might fail in obtaining satisfactory results in some real-world scenarios, especially when the captured images suffer from heavy rain that brings not only rain streaks but also a haze-like effect (caused by the accumulation of tiny raindrops). Different from most of the existing deep learning deraining methods that focus on handling rain streaks, we add a new variable to model the haze-like effect in a general model for rain, based on which a deep neural network is designed accordingly. Specifically, in our method, two branches are designed to handle rain streaks and the haze-like effect, respectively. The output of such branch structure is fed to an additional module to further enhance the performance. Three modules are trained jointly, leading to an end-to-end network, which supports a adjustment to the strength of removing the haze-like effect. Extensive experiments on several datasets show that our method outperforms several state-of-the-art methods in both objective assessment and visual quality. Yinglong Wang 0002, Dong Gong, Jie Yang 0002, Qinfeng Shi, Anton van den Hengel, Dehua Xie, Bing Zeng 0001 |
IEEE Trans. Multim. | 4 |
| 2020 | Gold Seeker: Information Gain From Policy Distributions for Goal-Oriented Vision-and-Langauge ReasoningabstractAs Computer Vision moves from passive analysis of pixels to active analysis of semantics, the breadth of information algorithms need to reason over has expanded significantly. One of the key challenges in this vein is the ability to identify the information required to make a decision, and select an action that will recover it. We propose a reinforcement-learning approach that maintains a distribution over its internal information, thus explicitly representing the ambiguity in what it knows, and needs to know, towards achieving its goal. Potential actions are then generated according to this distribution. For each potential action a distribution of the expected outcomes is calculated, and the value of the potential information gain assessed. The action taken is that which maximizes the potential information gain. We demonstrate this approach applied to two vision-and-language problems that have attracted significant recent interest, visual dialog and visual query generation. In both cases the method actively selects actions that will best reduce its internal uncertainty, and outperforms its competitors in achieving the goal of the challenge. Ehsan Abbasnejad, Iman Abbasnejad, Qi Wu 0001, Qinfeng Shi, Anton van den Hengel |
CVPR | 4 |
| 2020 | Counterfactual Vision and Language LearningabstractThe ongoing success of visual question answering methods has been somwehat surprising given that, at its most general, the problem requires understanding the entire variety of both visual and language stimuli. It is particularly remarkable that this success has been achieved on the basis of comparatively small datasets, given the scale of the problem. One explanation is that this has been accomplished partly by exploiting bias in the datasets rather than developing deeper multi-modal reasoning. This fundamentally limits the generalization of the method, and thus its practical applicability. We propose a method that addresses this problem by introducing counterfactuals in the training. In doing so we leverage structural causal models for counterfactual evaluation to formulate alternatives, for instance, questions that could be asked of the same image set. We show that simulating plausible alternative training data through this process results in better generalization. Ehsan Abbasnejad, Damien Teney, Amin Parvaneh, Qinfeng Shi, Anton van den Hengel |
CVPR | 4 |
| 2020 | Joint Learning of Social Groups, Individuals Action and Sub-group Activities in Videos
Mahsa Ehsanpour, Alireza Abedin Varamin, Fatemehsadat Saleh, Qinfeng Shi, Ian D. Reid 0001, Seyed Hamid Rezatofighi |
ECCV (9) | 4 |
| 2020 | Harmonic Structure-Based Neural Network Model for Music Pitch DetectionabstractPitched music notes are characterized by a harmonic structure. That is, the waveform of each note is composed of a number of sinusoids with different frequencies. Among these frequencies, the lowest is called the pitch or the fundamental frequency of the note, and all the others are integer multiples of the pitch and referred to as the harmonics of the note. In this paper we propose a neural network model for pitch detection that detects the active pitches present in a music recording. The model first uses conventional convolution whose receptive field is continuous to capture local frequency patterns in the spectrogram of the recording, and then uses sparse convolution with discontinuous receptive field to capture global harmonic frequency patterns. When the MAPS dataset is used both for training and for test, the model outperforms all existing models by large margins. When the MAESTRO dataset is used for training and the MAPS dataset is used for test, the model achieves the state-of-the-art performance as well. Lingqiao Liu, Qinfeng Shi |
ICMLA | 3 |
| 2020 | Enhancing Piano Transcription by Dilated ConvolutionabstractWhen detecting music pitch with deep learning, the most challenging issue is how to aggregate the multi-scale frequency information scattered in the spectrogram of a music signal. Traditionally, this aggregation was achieved via strided pooling and fully connected operations, or domain-specific sparse convolutions. In this paper we propose to use more general-purpose dilated convolution to this end. In particular, first we design a dilated convolutional acoustic model for piano pitch detection. This model leads existing acoustic models by large margins. Based on this model, we then design an automatic music transcription (AMT) system. When trained and tested on the MAPS dataset, this system outperforms existing AMT systems again by large margins. When trained and tested on the MAESTRO dataset, this system performs equally well as a state- of-the-art system, but has much better generalization capability. Lingqiao Liu, Qinfeng Shi |
ICMLA | 3 |
| 2020 | Counterfactual Vision-and-Language Navigation: Unravelling the UnseenabstractThe task of vision-and-language navigation (VLN) requires an agent to follow text instructions to find its way through simulated household environments. A prominent challenge is to train an agent capable of generalising to new environments at test time, rather than one that simply memorises trajectories and visual details observed during training. We propose a new learning strategy that learns both from observations and generated counterfactual environments. We describe an effective algorithm to generate counterfactual observations on the fly for VLN, as linear combinations of existing environments. Simultaneously, we encourage the agent's actions to remain stable between original and counterfactual environments through our novel training objective-effectively removing the spurious features that otherwise bias the agent. Our experiments show that this technique provides significant improvements in generalisation on benchmarks for Room-to-Room navigation and Embodied Question Answering. Amin Parvaneh, Ehsan Abbasnejad, Damien Teney, Qinfeng Shi, Anton van den Hengel |
NeurIPS | 4 |
| 2020 | Ghost Removal via Channel Attention in Exposure Fusion
Qingsen Yan, Bo Wang 0011, Xianjun Li, Qinfeng Shi, Zheng You, Yu Zhu 0004, Jinqiu Sun, Yanning Zhang 0001 |
Comput. Vis. Image Underst. | 6 |
| 2020 | GADE: A Generative Adversarial Approach to Density Estimation and its Applications
Ehsan Abbasnejad, Qinfeng Shi, Anton van den Hengel, Lingqiao Liu |
Int. J. Comput. Vis. | 2 |
| 2020 | Multi-way backpropagation for training compact deep neural networks
Jian Chen 0011, Anton van den Hengel, Qinfeng Shi, Mingkui Tan |
Neural Networks | 5 |
| 2020 | Model-Free Tracker for Multiple Objects Using Joint Appearance and Motion InferenceabstractModel-free tracking is a widely-accepted approach to track an arbitrary object in a video using a single frame annotation with no further prior knowledge about the object of interest. Extending this problem to track multiple objects is really challenging because: a) the tracker is not aware of the objects' type while trying to distinguish them from background (detection task), and b) The tracker needs to distinguish one object from other potentially similar objects (data association task) to generate stable trajectories. In order to track multiple arbitrary objects, most existing model-free tracking approaches rely on tracking each target individually by updating their appearance model independently. Therefore, in this scenario they often fail to perform well due to confusion between the appearance of similar objects, their sudden appearance changes and occlusion. To tackle this problem, we propose to use both appearance and motion models, and to learn them jointly using graphical models and deep neural networks features. We introduce an indicator variable to predict sudden appearance change and/or occlusion. When these happen, our model does not update the appearance model thus avoiding using the background and/or incorrect object to update the appearance of the object of interest mistakenly, and relies on our motion model to track. Moreover, we consider the correlation among all targets, and seek the joint optimal locations for all targets simultaneously as a graphical model inference problem. We learn the joint parameters for both appearance model and motion model in an online fashion under the framework of LaRank. Experiment results show that our method achieved superior performance compared to the competitive methods. Chongyu Liu, Rui Yao 0006, Seyed Hamid Rezatofighi, Ian D. Reid 0001, Qinfeng Shi |
IEEE Trans. Image Process. | 5 |
| 2020 | Deep HDR Imaging via A Non-Local NetworkabstractOne of the most challenging problems in reconstructing a high dynamic range (HDR) image from multiple low dynamic range (LDR) inputs is the ghosting artifacts caused by the object motion across different inputs. When the object motion is slight, most existing methods can well suppress the ghosting artifacts through aligning LDR inputs based on optical flow or detecting anomalies among them. However, they often fail to produce satisfactory results in practice, since the real object motion can be very large. In this study, we present a novel deep framework, termed NHDRRnet, which adopts an alternative direction and attempts to remove ghosting artifacts by exploiting the non-local correlation in inputs. In NHDRRnet, we first adopt an Unet architecture to fuse all inputs and map the fusion results into a low-dimensional deep feature space. Then, we feed the resultant features into a novel global non-local module which reconstructs each pixel by weighted averaging all the other pixels using the weights determined by their correspondences. By doing this, the proposed NHDRRnet is able to adaptively select the useful information (e.g., which are not corrupted by large motions or adverse lighting conditions) in the whole deep feature space to accurately reconstruct each pixel. In addition, we also incorporate a triple-pass residual module to capture more powerful local features, which proves to be effective in further boosting the performance. Extensive experiments on three benchmark datasets demonstrate the superiority of the proposed NDHRnet in terms of suppressing the ghosting artifacts in HDR reconstruction, especially when the objects have large motions. Qingsen Yan, Lei Zhang 0054, Yu Liu 0029, Yu Zhu 0004, Jinqiu Sun, Qinfeng Shi, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 6 |
| 2020 | Learning Distilled Graph for Large-Scale Social Network Data ClusteringabstractSpectral analysis is critical in social network analysis. As a vital step of the spectral analysis, the graph construction in many existing works utilizes content data only. Unfortunately, the content data often consists of noisy, sparse, and redundant features, which makes the resulting graph unstable and unreliable. In practice, besides the content data, social network data also contain link information, which provides additional information for graph construction. Some of previous works utilize the link data. However, the link data is often incomplete, which makes the resulting graph incomplete. To address these issues, we propose a novel Distilled Graph Clustering (DGC) method. It pursuits adistilled graphbased on both the content data and the link data. The proposed algorithm alternates between two steps: in the feature selection step, it finds the most representative feature subset w.r.t. an intermediate graph initialized with link data; in graph distillation step, the proposed method updates and refines the graph based on only the selected features. The final resulting graph, which is referred to as the distilled graph, is then utilized for spectral clustering on the large-scale social network data. Extensive experiments demonstrate the superiority of the proposed method. Wenhe Liu, Dong Gong, Mingkui Tan, Qinfeng Shi, Yi Yang 0001, Alex Hauptmann 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2020 | Fast and Low Memory Cost Matrix Factorization: Algorithm, Analysis, and Case StudyabstractMatrix factorization has been widely applied to various applications. With the fast development of storage and internet technologies, we have been witnessing a rapid increase of data. In this paper, we propose new algorithms for matrix factorization with the emphasis on efficiency. In addition, most existing methods of matrix factorization only consider a general smooth least square loss. Differently, many real-world applications have distinctive characteristics. As a result, different losses should be used accordingly. Therefore, it is beneficial to design new matrix factorization algorithms that are able to deal with both smooth and non-smooth losses. To this end, one needs to analyze the characteristics of target data and use the most appropriate loss based on the analysis. We particularly study two representative cases of low-rank matrix recovery, i.e., collaborative filtering for recommendation and high dynamic range imaging. To solve these two problems, we respectively propose a stage-wise matrix factorization algorithm by exploiting manifold optimization techniques. From our theoretical analysis, they are both are provably guaranteed to converge to a stationary point. Extensive experiments on recommender systems and high dynamic range imaging demonstrate the satisfactory performance and efficiency of our proposed method on large-scale real data. Yan Yan 0006, Mingkui Tan, Ivor W. Tsang, Yi Yang 0001, Qinfeng Shi, Chengqi Zhang |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2020 | Learning Deep Gradient Descent Optimization for Image DeconvolutionabstractAs an integral component of blind image deblurring, non-blind deconvolution removes image blur with a given blur kernel, which is essential but difficult due to the ill-posed nature of the inverse problem. The predominant approach is based on optimization subject to regularization functions that are either manually designed or learned from examples. Existing learning-based methods have shown superior restoration quality but are not practical enough due to their restricted and static model design. They solely focus on learning a prior and require to know the noise level for deconvolution. We address the gap between the optimization- and learning-based approaches by learning a universal gradient descent optimizer. We propose a recurrent gradient descent network (RGDN) by systematically incorporating deep neural networks into a fully parameterized gradient descent scheme. A hyperparameter-free update unit shared across steps is used to generate the updates from the current estimates based on a convolutional neural network. By training on diverse examples, the RGDN learns an implicit image prior and a universal update rule through recursive supervision. The learned optimizer can be repeatedly used to improve the quality of diverse degenerated observations. The proposed method possesses strong interpretability and high generalization. Extensive experiments on synthetic benchmarks and challenging real-world images demonstrate that the proposed deep optimization method is effective and robust to produce favorable results as well as practical for real-world image deblurring applications. Dong Gong, Zhen Zhang 0008, Qinfeng Shi, Anton van den Hengel, Chunhua Shen, Yanning Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Accurate Tensor Completion via Adaptive Low-Rank RepresentationabstractLow-rank representation-based approaches that assume low-rank tensors and exploit their low-rank structure with appropriate prior models have underpinned much of the recent progress in tensor completion. However, real tensor data only approximately comply with the low-rank requirement in most cases, viz., the tensor consists of low-rank (e.g., principle part) as well as non-low-rank (e.g., details) structures, which limit the completion accuracy of these approaches. To address this problem, we propose an adaptive low-rank representation model for tensor completion that represents low-rank and non-low-rank structures of a latent tensor separately in a Bayesian framework. Specifically, we reformulate the CANDECOMP/PARAFAC (CP) tensor rank and develop a sparsity-induced prior for the low-rank structure that can be used to determine tensor rank automatically. Then, the non-low-rank structure is modeled using a mixture of Gaussians prior that is shown to be sufficiently flexible and powerful to inform the completion process for a variety of real tensor data. With these two priors, we develop a Bayesian minimum mean-squared error estimate framework for inference. The developed framework can capture the important distinctions between low-rank and non-low-rank structures, thereby enabling more accurate model, and ultimately, completion. For various applications, compared with the state-of-the-art methods, the proposed model yields more accurate completion results. Lei Zhang 0054, Wei Wei 0008, Qinfeng Shi, Chunhua Shen, Anton van den Hengel, Yanning Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | A Generative Adversarial Density EstimatorabstractDensity estimation is a challenging unsupervised learning problem. Current maximum likelihood approaches for density estimation are either restrictive or incapable of producing high-quality samples. On the other hand, likelihood-free models such as generative adversarial networks, produce sharp samples without a density model. The lack of a density estimate limits the applications to which the sampled data can be put, however. We propose a Generative Adversarial Density Estimator, a density estimation approach that bridges the gap between the two. Allowing for a prior on the parameters of the model, we extend our density estimator to a Bayesian model where we can leverage the predictive variance to measure our confidence in the likelihood. Our experiments on challenging applications such as visual dialog where the density and the confidence in predictions are crucial shows the effectiveness of our approach. Ehsan Abbasnejad, Qinfeng Shi, Anton van den Hengel, Lingqiao Liu |
CVPR | 2 |
| 2019 | What's to Know? Uncertainty as a Guide to Asking Goal-Oriented QuestionsabstractOne of the core challenges in Visual Dialogue problems is asking the question that will provide the most useful information towards achieving the required objective. Encouraging an agent to ask the right questions is difficult because we don't know a-priori what information the agent will need to achieve its task, and we don't have an explicit model of what it knows already. We propose a solution to this problem based on a Bayesian model of the uncertainty in the implicit model maintained by the visual dialogue agent, and in the function used to select an appropriate output. By selecting the question that minimises the predicted regret with respect to this implicit model the agent actively reduces ambiguity. The Bayesian model of uncertainty also enables a principled method for identifying when enough information has been acquired, and an action should be selected. We evaluate our approach on two goal-oriented dialogue datasets, one for visual-based collaboration task and the other for a negotiation-based task. Our uncertainty-aware information-seeking model outperforms its counterparts in these two challenging problems. Ehsan Abbasnejad, Qi Wu 0001, Qinfeng Shi, Anton van den Hengel |
CVPR | 3 |
| 2019 | RGBD Based Dimensional Decomposition Residual Network for 3D Semantic Scene CompletionabstractRGB images differentiate from depth as they carry more details about the color and texture information, which can be utilized as a vital complement to depth for boosting the performance of 3D semantic scene completion (SSC). SSC is composed of 3D shape completion (SC) and semantic scene labeling while most of the existing approaches use depth as the sole input which causes the performance bottleneck. Moreover, the state-of-the-art methods employ 3D CNNs which have cumbersome networks and tremendous parameters. We introduce a light-weight Dimensional Decomposition Residual network (DDR) for 3D dense prediction tasks. The novel factorized convolution layer is effective for reducing the network parameters, and the proposed multi-scale fusion mechanism for depth and color image can improve the completion and segmentation accuracy simultaneously. Our method demonstrates excellent performance on two public datasets. Compared with the latest method SSCNet, we achieve 5.9% gains in SC-IoU and 5.7% gains in SSC-IOU, albeit with only 21% network parameters and 16.6% FLOPs employed compared with that of SSCNet. Jie Li 0040, Yu Liu 0029, Dong Gong, Qinfeng Shi, Xia Yuan, Chunxia Zhao, Ian D. Reid 0001 |
CVPR | 4 |
| 2019 | Variational Bayesian Dropout With a Hierarchical PriorabstractVariational dropout (VD) is a generalization of Gaussian dropout, which aims at inferring the posterior of network weights based on a log-uniform prior on them to learn these weights as well as dropout rate simultaneously. The log-uniform prior not only interprets the regularization capacity of Gaussian dropout in network training, but also underpins the inference of such posterior. However, the log-uniform prior is an improper prior (i.e., its integral is infinite), which causes the inference of posterior to be ill-posed, thus restricting the regularization performance of VD. To address this problem, we present a new generalization of Gaussian dropout, termed variational Bayesian dropout (VBD), which turns to exploit a hierarchical prior on the network weights and infer a new joint posterior. Specifically, we implement the hierarchical prior as a zero-mean Gaussian distribution with variance sampled from a uniform hyper-prior. Then, we incorporate such a prior into inferring the joint posterior over network weights and the variance in the hierarchical prior, with which both the network training and dropout rate estimation can be cast into a joint optimization problem. More importantly, the hierarchical prior is a proper prior which enables the inference of posterior to be well-posed. In addition, we further show that the proposed VBD can be seamlessly applied to network compression. Experiments on classification and network compression demonstrate the superior performance of the proposed VBD in regularizing network training. Yuhang Liu 0002, Wenyong Dong, Lei Zhang 0054, Dong Gong, Qinfeng Shi |
CVPR | 5 |
| 2019 | Attention-Guided Network for Ghost-Free High Dynamic Range ImagingabstractGhosting artifacts caused by moving objects or misalignments is a key challenge in high dynamic range (HDR) imaging for dynamic scenes. Previous methods first register the input low dynamic range (LDR) images using optical flow before merging them, which are error-prone and cause ghosts in results. A very recent work tries to bypass optical flows via a deep network with skip-connections, however, which still suffers from ghosting artifacts for severe movement. To avoid the ghosting from the source, we propose a novel attention-guided end-to-end deep neural network (AHDRNet) to produce high-quality ghost-free HDR images. Unlike previous methods directly stacking the LDR images or features for merging, we use attention modules to guide the merging according to the reference image. The attention modules automatically suppress undesired components caused by misalignments and saturation and enhance desirable fine details in the non-reference images. In addition to the attention model, we use dilated residual dense block (DRDB) to make full use of the hierarchical features and increase the receptive field for hallucinating the missing details. The proposed AHDRNet is a non-flow-based method, which can also avoid the artifacts generated by optical-flow estimation error. Experiments on different datasets show that the proposed AHDRNet can achieve state-of-the-art quantitative and qualitative results. Qingsen Yan, Dong Gong, Qinfeng Shi, Anton van den Hengel, Chunhua Shen, Ian D. Reid 0001, Yanning Zhang 0001 |
CVPR | 3 |
| 2019 | New Convex Relaxations for MRF Inference With Unknown GraphsabstractTreating graph structures of Markov random fields as unknown and estimating them jointly with labels have been shown to be useful for modeling human activity recognition and other related tasks. We propose two novel relaxations for solving this problem. The first is a linear programming (LP) relaxation, which is provably tighter than the existing LP relaxation. The second is a non-convex quadratic programming (QP) relaxation, which admits an efficient concave-convex procedure (CCCP). The CCCP algorithm is initialized by solving a convex QP relaxation of the problem, which is obtained by modifying the diagonal of the matrix that specifies the non-convex QP relaxation. We show that our convex QP relaxation is optimal in the sense that it minimizes the L1 norm of the diagonal modification vector. While the convex QP relaxation is not as tight as the existing and the new LP relaxations, when used in conjunction with the CCCP algorithm for the non-convex QP relaxation, it provides accurate solutions. We demonstrate the efficacy of our new relaxations for both synthetic data and human activity recognition. Zhenhua Wang 0003, Qinfeng Shi, M. Pawan Kumar, Jianhua Zhang 0002 |
ICCV | 3 |
| 2019 | Exploiting Stereo Sound Channels to Boost Performance of Neural Network-Based Music TranscriptionabstractIn recent years deep learning begins to show great potential for automatic music transcription that reproduces MIDI-like music composition information, such as note pitches and onset and offset times, from music recordings. In the literature without exception the two stereo sound channels coming with music recordings were averaged into a single channel to alleviate the computation overhead, which, from an entropy standpoint, definitely sacrifices information. In this paper we propose a method to properly combine the two sound channels for deep learning-based pitch detection. In particular, through modifying the loss function the network is forced to focus on the worse performing sound channel. This method achieves start-of-the-art frame-wise pitch detection performance on the MAPS dataset. Lingqiao Liu, Qinfeng Shi |
ICMLA | 3 |
| 2019 | Adaptive Neuro-Surrogate-Based Optimisation Method for Wave Energy Converters Placement Optimisation
Mehdi Neshat, Ehsan Abbasnejad, Qinfeng Shi, Bradley Alexander, Markus Wagner 0007 |
ICONIP (2) | 3 |
| 2019 | SparseSense: Human Activity Recognition from Highly Sparse Sensor Data-streams Using Set-based Neural NetworksabstractBatteryless or so called passive wearables are providing new and innovative methods for human activity recognition (HAR), especially in healthcare applications for older people. Passive sensors are low cost, lightweight, unobtrusive and desirably disposable; attractive attributes for healthcare applications in hospitals and nursing homes. Despite the compelling propositions for sensing applications, the data streams from these sensors are characterised by high sparsity---the time intervals between sensor readings are irregular while the number of readings per unit time are often limited. In this paper, we rigorously explore the problem of learning activity recognition models from temporally sparse data. We describe how to learn directly from sparse data using a deep learning paradigm in an end-to-end manner. We demonstrate significant classification performance improvements on real-world passive sensor datasets from older people over the state-of-the-art deep learning human activity recognition models. Further, we provide insights into the model's behaviour through complementary experiments on a benchmark dataset and visualisation of the learned activity feature spaces. Alireza Abedin Varamin, Seyed Hamid Rezatofighi, Qinfeng Shi, Damith Chinthana Ranasinghe |
IJCAI | 3 |
| 2019 | Multi-Scale Dense Networks for Deep High Dynamic Range ImagingabstractGenerating a high dynamic range (HDR) image from a set of sequential exposures is a challenging task for dynamic scenes. The most common approaches are aligning the input images to a reference image before merging them into an HDR image, but artifacts often appear in cases of large scene motion. The state-of-the-art method using deep learning can solve this problem effectively. In this paper, we propose a novel deep convolutional neural network to generate HDR, which attempts to produce more vivid images. The key idea of our method is using the coarse-to-fine scheme to gradually reconstruct the HDR image with the multi-scale architecture and residual network. By learning the relative changes of inputs and ground truth, our method can produce not only artificial free image but also restore missing information. Furthermore, we compare to existing methods for HDR reconstruction, and show high-quality results from a set of low dynamic range (LDR) images. We evaluate the results in qualitative and quantitative experiments, our method consistently produces excellent results than existing state-of-the-art approaches in challenging scenes. Qingsen Yan, Dong Gong, Qinfeng Shi, Jinqiu Sun, Ian D. Reid 0001, Yanning Zhang 0001 |
WACV | 4 |
| 2019 | Accurate imagery recovery using a multi-observation patch model
Lei Zhang 0054, Wei Wei 0008, Qinfeng Shi, Chunhua Shen, Anton van den Hengel, Yanning Zhang 0001 |
Inf. Sci. | 3 |
| 2019 | One-step adaptive markov random field for structured compressive sensing
Suwichaya Suwanwimolkul, Lei Zhang 0054, Damith Chinthana Ranasinghe, Qinfeng Shi |
Signal Process. | 4 |
| 2019 | Semantics-Aware Visual Object TrackingabstractIn this paper, we propose a semantics-aware visual object tracking method, which introduces semantics into the tracking procedure and extends the model of an object with explicit semantics prior to enhancing the robustness of three key aspects of the tracking framework, i.e., appearance model, search scheme, and scale adaptation. We first present a semantic object proposal generation method for video sequences to generate high-quality category-oriented object proposals. Then, a hybrid semantics-aware tracking algorithm with semantic compatibility is proposed. This algorithm takes full advantages of globally sparse semantic object proposal prediction and locally dense prediction with a template model and semantic distractor-aware color appearance model. Furthermore, we propose to exploit semantics to localize object accurately via an energy minimization framework-based scale adaptation method, which jointly integrates dense location prior, instance-specific color, and category-specific semantic information. Extensive experiments are conducted on two widely used benchmarks, and the results demonstrate that our method achieves the state-of-the-art performance. Rui Yao 0006, Guosheng Lin, Chunhua Shen, Yanning Zhang 0001, Qinfeng Shi |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2019 | MPTV: Matching Pursuit-Based Total Variation Minimization for Image DeconvolutionabstractTotal variation (TV) regularization has proven effective for a range of computer vision tasks through its preferential weighting of sharp image edges. Existing TV-based methods, however, often suffer from the over-smoothing issue and solution bias caused by the homogeneous penalization. In this paper, we consider addressing these issues by applying inhomogeneous regularization on different image components. We formulate the inhomogeneous TV minimization problem as a convex quadratic constrained linear programming problem. Relying on this new model, we propose a matching pursuit-based total variation minimization method (MPTV), specifically for image deconvolution. The proposed MPTV method is essentially a cutting-plane method that iteratively activates a subset of nonzero image gradients and then solves a subproblem focusing on those activated gradients only. Compared with existing methods, the MPTV is less sensitive to the choice of the trade-off parameter between data fitting and regularization. Moreover, the inhomogeneity of MPTV alleviates the over-smoothing and ringing artifacts and improves the robustness to errors in blur kernel. Extensive experiments on different tasks demonstrate the superiority of the proposed method over the current state of the art. Dong Gong, Mingkui Tan, Qinfeng Shi, Anton van den Hengel, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2019 | An Adaptive Markov Random Field for Structured Compressive SensingabstractExploiting intrinsic structures in sparse signals underpins the recent progress in compressive sensing (CS). The key for exploiting such structures is to achieve two desirable properties: generality (i.e., the ability to fit a wide range of signals with diverse structures) and adaptability (i.e., being adaptive to a specific signal). Most existing approaches, however, often only achieve one of these two properties. In this study, we propose a novel adaptive Markov random field sparsity prior for CS, which not only is able to capture a broad range of sparsity structures, but also can adapt to each sparse signal through refining the parameters of the sparsity prior with respect to the compressed measurements. To maximize the adaptability, we also propose a new sparse signal estimation where the sparse signals, support, noise and signal parameter estimation are unified into a variational optimization problem, which can be effectively solved with an alternative minimization scheme. Extensive experiments on three real-world datasets demonstrate the effectiveness of the proposed method in recovery accuracy, noise tolerance, and runtime. Suwichaya Suwanwimolkul, Lei Zhang 0054, Dong Gong, Zhen Zhang 0008, Chao Chen 0012, Damith Chinthana Ranasinghe, Qinfeng Shi |
IEEE Trans. Image Process. | 7 |
| 2019 | Auto-Embedding Generative Adversarial Networks For High Resolution Image SynthesisabstractGenerating images via a generative adversarial network (GAN) has attracted much attention recently. However, most of the existing GAN-based methods can only produce lowresolution images of limited quality. Directly generating highresolution images using GANs is nontrivial, and often produces problematic images with incomplete objects. To address this issue, we develop a novel GAN called auto-embedding generative adversarial network, which simultaneously encodes the global structure features and captures the fine-grained details. In our network, we use an autoencoder to learn the intrinsic high-level structure of real images and design a novel denoiser network to provide photo-realistic details for the generated images. In the experiments, we are able to produce 512 × 512 images of promising quality directly from the input noise. The resultant images exhibit better perceptual photo-realism, that is, with sharper structure and richer details, than other baselines on several datasets, including Oxford-102 Flowers, Caltech-UCSD Birds (CUB), High-Quality Large-scale CelebFaces Attributes (CelebAHQ), Large-scale Scene Understanding (LSUN), and ImageNet. Qi Chen 0014, Jian Chen 0011, Qingyao Wu, Qinfeng Shi, Mingkui Tan |
IEEE Trans. Multim. | 5 |
| 2018 | Joint Learning of Set Cardinality and State DistributionabstractWe present a novel approach for learning to predict sets using deep learning. In recent years, deep neural networks have shown remarkable results in computer vision, natural language processing and other related problems. Despite their success,traditional architectures suffer from a serious limitation in that they are built to deal with structured input and output data,i.e. vectors or matrices. Many real-world problems, however, are naturally described as sets, rather than vectors. Existing techniques that allow for sequential data, such as recurrent neural networks, typically heavily depend on the input and output order and do not guarantee a valid solution. Here, we derive in a principled way, a mathematical formulation for set prediction where the output is permutation invariant. In particular, our approach jointly learns both the cardinality and the state distribution of the target set. We demonstrate the validity of our method on the task of multi-label image classification and achieve a new state of the art on the PASCAL VOC and MS COCO datasets. Seyed Hamid Rezatofighi, Anton Milan, Qinfeng Shi, Anthony R. Dick, Ian D. Reid 0001 |
AAAI | 3 |
| 2018 | Active Learning from Noisy Tagged Images
Ehsan Abbasnejad, Anthony R. Dick, Qinfeng Shi, Anton van den Hengel |
BMVC | 3 |
| 2018 | Deblurring Natural Image Using Super-Gaussian Fields
Yuhang Liu 0002, Wenyong Dong, Dong Gong, Lei Zhang 0054, Qinfeng Shi |
ECCV (1) | 5 |
| 2018 | Seeing Deeply and Bidirectionally: A Deep Learning Approach for Single Image Reflection Removal
Jie Yang 0002, Dong Gong, Lingqiao Liu, Qinfeng Shi |
ECCV (3) | 4 |
| 2018 | Deep Auto-Set: A Deep Auto-Encoder-Set Network for Activity Recognition Using WearablesabstractAutomatic recognition of human activities from time-series sensor data (referred to as HAR) is a growing area of research in ubiquitous computing. Most recent research in the field adopts supervised deep learning paradigms to automate extraction of intrinsic features from raw signal inputs and addresses HAR as a multi-class classification problem where detecting a single activity class within the duration of a sensory data segment suffices. However, due to the innate diversity of human activities and their corresponding duration, no data segment is guaranteed to contain sensor recordings of a single activity type. In this paper, we express HAR more naturally as a set prediction problem where the predictions are sets of ongoing activity elements with unfixed and unknown cardinality. For the first time, we address this problem by presenting a novel HAR approach that learns to output activity sets using deep neural networks. Moreover, motivated by the limited availability of annotated HAR datasets as well as the unfortunate immaturity of existing unsupervised systems, we complement our supervised set learning scheme with a prior unsupervised feature learning process that adopts convolutional auto-encoders to exploit unlabeled data. The empirical experiments on two widely adopted HAR datasets demonstrate the substantial improvement of our proposed methodology over the baseline models. Alireza Abedin Varamin, Ehsan Abbasnejad, Qinfeng Shi, Damith Chinthana Ranasinghe, Seyed Hamid Rezatofighi |
MobiQuitous | 3 |
| 2018 | Cluster Sparsity Field: An Internal Hyperspectral Imagery Prior for Reconstruction
Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001, Chunhua Shen, Anton van den Hengel, Qinfeng Shi |
Int. J. Comput. Vis. | 6 |
| 2018 | Efficient dense labelling of human activity sequences from wearables using fully convolutional networks
Rui Yao 0006, Guosheng Lin, Qinfeng Shi, Damith Chinthana Ranasinghe |
Pattern Recognit. | 3 |
| 2017 | MPGL: An Efficient Matching Pursuit Method for Generalized LASSOabstractUnlike traditional LASSO enforcing sparsity on the variables, Generalized LASSO (GL) enforces sparsity on a linear transformation of the variables, gaining flexibility and success in many applications. However, many existing GL algorithms do not scale up to high-dimensional problems, and/or only work well for a specific choice of the transformation. We propose an efficient Matching Pursuit Generalized LASSO (MPGL) method, which overcomes these issues, and is guaranteed to converge to a global optimum. We formulate the GL problem as a convex quadratic constrained linear programming (QCLP) problem and tailor-make a cutting plane method. More specifically, our MPGL iteratively activates a subset of nonzero elements of the transformed variables, and solves a subproblem involving only the activated elements thus gaining significant speed-up. Moreover, MPGL is less sensitive to the choice of the trade-off hyper-parameter between data fitting and regularization, and mitigates the long-standing hyper-parameter tuning issue in many existing methods. Experiments demonstrate the superior efficiency and accuracy of the proposed method over the state-of-the-arts in both classification and image processing tasks. Dong Gong, Mingkui Tan, Yanning Zhang 0001, Anton van den Hengel, Qinfeng Shi |
AAAI | 5 |
| 2017 | Solving Constrained Combinatorial Optimisation Problems via MAP Inference without High-Order PenaltiesabstractSolving constrained combinatorial optimisation problems via MAP inference is often achieved by introducing extra potential functions for each constraint. This can result in very high order potentials, e.g. a 2nd-order objective with pairwise potentials and a quadratic constraint over all N variables would correspond to an unconstrained objective with an order-N potential. This limits the practicality of such an approach, since inference with high order potentials is tractable only for a few special classes of functions. We propose an approach which is able to solve constrained combinatorial problems using belief propagation without increasing the order. For example, in our scheme the 2nd-order problem above remains order 2 instead of order N. Experiments on applications ranging from foreground detection, image reconstruction, quadratic knapsack, and the M-best solutions problem demonstrate the effectiveness and efficiency of our method. Moreover, we show several situations in which our approach outperforms commercial solvers like CPLEX and others designed for specific constrained MAP inference problems. Zhen Zhang 0008, Qinfeng Shi, Julian J. McAuley, Wei Wei 0008, Yanning Zhang 0001, Rui Yao 0006, Anton van den Hengel |
AAAI | 2 |
| 2017 | From Motion Blur to Motion Flow: A Deep Learning Solution for Removing Heterogeneous Motion BlurabstractRemoving pixel-wise heterogeneous motion blur is challenging due to the ill-posed nature of the problem. The predominant solution is to estimate the blur kernel by adding a prior, but extensive literature on the subject indicates the difficulty in identifying a prior which is suitably informative, and general. Rather than imposing a prior based on theory, we propose instead to learn one from the data. Learning a prior over the latent image would require modeling all possible image content. The critical observation underpinning our approach, however, is that learning the motion flow instead allows the model to focus on the cause of the blur, irrespective of the image content. This is a much easier learning task, but it also avoids the iterative process through which latent image priors are typically applied. Our approach directly estimates the motion flow from the blurred image through a fully-convolutional deep neural network (FCN) and recovers the unblurred image from the estimated motion flow. Our FCN is the first universal end-to-end mapping from the blurred image to the dense motion flow. To train the FCN, we simulate motion flows to generate synthetic blurred-image-motion-flow pairs thus avoiding the need for human labeling. Extensive experiments on challenging realistic blurred images demonstrate that the proposed method outperforms the state-of-the-art. Dong Gong, Jie Yang 0002, Lingqiao Liu, Yanning Zhang 0001, Ian D. Reid 0001, Chunhua Shen, Anton van den Hengel, Qinfeng Shi |
CVPR | 8 |
| 2017 | Self-Paced Kernel Estimation for Robust Blind Image DeblurringabstractThe challenge in blind image deblurring is to remove the effects of blur with limited prior information about the nature of the blur process. Existing methods often assume that the blur image is produced by linear convolution with additive Gaussian noise. However, including even a small number of outliers to this model in the kernel estimation process can significantly reduce the resulting image quality. Previous methods mainly rely on some simple but unreliable heuristics to identify outliers for kernel estimation. Rather than attempt to identify outliers to the model a priori, we instead propose to sequentially identify inliers, and gradually incorporate them into the estimation process. The self-paced kernel estimation scheme we propose represents a generalization of existing self-paced learning approaches, in which we gradually detect and include reliable inlier pixel sets in a blurred image for kernel estimation. Moreover, we automatically activate a subset of significant gradients w.r.t. the reliable inlier pixels, and then update the intermediate sharp image and the kernel accordingly. Experiments on both synthetic data and real-world images with various kinds of outliers demonstrate the effectiveness and robustness of the proposed method compared to the stateof- the-art methods. Dong Gong, Mingkui Tan, Yanning Zhang 0001, Anton van den Hengel, Qinfeng Shi |
ICCV | 5 |
| 2017 | Dynamic Programming Bipartite Belief Propagation For Hyper Graph MatchingabstractHyper graph matching problems have drawn attention recently due to their ability to embed higher order relations between nodes. In this paper, we formulate hyper graph matching problems as constrained MAP inference problems in graphical models. Whereas previous discrete approaches introduce several global correspondence vectors, we introduce only one global correspondence vector, but several local correspondence vectors. This allows us to decompose the problem into a (linear) bipartite matching problem and several belief propagation sub-problems. Bipartite matching can be solved by traditional approaches, while the belief propagation sub-problem is further decomposed as two sub-problems with optimal substructure. Then a newly proposed dynamic programming procedure is used to solve the belief propagation sub-problem. Experiments show that the proposed methods outperform state-of-the-art techniques for hyper graph matching. Zhen Zhang 0008, Julian J. McAuley, Wei Wei 0008, Yanning Zhang 0001, Qinfeng Shi |
IJCAI | 6 |
| 2017 | A hierarchical model for recognizing alarming states in a batteryless sensor alarm intervention for preventing falls in older people
Roberto Luis Shinmoto Torres, Qinfeng Shi, Anton van den Hengel, Damith Chinthana Ranasinghe |
Pervasive Mob. Comput. | 2 |
| 2017 | Part-Based Robust Tracking Using Online Latent Structured LearningabstractDespite many advances made in the area, deformable targets and partial occlusions continue to represent key problems in visual tracking. Structured learning has shown good results when applied to tracking whole targets, but applying this approach to a part-based target model is complicated by the need to model the relationships between parts and to avoid lengthy initialization processes. We thus propose a method that models the unknown parts by using latent variables. In doing so, we extend the online algorithm Pegasos to the structured prediction case (i.e., predicting the location of the bounding boxes) with latent part variables. We also incorporate the very recently proposed spatial constraints to preserve distances between parts. To better estimate the parts, and to avoid overfitting caused by the extra model complexity/capacity introduced by the parts, we propose a two-stage training process based on the primal rather than the dual form. We then show that the method outperforms the state of the arts in extensive experiments. Rui Yao 0006, Qinfeng Shi, Chunhua Shen, Yanning Zhang 0001, Anton van den Hengel |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Learning Sparse Confidence-Weighted Classifier on Very High Dimensional DataabstractConfidence-weighted (CW) learning is a successful online learning paradigm which maintains a Gaussian distribution over classifier weights and adopts a covariancematrix to represent the uncertainties of the weight vectors. However, there are two deficiencies in existing full CW learning paradigms, these being the sensitivity to irrelevant features, and the poor scalability to high dimensional data due to the maintenance of the covariance structure. In this paper, we begin by presenting an online-batch CW learning scheme, and then present a novel paradigm to learn sparse CW classifiers. The proposed paradigm essentially identifies feature groups and naturally builds a block diagonal covariance structure, making it very suitable for CW learning over very high-dimensional data.Extensive experimental results demonstrate the superior performance of the proposed methods over state-of-the-art counterparts on classification and feature selection tasks. Mingkui Tan, Yan Yan 0006, Li Wang 0033, Anton van den Hengel, Ivor W. Tsang, Qinfeng Shi |
AAAI | 6 |
| 2016 | Efficient Orthogonal Non-negative Matrix Factorization over Stiefel ManifoldabstractOrthogonal Non-negative Matrix Factorization (ONMF) approximates a data matrix X by the product of two lower dimensional factor matrices: X -- UVT, with one of them orthogonal. ONMF has been widely applied for clustering, but it often suffers from high computational cost due to the orthogonality constraint. In this paper, we propose a method, called Nonlinear Riemannian Conjugate Gradient ONMF (NRCG-ONMF), which updates U and V alternatively and preserves the orthogonality of U while achieving fast convergence speed. Specifically, in order to update U, we develop a Nonlinear Riemannian Conjugate Gradient (NRCG) method on the Stiefel manifold using Barzilai-Borwein (BB) step size. For updating V, we use a closed-form solution under non-negativity constraint. Extensive experiments on both synthetic and real-world data sets show consistent superiority of our method over other approaches in terms of orthogonality preservation, convergence speed and clustering performance. Wei Zhang 0098, Mingkui Tan, Quan Z. Sheng, Lina Yao 0001, Qinfeng Shi |
CIKM | 5 |
| 2016 | Blind Image Deconvolution by Automatic Gradient ActivationabstractBlind image deconvolution is an ill-posed inverse problem which is often addressed through the application of appropriate prior. Although some priors are informative in general, many images do not strictly conform to this, leading to degraded performance in the kernel estimation. More critically, real images may be contaminated by nonuniform noise such as saturation and outliers. Methods for removing specific image areas based on some priors have been proposed, but they operate either manually or by defining fixed criteria. We show here that a subset of the image gradients are adequate to estimate the blur kernel robustly, no matter the gradient image is sparse or not. We thus introduce a gradient activation method to automatically select a subset of gradients of the latent image in a cutting-plane-based optimization scheme for kernel estimation. No extra assumption is used in our model, which greatly improves the accuracy and flexibility. More importantly, the proposed method affords great convenience for handling noise and outliers. Experiments on both synthetic data and real-world images demonstrate the effectiveness and robustness of the proposed method in comparison with the state-of-the-art methods. Dong Gong, Mingkui Tan, Yanning Zhang 0001, Anton van den Hengel, Qinfeng Shi |
CVPR | 5 |
| 2016 | Joint Probabilistic Matching Using m-Best SolutionsabstractMatching between two sets of objects is typically approached by finding the object pairs that collectively maximize the joint matching score. In this paper, we argue that this single solution does not necessarily lead to the optimal matching accuracy and that general one-to-one assignment problems can be improved by considering multiple hypotheses before computing the final similarity measure. To that end, we propose to utilize the marginal distributionsfor each entity. Previously, this idea has been neglected mainly because exact marginalization is intractable due to a combinatorial number of all possible matching permutations. Here, we propose a generic approach to efficiently approximate the marginal distributions by exploiting the m-best solutions of the original problem. This approach not only improves the matching solution, but also provides more accurate ranking of the results, because of the extra information included in the marginal distribution. We validate our claim on two distinct objectives: (i) person re-identification and temporal matching modeled as an integer linear program, and (ii) feature point matching using a quadratic cost function. Our experiments confirm that marginalization indeed leads to superior performance compared to the single (nearly) optimal solution, yielding state-of-the-art results in both applications on standard benchmarks. Seyed Hamid Rezatofighi, Anton Milan, Zhen Zhang 0008, Qinfeng Shi, Anthony R. Dick, Ian D. Reid 0001 |
CVPR | 4 |
| 2016 | Proximal Riemannian Pursuit for Large-Scale Trace-Norm MinimizationabstractTrace-norm regularization plays an important role in many areas such as computer vision and machine learning. When solving general large-scale trace-norm regularized problems, existing methods may be computationally expensive due to many high-dimensional truncated singular value decompositions (SVDs) or the unawareness of matrix ranks. In this paper, we propose a proximal Riemannian pursuit (PRP) paradigm which addresses a sequence of trace-norm regularized subproblems defined on nonlinear matrix varieties. To address the subproblem, we extend the proximal gradient method on vector space to nonlinear matrix varieties, in which the SVDs of intermediate solutions are maintained by cheap low-rank QR decompositions, therefore making the proposed method more scalable. Empirical studies on several tasks, such as matrix completion and low-rank representation based subspace clustering, demonstrate the competitive performance of the proposed paradigms over existing methods. Mingkui Tan, Shijie Xiao, Junbin Gao, Dong Xu 0001, Anton van den Hengel, Qinfeng Shi |
CVPR | 6 |
| 2016 | Pairwise Matching through Max-Weight Bipartite Belief PropagationabstractFeature matching is a key problem in computer vision and pattern recognition. One way to encode the essential interdependence between potential feature matches is to cast the problem as inference in a graphical model, though recently alternatives such as spectral methods, or approaches based on the convex-concave procedure have achieved the state-of-the-art. Here we revisit the use of graphical models for feature matching, and propose a belief propagation scheme which exhibits the following advantages: (1) we explicitly enforce one-to-one matching constraints, (2) we offer a tighter relaxation of the original cost function than previous graphical-model-based approaches, and (3) our sub-problems decompose into max-weight bipartite matching, which can be solved efficiently, leading to orders-of-magnitude reductions in execution time. Experimental results show that the proposed algorithm produces results superior to those of the current state-of-the-art. Zhen Zhang 0008, Qinfeng Shi, Julian J. McAuley, Wei Wei 0008, Yanning Zhang 0001, Anton van den Hengel |
CVPR | 2 |
| 2016 | Cluster Sparsity Field for Hyperspectral Imagery Denoising
Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001, Chunhua Shen, Anton van den Hengel, Qinfeng Shi |
ECCV (5) | 6 |
| 2016 | Dictionary Learning for Promoting Structured Sparsity in Hyperspectral Compressive SensingabstractThe ability to accurately represent a hyperspectral image (HSI) as a combination of a small number of elements from an appropriate dictionary underpins much of the recent progress in hyperspectral compressive sensing (HCS). Preserving structure in the sparse representation is critical to achieving an accurate reconstruction but has thus far only been partially exploited because existing methods assume a predefined dictionary. To address this problem, a structured sparsity-based hyperspectral blind compressive sensing method is presented in this study. For the reconstructed HSI, a data-adaptive dictionary is learned directly from its noisy measurements, which promotes the underlying structured sparsity and obviously improves reconstruction accuracy. Specifically, a fully structured dictionary prior is first proposed to jointly depict the structure in each dictionary atom as well as the correlation between atoms, where the magnitude of each atom is also regularized. Then, a reweighted Laplace prior is employed to model the structured sparsity in the representation of the HSI. Based on these two priors, a unified optimization framework is proposed to learn both the dictionary and sparse representation from the measurements by alternatively optimizing two separate latent variable Bayes models. With the learned dictionary, the structured sparsity of HSIs can be well described by the reweighted Laplace prior. In addition, both the learned dictionary and sparse representation are robust to noise corruption in the measurements. Extensive experiments on three hyperspectral data sets demonstrate that the proposed method outperforms several state-of-the-art HCS methods in terms of the reconstruction accuracy achieved. Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001, Chunhua Shen, Anton van den Hengel, Qinfeng Shi |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2015 | Learning graph structure for multi-label image classification via clique generationabstractExploiting label dependency for multi-label image classification can significantly improve classification performance. Probabilistic Graphical Models are one of the primary methods for representing such dependencies. The structure of graphical models, however, is either determined heuristically or learned from very limited information. Moreover, neither of these approaches scales well to large or complex graphs. We propose a principled way to learn the structure of a graphical model by considering input features and labels, together with loss functions. We formulate this problem into a max-margin framework initially, and then transform it into a convex programming problem. Finally, we propose a highly scalable procedure that activates a set of cliques iteratively. Our approach exhibits both strong theoretical properties and a significant performance improvement over state-of-the-art methods on both synthetic and real-world data sets. Mingkui Tan, Qinfeng Shi, Anton van den Hengel, Chunhua Shen, Junbin Gao, Fuyuan Hu, Zhen Zhang 0008 |
CVPR | 2 |
| 2015 | Joint Probabilistic Data Association RevisitedabstractIn this paper, we revisit the joint probabilistic data association (JPDA) technique and propose a novel solution based on recent developments in finding the m-best solutions to an integer linear program. The key advantage of this approach is that it makes JPDA computationally tractable in applications with high target and/or clutter density, such as spot tracking in fluorescence microscopy sequences and pedestrian tracking in surveillance footage. We also show that our JPDA algorithm embedded in a simple tracking framework is surprisingly competitive with state-of-the-art global tracking methods in these two applications, while needing considerably less processing time. Seyed Hamid Rezatofighi, Anton Milan, Zhen Zhang 0008, Qinfeng Shi, Anthony R. Dick, Ian D. Reid 0001 |
ICCV | 4 |
| 2015 | Hyperspectral Compressive Sensing Using Manifold-Structured Sparsity PriorabstractTo reconstruct hyperspectral image (HSI) accurately from a few noisy compressive measurements, we present a novel manifold-structured sparsity prior based hyperspectral compressive sensing (HCS) method in this study. A matrix based hierarchical prior is first proposed to represent the spectral structured sparsity and spatial unknown manifold structure of HSI simultaneously. Then, a latent variable Bayes model is introduced to learn the sparsity prior and estimate the noise jointly from measurements. The learned prior can fully represent the inherent 3D structure of HSI and regulate its shape based on the estimated noise level. Thus, with this learned prior, the proposed method improves the reconstruction accuracy significantly and shows strong robustness to unknown noise in HCS. Experiments on four real hyperspectral datasets show that the proposed method outperforms several state-of-the-art methods on the reconstruction accuracy of HSI. Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001, Fei Li 0011, Chunhua Shen, Qinfeng Shi |
ICCV | 6 |
| 2015 | Scalable Maximum Margin Matrix Factorization by Active Riemannian Subspace Search
Yan Yan 0006, Mingkui Tan, Ivor W. Tsang, Yi Yang 0001, Chengqi Zhang, Qinfeng Shi |
IJCAI | 6 |
| 2015 | Image-Based Recommendations on Styles and SubstitutesabstractHumans inevitably develop a sense of the relationships between objects, some of which are based on their appearance. Some pairs of objects might be seen as being alternatives to each other (such as two pairs of jeans), while others may be seen as being complementary (such as a pair of jeans and a matching shirt). This information guides many of the choices that people make, from buying clothes to their interactions with each other. We seek here to model this human sense of the relationships between objects based on their appearance. Our approach is not based on fine-grained modeling of user annotations but rather on capturing the largest dataset possible and developing a scalable method for uncovering human notions of the visual relationships within. We cast this as a network inference problem defined on graphs of related images, and provide a large-scale dataset for the training and evaluation of the same. The system we develop is capable of recommending which clothes and accessories will go well together (and which will not), amongst a host of other applications. Julian J. McAuley, Christopher Targett, Qinfeng Shi, Anton van den Hengel |
SIGIR | 3 |
| 2015 | A Hybrid Loss for Multiclass and Structured PredictionabstractWe propose a novel hybrid loss for multiclass and structured prediction problems that is a convex combination of a log loss for Conditional Random Fields (CRFs) and a multiclass hinge loss for Support Vector Machines (SVMs). We provide a sufficient condition for when the hybrid loss is Fisher consistent for classification. This condition depends on a measure of dominance between labels-specifically, the gap between the probabilities of the best label and the second best label. We also prove Fisher consistency is necessary for parametric consistency when learning models such as CRFs. We demonstrate empirically that the hybrid loss typically performs least as well as-and often better than-both of its constituent losses on a variety of tasks, such as human action recognition. In doing so we also provide an empirical comparison of the efficacy of probabilistic and margin based approaches to multiclass and structured prediction. Qinfeng Shi, Mark D. Reid, Tibério S. Caetano, Anton van den Hengel, Zhenhua Wang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Worst Case Linear Discriminant Analysis as Scalable Semidefinite Feasibility ProblemsabstractIn this paper, we propose an efficient semidefinite programming (SDP) approach to worst case linear discriminant analysis (WLDA). Compared with the traditional LDA, WLDA considers the dimensionality reduction problem from the worst case viewpoint, which is in general more robust for classification. However, the original problem of WLDA is non-convex and difficult to optimize. In this paper, we reformulate the optimization problem of WLDA into a sequence of semidefinite feasibility problems. To efficiently solve the semidefinite feasibility problems, we design a new scalable optimization method with a quasi-Newton method and eigen-decomposition being the core components. The proposed method is orders of magnitude faster than standard interior-point SDP solvers. Experiments on a variety of classification problems demonstrate that our approach achieves better performance than standard LDA. Our method is also much faster and more scalable than standard interior-point SDP solvers-based WLDA. The computational complexity for an SDP with m constraints and matrices of size d by d is roughly reduced from O(m(3)+md(3)+m(2)d(2)) to O(d(3)) (m>d in our case). Hui Li 0031, Chunhua Shen, Anton van den Hengel, Qinfeng Shi |
IEEE Trans. Image Process. | 4 |
| 2015 | Hashing on Nonlinear ManifoldsabstractLearning-based hashing methods have attracted considerable attention due to their ability to greatly increase the scale at which existing algorithms may operate. Most of these methods are designed to generate binary codes preserving the Euclidean similarity in the original space. Manifold learning techniques, in contrast, are better able to model the intrinsic structure embedded in the original high-dimensional data. The complexities of these models, and the problems with out-of-sample data, have previously rendered them unsuitable for application to large-scale embedding, however. In this paper, how to learn compact binary embeddings on their intrinsic manifolds is considered. In order to address the above-mentioned difficulties, an efficient, inductive solution to the out-of-sample data problem, and a process by which nonparametric manifold learning may be used as the basis of a hashing method are proposed. The proposed approach thus allows the development of a range of new hashing techniques exploiting the flexibility of the wide variety of manifold learning approaches available. It is particularly shown that hashing on the basis of t-distributed stochastic neighbor embedding outperforms state-of-the-art hashing methods on large-scale benchmark data sets, and is very effective for image classification with very short code lengths. It is shown that the proposed framework can be further improved, for example, by minimizing the quantization error with learned orthogonal rotations without much computation overhead. In addition, a supervised inductive manifold hashing framework is developed by incorporating the label information, which is shown to greatly advance the semantic retrieval performance. Fumin Shen, Chunhua Shen, Qinfeng Shi, Anton van den Hengel, Zhenmin Tang, Heng Tao Shen |
IEEE Trans. Image Process. | 3 |
| 2015 | Dual Graph Regularized Latent Low-Rank Representation for Subspace ClusteringabstractLow-rank representation (LRR) has received considerable attention in subspace segmentation due to its effectiveness in exploring low-dimensional subspace structures embedded in data. To preserve the intrinsic geometrical structure of data, a graph regularizer has been introduced into LRR framework for learning the locality and similarity information within data. However, it is often the case that not only the high-dimensional data reside on a non-linear low-dimensional manifold in the ambient space, but also their features lie on a manifold in feature space. In this paper, we propose a dual graph regularized LRR model (DGLRR) by enforcing preservation of geometric information in both the ambient space and the feature space. The proposed method aims for simultaneously considering the geometric structures of the data manifold and the feature manifold. Furthermore, we extend the DGLRR model to include non-negative constraint, leading to a parts-based representation of data. Experiments are conducted on several image data sets to demonstrate that the proposed method outperforms the state-of-the-art approaches in image clustering. Ming Yin 0002, Junbin Gao, Zhouchen Lin, Qinfeng Shi, Yi Guo 0001 |
IEEE Trans. Image Process. | 4 |
| 2014 | Fast Supervised Hashing with Decision Trees for High-Dimensional DataabstractSupervised hashing aims to map the original features to compact binary codes that are able to preserve label based similarity in the Hamming space. Non-linear hash functions have demonstrated their advantage over linear ones due to their powerful generalization capability. In the literature, kernel functions are typically used to achieve non-linearity in hashing, which achieve encouraging retrieval perfor- mance at the price of slow evaluation and training time. Here we propose to use boosted decision trees for achieving non-linearity in hashing, which are fast to train and evaluate, hence more suitable for hashing with high dimensional data. In our approach, we first propose sub-modular formulations for the hashing binary code inference problem and an efficient GraphCut based block search method for solving large-scale inference. Then we learn hash func- tions by training boosted decision trees to fit the binary codes. Experiments demonstrate that our proposed method significantly outperforms most state-of-the-art methods in retrieval precision and training time. Especially for high- dimensional data, our method is orders of magnitude faster than many methods in terms of training time. Guosheng Lin, Chunhua Shen, Qinfeng Shi, Anton van den Hengel, David Suter |
CVPR | 3 |
| 2014 | RandomBoost: Simplified Multiclass Boosting Through RandomizationabstractWe propose a novel boosting approach to multiclass classification problems, in which multiple classes are distinguished by a set of random projection matrices in essence. The approach uses random projections to alleviate the proliferation of binary classifiers typically required to perform multiclass classification. The result is a multiclass classifier with a single vector-valued parameter, irrespective of the number of classes involved. Two variants of this approach are proposed. The first method randomly projects the original data into new spaces, while the second method randomly projects the outputs of learned weak classifiers. These methods are not only conceptually simple but also effective and easy to implement. A series of experiments on synthetic, machine learning, and visual recognition data sets demonstrate that our proposed methods could be compared favorably with existing multiclass boosting algorithms in terms of both the convergence rate and classification accuracy. Sakrapee Paisitkriangkrai, Chunhua Shen, Qinfeng Shi, Anton van den Hengel |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2013 | Inductive Hashing on ManifoldsabstractLearning based hashing methods have attracted considerable attention due to their ability to greatly increase the scale at which existing algorithms may operate. Most of these methods are designed to generate binary codes that preserve the Euclidean distance in the original space. Manifold learning techniques, in contrast, are better able to model the intrinsic structure embedded in the original high-dimensional data. The complexity of these models, and the problems with out-of-sample data, have previously rendered them unsuitable for application to large-scale embedding, however. In this work, we consider how to learn compact binary embeddings on their intrinsic manifolds. In order to address the above-mentioned difficulties, we describe an efficient, inductive solution to the out-of-sample data problem, and a process by which non-parametric manifold learning may be used as the basis of a hashing method. Our proposed approach thus allows the development of a range of new hashing techniques exploiting the flexibility of the wide variety of manifold learning approaches available. We particularly show that hashing on the basis of t-SNE [29] outperforms state-of-the-art hashing methods on large-scale benchmark datasets, and is very effective for image classification with very short code lengths. Fumin Shen, Chunhua Shen, Qinfeng Shi, Anton van den Hengel, Zhenmin Tang |
CVPR | 3 |
| 2013 | Bilinear Programming for Human Activity Recognition with Unknown MRF GraphsabstractMarkov Random Fields (MRFs) have been successfully applied to human activity modelling, largely due to their ability to model complex dependencies and deal with local uncertainty. However, the underlying graph structure is often manually specified, or automatically constructed by heuristics. We show, instead, that learning an MRF graph and performing MAP inference can be achieved simultaneously by solving a bilinear program. Equipped with the bilinear program based MAP inference for an unknown graph, we show how to estimate parameters efficiently and effectively with a latent structural SVM. We apply our techniques to predict sport moves (such as serve, volley in tennis) and human activity in TV episodes (such as kiss, hug and Hi-Five). Experimental results show the proposed method outperforms the state-of-the-art. Zhenhua Wang 0003, Qinfeng Shi, Chunhua Shen, Anton van den Hengel |
CVPR | 2 |
| 2013 | Part-Based Visual Tracking with Online Latent Structural LearningabstractDespite many advances made in the area, deformable targets and partial occlusions continue to represent key problems in visual tracking. Structured learning has shown good results when applied to tracking whole targets, but applying this approach to a part-based target model is complicated by the need to model the relationships between parts, and to avoid lengthy initialisation processes. We thus propose a method which models the unknown parts using latent variables. In doing so we extend the online algorithm pegasos to the structured prediction case (i.e., predicting the location of the bounding boxes) with latent part variables. To better estimate the parts, and to avoid over-fitting caused by the extra model complexity/capacity introduced by the parts, we propose a two-stage training process, based on the primal rather than the dual form. We then show that the method outperforms the state-of-the-art (linear and non-linear kernel) trackers. Rui Yao 0006, Qinfeng Shi, Chunhua Shen, Yanning Zhang 0001, Anton van den Hengel |
CVPR | 2 |
| 2013 | Evaluation of Wearable Sensor Tag Data Segmentation Approaches for Real Time Activity Classification in Elderly
Roberto Luis Shinmoto Torres, Damith Chinthana Ranasinghe, Qinfeng Shi |
MobiQuitous | 3 |
| 2012 | Non-sparse linear representations for visual tracking with online reservoir metric learningabstractMost sparse linear representation-based trackers need to solve a computationally expensive li-regularized optimization problem. To address this problem, we propose a visual tracker based on non-sparse linear representations, which admit an efficient closed-form solution without sacrificing accuracy. Moreover, in order to capture the correlation information between different feature dimensions, we learn a Mahalanobis distance metric in an online fashion and incorporate the learned metric into the optimization problem for obtaining the linear representation. We show that online metric learning using proximity comparison significantly improves the robustness of the tracking, especially on those sequences exhibiting drastic appearance changes. Furthermore, in order to prevent the unbounded growth in the number of training samples for the metric learning, we design a time-weighted reservoir sampling method to maintain and update limited-sized foreground and background sample buffers for balancing sample diversity and adaptability. Experimental results on challenging videos demonstrate the effectiveness and robustness of the proposed tracker. Xi Li 0001, Chunhua Shen, Qinfeng Shi, Anthony R. Dick, Anton van den Hengel |
CVPR | 3 |
| 2012 | Robust Tracking with Weighted Online Structured Learning
Rui Yao 0006, Qinfeng Shi, Chunhua Shen, Yanning Zhang 0001, Anton van den Hengel |
ECCV (3) | 2 |
| 2012 | Is margin preserved after random projection?
Qinfeng Shi, Chunhua Shen, Rhys Hill, Anton van den Hengel |
ICML | 1 |
| 2012 | Dimensionality reduction via compressive sensing
Junbin Gao, Qinfeng Shi, Tibério S. Caetano |
Pattern Recognit. Lett. | 2 |
| 2011 | Real-time visual tracking using compressive sensingabstractThe ℓ1tracker obtains robustness by seeking a sparse representation of the tracking object via ℓ1norm minimization. However, the high computational complexity involved in the ℓ1tracker may hamper its applications in real-time processing scenarios. Here we propose Real-time Com-pressive Sensing Tracking (RTCST) by exploiting the signal recovery power of Compressive Sensing (CS). Dimensionality reduction and a customized Orthogonal Matching Pursuit (OMP) algorithm are adopted to accelerate the CS tracking. As a result, our algorithm achieves a realtime speed that is up to 5,000 times faster than that of the ℓ1tracker. Meanwhile, RTCST still produces competitive (sometimes even superior) tracking accuracy compared to the ℓ1tracker. Furthermore, for a stationary camera, a refined tracker is designed by integrating a CS-based background model (CSBM) into tracking. This CSBM-equipped tracker, termed RTCST-B, outperforms most state-of-the-art trackers in terms of both accuracy and robustness. Finally, our experimental results on various video sequences, which are verified by a new metric - Tracking Success Probability (TSP), demonstrate the excellence of the proposed algorithms. Chunhua Shen, Qinfeng Shi |
CVPR | 3 |
| 2011 | Is face recognition really a Compressive Sensing problem?abstractCompressive Sensing has become one of the standard methods of face recognition within the literature. We show, however, that the sparsity assumption which underpins much of this work is not supported by the data. This lack of sparsity in the data means that compressive sensing approach cannot be guaranteed to recover the exact signal, and therefore that sparse approximations may not deliver the robustness or performance desired. In this vein we show that a simple ℓ2approach to the face recognition problem is not only significantly more accurate than the state-of-the-art approach, it is also more robust, and much faster. These results are demonstrated on the publicly available YaleB and AR face datasets but have implications for the application of Compressive Sensing more broadly. Qinfeng Shi, Anders P. Eriksson, Anton van den Hengel, Chunhua Shen |
CVPR | 1 |
| 2011 | Human Action Segmentation and Recognition Using Discriminative Semi-Markov Models
Qinfeng Shi, Li Cheng 0001, Li Wang 0033, Alexander J. Smola |
Int. J. Comput. Vis. | 1 |
| 2010 | Rapid face recognition using hashingabstractWe propose a face recognition approach based on hashing. The approach yields comparable recognition rates with the random ℓ1approach, which is considered the state-of-the-art. But our method is much faster: it is up to 150 times faster than on the YaleB dataset. We show that with hashing, the sparse representation can be recovered with a high probability because hashing preserves the restrictive isometry property. Moreover, we present a theoretical analysis on the recognition rate of the proposed hashing approach. Experiments show a very competitive recognition rate and significant speedup compared with the state-of-the-art. Qinfeng Shi, Chunhua Shen |
CVPR | 1 |
| 2009 | Discriminative Maximum Margin Image Object Categorization with Exact InferenceabstractCategorizing multiple objects in images is essentially a structured prediction problem: the label of an object is in general dependent on the labels of other objects in the image. We explicitly model object dependencies in a sparse graphical topology induced by the adjacency of objects in the image, which benefits inference, and then use maximum margin principle to learn the model discriminatively. Moreover, we propose a novel exact inference method, which is used in training to find the most violated constraint required by cutting plane method. A slightly modified inference method is used in testing when the target labels are unseen. Experiment results on both synthetic and real datasets demonstrate the improvement of the proposed approach over the state-of-the-art methods. Qinfeng Shi, Luping Zhou, Li Cheng 0001, Dale Schuurmans |
ICIG | 1 |
| 2009 | Hash Kernels for Structured Data
Qinfeng Shi, James Petterson, Gideon Dror, John Langford 0001, Alexander J. Smola, S. V. N. Vishwanathan |
J. Mach. Learn. Res. | 1 |
| 2008 | Discriminative human action segmentation and recognition using semi-Markov modelabstractGiven an input video sequence of one person conducting a sequence of continuous actions, we consider the problem of jointly segmenting and recognizing actions. We propose a discriminative approach to this problem under a semi-Markov model framework, where we are able to define a set of features over input-output space that captures the characteristics on boundary frames, action segments and neighboring action segments, respectively. In addition, we show that this method can also be used to recognize the person who performs in this video sequence. A Viterbi-like algorithm is devised to help efficiently solve the induced optimization problem. Experiments on a variety of datasets demonstrate the effectiveness of the proposed method. Qinfeng Shi, Li Wang 0033, Li Cheng 0001, Alexander J. Smola |
CVPR | 1 |
| 2007 | Semi-Markov Models for Sequence Segmentation
Qinfeng Shi, Yasemin Altun, Alexander J. Smola, S. V. N. Vishwanathan |
EMNLP-CoNLL | 1 |