Nick Barnes

dblp:41/2904 · DBLP profile ↗
← Back
128ranked-venue papers
11as first author
48since 2021 · last 2026
0000-0002-9343-9535ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 92 · 9 first-author · 34 since 2021Graphics, computer vision, multimedia, augmented reality and games · 72 · 3 first-author · 28 since 2021Systems, architecture and hardware · 11 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 8 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 AusSmoke meets MultiNatSmoke: a fully-labelled diverse smoke segmentation dataset
Hongjin Zhao, Gao Zhu, Ge-Peng Ji, Marta Yebra, Nick Barnes
WACV7
2026 SmokeBench: Evaluating Multimodal Large Language Models for Wildfire Smoke Detection
Tianye Qi, Nick Barnes
WACV3
2026 False Alarm Rectification for Early Smoke Segmentation
Hongjin Zhao, Ge-Peng Ji, Nick Barnes
WACV4
2026 DermEVAL: A Dermatologist-Reviewed Benchmark for Multimodal Large Language Models
abstract
Clinical photographs play a vital role in conversational computer-aided diagnosis, particularly in dermatology. However, existing skin disease benchmarks contain limitations like insufficient dataset size, the sole presence of categorical labels, the lack of expert inspections, and limited diversity in annotations. To address these shortcomings, we introduce DermEVAL, a large-scale benchmark specifically designed to evaluate the performance of Multimodal Large Language Models (MLLMs) in dermatology. Our benchmark includes image-text pairs depicting 16 distinct skin diseases, featuring a total of 11,347 representative images drawn from various dermatological datasets, carefully selected and annotated with the guidance of dermatologists. DermEVAL enables two primary tasks: visual question answering (VQA) and medical report generation (MRG), designed to simulate real-world medical diagnostics. We evaluate the performance of MLLMs in dermatology using multiple metrics, including traditional metrics and GPT-4V-based assessments. Our results indicate that accurately diagnosing skin diseases remains challenging for state-of-the-art MLLMs. We also demonstrate that fine-tuning MLLMs using DermEVAL significantly improves their performance on dermatology-related image–text tasks.
Hongjin Zhao, Zhenyue Qin, Ge-Peng Ji, Tom Gedeon, Nick Barnes
WACV7
2026 PointCaM: Cut-and-Mix for open-set point cloud learning
Shi Qiu 0001, Weihao Li 0005, Saeed Anwar, Mehrtash Harandi, Nick Barnes, Lars Petersson
Comput. Vis. Image Underst.6
2025 Open Set Label Shift with Test Time Out-of-Distribution Reference
abstract
Open set label shift (OSLS) occurs when label distributions change from a source to a target distribution, and the target distribution has an additional out-of-distribution (OOD) class. In this work, we build estimators for both source and target open set label distributions using a source domain in-distribution (ID) classifier and an ID/OOD classifier. With reasonable assumptions on the ID/OOD classifier, the estimators are assembled into a sequence of three stages: 1) an estimate of the source label distribution of the OOD class, 2) an EM algorithm for Maximum Likelihood estimates (MLE) of the target label distribution, and 3) an estimate of the target label distribution of OOD class under relaxed assumptions on the OOD classifier. The sampling errors of estimates in 1) and 3) are quantified with a concentration inequality. The estimation result allows us to correct the ID classifier trained on the source distribution to the target distribution without retraining. Experiments on a variety of open set label shift settings demonstrate the effectiveness of our model. Our code is available at https://github.com/ChangkunYe/OpenSetLabelShift.
Changkun Ye, Russell Tsuchida, Lars Petersson, Nick Barnes
CVPR4
2025 A Comprehensive Overview of Large Language Models
abstract
Large Language Models (LLMs) have recently demonstrated remarkable capabilities in natural language processing tasks and beyond. This success of LLMs has led to a large influx of research contributions in this direction. These works encompass diverse topics such as architectural innovations, better training strategies, context length improvements, fine-tuning, multimodal LLMs, robotics, datasets, benchmarking, efficiency, and more. With the rapid development of techniques and regular breakthroughs in LLM research, it has become considerably challenging to perceive the bigger picture of the advances in this direction. Considering the rapidly emerging plethora of literature on LLMs, it is imperative that the research community is able to benefit from a concise yet comprehensive overview of the recent developments in this field. This article provides an overview of the literature on a broad range of LLM-related concepts. Our self-contained comprehensive overview of LLMs discusses relevant background concepts along with covering the advanced topics at the frontier of research in LLMs. This review article is intended to provide not only a systematic survey but also a quick, comprehensive reference for the researchers and practitioners to draw insights from extensive, informative summaries of the existing works to advance the LLM research.
Humza Naveed, Asad Ullah Khan, Shi Qiu 0001, Saeed Anwar, Muhammad Usman 0010, Naveed Akhtar, Nick Barnes, Ajmal Mian
ACM Trans. Intell. Syst. Technol.8
2025 Attention-Based Real Image Restoration
abstract
Deep convolutional neural networks perform better on images containing spatially invariant degradations, also known as synthetic degradations; however, their performance is limited on real-degraded photographs and requires multiple-stage network modeling. To advance the practicability of restoration algorithms, this article proposes a novel single-stage blind real image restoration network ( Net) by employing a modular architecture. We use a residual on the residual structure to ease low-frequency information flow and apply feature attention to exploit the channel dependencies. Furthermore, the evaluation in terms of quantitative metrics and visual quality for four restoration tasks, i.e., denoising, super-resolution, raindrop removal, and JPEG compression on 11 real degraded datasets against more than 30 state-of-the-art algorithms, demonstrates the superiority of our Net. We also present the comparison on three synthetically generated degraded datasets for denoising to showcase our method's capability on synthetics denoising. The codes, trained models, and results are available on https://github.com/saeed-anwar/R2Net.
Saeed Anwar, Nick Barnes, Lars Petersson
IEEE Trans. Neural Networks Learn. Syst.2
2025 Polarity Loss: Improving Visual-Semantic Alignment for Zero-Shot Detection
abstract
Conventional object detection models require large amounts of training data. In comparison, humans can recognize previously unseen objects by merely knowing their semantic description. To mimic similar behavior, zero-shot object detection (ZSD) aims to recognize and localize "unseen" object instances by using only their semantic information. The model is first trained to learn the relationships between visual and semantic domains for seen objects, later transferring the acquired knowledge to totally unseen objects. This setting gives rise to the need for correct alignment between visual and semantic concepts so that the unseen objects can be identified using only their semantic attributes. In this article, we propose a novel loss function called "polarity loss" that promotes correct visual-semantic alignment for an improved ZSD. On the one hand, it refines the noisy semantic embeddings via metric learning on a "semantic vocabulary" of related concepts to establish a better synergy between visual and semantic domains. On the other hand, it explicitly maximizes the gap between positive and negative predictions to achieve better discrimination between seen, unseen, and background objects. Our approach is inspired by embodiment theories in cognitive science that claim human semantic understanding to be grounded in past experiences (seen objects), related linguistic concepts (word vocabulary), and visual perception (seen/unseen object images). We conduct extensive evaluations on the Microsoft Common Objects in Context (MS-COCO) and Pascal Visual Object Classes (VOC) datasets, showing significant improvements over state of the art. Our code and evaluation protocols available at: https://github.com/salman-h-khan/PL-ZSD_Release.
Shafin Rahman, Salman Khan 0001, Nick Barnes
IEEE Trans. Neural Networks Learn. Syst.3
2024 Self-Calibrating Vicinal Risk Minimisation for Model Calibration
abstract
Model calibration, measuring the alignment between the prediction accuracy and model confidence, is an important metric reflecting model trustworthiness. Existing dense binary classification methods, without proper regularisation of model confidence, are prone to being over-confident. To calibrate Deep Neural Networks (DNNs), we propose a SelfCalibrating Vicinal Risk Minimisation (SCVRM) that explores the vicinity space of labeled data, where vicinal images that are farther away from labeled images adopt the groundtruth label with decreasing label confidence. We prove that in the logistic regression problem, SCVRM can be seen as a Vicinal Risk Minimisation plus a regularisation term that penalises the over-confident predictions. In practical implementation, SCVRM is approximated using Monte Carlo sampling that samples additional augmented training images and labels from the vicinal distributions. Experimental results demonstrate that SCVRM can signifi-cantly enhance model calibration for different dense classification tasks on both in-distribution and out-of-distribution data. Code is available at https://github.com/Carlisle-Liu/SCVRM.
Changkun Ye, Ruikai Cui, Nick Barnes
CVPR4
2024 LAM3D: Large Image-Point Clouds Alignment Model for 3D Reconstruction from Single Image
abstract
Large Reconstruction Models have made significant strides in the realm of automated 3D content generation from single or multiple input images. Despite their success, these models often produce 3D meshes with geometric inaccuracies, stemming from the inherent challenges of deducing 3D shapes solely from image data. In this work, we introduce a novel framework, the Large Image and Point Cloud Alignment Model (LAM3D), which utilizes 3D point cloud data to enhance the fidelity of generated 3D meshes. Our methodology begins with the development of a point-cloud-based network that effectively generates precise and meaningful latent tri-planes, laying the groundwork for accurate 3D mesh reconstruction. Building upon this, our Image-Point-Cloud Feature Alignment technique processes a single input image, aligning to the latent tri-planes to imbue image features with robust 3D information. This process not only enriches the image features but also facilitates the production of high-fidelity 3D meshes without the need for multi-view input, significantly reducing geometric distortions. Our approach achieves state-of-the-art high-fidelity 3D mesh reconstruction from a single image in just 6 seconds, and experiments on various datasets demonstrate its effectiveness.
Ruikai Cui, Xibin Song, Weixuan Sun, Senbo Wang, Weizhe Liu, Shenzhou Chen, Taizhang Shang, Yang Li 0193, Nick Barnes, Hongdong Li, Pan Ji
NeurIPS9
2024 Label Shift Estimation for Class-Imbalance Problem: A Bayesian Approach
abstract
As a type of distribution shift, label shift occurs when the source and target domains have different label distributions $\mathbb{P}(Y)$ but identical conditional distributions of data given labels $\mathbb{P}(X|Y)$. Under a Bayesian framework, we propose a novel Maximum A Posteriori (MAP) model and a novel posterior sampling model for the label shift problem. We prove the MAP objective admits a unique optimum and derive an EM algorithm that converges to the global optimum. We propose a novel Adaptive Prior Learning (APL) model to adaptively select prior parameters given data. We use the Markov Chain Monte Carlo (MCMC) method in our posterior sampling model to estimate and correct for label shift. Our methods can effectively resolve class imbalance problems on large-scale datasets without fine-tuning the classifier. Experiments show that our model outperforms existing methods on a variety of label shift settings. Our code is available at https://github.com/ChangkunYe/MAPLS/.
Changkun Ye, Russell Tsuchida, Lars Petersson, Nick Barnes
WACV4
2024 Synergizing triple attention with depth quality for RGB-D salient object detection
Peipei Song, Peiyan Zhong, Jing Zhang 0052, Piotr Koniusz, Feng Duan 0006, Nick Barnes
Neurocomputing7
2024 Measuring and Modeling Uncertainty Degree for Monocular Depth Estimation
abstract
Effectively measuring and modeling the reliability of a trained model is essential to the real-world deployment of monocular depth estimation (MDE) models. However, the intrinsic ill-posedness and ordinal-sensitive nature of MDE pose major challenges to the estimation of uncertainty degree of the trained models. On the one hand, utilizing current uncertainty modeling methods may increase memory consumption and usually take more time. On the other hand, measuring the uncertainty based on model accuracy can also be problematic, where uncertainty reliability and prediction accuracy are not well decoupled. In this paper, we propose to model the uncertainty of MDE models from the perspective of the inherent probability distributions originating from the depth probability volume and its extensions, and to assess it more fairly with more comprehensive metrics. By simply introducing additional training regularization terms, our model, with surprisingly simple formations and without requiring extra modules or multiple inferences, can provide uncertainty estimations with state-of-the-art reliability, and can be further improved when combined with ensemble or sampling methods. A series of experiments demonstrate the effectiveness of our methods. Code and results are available at https://github.com/npucvr/MDEUncertainty.
Mochu Xiang, Jing Zhang 0052, Nick Barnes, Yuchao Dai
IEEE Trans. Circuits Syst. Video Technol.3
2024 Weakly-Supervised Contrastive Learning for Unsupervised Object Discovery
abstract
Unsupervised object discovery (UOD) refers to the task of discriminating the whole region of objects from the background within a scene without relying on labeled datasets, which benefits the task of bounding-box-level localization and pixel-level segmentation. This task is promising due to its ability to discover objects in a generic manner. We roughly categorize existing techniques into two main directions, namely the generative solutions based on image resynthesis, and the clustering methods based on self-supervised models. We have observed that the former heavily relies on the quality of image reconstruction, while the latter shows limitations in effectively modeling semantic correlations. To directly target at object discovery, we focus on the latter approach and propose a novel solution by incorporating weakly-supervised contrastive learning (WCL) to enhance semantic information exploration. We design a semantic-guided self-supervised learning model to extract high-level semantic features from images, which is achieved by fine-tuning the feature encoder of a self-supervised model, namely DINO, via WCL. Subsequently, we introduce Principal Component Analysis (PCA) to localize object regions. The principal projection direction, corresponding to the maximal eigenvalue, serves as an indicator of the object region(s). Extensive experiments on benchmark unsupervised object discovery datasets demonstrate the effectiveness of our proposed solution. The source code and experimental results are publicly available via our project page at https://github.com/npucvr/WSCUOD.git.
Yunqiu Lv, Jing Zhang 0052, Nick Barnes, Yuchao Dai
IEEE Trans. Image Process.3
2023 Learning Audio-Visual Source Localization via False Negative Aware Contrastive Learning
abstract
Self-supervised audio-visual source localization aims to locate sound-source objects in video frames without extra annotations. Recent methods often approach this goal with the help of contrastive learning, which assumes only the audio and visual contents from the same video are positive samples for each other. However, this assumption would suffer from false negative samples in real-world training. For example, for an audio sample, treating the frames from the same audio class as negative samples may mislead the model and therefore harm the learned representations (e.g., the audio of a siren wailing may reasonably correspond to the ambulances in multiple images). Based on this observation, we propose a new learning strategy named False Negative Aware Contrastive (FNAC) to mitigate the problem of misleading the training with such false negative samples. Specifically, we utilize the intra-modal similarities to identify potentially similar samples and construct corresponding adjacency matrices to guide contrastive learning. Further, we propose to strengthen the role of true negative samples by explicitly leveraging the visual features of sound sources to facilitate the differentiation of authentic sounding source regions. FNAC achieves state-of-the-art performances on Flickr-SoundNet, VGG-Sound, and AVSBench, which demonstrates the effectiveness of our method in mitigating the false negative issue. The code is available at https://github.com/OpenNLPLab/FNAC_AVL.
Weixuan Sun, Zheyuan Liu 0002, Yiran Zhong, Tianpeng Feng, Yandong Guo, Nick Barnes
CVPR9
2023 P2C: Self-Supervised Point Cloud Completion from Single Partial Clouds
abstract
Point cloud completion aims to recover the complete shape based on a partial observation. Existing methods require either complete point clouds or multiple partial observations of the same object for learning. In contrast to previous approaches, we present Partial2Complete (P2C), the first self-supervised framework that completes point cloud objects using training samples consisting of only a single incomplete point cloud per object. Specifically, our framework groups incomplete point clouds into local patches as input and predicts masked patches by learning prior information from different partial objects. We also propose Region-Aware Chamfer Distance to regularize shape mismatch without limiting completion capability, and devise the Normal Consistency Constraint to incorporate a local planarity assumption, encouraging the recovered shape surface to be continuous and complete. In this way, P2C no longer needs multiple observations or complete point clouds as ground truth. Instead, structural cues are learned from a category-specific dataset to complete partial point clouds of objects. We demonstrate the effectiveness of our approach on both synthetic ShapeNet data and real-world ScanNet data, showing that P2C produces comparable results to methods trained with complete shapes, and outperforms methods learned with multiple partial observations. Code is available at https://github.com/CuiRuikai/Partial2Complete.
Ruikai Cui, Shi Qiu 0001, Saeed Anwar, Jiawei Liu 0005, Chaoyue Xing, Jing Zhang 0052, Nick Barnes
ICCV7
2023 Model Calibration in Dense Classification with Adaptive Label Perturbation
abstract
For safety-related applications, it is crucial to produce trustworthy deep neural networks whose prediction is associated with confidence that can represent the likelihood of correctness for subsequent decision-making. Existing dense binary classification models are prone to being over-confident. To improve model calibration, we propose Adaptive Stochastic Label Perturbation (ASLP) which learns a unique label perturbation level for each training image. ASLP employs our proposed Self-Calibrating Binary Cross Entropy (SC-BCE) loss, which unifies label perturbation processes including stochastic approaches (like DisturbLabel), and label smoothing, to correct calibration while maintaining classification rates. ASLP follows Maximum Entropy Inference of classic statistical mechanics to maximise prediction entropy with respect to missing information. It performs this while: (1) preserving classification accuracy on known data as a conservative solution, or (2) specifically improves model calibration degree by minimising the gap between the prediction accuracy and expected confidence of the target training label. Extensive results demonstrate that ASLP can significantly improve calibration degrees of dense binary classification models on both in-distribution and out-of-distribution data. The code is available on https://github.com/Carlisle-Liu/ASLP.
Jiawei Liu 0005, Changkun Ye, Shan Wang 0010, Ruikai Cui, Jing Zhang 0052, Kaihao Zhang, Nick Barnes
ICCV7
2023 A Multimodal Hierarchical Variational Autoencoder for Saliency Detection
abstract
Existing multimodal Salient Object Detection (SOD) methods do not generalize well for more complex and scalable multimodal learning scenarios. In this paper, we propose a multimodal hierarchical variational auto-encoder for generalized multimodal SOD. By introducing joint inference factorization methods in multimodal VAEs, our model is scalable to partially missing modality data. A latent hierarchy is proposed which enhances expressiveness in latent space and allows multi-level interaction between features across modalities. By further exploring the latent hierarchy, we provide intuitive uncertainty visualizations and observe that the main source of uncertainty in SOD derives from the lower-level features. Based on this, we propose a simple yet effective sampling-based confidence estimation method, that brings robustness when encountering untrustworthy modality data for inference. Extensive experimental analysis illustrate that our model can satisfy crucial properties that make a desirable multimodal SOD framework.
Nick Barnes
IJCNN3
2023 PnP-3D: A Plug-and-Play for 3D Point Clouds
abstract
With the help of the deep learning paradigm, many point cloud networks have been invented for visual analysis. However, there is great potential for development of these networks since the given information of point cloud data has not been fully exploited. To improve the effectiveness of existing networks in analyzing point cloud data, we propose a plug-and-play module, PnP-3D, aiming to refine the fundamental point cloud feature representations by involving more local context and global bilinear response from explicit 3D space and implicit feature space. To thoroughly evaluate our approach, we conduct experiments on three standard point cloud analysis tasks, including classification, semantic segmentation, and object detection, where we select three state-of-the-art networks from each task for evaluation. Serving as a plug-and-play module, PnP-3D can significantly boost the performances of established networks. In addition to achieving state-of-the-art results on four widely used point cloud benchmarks, we present comprehensive ablation studies and visualizations to demonstrate our approach's advantages. The code will be available at https://github.com/ShiQiu0419/pnp-3d.
Shi Qiu 0001, Saeed Anwar, Nick Barnes
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Vicinity Vision Transformer
abstract
Vision transformers have shown great success on numerous computer vision tasks. However, their central component, softmax attention, prohibits vision transformers from scaling up to high-resolution images, due to both the computational complexity and memory footprint being quadratic. Linear attention was introduced in natural language processing (NLP) which reorders the self-attention mechanism to mitigate a similar issue, but directly applying existing linear attention to vision may not lead to satisfactory results. We investigate this problem and point out that existing linear attention methods ignore an inductive bias in vision tasks, i.e., 2D locality. In this article, we propose Vicinity Attention, which is a type of linear attention that integrates 2D locality. Specifically, for each image patch, we adjust its attention weight based on its 2D Manhattan distance from its neighbouring patches. In this case, we achieve 2D locality in a linear complexity where the neighbouring image patches receive stronger attention than far away patches. In addition, we propose a novel Vicinity Attention Block that is comprised of Feature Reduction Attention (FRA) and Feature Preserving Connection (FPC) in order to address the computational bottleneck of linear attention approaches, including our Vicinity Attention, whose complexity grows quadratically with respect to the feature dimension. The Vicinity Attention Block computes attention in a compressed feature space with an extra skip connection to retrieve the original feature distribution. We experimentally validate that the block further reduces computation without degenerating the accuracy. Finally, to validate the proposed methods, we build a linear vision transformer backbone named Vicinity Vision Transformer (VVT). Targeting general vision tasks, we build VVT in a pyramid structure with progressively reduced sequence length. We perform extensive experiments on CIFAR-100, ImageNet-1 k, and ADE20 K datasets to validate the effectiveness of our method. Our method has a slower growth rate in terms of computational overhead than previous transformer-based and convolution-based networks when the input resolution increases. In particular, our approach achieves state-of-the-art image classification accuracy with 50% fewer parameters than previous approaches.
Weixuan Sun, Zhen Qin 0003, Yi Zhang 0137, Kaihao Zhang, Nick Barnes, Stanley T. Birchfield, Lingpeng Kong, Yiran Zhong
IEEE Trans. Pattern Anal. Mach. Intell.7
2023 An Energy-Based Prior for Generative Saliency
abstract
We propose a novel generative saliency prediction framework that adopts an informative energy-based model as a prior distribution. The energy-based prior model is defined on the latent space of a saliency generator network that generates the saliency map based on a continuous latent variables and an observed image. Both the parameters of saliency generator and the energy-based prior are jointly trained via Markov chain Monte Carlo-based maximum likelihood estimation, in which the sampling from the intractable posterior and prior distributions of the latent variables are performed by Langevin dynamics. With the generative saliency model, we can obtain a pixel-wise uncertainty map from an image, indicating model confidence in the saliency prediction. Different from existing generative models, which define the prior distribution of the latent variables as a simple isotropic Gaussian distribution, our model uses an energy-based informative prior which can be more expressive in capturing the latent space of the data. With the informative energy-based prior, we extend the Gaussian distribution assumption of generative models to achieve a more representative distribution of the latent space, leading to more reliable uncertainty estimation. We apply the proposed frameworks to both RGB and RGB-D salient object detection tasks with both transformer and convolutional neural network backbones. We further propose an adversarial learning algorithm and a variational inference algorithm as alternatives to train the proposed generative framework. Experimental results show that our generative saliency model with an energy-based prior can achieve not only accurate saliency predictions but also reliable uncertainty maps that are consistent with human perception.
Jing Zhang 0052, Jianwen Xie, Nick Barnes, Ping Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Toward Deeper Understanding of Camouflaged Object Detection
abstract
Preys in the wild evolve to be camouflaged to avoid being recognized by predators. In this way, camouflage acts as a key defence mechanism across species that is critical to survival. To detect and segment the whole scope of a camouflaged object, camouflaged object detection (COD) is introduced as a binary segmentation task, with the binary ground truth camouflage map indicating the exact regions of the camouflaged objects. In this paper, we revisit this task and argue that the binary segmentation setting fails to fully understand the concept of camouflage. We find that explicitly modeling the conspicuousness of camouflaged objects against their particular backgrounds can not only lead to a better understanding about camouflage, but also provide guidance to designing more sophisticated camouflage techniques. Furthermore, we observe that it is some specific parts of camouflaged objects that make them detectable by predators. With the above understanding about camouflaged objects, we present the first triple-task learning framework to simultaneouslylocalize, segment, and rankcamouflaged objects, indicating the conspicuousness level of camouflage. As no corresponding datasets exist for either the localization model or the ranking model, we generate localization maps with an eye tracker, which are then processed according to the instance level labels to generate our ranking-based training and testing dataset. We also contribute the largest COD testing set to comprehensively analyse performance of the COD models. Experimental results show that our triple-task learning framework achieves new state-of-the-art, leading to a more explainable COD network. Our code, data, and results are available at:https://github.com/JingZhang617/COD-Rank-Localize-and-Segment.
Yunqiu Lv, Jing Zhang 0052, Yuchao Dai, Aixuan Li, Nick Barnes, Deng-Ping Fan
IEEE Trans. Circuits Syst. Video Technol.5
2022 Transmission-Guided Bayesian Generative Model for Smoke Segmentation
abstract
Smoke segmentation is essential to precisely localize wildfire so that it can be extinguished in an early phase. Although deep neural networks have achieved promising results on image segmentation tasks, they are prone to be overconfident for smoke segmentation due to its non-rigid shape and transparent appearance. This is caused by both knowledge level uncertainty due to limited training data for accurate smoke segmentation and labeling level uncertainty representing the difficulty in labeling ground-truth. To effectively model the two types of uncertainty, we introduce a Bayesian generative model to simultaneously estimate the posterior distribution of model parameters and its predictions. Further, smoke images suffer from low contrast and ambiguity, inspired by physics-based image dehazing methods, we design a transmission-guided local coherence loss to guide the network to learn pair-wise relationships based on pixel distance and the transmission feature. To promote the development of this field, we also contribute a high-quality smoke segmentation dataset, SMOKE5K, consisting of 1,400 real and 4,000 synthetic images with pixel-wise annotation. Experimental results on benchmark testing datasets illustrate that our model achieves both accurate predictions and reliable uncertainty maps representing model ignorance about its prediction. Our code and dataset are publicly available at: https://github.com/redlessme/Transmission-BVM.
Siyuan Yan, Jing Zhang 0052, Nick Barnes
AAAI3
2022 Energy-Based Generative Cooperative Saliency Prediction
abstract
Conventional saliency prediction models typically learn a deterministic mapping from an image to its saliency map, and thus fail to explain the subjective nature of human attention. In this paper, to model the uncertainty of visual saliency, we study the saliency prediction problem from the perspective of generative models by learning a conditional probability distribution over the saliency map given an input image, and treating the saliency prediction as a sampling process from the learned distribution. Specifically, we propose a generative cooperative saliency prediction framework, where a conditional latent variable model~(LVM) and a conditional energy-based model~(EBM) are jointly trained to predict salient objects in a cooperative manner. The LVM serves as a fast but coarse predictor to efficiently produce an initial saliency map, which is then refined by the iterative Langevin revision of the EBM that serves as a slow but fine predictor. Such a coarse-to-fine cooperative saliency prediction strategy offers the best of both worlds. Moreover, we propose a ``cooperative learning while recovering" strategy and apply it to weakly supervised saliency prediction, where saliency annotations of training images are partially observed. Lastly, we find that the learned energy function in the EBM can serve as a refinement module that can refine the results of other pre-trained saliency prediction models. Experimental results show that our model can produce a set of diverse and plausible saliency maps of an image, and obtain state-of-the-art performance in both fully supervised and weakly supervised saliency prediction tasks.
Jing Zhang 0052, Jianwen Xie, Zilong Zheng, Nick Barnes
AAAI4
2022 PU-Transformer: Point Cloud Upsampling Transformer
Shi Qiu 0001, Saeed Anwar, Nick Barnes
ACCV (1)3
2022 Energy-Based Residual Latent Transport for Unsupervised Point Cloud Completion
Ruikai Cui, Shi Qiu 0001, Saeed Anwar, Jing Zhang 0052, Nick Barnes
BMVC5
2022 Robust normalizing flows using Bernstein-type polynomials
Sameera Ramasinghe, Kasun Fernando, Salman Khan 0001, Nick Barnes
BMVC4
2022 The Devil in Linear Transformer
abstract
Linear transformers aim to reduce the quadratic space-time complexity of vanilla transformers.However, they usually suffer from degraded performances on various tasks and corpora.In this paper, we examine existing kernel-based linear transformers and identify two key issues that lead to such performance gaps: 1) unbounded gradients in the attention computation adversely impact the convergence of linear transformer models; 2) attention dilution which trivially distributes attention scores over long sequences while neglecting neighbouring structures.To address these issues, we first identify that the scaling of attention matrices is the devil in unbounded gradients, which turns out unnecessary in linear attention as we show theoretically and empirically.To this end, we propose a new linear attention that replaces the scaling operation with a normalization to stabilize gradients.For the issue of attention dilution, we leverage a diagonal attention to confine attention to only neighbouring tokens in early layers.Benefiting from the stable gradients and improved attention, our new linear transformer model, TRANSNORMER, demonstrates superior performance on text classification and language modeling tasks, as well as on the challenging Long-Range Arena benchmark, surpassing vanilla transformer and existing linear variants by a clear margin while being significantly more space-time efficient.The code is available at TRANSNORMER.
Zhen Qin 0003, Xiaodong Han, Weixuan Sun, Dongxu Li 0003, Lingpeng Kong, Nick Barnes, Yiran Zhong
EMNLP6
2022 Multi-Modal Transformer for RGB-D Salient Object Detection
abstract
The main focus of existing RGB-D salient object detection models is achieving effective multi-modal fusion. Due to the limited receptive field of conventional convolutional neural networks (CNNs), CNN-based multi-modal fusion strategies fail to extensively model the correlation between the two modalities (appearance information from the RGB image and geometric information from the depth data). Given the success of transformer networks for long-range dependency modeling, we investigate multi-modal transformer networks for RGB-D salient object detection. Specifically, a transformer-based multi-modal fusion module is presented to effectively fuse appearance features and geometric features. Experimental results on six challenging benchmark RGB-D salient object detection datasets demonstrate the effectiveness of our approach.
Peipei Song, Jing Zhang 0052, Piotr Koniusz, Nick Barnes
ICIP4
2022 Efficient Gaussian Process Model on Class-Imbalanced Datasets for Generalized Zero-Shot Learning
abstract
Zero-Shot Learning (ZSL) models aim to classify object classes that are not seen during the training process. However, the problem of class imbalance is rarely discussed, despite its presence in several ZSL datasets. In this paper, we propose a Neural Network model that learns a latent feature embedding and a Gaussian Process (GP) regression model that predicts latent feature prototypes of unseen classes. A calibrated classifier is then constructed for ZSL and Generalized ZSL tasks. Our Neural Network model is trained efficiently with a simple training strategy that mitigates the impact of class-imbalanced training data. The model has an average training time of 5 minutes and can achieve state-of-the-art (SOTA) performance on imbalanced ZSL benchmark datasets like AWA2, AWA1 and APY, while having relatively good performance on the SUN and CUB datasets.
Changkun Ye, Nick Barnes, Lars Petersson, Russell Tsuchida
ICPR2
2022 Modeling Aleatoric Uncertainty for Camouflaged Object Detection
abstract
Aleatoric uncertainty captures noise within the observations. For camouflaged object detection, due to similar appearance of the camouflaged foreground and the back-ground, it’s difficult to obtain highly accurate annotations, especially annotations around object boundaries. We argue that training directly with the "noisy" camouflage map may lead to a model of poor generalization ability. In this paper, we introduce an explicitly aleatoric uncertainty estimation technique to represent predictive uncertainty due to noisy labeling. Specifically, we present a confidence-aware camouflaged object detection (COD) framework using dynamic supervision to produce both an accurate camouflage map and a reliable "aleatoric uncertainty". Different from existing techniques that produce deterministic prediction following the point estimation pipeline, our framework formalises aleatoric uncertainty as probability distribution over model output and the input image. We claim that, once trained, our confidence estimation network can evaluate the pixel-wise accuracy of the prediction without relying on the ground truth camouflage map. Extensive results illustrate the superior performance of the proposed model in explaining the camouflage prediction. Our codes are available at https://github.com/Carlisle-Liu/OCENet
Jiawei Liu 0005, Jing Zhang 0052, Nick Barnes
WACV3
2022 Inferring the Class Conditional Response Map for Weakly Supervised Semantic Segmentation
abstract
Image-level weakly supervised semantic segmentation (WSSS) relies on class activation maps (CAMs) for pseudo labels generation. As CAMs only highlight the most discriminative regions of objects, the generated pseudo labels are usually unsatisfactory to serve directly as supervision. To solve this, most existing approaches follow a multi-training pipeline to refine CAMs for better pseudo-labels, which includes: 1) re-training the classification model to generate CAMs; 2) post-processing CAMs to obtain pseudo labels; and 3) training a semantic segmentation model with the obtained pseudo labels. However, this multi-training pipeline requires complicated adjustment and additional time. To address this, we propose a class-conditional inference strategy and an activation aware mask refinement loss function to generate better pseudo labels without retraining the classifier. The class conditional inference-time approach is presented to separately and iteratively reveal the classification network’s hidden object activation to generate more complete response maps. Further, our activation aware mask refinement loss function introduces a novel way to exploit saliency maps during segmentation training and refine the foreground object masks without suppressing background objects. Our method achieves superior WSSS results without requiring re-training of the classifier. https://github.com/weixuansun/InferCam
Weixuan Sun, Jing Zhang 0052, Nick Barnes
WACV3
2022 Densely Residual Laplacian Super-Resolution
abstract
Super-Resolution convolutional neural networks have recently demonstrated high-quality restoration for single images. However, existing algorithms often require very deep architectures and long training times. Furthermore, current convolutional neural networks for super-resolution are unable to exploit features at multiple scales and weigh them equally or at only static scale only, limiting their learning capability. In this exposition, we present a compact and accurate super-resolution algorithm, namely, densely residual laplacian network (DRLN). The proposed network employs cascading residual on the residual structure to allow the flow of low-frequency information to focus on learning high and mid-level features. In addition, deep supervision is achieved via the densely concatenated residual blocks settings, which also helps in learning from high-level complex features. Moreover, we propose Laplacian attention to model the crucial features to learn the inter and intra-level dependencies between the feature maps. Furthermore, comprehensive quantitative and qualitative evaluations on low-resolution, noisy low-resolution, and real historical image benchmark datasets illustrate that our DRLN algorithm performs favorably against the state-of-the-art methods visually and accurately.
Saeed Anwar, Nick Barnes
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Uncertainty Inspired RGB-D Saliency Detection
abstract
We propose the first stochastic framework to employ uncertainty for RGB-D saliency detection by learning from the data labeling process. Existing RGB-D saliency detection models treat this task as a point estimation problem by predicting a single saliency map following a deterministic learning pipeline. We argue that, however, the deterministic solution is relatively ill-posed. Inspired by the saliency data labeling process, we propose a generative architecture to achieve probabilistic RGB-D saliency detection which utilizes a latent variable to model the labeling variations. Our framework includes two main models: 1) a generator model, which maps the input image and latent variable to stochastic saliency prediction, and 2) an inference model, which gradually updates the latent variable by sampling it from the true or approximate posterior distribution. The generator model is an encoder-decoder saliency network. To infer the latent variable, we introduce two different solutions: i) a Conditional Variational Auto-encoder with an extra encoder to approximate the posterior distribution of the latent variable; and ii) an Alternating Back-Propagation technique, which directly samples the latent variable from the true posterior distribution. Qualitative and quantitative results on six challenging RGB-D benchmark datasets show our approach's superior performance in learning the distribution of saliency maps. The source code is publicly available via our project page: https://github.com/JingZhang617/UCNet.
Jing Zhang 0052, Deng-Ping Fan, Yuchao Dai, Saeed Anwar, Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Nick Barnes
IEEE Trans. Pattern Anal. Mach. Intell.7
2022 Geometric Back-Projection Network for Point Cloud Classification
abstract
As the basic task of point cloud analysis, classification is fundamental but always challenging. To address some unsolved problems of existing methods, we propose a network that captures geometric features of point clouds for better representations. To achieve this, on the one hand, we enrich the geometric information of points in low-level 3D space explicitly. On the other hand, we apply CNN-based structures in high-level feature spaces to learn local geometric context implicitly. Specifically, we leverage an idea of error-correcting feedback structure to capture the local features of point clouds comprehensively. Furthermore, an attention module based on channel affinity assists the feature map to avoid possible redundancy by emphasizing its distinct channels. The performance on both synthetic and real-world point clouds datasets demonstrate the superiority and applicability of our network. Comparing with other state-of-the-art methods, our approach balances accuracy and efficiency.
Shi Qiu 0001, Saeed Anwar, Nick Barnes
IEEE Trans. Multim.3
2021 Simultaneously Localize, Segment and Rank the Camouflaged Objects
abstract
Camouflage is a key defence mechanism across species that is critical to survival. Common strategies for camouflage include background matching, imitating the color and pattern of the environment, and disruptive coloration, disguising body outlines [37]. Camouflaged object detection (COD) aims to segment camouflaged objects hiding in their surroundings. Existing COD models are built upon binary ground truth to segment the camouflaged objects without illustrating the level of camouflage. In this paper, we revisit this task and argue that explicitly modeling the conspicuousness of camouflaged objects against their particular backgrounds can not only lead to a better understanding about camouflage and evolution of animals, but also provide guidance to design more sophisticated camouflage techniques. Furthermore, we observe that it is some specific parts of the camouflaged objects that make them detectable by predators. With the above understanding about camouflaged objects, we present the first ranking based COD network (Rank-Net) to simultaneously localize, segment and rank camouflaged objects. The localization model is proposed to find the discriminative regions that make the camouflaged object obvious. The segmentation model segments the full scope of the camouflaged objects. Further, the ranking model infers the detectability of different camouflaged objects. Moreover, we contribute a large COD testing set to evaluate the generalization ability of COD models. Experimental results show that our model achieves new state-of-the-art, leading to a more interpretable COD network1.
Yunqiu Lv, Jing Zhang 0052, Yuchao Dai, Aixuan Li, Bowen Liu 0012, Nick Barnes, Deng-Ping Fan
CVPR6
2021 Semantic Segmentation for Real Point Cloud Scenes via Bilateral Augmentation and Adaptive Fusion
abstract
Given the prominence of current 3D sensors, a fine-grained analysis on the basic point cloud data is worthy of further investigation. Particularly, real point cloud scenes can intuitively capture complex surroundings in the real world, but due to 3D data’s raw nature, it is very challenging for machine perception. In this work, we concentrate on the essential visual task, semantic segmentation, for large-scale point cloud data collected in reality. On the one hand, to reduce the ambiguity in nearby points, we augment their local context by fully utilizing both geometric and semantic features in a bilateral structure. On the other hand, we comprehensively interpret the distinctness of the points from multiple resolutions and represent the feature map following an adaptive fusion method at point-level for accurate semantic segmentation. Further, we provide specific ablation studies and intuitive visualizations to validate our key modules. By comparing with state-of-the-art networks on three different benchmarks, we demonstrate the effectiveness of our network.
Shi Qiu 0001, Saeed Anwar, Nick Barnes
CVPR3
2021 Weakly Supervised Video Salient Object Detection
abstract
Significant performance improvement has been achieved for fully-supervised video salient object detection with the pixel-wise labeled training datasets, which are time-consuming and expensive to obtain. To relieve the burden of data annotation, we present the first weakly super-vised video salient object detection model based on relabeled “fixation guided scribble annotations”. Specifically, an "Appearance-motion fusion module" and bidirectional ConvLSTM based framework are proposed to achieve effective multi-modal learning and long-term temporal context modeling based on our new weak annotations. Further, we design a novel foreground-background similarity loss to further explore the labeling similarity across frames. A weak annotation boosting strategy is also introduced to boost our model performance with a new pseudo-label generation technique. Extensive experimental results on six benchmark video saliency detection datasets illustrate the effectiveness of our solution1.
Wangbo Zhao, Jing Zhang 0052, Long Li 0008, Nick Barnes, Nian Liu 0002, Junwei Han 0001
CVPR4
2021 RGB-D Saliency Detection via Cascaded Mutual Information Minimization
abstract
Existing RGB-D saliency detection models do not explicitly encourage RGB and depth to achieve effective multi-modal learning. In this paper, we introduce a novel multistage cascaded learning framework via mutual information minimization to explicitly model the multi-modal information between RGB image and depth data. Specifically, we first map the feature of each mode to a lower dimensional feature vector, and adopt mutual information minimization as a regularizer to reduce the redundancy between appearance features from RGB and geometric features from depth. We then perform multi-stage cascaded learning to impose the mutual information minimization constraint at every stage of the network. Extensive experiments on benchmark RGB-D saliency datasets illustrate the effectiveness of our framework. Further, to prosper the development of this field, we contribute the largest (7× larger than NJU2K) COME15K dataset, which contains 15,625 image pairs with high quality polygon-/scribble-/object-/instance-/rank-level annotations. Based on these rich labels, we additionally construct four new benchmarks with strong baselines and observe some interesting phenomena, which can motivate future model design. Source code and dataset are available at https://github.com/JingZhang617/cascaded_rgbd_sod.
Jing Zhang 0052, Deng-Ping Fan, Yuchao Dai, Xin Yu 0002, Yiran Zhong, Nick Barnes, Ling Shao 0001
ICCV6
2021 Conditional Generative Modeling via Learning the Latent Space
Sameera Ramasinghe, Kanchana Ranasinghe, Salman Khan 0001, Nick Barnes, Stephen Gould
ICLR4
2021 Learning structure-aware semantic segmentation with image-level supervision
abstract
Compared with expensive pixel-wise annotations, image-level labels make it possible to learn semantic segmentation in a weakly-supervised manner. Within this pipeline, the class activation map (CAM) is obtained and further processed to serve as a pseudo label to train the semantic segmentation model in a fully-supervised manner. In this paper, we argue that the lost structure information in CAM limits its application in downstream semantic segmentation, leading to deteriorated predictions. Furthermore, the inconsistent class activation scores inside the same object contradicts the common sense that each region of the same object should belong to the same semantic category. To produce sharp prediction with structure information, we introduce an auxiliary semantic boundary detection module, which penalizes the deteriorated predictions. Furthermore, we adopt smoothness loss to encourage prediction inside the object to be consistent. Experimental results on the PASCAL-VOC dataset illustrate the effectiveness of the proposed solution.
Jiawei Liu 0005, Jing Zhang 0052, Yicong Hong, Nick Barnes
IJCNN4
2021 Recursive Training for Zero-Shot Semantic Segmentation
abstract
General purpose semantic segmentation relies on a backbone CNN network to extract discriminative features that help classify each image pixel into a ‘seen’ object class (i.e., the object classes available during training) or a background class. Zero-shot semantic segmentation is a challenging task that requires a computer vision model to identify image pixels belonging to an object class which it has never seen before. Equipping a general purpose semantic segmentation model to separate image pixels of ‘unseen’ classes from the background remains an open challenge. Some recent models have approached this problem by fine-tuning the final pixel classification layer of a semantic segmentation model for a Zero-Shot setting, but struggle to learn discriminative features due to the lack of supervision. We propose a recursive training scheme to supervise the retraining of a semantic segmentation model for a zero-shot setting using a pseudo-feature representation. To this end, we propose a Zero-Shot Maximum Mean Discrepancy (ZS-MMD) loss that weighs high confidence outputs of the pixel classification layer as a pseudo-feature representation, and feeds it back to the generator. By closing-the-loop on the generator end, we provide supervision during retraining that in turn helps the model learn a more discriminative feature representation for ‘unseen’ classes. We show that using our recursive training and ZS-MMD loss, our proposed model achieves state-of-the-art performance on the Pascal-VOC 2012 dataset and Pascal-Context dataset.
Moshiur R. Farazi, Nick Barnes
IJCNN3
2021 Rethinking conditional GAN training: An approach using geometrically structured latent manifolds
abstract
Conditional GANs (cGAN), in their rudimentary form, suffer from critical drawbacks such as the lack of diversity in generated outputs and distortion between the latent and output manifolds. Although efforts have been made to improve results, they can suffer from unpleasant side-effects such as the topology mismatch between latent and output spaces. In contrast, we tackle this problem from a geometrical perspective and propose a novel training mechanism that increases both the diversity and the visual quality of a vanilla cGAN, by systematically encouraging a bi-lipschitz mapping between the latent and the output manifolds. We validate the efficacy of our solution on a baseline cGAN (i.e., Pix2Pix) which lacks diversity, and show that by only modifying its training mechanism (i.e., with our proposed Pix2Pix-Geo), one can achieve more diverse and realistic outputs on a broad set of image-to-image translation tasks.
Sameera Ramasinghe, Moshiur R. Farazi, Salman Khan 0001, Nick Barnes, Stephen Gould
NeurIPS4
2021 Learning Generative Vision Transformer with Energy-Based Latent Space for Saliency Prediction
abstract
Vision transformer networks have shown superiority in many computer vision tasks. In this paper, we take a step further by proposing a novel generative vision transformer with latent variables following an informative energy-based prior for salient object detection. Both the vision transformer network and the energy-based prior model are jointly trained via Markov chain Monte Carlo-based maximum likelihood estimation, in which the sampling from the intractable posterior and prior distributions of the latent variables are performed by Langevin dynamics. Further, with the generative vision transformer, we can easily obtain a pixel-wise uncertainty map from an image, which indicates the model confidence in predicting saliency from the image. Different from the existing generative models which define the prior distribution of the latent variables as a simple isotropic Gaussian distribution, our model uses an energy-based informative prior which can be more expressive to capture the latent space of the data. We apply the proposed framework to both RGB and RGB-D salient object detection tasks. Extensive experimental results show that our framework can achieve not only accurate saliency predictions but also meaningful uncertainty maps that are consistent with the human perception.
Jing Zhang 0052, Jianwen Xie, Nick Barnes, Ping Li 0001
NeurIPS3
2021 Dense-Resolution Network for Point Cloud Classification and Segmentation
abstract
Point cloud analysis is attracting attention from Artificial Intelligence research since it can be widely used in applications such as robotics, Augmented Reality, self-driving. However, it is always challenging due to irregularities, unorderedness, and sparsity. In this article, we propose a novel network named Dense-Resolution Network (DRNet) for point cloud analysis. Our DRNet is designed to learn local point features from the point cloud in different resolutions. In order to learn local point groups more effectively, we present a novel grouping method for local neighborhood searching and an error-minimizing module for capturing local features. In addition to validating the network on widely used point cloud segmentation and classification benchmarks, we also test and visualize the performance of the components. Comparing with other state-of-the-art methods, our network shows superiority on ModelNet40, ShapeNet synthetic and ScanObjectNN real point cloud datasets.
Shi Qiu 0001, Saeed Anwar, Nick Barnes
WACV3
2021 Learning Saliency From Single Noisy Labelling: A Robust Model Fitting Perspective
abstract
The advances made in predicting visual saliency using deep neural networks come at the expense of collecting large-scale annotated data. However, pixel-wise annotation is labor-intensive and overwhelming. In this paper, we propose to learn saliency prediction from a single noisy labelling, which is easy to obtain (e.g., from imperfect human annotation or from unsupervised saliency prediction methods). With this goal, we address a natural question: Can we learn saliency prediction while identifying clean labels in a unified framework? To answer this question, we call on the theory of robust model fitting and formulate deep saliency prediction from a single noisy labelling as robust network learning and exploit model consistency across iterations to identify inliers and outliers (i.e., noisy labels). Extensive experiments on different benchmark datasets demonstrate the superiority of our proposed framework, which can learn comparable saliency prediction with state-of-the-art fully supervised saliency methods. Furthermore, we show that simply by treating ground truth annotations as noisy labelling, our framework achieves tangible improvements over state-of-the-art methods.
Jing Zhang 0052, Yuchao Dai, Tong Zhang 0023, Mehrtash Harandi, Nick Barnes, Richard I. Hartley
IEEE Trans. Pattern Anal. Mach. Intell.5
2021 Accuracy vs. complexity: A trade-off in visual question answering models
Moshiur R. Farazi, Salman Khan 0001, Nick Barnes
Pattern Recognit.3
2020 Improved Visual-Semantic Alignment for Zero-Shot Object Detection
abstract
Zero-shot object detection is an emerging research topic that aims to recognize and localize previously ‘unseen’ objects. This setting gives rise to several unique challenges, e.g., highly imbalanced positive vs. negative instance ratio, proper alignment between visual and semantic concepts and the ambiguity between background and unseen classes. Here, we propose an end-to-end deep learning framework underpinned by a novel loss function that handles class-imbalance and seeks to properly align the visual and semantic cues for improved zero-shot learning. We call our objective the ‘Polarity loss’ because it explicitly maximizes the gap between positive and negative predictions. Such a margin maximizing formulation is not only important for visual-semantic alignment but it also resolves the ambiguity between background and unseen objects. Further, the semantic representations of objects are noisy, thus complicating the alignment between visual and semantic domains. To this end, we perform metric learning using a ‘Semantic vocabulary’ of related concepts that refines the noisy semantic embeddings and establishes a better synergy between visual and semantic domains. Our approach is inspired by the embodiment theories in cognitive science, that claim human semantic understanding to be grounded in past experiences (seen objects), related linguistic concepts (word vocabulary) and the visual perception (seen/unseen object images). Our extensive results on MS-COCO and Pascal VOC datasets show significant improvements over state of the art.1
Shafin Rahman, Salman Khan 0001, Nick Barnes
AAAI3
2020 Any-Shot Object Detection
Shafin Rahman, Salman Khan 0001, Nick Barnes, Fahad Shahbaz Khan
ACCV (3)3
2020 3D Guided Weakly Supervised Semantic Segmentation
Weixuan Sun, Jing Zhang 0052, Nick Barnes
ACCV (1)3
2020 From Depth What Can You See? Depth Completion via Auxiliary Image Reconstruction
abstract
Depth completion recovers dense depth from sparse measurements, e.g., LiDAR. Existing depth-only methods use sparse depth as the only input. However, these methods may fail to recover semantics consistent boundaries, or small/thin objects due to 1) the sparse nature of depth points and 2) the lack of images to provide semantic cues. This paper continues this line of research and aims to overcome the above shortcomings. The unique design of our depth completion model is that it simultaneously outputs a reconstructed image and a dense depth map. Specifically, we formulate image reconstruction from sparse depth as an auxiliary task during training that is supervised by the unlabelled gray-scale images. During testing, our system accepts sparse depth as the only input, i.e., the image is not required. Our design allows the depth completion network to learn complementary image features that help to better understand object structures. The extra supervision incurred by image reconstruction is minimal, because no annotations other than the image are needed. We evaluate our method on the KITTI depth completion benchmark and show that depth completion can be significantly improved via the auxiliary supervision of image reconstruction. Our algorithm consistently outperforms depth-only methods and is also effective for indoor scenes like NYUv2.
Kaiyue Lu, Nick Barnes, Saeed Anwar, Liang Zheng 0001
CVPR2
2020 UC-Net: Uncertainty Inspired RGB-D Saliency Detection via Conditional Variational Autoencoders
abstract
In this paper, we propose the first framework (UCNet) to employ uncertainty for RGB-D saliency detection by learning from the data labeling process. Existing RGB-D saliency detection methods treat the saliency detection task as a point estimation problem, and produce a single saliency map following a deterministic learning pipeline. Inspired by the saliency data labeling process, we propose probabilistic RGB-D saliency detection network via conditional variational autoencoders to model human annotation uncertainty and generate multiple saliency maps for each input image by sampling in the latent space. With the proposed saliency consensus process, we are able to generate an accurate saliency map based on these multiple predictions. Quantitative and qualitative evaluations on six challenging benchmark datasets against 18 competing algorithms demonstrate the effectiveness of our approach in learning the distribution of saliency maps, leading to a new state-of-the-art in RGB-D saliency detection.
Jing Zhang 0052, Deng-Ping Fan, Yuchao Dai, Saeed Anwar, Fatemehsadat Saleh, Tong Zhang 0023, Nick Barnes
CVPR7
2020 Reducing the Sim-to-Real Gap for Event Cameras
Timo Stoffregen, Cedric Scheerlinck, Davide Scaramuzza 0001, Tom Drummond, Nick Barnes, Lindsay Kleeman, Robert E. Mahony
ECCV (27)5
2020 Learning Noise-Aware Encoder-Decoder from Noisy Labels by Alternating Back-Propagation for Saliency Detection
Jing Zhang 0052, Jianwen Xie, Nick Barnes
ECCV (17)3
2020 Question-Agnostic Attention for Visual Question Answering
abstract
Visual Question Answering (VQA) models employ attention mechanisms to discover image locations that are most relevant for answering a specific question. For this purpose, several multimodal fusion strategies have been proposed, ranging from relatively simple operations (e.g. linear sum) to more complex ones (e.g. Block [1]. The resulting multimodal representations define an intermediate feature space for capturing the interplay between visual and semantic features, that is helpful in selectively focusing on image content. In this paper, we propose a question-agnostic attention mechanism that is complementary to the existing question-dependent attention mechanisms. Our proposed model parses object instances to obtain an `object map' and applies this map on the visual features to generate Question-Agnostic Attention (QAA) features. In contrast to question-dependent attention approaches that are learned end-to-end, the proposed QAA does not involve question-specific training, and can be easily included in almost any existing VQA model as a generic light-weight pre-processing step, thereby adding minimal computation overhead for training. Further, when used in complement with the question-dependent attention, the QAA allows the model to focus on the regions containing objects that might have been overlooked by the learned attention representation. Through extensive evaluation on VQAv1, VQAv2 and TDIUC datasets, we show that incorporating complementary QAA allows state-of-the-art VQA models to perform better, and provides significant boost to simplistic VQA models, enabling them to performance on par with highly sophisticated fusion strategies.
Moshiur R. Farazi, Salman Khan 0001, Nick Barnes
ICPR3
2020 Learned and Hand-crafted Feature Fusion in Unit Ball for 3D Object Classification
abstract
Convolution is an effective technique that can be used to obtain abstract feature representations using hierarchical layers in deep networks. However, performing convolution in non-Euclidean topological spaces such as the unit ball (B 3 ) is still an under-explored problem. In this paper, we propose a light-weight experimental architecture for 3D object classification, that operates in B 3 . The proposed network utilizes both hand-crafted and learned features, and uses capsules in the penultimate layer to disentangle 3D shape features through pose and view equivariance. It simultaneously maintains an intrinsic co-ordinate frame, where mutual relationships between object parts are preserved. Furthermore, we show that the optimal view angles for extracting patterns from 3D objects depend on its shape and achieve compelling results with a relatively shallow network, compared to the state-of-the-art.
Sameera Ramasinghe, Salman Khan 0001, Nick Barnes
ICPRAM3
2020 Spectral-GANs for High-Resolution 3D Point-cloud Generation
abstract
Point-clouds are a popular choice for robotics and computer vision tasks due to their accurate shape description and direct acquisition from range-scanners. This demands the ability to synthesize and reconstruct high-quality point-clouds. Current deep generative models for 3D data generally work on simplified representations (e.g., voxelized objects) and cannot deal with the inherent redundancy and irregularity in point-clouds. A few recent efforts on 3D point-cloud generation offer limited resolution and their complexity grows with the increase in output resolution. In this paper, we develop a principled approach to synthesize 3D point-clouds using a spectral-domain Generative Adversarial Network (GAN). Our spectral representation is highly structured and allows us to disentangle various frequency bands such that the learning task is simplified for a GAN model. As compared to spatial-domain generative approaches, our formulation allows us to generate high-resolution point-clouds with minimal computational overhead. Furthermore, we propose a fully differentiable block to transform from the spectral to the spatial domain and back, thereby allowing us to integrate knowledge from well-established spatial models. We demonstrate that Spectral-GAN performs well for point-cloud generation task. Additionally, it can learn a highly discriminative representation in an unsupervised fashion and can be used to accurately reconstruct 3D objects. Our codes are available at https://github.com/samgregoost/Spectral-GAN/.
Sameera Ramasinghe, Salman Khan 0001, Nick Barnes, Stephen Gould
IROS3
2020 OfGAN: Realistic Rendition of Synthetic Colonoscopy Videos
Jiabo Xu, Saeed Anwar, Nick Barnes, Florian Grimpen, Olivier Salvado, Stuart Anderson 0004, Mohammad Ali Armin
MICCAI (3)3
2020 Blended Convolution and Synthesis for Efficient Discrimination of 3D Shapes
abstract
Existing models for shape analysis directly learn feature representations on 3D point clouds. We argue that 3D point clouds are highly redundant and hold irregular (permutation-invariant) structure, which makes it difficult to achieve inter-class discrimination efficiently. In this paper, we propose a two-pronged solution to this problem that is seamlessly integrated in a single blended convolution and synthesis layer. This fully differentiable layer performs two critical tasks in succession. In the first step, it projects the input 3D point clouds into a latent 3D space to synthesize a highly compact and inter-class discriminative point cloud representation. Since, 3D point clouds do not follow a Euclidean topology, standard 2/3D convolutional neural networks offer limited representation capability. Therefore, in the second step, we propose a novel 3D convolution operator functioning inside the unit ball to extract useful volumetric features. We derive formulae to achieve both translation and rotation of our novel convolution kernels. Finally, using the proposed techniques we present an extremely light-weight, end-to-end architecture that achieves compelling results on 3D shape recognition and retrieval.
Sameera Ramasinghe, Salman Khan 0001, Nick Barnes, Stephen Gould
WACV3
2020 Fast Image Reconstruction with an Event Camera
abstract
Event cameras are powerful new sensors able to capture high dynamic range with microsecond temporal resolution and no motion blur. Their strength is detecting brightness changes (called events) rather than capturing direct brightness images; however, algorithms can be used to convert events into usable image representations for applications such as classification. Previous works rely on hand-crafted spatial and temporal smoothing techniques to reconstruct images from events. State-of-the-art video reconstruction has recently been achieved using neural networks that are large (10M parameters) and computationally expensive, requiring 30ms for a forward-pass at 640 × 480 resolution on a modern GPU. We propose a novel neural network architecture for video reconstruction from events that is smaller (38k vs. 10M parameters) and faster (10ms vs. 30ms) than state-of-the-art with minimal impact to performance.
Cedric Scheerlinck, Henri Rebecq, Daniel Gehrig, Nick Barnes, Robert E. Mahony, Davide Scaramuzza 0001
WACV4
2020 Representation Learning on Unit Ball with 3D Roto-translational Equivariance
Sameera Ramasinghe, Salman Khan 0001, Nick Barnes, Stephen Gould
Int. J. Comput. Vis.3
2020 From known to the unknown: Transferring knowledge to answer questions about novel visual and semantic concepts
Moshiur R. Farazi, Salman Khan 0001, Nick Barnes
Image Vis. Comput.3
2020 Deep0Tag: Deep Multiple Instance Learning for Zero-Shot Image Tagging
abstract
Zero-shot learning aims to perform visual reasoning about unseen objects. In-line with the success of deep learning on object recognition problems, several end-to-end deep models for zero-shot recognition have been proposed in the literature. These models are successful in predicting a single unseen label given an input image but do not scale to cases where multiple unseen objects are present. Here, we focus on the challenging problem of zero-shot image tagging, where multiple labels are assigned to an image, that may relate to objects, attributes, actions, events, and scene type. Discovery of these scene concepts requires the ability to process multi-scale information. To encompass global as well as local image details, we propose an automatic approach to locate relevant image patches and model image tagging within the Multiple Instance Learning (MIL) framework. To the best of our knowledge, we propose the first end-to-end trainable deep MIL framework for the multi-label zero-shot tagging problem. We explore several alternatives for instance-level evidence aggregation and perform an extensive ablation study to identify the optimal pooling strategy. Due to its novel design, the proposed framework has several interesting features: 1) unlike previous deep MIL models, it does not use any off-line procedure (e.g., Selective Search or EdgeBoxes) for bag generation. 2) During test time, it can process any number of unseen labels given their semantic embedding vectors. 3) Using only image-level seen labels as weak annotation, it can produce a localized bounding box for each predicted label. We experiment with the large-scale NUS-WIDE and MS-COCO datasets and achieve superior performance across conventional, zero-shot, and generalized zero-shot tagging tasks.
Shafin Rahman, Salman Khan 0001, Nick Barnes
IEEE Trans. Multim.3
2019 Unsupervised Primitive Discovery for Improved 3D Generative Modeling
Salman Khan 0001, Yulan Guo, Munawar Hayat, Nick Barnes
CVPR4
2019 Real Image Denoising With Feature Attention
abstract
Deep convolutional neural networks perform better on images containing spatially invariant noise (synthetic noise); however, its performance is limited on real-noisy photographs and requires multiple stage network modeling. To advance the practicability of the denoising algorithms, this paper proposes a novel single-stage blind real image denoising network (RIDNet) by employing a modular architecture. We use residual on the residual structure to ease the flow of low-frequency information and apply feature attention to exploit the channel dependencies. Furthermore, the evaluation in terms of quantitative metrics and visual quality on three synthetic and four real noisy datasets against 19 state-of-the-art algorithms demonstrate the superiority of our RIDNet.
Saeed Anwar, Nick Barnes
ICCV2
2019 Transductive Learning for Zero-Shot Object Detection
abstract
Zero-shot object detection (ZSD) is a relatively unexplored research problem as compared to the conventional zero-shot recognition task. ZSD aims to detect previously unseen objects during inference. Existing ZSD works suffer from two critical issues: (a) large domain-shift between the source (seen) and target (unseen) domains since the two distributions are highly mismatched. (b) the learned model is biased against unseen classes, therefore in generalized ZSD settings, where both seen and unseen objects co-occur during inference, the learned model tends to misclassify unseen to seen categories. This brings up an important question: How effectively can a transductive setting address the aforementioned problems? To the best of our knowledge, we are the first to propose a transductive zero-shot object detection approach that convincingly reduces the domain-shift and model-bias against unseen classes. Our approach is based on a self-learning mechanism that uses a novel hybrid pseudo-labeling technique. It progressively updates learned model parameters by associating unlabeled data samples to their corresponding classes. During this process, our technique makes sure that knowledge that was previously acquired on the source domain is not forgotten. We report significant 'relative' improvements of 34.9% and 77.1% in terms of mAP and recall rates over the previous best inductive models on MSCOCO dataset.
Shafin Rahman, Salman Khan 0001, Nick Barnes
ICCV3
2018 Continuous-Time Intensity Estimation Using Event Cameras
Cedric Scheerlinck, Nick Barnes, Robert E. Mahony
ACCV (5)2
2018 Deep Texture and Structure Aware Filtering Network for Image Smoothing
Kaiyue Lu, Shaodi You, Nick Barnes
ECCV (4)3
2018 Adversarial Training of Variational Auto-Encoders for High Fidelity Image Generation
abstract
Variational auto-encoders (VAEs) provide an attractive solution to image generation problem. However, they tend to produce blurred and over-smoothed images due to their dependence on pixel-wise reconstruction loss. This paper introduces a new approach to alleviate this problem in the VAE based generative models. Our model simultaneously learns to match the data, reconstruction loss and the latent distributions of real and fake images to improve the quality of generated samples. To compute the loss distributions, we introduce an auto-encoder based discriminator model which allows an adversarial learning procedure. The discriminator in our model also provides perceptual guidance to the VAE by matching the learned similarity metric of the real and fake samples in the latent space. To stabilize the overall training process, our model uses an error feedback approach to maintain the equilibrium between competing networks in the model. Our experiments show that the generated samples from our proposed model exhibit a diverse set of attributes and facial expressions and scale up to highresolution images very well.
Salman Khan 0001, Munawar Hayat, Nick Barnes
WACV3
2018 3-D Shape Matching and Non-Rigid Correspondence for Hippocampi Based on Markov Random Fields
abstract
The purpose of this paper is to recover dense correspondence between non-rigid shapes for anatomical objects, which is a key element of disease diagnosis and analysis. We proposed a shape matching framework based on Markov random fields to obtain non-rigid correspondence. We constructed an energy function by summing up two terms where one was a unary term and the other was a binary term. By using this formulation, shape matching was represented as an energy function minimisation problem. Loopy belief propagation (LBP) was then used to minimize the energy function. We adopted a new sparse update technique for LBP update to increase computational efficiency. At the same time, we also proposed to use a novel clamping technique, an expectation-maximization (EM) like approach, to enhance matching accuracy. Experiments with the hippocampal data from OASIS and PATH showed that the sparse update was 160 times faster than standard BP. By iteratively running the EM-like clamping procedure, we were able to obtain high quality non-rigid correspondence results to achieve 97% matching rate between two hippocampi. Our shape matching based approach overcomes the flip problem of first-order ellipsoid and does not assume pre-alignment unlike iterative closest point.
Pengdong Xiao, Nick Barnes, Tibério S. Caetano
IEEE Trans. Image Process.2
2016 Local Background Enclosure for RGB-D Salient Object Detection
abstract
Recent work in salient object detection has considered the incorporation of depth cues from RGB-D images. In most cases, depth contrast is used as the main feature. However, areas of high contrast in background regions cause false positives for such methods, as the background frequently contains regions that are highly variable in depth. Here, we propose a novel RGB-D saliency feature. Local Background Enclosure (LBE) captures the spread of angular directions which are background with respect to the candidate region and the object that it is part of. We show that our feature improves over state-of-the-art RGB-D saliency approaches as well as RGB methods on the RGBD1000 and NJUDS2000 datasets.
David Feng 0002, Nick Barnes, Shaodi You, Chris McCarthy
CVPR2
2016 Learning Hough Transform with Latent Structures for Joint Object Detection and Pose Estimation
Xuming He 0001, Nick Barnes, Mingwen Wang 0001
MMM (2)3
2016 Efficient transductive semantic segmentation
abstract
Semantically describing the contents of images is one of the classical problems of computer vision. With huge numbers of images being made available daily, there is increasing interest in methods for semantic pixel labelling that exploit large image sets. Graph transduction provides a framework for the flexible inclusion of labeled data that can be exploited in the classification of unlabeled samples without requiring a trained classifier. Unfortunately, current approaches lack the scalability to tackle the joint segmentation of large image sets. Here we introduce an efficient flexible graph transduction approach to semantic segmentation that allows simple and efficient leveraging of large image sets without requiring separate computation of unary potentials, or a trained classifier. We demonstrate that this technique can handle far larger graphs than previous methods, and that results continue to improve as more labeled images are made available. Furthermore, we show that the method is able to benefit from dense or sparse unary labels when they are available.
José M. Álvarez 0004, Mathieu Salzmann, Nick Barnes
WACV3
2016 Semantic labeling for prosthetic vision
Lachlan Horne, José M. Álvarez 0004, Chris McCarthy, Mathieu Salzmann, Nick Barnes
Comput. Vis. Image Underst.5
2016 Exploiting Large Image Sets for Road Scene Parsing
abstract
There is an increasing interest in exploiting multiple images for scene understanding, with great progress in areas such as cosegmentation and video segmentation. Jointly analyzing the images in a large set offers the opportunity to exploit a greater source of information than when considering a single image on its own. However, this also yields challenges since, to effectively exploit all the available information, the resulting methods need to consider not just local connections, but efficiently analyze similarity between all pairs of pixels within and across all the images. In this paper, we propose to model an image set as a fully connected pairwise Conditional Random Field (CRF) defined over the image pixels, or superpixels, with Gaussian edge potentials. We show that this lets us co-label the images of a large set efficiently, thus yielding increased accuracy at no additional computational cost compared to sequential labeling of the images. Furthermore, we extend our framework to incorporate temporal dependence, thus effectively encompassing video segmentation as a special case of our approach, as well as to modeling label dependence over larger image regions. Our experimental evaluation demonstrates that our framework lets us handle over 10 000 images in a matter of seconds.
José M. Álvarez 0004, Mathieu Salzmann, Nick Barnes
IEEE Trans. Intell. Transp. Syst.3
2015 Efficient scene parsing by sampling unary potentials in a fully-connected CRF
abstract
Efficient, fully-connected CRF inference enables fast semantic labelling of images. However, this requires high-quality unary potentials to be computed, which is currently time-consuming. While some recent work attempts to address this issue by only computing a subset of unary potentials, a need remains for a simple, fast way to decide which unary potentials should be computed, without sacrificing accuracy. In particular, for embedded applications, a method which avoids time or memory-intensive operations is desired. In this paper, we introduce an approach to selecting good locations to compute unary potentials. We implement an efficient morphological approach to select a small proportion of pixel locations where unary potentials will be calculated. The speed of our labelling method allows us to directly search a large parameter space to optimize our method for a given task. We show that our method can achieve comparable accuracy to what can be achieved when all unary potentials are calculated, with significant time saving. Furthermore, we show that it is possible to tune our method to yield improved accuracy for certain classes of interest. We demonstrate this over multiple datasets representing challenging applications for our approach.
Lachlan Horne, José M. Álvarez 0004, Mathieu Salzmann, Nick Barnes
Intelligent Vehicles Symposium4
2015 Motion Segmentation of Truncated Signed Distance Function Based Volumetric Surfaces
abstract
Truncated signed distance function (TSDF) based volumetric surface reconstructions of static environments can be readily acquired using recent RGB-D camera based mapping systems. If objects in the environment move then a previously obtained TSDF reconstruction is no longer current. Handling this problem requires segmenting moving objects from the reconstruction. To this end, we present a novel solution to the motion segmentation of TSDF volumes. The segmentation problem is cast as CRF-based MAP inference in the voxel space. We propose: a novel data term by solving sparse multi-body motion segmentation and computing likelihoods for each motion label in the RGB-D image space, and, a novel pairwise term based on gradients of the TSDF volume. Experimental evaluation shows that the proposed approach achieves successful segmentations on reconstructions acquired with Kinect Fusion. Unlike the existing solutions which only work if the objects move completely from their initially occupied spaces, the proposed method permits segmentation of objects when they start to move.
Samunda Perera, Nick Barnes, Xuming He 0001, Shahram Izadi, Pushmeet Kohli, Ben Glocker
WACV2
2014 Importance weighted image enhancement for prosthetic vision: An augmentation framework
abstract
Augmentations to enhance perception in prosthetic vision (also known as bionic eyes) have the potential to improve functional outcomes significantly for implantees. In current (and near-term) im-plantable electrode arrays resolution and dynamic range are highly constrained in comparison to images from modern cameras that can be head mounted. In this paper, we propose a novel, generally applicable adaptive contrast augmentation framework for prosthetic vision that addresses the specific perceptual needs of low resolution and low dynamic range displays. The scheme accepts an externally defined pixel-wise weighting of importance describing features of the image to enhance in the output dynamic range. Our approach explicitly incorporates the logarithmic scaling of enhancement required in human visual perception to ensure perceivability of all contrast augmentations. It requires no pre-existing contrast, and thus extends previous work in local contrast enhancement to a formulation for general image augmentation. We demonstrate the generality of our augmentation scheme for scene structure and looming object enhancement using simulated prosthetic vision.
Chris McCarthy, Nick Barnes
ISMAR2
2014 Large-scale semantic co-labeling of image sets
abstract
As evidenced by video segmentation and cosegmentation approaches, exploiting multiple images is key to the success of visual scene understanding. With the availability of increasingly large sets of images, there is a clear need for methods that can efficiently analyze the similarities and structure across huge numbers of image pixels. Furthermore, to make effective use of this data, these similarities should not just be considered locally between neighboring pixels, but between all pairs of pixels across all images. In this paper, we tackle this challenging scenario by introducing a semantic co-labeling approach that performs efficient inference in a fully-connected CRF defined over the pixels, or superpixels, of an image set. Our experimental evaluation demonstrates that our approach yields improved accuracy while coming at no additional computation cost compared to performing segmentation sequentially on individual images. Furthermore, our formulation lets us perform inference over ten thousand images in a matter of seconds.
José M. Álvarez 0004, Mathieu Salzmann, Nick Barnes
WACV3
2014 Data-driven road detection
abstract
In this paper, we tackle the problem of road detection from RGB images. In particular, we follow a data-driven approach to segmenting the road pixels in an image. To this end, we introduce two road detection methods: A top-down approach that builds an image-level road prior based on the traffic pattern observed in an input image, and a bottom-up technique that estimates the probability that an image superpixel belongs to the road surface in a nonparametric manner. Both our algorithms work on the principle of label transfer in the sense that the road prior is directly constructed from the ground-truth segmentations of training images. Our experimental evaluation on four different datasets shows that this approach outperforms existing top-down and bottom-up techniques, and is key to the robustness of road detection algorithms to the dataset bias.
José M. Álvarez 0004, Mathieu Salzmann, Nick Barnes
WACV3
2013 Learning Structured Hough Voting for Joint Object Detection and Occlusion Reasoning
abstract
We propose a structured Hough voting method for detecting objects with heavy occlusion in indoor environments. First, we extend the Hough hypothesis space to include both object location and its visibility pattern, and design a new score function that accumulates votes for object detection and occlusion prediction. In addition, we explore the correlation between objects and their environment, building a depth-encoded object-context model based on RGB-D data. Particularly, we design a layered context representation and allow image patches from both objects and backgrounds voting for the object hypotheses. We demonstrate that using a data-driven 2.1D representation we can learn visual codebooks with better quality, and more interpretable detection results in terms of spatial relationship between objects and viewer. We test our algorithm on two challenging RGB-D datasets with significant occlusion and intraclass variation, and demonstrate the superior performance of our method.
Tao Wang 0047, Xuming He 0001, Nick Barnes
CVPR3
2013 An overview of vision processing in implantable prosthetic vision
abstract
Electrically stimulating prosthetic vision devices offer a potential therapy to blind individuals. There are currently two multi-centre trials of devices by Second Sight Medical Products, and by Zrenner's group at University of Tuebingen. In Australia, Bionic Vision Australia has a retinal implant trial with three patients. Current implants provide restricted information for implantees, and some limitations are likely to remain in the future. To provide a substantial benefit to individual's abilities to perform key tasks such as orientation and mobility, activities of daily living, reading and face recognition there is much work to be done. Vision processing's role is to ensure the key visual information is available to undertake tasks given these limitations. This paper frames the background and challenges in vision processing for implantable prosthetic vision, and gives an overview of recent work.
Nick Barnes
ICIP1
2013 Glass object segmentation by label transfer on joint depth and appearance manifolds
abstract
We address the glass object localization problem with a RGB-D camera. Our approach uses a nonparametric, data-driven label transfer scheme for local glass boundary estimation. A weighted voting scheme based on a joint feature manifold is adopted to integrate depth and appearance cues, and we learn a distance metric on the depth-encoded feature manifold. Local boundary evidence is then integrated into a MRF framework for spatially coherent glass object detection and segmentation. The efficacy of our approach is verified on a challenging RGB-D glass dataset where we obtained a clear improvement over the state-of-the-art both in terms of accuracy and speed.
Tao Wang 0047, Xuming He 0001, Nick Barnes
ICIP3
2013 Learning appearance models for road detection
abstract
We introduce an approach to image-based road detection that exploits the availability of unannotated training images to learn an appearance model. Our approach allows us to remove the standard assumption that the lower part of the input image belongs to the road surface, which does not always hold and often yields strongly biased appearance models. Instead, we exploit this assumption in the training images, which yields a much more general appearance model. We then use the learned model to classify the pixels of an input image as road or background without requiring any assumptions about this image. Our experimental evaluation shows the benefits of our approach over existing methods in challenging real-world driving scenarios.
José M. Álvarez 0004, Mathieu Salzmann, Nick Barnes
Intelligent Vehicles Symposium3
2012 Maximal Cliques Based Rigid Body Motion Segmentation with a RGB-D Camera
Samunda Perera, Nick Barnes
ACCV (2)2
2012 Glass object localization by joint inference of boundary and depth
Tao Wang 0047, Xuming He 0001, Nick Barnes
ICPR3
2012 The role of computer vision in prosthetic vision
abstract
The cost of vision loss worldwide has been estimated at nearly $3 trillion (http://www.amdalliance.org/cost-of-blindness.html). Non-preventable diseases cause a significant proportion of blindness in developed nations and will become more prevalent as people live longer. Prosthetic vision technologies including retinal implants will play an important therapeutic role. Retinal implants convert an input image stream to visual percepts via stimulation of the retina. This paper highlights some barriers to restoring functional human vision for current generation visual prosthetic devices that computer vision can help overcome. Such computer vision is interactive, aiming to restore function including visuo-motor tasks and recognition.
Nick Barnes
Image Vis. Comput.1
2012 A Unified Strategy for Landing and Docking Using Spherical Flow Divergence
abstract
We present a new visual control input from optical flow divergence enabling the design of novel, unified control laws for docking and landing. While divergence-based time-to-contact estimation is well understood, the use of divergence in visual control currently assumes knowledge of surface orientation, and/or egomotion. There exists no directly observable visual cue capable of supporting approaches to surfaces of arbitrary orientation under general motion. Central to our measure is the use of the maximum flow field divergence on the view sphere (max-div). We prove kinematic properties governing the location of max-div, and show that max-div provides a temporal measure of proximity. From this, we contribute novel control laws for regulating both approach velocity and angle of approach toward planar surfaces of arbitrary orientation, without structure-from-motion recovery. The strategy is tested in simulation, over real image sequences and in closed-loop control of docking/landing maneuvers on a mobile platform.
Chris McCarthy, Nick Barnes
IEEE Trans. Pattern Anal. Mach. Intell.2
2012 Fast and Robust Object Detection Using Asymmetric Totally Corrective Boosting
abstract
Boosting-based object detection has received significant attention recently. In this paper, we propose totally corrective asymmetric boosting algorithms for real-time object detection. Our algorithms differ from Viola and Jones' detection framework in two ways. Firstly, our boosting algorithms explicitly optimize asymmetric loss of objectives, while AdaBoost used by Viola and Jones optimizes a symmetric loss. Secondly, by carefully deriving the Lagrange duals of the optimization problems, we design more efficient boosting in that the coefficients of the selected weak classifiers are updated in a totally corrective fashion, in contrast to the stagewise optimization commonly used by most boosting algorithms. Column generation is employed to solve the proposed optimization problems. Unlike conventional boosting, the proposed boosting algorithms are able to de-select those irrelevant weak classifiers in the ensemble while training a classification cascade. This results in improved detection performance as well as fewer weak classifiers in the learned strong classifier. Compared with AsymBoost of Viola and Jones, our proposed asymmetric boosting is nonheuristic and the training procedure is much simpler. Experiments on face and pedestrian detection demonstrate that our methods have superior detection performance than some of the state-of-the-art object detectors.
Peng Wang 0015, Chunhua Shen, Nick Barnes
IEEE Trans. Neural Networks Learn. Syst.3
2010 Totally-Corrective Multi-class Boosting
Zhihui Hao, Chunhua Shen, Nick Barnes
ACCV (4)3
2010 Surface Extraction from Iso-disparity Contours
Chris McCarthy, Nick Barnes
ACCV (4)2
2010 Asymmetric Totally-Corrective Boosting for Real-Time Object Detection
Peng Wang 0015, Chunhua Shen, Nick Barnes, Zhang Ren
ACCV (1)3
2010 Hippocampal Shape Classification Using Redundancy Constrained Feature Selection
Luping Zhou, Lei Wang 0001, Chunhua Shen, Nick Barnes
MICCAI (2)4
2010 Estimation of the epipole using optical flow at antipodal points
John Lim, Nick Barnes
Comput. Vis. Image Underst.2
2010 Estimating Relative Camera Motion from the Antipodal-Epipolar Constraint
abstract
This paper introduces a novel antipodal-epipolar constraint on relative camera motion. By using antipodal points, which are available in large Field-of-View cameras, the translational and rotational motions of a camera are geometrically decoupled, allowing them to be separately estimated as two problems in smaller dimensions. We present a new formulation based on discrete camera motions, which works over a larger range of motions compared to previous differential techniques using antipodal points. The use of our constraints is demonstrated with two robust and practical algorithms, one based on RANSAC and the other based on Hough-like voting. As an application of the motion decoupling property, we also present a new structure-from-motion algorithm that does not require explicitly estimating rotation (it uses only the translation found with our methods). Finally, experiments involving simulations and real image sequences will demonstrate that our algorithms perform accurately and robustly, with some advantages over the state-of-the-art.
John Lim, Nick Barnes, Hongdong Li
IEEE Trans. Pattern Anal. Mach. Intell.2
2010 The regular polygon detector
Nick Barnes, Gareth Loy, David Shaw
Pattern Recognit.1
2009 Rotation Averaging with Application to Camera-Rig Calibration
Yuchao Dai, Jochen Trumpf, Hongdong Li, Nick Barnes, Richard I. Hartley
ACCV (2)4
2009 Identifying Anatomical Shape Difference by Regularized Discriminative Direction
abstract
Identifying the shape difference between two groups of anatomical objects is important for medical image analysis and computer-aided diagnosis. A method called "discriminative direction" in the literature has been proposed to solve this problem. In that method, the shape difference between groups is identified by deforming a shape along the discriminative direction. This paper conducts a thorough study about inferring this discriminative direction in an efficient and accurate way. First, finding the discriminative direction is reformulated as a preimage problem in kernel-based learning. This provides a complementary but conceptually simpler solution than the previous method. More importantly, we find that a shape deforming along the original discriminative direction cannot faithfully maintain its anatomical correctness. This unnecessarily introduces spurious shape differences and leads to inaccurate analysis. To overcome this problem, this paper further proposes a regularized discriminative direction by requiring a shape to conform to its underlying distribution when it deforms. Two different approaches are developed to impose the regularization, one from the perspective of probability distributions and the other from a geometric point of view, and their relationship is discussed. After verifying their superior performance through controlled experiments, we apply the proposed methods to detecting and localizing the hippocampal shape difference between sexes. We get results consistent with other independent research, providing a more compact representation of the shape difference compared with the established discriminative direction method.
Luping Zhou, Richard I. Hartley, Lei Wang 0001, Paulette Lieby, Nick Barnes
IEEE Trans. Medical Imaging5
2008 Directions of egomotion from antipodal points
abstract
We present a novel geometrical constraint on the egomotion of a single, moving camera. Using a camera with a large field-of-view (FOV), the optical flow measured at a single pair of antipodal points on the image sphere constrains the set of all possible camera motion directions to a subset region. By considering the flow at many such antipodal point pairs, it is shown that the intersection of all subset regions arising from each pair yields an estimate on the directions of motion. These antipodal point constraints rely on the geometrical properties of using a spherical representation of the image as well as the larger information content available from a large FOV. An algorithm using these constraints was implemented and tested on both simulated and real images. Results show comparable performance to the state of the art in the presence of noise and outliers whilst processing in constant time.
John Lim, Nick Barnes
CVPR2
2008 Regularized Discriminative Direction for Shape Difference Analysis
Luping Zhou, Richard I. Hartley, Lei Wang 0001, Paulette Lieby, Nick Barnes
MICCAI (1)5
2008 Real-Time Speed Sign Detection Using the Radial Symmetry Detector
abstract
Algorithms for classifying road signs have a high computational cost per pixel processed. A detection stage that has a lower computational cost can facilitate real-time processing. Various authors have used shape and color-based detectors. Shape-based detectors have an advantage under variable lighting conditions and sign deterioration that, although the apparent color may change, the shape is preserved. In this paper, we present the radial symmetry detector for detecting speed signs. We evaluate the detector itself in a system that is mounted within a road vehicle. We also evaluate its performance that is integrated with classification over a series of sequences from roads around Canberra and demonstrate it while running online in our road vehicle. We show that it can detect signs with high reliability in real time. We examine the internal parameters of the algorithm to adapt it to road sign detection. We demonstrate the stability of the system under the variation of these parameters and show computational speed gains through their tuning. The detector is demonstrated to work under a wide variety of visual conditions.
Nick Barnes, Alexander Zelinsky, Luke Fletcher
IEEE Trans. Intell. Transp. Syst.1
2008 A Robust Docking Strategy for a Mobile Robot Using Flow Field Divergence
abstract
We present a robust strategy for docking a mobile robot in close proximity with an upright surface using optical flow field divergence and proportional feedback control. Unlike previous approaches, we achieve this without the need for explicit segmentation of features in the image, and using complete gradient-based optical flow estimation (i.e., no affine models) in the optical flow computation. A key contribution is the development of an algorithm to compute the flow field divergence, or time-to-contact, in a manner that is robust to small rotations of the robot during ego-motion. This is done by tracking thefocusofexpansionof the flow field and using this to compensate for ego rotation of the image. The control law used is a simple proportional feedback, using the unfiltered flow field divergence as an input, for a dynamic vehicle model. Closed-loop stability analysis of docking under the proposed feedback is provided. Performance of the flow field divergence algorithm is demonstrated using offboard natural image sequences, and the performance of the closed-loop system is experimentally demonstrated by control of a mobile robot approaching a wall.
Chris McCarthy, Nick Barnes, Robert E. Mahony
IEEE Trans. Robotics2
2007 An MRF and Gaussian Curvature Based Shape Representation for Shape Matching
abstract
Matching and registration of shapes is a key issue in Computer Vision, Pattern Recognition, and Medical Image Analysis. This paper presents a shape representation framework based on Gaussian curvature and Markov random fields (MRFs) for the purpose of shape matching. The method is based on a surface mesh model in R3, which is projected into a two-dimensional space and there modeled as an extended boundary closed Markov random field. The surface is homeomorphic to S2. The MRF encodes in the nodes entropy features of the corresponding similarities based on Gaussian curvature, and in the edges the spatial consistency of the meshes. Correspondence between two surface meshes is then established by performing probabilistic inference on the MRF via Gibbs sampling. The technique combines both geometric, topological, and probabilistic information, which can be used to represent shapes in three dimensional space, and can be generalized to higher dimensional spaces. As a result, the representation can be used for shape matching, registration, and statistical shape analysis.
Pengdong Xiao, Nick Barnes, Tibério S. Caetano, Paulette Lieby
CVPR2
2007 Estimation of the Epipole using Optical Flow at Antipodal Points
abstract
This paper develops an algorithm for estimating the epipole or direction of translation of a moving monocular observer. To this end, we use constraints arising from two points that are antipodal on the image sphere. The antipodal point condition is necessary for decoupling rotation from translation. One such pair of points constrains the epipole to lie on a plane, and using two pairs of points, we have two such planes. The intersection of these two planes gives an estimate of the epipole. This means we require image motion measurements at two pairs of antipodal points to obtain an estimate. Repeating this will yield a set of possible solutions and a variety of methods could be applied to obtain a robust and refined estimate from this set. One robust and simple method is chosen for illustrative purposes and results on real images are shown. With real sequences, results of below 2deg error in the estimate of the epipole can be obtained. Since antipodal points on an image sphere are required, this algorithm must use some kind of omnidirectional or large field-of-view (FOV) sensor.
John Lim, Nick Barnes
ICCV2
2007 Real Time Biologically-Inspired Depth Maps from Spherical Flow
abstract
We present a strategy for generating real-time relative depth maps of an environment from optical flow, under general motion. We achieve this using an insect-inspired hemispherical fish-eye sensor with 190 degree FOV, and a de-rotated optical flow field. The de-rotation algorithm applied is based on the theoretical work of Nelson and Aloimonos (1988). From this we obtain the translational component of motion, and construct full relative depth maps on the sphere. We examine the robustness of this strategy in both simulation and real-world experiments, for a variety of environmental scenarios. To our knowledge, this is the first demonstrated implementation of the Nelson and Aloimonos algorithm working in real-time, over real image sequences. In addition, we apply this algorithm to the real-time recovery of full relative depth maps. These preliminary results demonstrate the feasibility of this approach for closed-loop control of a robot.
Chris McCarthy, Nick Barnes, Mandyam V. Srinivasan
ICRA2
2007 A Study of Hippocampal Shape Difference Between Genders by Efficient Hypothesis Test and Discriminative Deformation
Luping Zhou, Richard I. Hartley, Paulette Lieby, Nick Barnes, Kaarin Anstey, Nicolas Cherbuin, Perminder S. Sachdev
MICCAI (1)4
2007 MAP ZDF segmentation and tracking using active stereo vision: Hand tracking case study
Andrew Dankers, Nick Barnes, Alexander Zelinsky
Comput. Vis. Image Underst.2
2006 A Robust Docking Strategy for a Mobile Robot using Flow Field Divergence
abstract
We present a robust strategy for docking a mobile robot in close proximity with an upright surface using optical flow field divergence. Unlike previous approaches, we achieve this without the need for explicit segmentation of the surface in the image, and using complete optical estimation (i.e. no affine models) in the control loop. A simple proportional control law is used to regulate the vehicle's velocity, using only the raw, unfiltered flow divergence as input. Central to the robustness of our approach is the derivation of a time-to-contact estimator that accounts for small rotations of the robot during ego-motion. We present both analytical and experimental results showing that through tracking of the focus of expansion to a looming surface, we may compensate for such rotations, thereby significantly improving the robustness of the time-to-contact estimate. This is demonstrated using an off-board natural image sequence, and in closed-loop control of a mobile robot
Chris McCarthy, Nick Barnes
IROS2
2005 Regular Polygon Detection
abstract
This paper describes a new robust regular polygon detector. The regular polygon transform is posed as a mixture of regular polygons in a five dimensional space. Given the edge structure of an image, we derive the a posteriori probability for a mixture of regular polygons, and thus the probability density function for the appearance of a mixture of regular polygons. Likely regular polygons can be isolated quickly by discretising and collapsing the search space into three dimensions. The remaining dimensions may be efficiently recovered subsequently using maximum likelihood at the locations of the most likely polygons in the subspace. This leads to an efficient algorithm. Also the a posteriori formulation facilitates inclusion of additional a priori information leading to real-time application to road sign detection. The use of gradient information also reduces noise compared to existing approaches such as the generalised Hough transform. Results are presented for images with noise to show stability. The detector is also applied to two separate applications: real-time road sign detection for on-line driver assistance; and feature detection, recovering stable features in rectilinear environments.
Nick Barnes, Gareth Loy, David Shaw, Antonio Robles-Kelly
ICCV1
2005 Improved Signal To Noise Ratio And Computational Speed For Gradient-Based Detection Algorithms
abstract
Image gradient-based feature detectors offer great advantages over their standard edge-only equivalents. In driver support systems research, the radial symmetry detection algorithm has given real-time results for speed sign recognition. The regular polygon detector is a scan line algorithm for these features facilitating recognition of other road signs such as stop and give way signs. Radial symmetry has also been applied to real-time face detection, and the polygon detector is showing promising results as a feature detector for SLAM. However, gradient-based feature detection is more sensitive to noise than standard edge-based algorithms. As the total gradient magnitude at a pixel decreases, the component of the gradient at that point that arises from image noise increases. When a pixel votes in its gradient direction out to an extended radius, its position is more likely to be inaccurate if the gradient magnitude is low. In this paper, we analyse the performance of the radial symmetry and regular polygon detector algorithms under changes to the threshold on gradient magnitude. We show that the number of pixels correctly voting on a circle is not greatly reduced by thresholds that decrease the total number of pixels that vote in the image to 20%. This greatly reduces the noise component in the image, with only slight impact on the signal. This improves the performance, particularly for the regular polygon detector where the voting mechanism is complex and constitutes a large amount of the processing per pixel. This facilitates a real-time implementation, which is presented here.
Nick Barnes
ICRA1
2005 A Sign Reading Driver Assistance System Using Eye Gaze
abstract
Cars are becoming, in effect, a robotic system with an embedded human. It is not possible to know what the driver is thinking. We can, however, monitor their gaze and compare it with information in their view-field to make an inference. In this paper we present a complete system that reads speed signs in real-time, compares the driver gaze, and provides immediate feedback if the sign has been missed by the driver. This paper focuses on correlating measures of driver gaze direction with the position of signs in the road scene and improving recognition of signs through image enhancement.
Luke Fletcher, Lars Petersson, Nick Barnes, David J. Austin, Alexander Zelinsky
ICRA3
2004 Performance of Optical Flow Techniques for Indoor Navigation with a Mobile Robot
abstract
We present a comparison of four optical flow methods and three spatio-temporal filters for mobile robot navigation in corridor-like environments. Previous comparisons of optical flow methods have evaluated performance only in terms of accuracy and/or efficiency, and typically in isolation. These comparisons are inadequate for addressing applicability to continuous, real-time operation as part of a robot control loop. We emphasise the need for comparisons that consider the context of a system, and that are confirmed by in-system results. To this end, we give results for on and off-board trials of two biologically inspired behaviours: corridor centring and visual odometry. Our results show the best in-system performances are achieved using Lucas and Kanade's gradient-based method in combination with a recursive temporal filter. Results for traditionally used Gaussian filters indicate that long latencies significantly impede performance for real-time tasks in the control loop.
Chris McCarthy, Nick Barnes
ICRA2
2004 An Interactive Driver Assistance System Monitoring the Scene in and out of the Vehicle
abstract
This paper presents a framework for interactive driver assistance systems including techniques for fast speed sign detection and classification, car detection and tracking, and lane departure warning. In addition, the driver's actions are monitored. The integrated system uses information extracted from the road scene (speed signs, position within the lane, relative position to other cars, etc.) together with information about the driver's state such as eye gaze and head pose, to issue adequate warnings. A touch screen monitor presents relevant information and allows the driver to interact with the system. The research is focused around robust on-line algorithms. Initial results of online speed sign detection and car tracking are presented in the context of a driver assistance system.
Lars Petersson, Luke Fletcher, Nick Barnes, Alexander Zelinsky
ICRA3
2004 Fast Sum of Absolute Differences Visual Landmark Detector
abstract
This paper presents various optimisation that can be applied to the sum of absolute differences (SAD) correlation algorithm for automated landmark detection. This has applications in mobile robotic navigation and mapping. We show how some assumptions about the environment and the generic form of strong landmarks selected by the SAD correlation algorithm have led to the development of an algorithm to enable near real tune selection of strong landmarks from visual information. The landmarks that have been selected from a series of frames using our optimisation are shown to be stable through the image sequence, demonstration the scale invariance of the landmarks that are selected by the SAD correlation algorithm.
Craig Watman, David J. Austin, Nick Barnes, Gary Overett, Simon Thompson 0002
ICRA3
2004 Fast shape-based road sign detection for a driver assistance system
abstract
A new method is presented for detecting triangular, square and octagonal road signs efficiently and robustly. The method uses the symmetric nature of these shapes, together with the pattern of edge orientations exhibited by equiangular polygons with a known number of sides, to establish possible shape centroid locations in the image. This approach is invariant to in-plane rotation and returns the location and size of the shape detected. Results on still images show a detection rate of over 95%. The method is efficient enough for real-time applications, such as on-board-vehicle sign detection.
Gareth Blake Loy, Nick Barnes
IROS2
2004 Embodied categorisation for vision-guided mobile robots
Nick Barnes
Pattern Recognit.1
2003 Particle attraction localisation
abstract
In this paper, we present an original method for Bayesian localisation based on particle approximation. Our method overcomes a majority of problems inherent in previous Kalman filter and Bayesian approaches, including the recent Monte Carlo localisation methods. The algorithm converges quickly to any desired precision. It does not over-converge in the case of highly accurate sensor data and thus does not require a mixture-based approach. Also, the algorithm recovers well from random repositioning. These benefits are not hindered by computation which can be performed in real time on low powered processors. Further, the algorithm is intuitive and easy to implement. This algorithm is evaluated in simulation and has been applied to our entrant in the Sony four legged league of RoboCup, where it has been tested over many hours of international competition.
Damien George, Nick Barnes
IROS2
2003 Towards an efficient optimal trajectory planner for multiple mobile robots
abstract
In this paper, we present a real-time algorithm that plans mostly optimal trajectories for multiple mobile robots in a dynamic environment. This approach combines the use of a Delaunay triangulation to discretise the environment, a novel efficient use of the A* search method, and a novel cubic spline representation for a robot trajectory that meets the kinematic and dynamic constraints of the robot. We show that for complex environments the shortest-distance path is not always the shortest-time path due to these constraints. The algorithm has been implemented on real robots, and we present experimental results in cluttered environments.
Jason Thomas, Alan Blair 0001, Nick Barnes
IROS3
2003 Knowledge-Based Autonomous Dynamic Colour Calibration
Daniel Cameron, Nick Barnes
RoboCup2
2003 On-Board Vision Using Visual-Servoing for RoboCup F-180 League Mobile Robots
Paul Lee, Tim Dean, Andrew Yap, Dariusz Walter, Leslie J. Kitchen, Nick Barnes
RoboCup6
2002 Towards Real-Time Strategic Teamwork: A RoboCup Case Study
Kenichi Yoshimura, Nick Barnes, Ralph Rönnquist, Liz Sonenberg
RoboCup2
2001 Fuzzy Control for Active Perceptual Docking
abstract
Demonstrates fuzzy control of heading direction for a mobile robot. The robot is engaged in a docking maneuver guided by an active perceptual behaviour. In order to dock, the robot must control its heading direction to move directly towards a target object. Previously, we have developed a technique in which the robot fixates on the desired target, and corrects its heading direction based on information derived from optical flow data from a log-polar camera. However, the optical flow data derived is noisy. Thus, when robot direction control is based directly on the optical flow data the resulting path is erratic. A smooth path is more direct and so faster, leads to less strain on motors, and simplifies fixation. Further, constant change in direction can result in wheel-slippage, leading to errors in odometry. The results show a significant smoothing of the robot path in simulation.
Nick Barnes
FUZZ-IEEE1
2001 RoboMutts++
Kate Clarke, Stephen Dempster, Ian Falcao, Bronwen Jones, Daniel Rudolph, Alan Blair 0001, Chris McCarthy, Dariusz Walter, Nick Barnes
RoboCup9
2000 Direction Control for an Active Docking Behaviour Based on the Rotational Component of Log-Polar Optic Flow
Nick Barnes, Giulio Sandini
ECCV (2)1
2000 RoboMutts
Robert Sim, Paul Russo, Andrew Grahm, Andrew Blair, Nick Barnes, Alan Blair 0001
RoboCup5
2000 Vision Guided Circumnavigating Autonomous Robots
abstract
We present a system for vision guided autonomous circumnavigation, allowing a mobile robot to navigate safely around objects of arbitrary pose, and avoid obstacles. The system performs model-based object recognition from an intensity image. By enabling robots to recognize and navigate with respect to particular objects, this system empowers robots to perform deterministic actions on specific objects, rather than general exploration and navigation as emphasized in much of the current literature. This paper describes a fully integrated system, and, in particular, introduces canonical-views. Further, we derive a direct algebraic method for finding object pose and position for the four-dimensional case of a ground-based robot with uncalibrated vertical movement of its camera. Vision for mobile robots can be treated as a very different problem to traditional computer vision, as mobile robots have a characteristic perspective, and there is a causal relation between robot actions and view changes. Canonical-views are a novel, active object representation designed specifically to take advantage of the constraints of the robot navigation problem to allow efficient recognition and navigation.
Nick Barnes
Int. J. Pattern Recognit. Artif. Intell.1
1999 Knowledge-Based Shape-From Shading
abstract
In this paper, we study the problem of recovering approximate shape from the shading of a three-dimensional object in a single image when knowledge about the object is available. The application of knowledge-based methods to low-level image processing tasks will help overcome problems that arise from processing images using a pixel-based approach. Shape-from-shading has generally been approached by precognitive vision methods where a standard operator is applied to the image based on assumptions about the imaging process and generic properties of what appears. This paper explores some advantages of applying knowledge and hypotheses about what appears in the image. The knowledge and hypotheses used here come from domain knowledge and edge-matching. Specifically, we are able to find solutions to some problems that cannot be solved by other methods and gain advantages in terms of computation speed over similar approaches. Further, we can fully automate the derivation of the approximate shape of an object. This paper demonstrates the efficacy of using knowledge in the basic operation of an early vision operator, and so introduces a new paradigm for computer vision that may be applied to other early vision operators.
Nick Barnes
Int. J. Pattern Recognit. Artif. Intell.1