Neil D. B. Bruce

dblp:00/5848 · DBLP profile ↗
← Back
47ranked-venue papers
12as first author
11since 2021 · last 2025
0000-0002-5710-1107ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 8 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 8 · 2 first-authorDatabases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Deep Learning-Based Segmentation for Mapping Backyard Poultry in Canada
abstract
Backyard poultry operations pose a potential risk of avian influenza transmission, a disease with severe economic consequences due to flock culling and trade restrictions. Existing surveillance efforts rely primarily on data from registered farms, often overlooking unregistered small-scale flocks that lack biosecurity and are more exposed to wild birds, a known reservoir for the virus. This creates a gap in the monitoring of avian influenza. This study proposes a deep learning-based approach specifically designed to detect backyard operations using high resolution satellite imagery. Although previous studies have applied satellite imagery and deep learning techniques to detect commercial poultry farms and large-scale livestock operations, these approaches have not been extended to backyard poultry detection. Our method addresses the challenge of identifying small, irregular and often unregistered backyard setups. We employ a fully convolutional network (FCN) with ResNet-50 backbone to perform binary semantic segmentation. The model achieved an accuracy of 81.13%, precision of 78.92%, recall of 84.96%, and an F1 score of 81.83%, outperforming other models on our dataset.
Mina Khoshbazm Farimani, Neil D. B. Bruce, Shayan Sharif, Rozita Dara 0001
ICTAI2
2025 Leveraging social media and google trends to identify waves of avian influenza outbreaks in USA and Canada
abstract
Avian Influenza Virus (AIV) poses significant threats to the poultry industry, humans, domestic animals, and wildlife health worldwide. Monitoring this infectious disease is important for rapid and effective response to potential outbreaks. Conventional avian influenza surveillance systems have exhibited limitations in providing timely alerts for potential outbreaks. This study aimed to examine the idea of using online activity on social media, and Google searches to improve the identification of AIV in the early stage of an outbreak in a region. To this end, to evaluate the feasibility of this approach, we collected historical data on online user activities from X (formerly known as Twitter) and Google Trends and assessed the statistical correlation of activities in a region with the AIV outbreak officially reported case numbers. In order to mitigate the effect of the noisy content on the outbreak identification process, large language models were utilized to filter out the relevant online activity on X that could be indicative of an outbreak. Additionally, we conducted trend analysis on the selected internet-based data sources in terms of their timeliness and statistical significance in identifying AIV outbreaks. Moreover, we performed an ablation study using autoregsressive forecasting models to identify the contribution of X and Google Trends in predicting AIV outbreaks. The experimental findings illustrate that online activity on social media and search engine trends can detect avian influenza outbreaks, providing alerts earlier compared to official reports. This study suggests that real-time analysis of social media outlets and Google search trends can be used in avian influenza outbreak early warning systems, supporting epidemiologists and animal health professionals in informed decision-making.
Marzieh Soltani, Rozita Dara 0001, Zvonimir Poljak, Caroline Dubé, Neil D. B. Bruce, Shayan Sharif
Expert Syst. Appl.5
2025 Quantifying and Learning Static vs. Dynamic Information in Deep Spatiotemporal Networks
abstract
There is limited understanding of the information captured by deep spatiotemporal models in their intermediate representations. For example, while evidence suggests that action recognition algorithms are heavily influenced by visual appearance in single frames, no quantitative methodology exists for evaluating such static bias in the latent representation compared to bias toward dynamics. We tackle this challenge by proposing an approach for quantifying the static and dynamic biases of any spatiotemporal model, and apply our approach to three tasks, action recognition, automatic video object segmentation (AVOS) and video instance segmentation (VIS). Our key findings are: (i) Most examined models are biased toward static information. (ii) Some datasets that are assumed to be biased toward dynamics are actually biased toward static information. (iii) Individual channels in an architecture can be biased toward static, dynamic or jointly encode a combination static and dynamic information. (iv) Most models converge to their culminating biases in the first half of training. We then explore how these biases affect performance on dynamically biased datasets. For action recognition, we propose StaticDropout, a semantically guided dropout that debiases a model from static information toward dynamics. For AVOS, we design a better combination of fusion and cross connection layers compared with previous architectures.
Matthew Kowal, Mennatullah Siam, Md. Amirul Islam, Neil D. B. Bruce, Richard P. Wildes, Konstantinos G. Derpanis
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Position, Padding and Predictions: A Deeper Look at Position Information in CNNs
Md. Amirul Islam, Matthew Kowal, Sen Jia 0002, Konstantinos G. Derpanis, Neil D. B. Bruce
Int. J. Comput. Vis.5
2023 SegMix: Co-occurrence Driven Mixup for Semantic Segmentation and Adversarial Robustness
Md. Amirul Islam, Matthew Kowal, Konstantinos G. Derpanis, Neil D. B. Bruce
Int. J. Comput. Vis.4
2022 Maximizing Mutual Shape Information
Md. Amirul Islam, Matthew Kowal, Patrick Esser, Björn Ommer, Konstantinos G. Derpanis, Neil D. B. Bruce
BMVC6
2022 A Deeper Dive Into What Deep Spatiotemporal Networks Encode: Quantifying Static vs. Dynamic Information
abstract
Deep spatiotemporal models are used in a variety of computer vision tasks, such as action recognition and video object segmentation. Currently, there is a limited understanding of what information is captured by these models in their intermediate representations. For example, while it has been observed that action recognition algorithms are heavily influenced by visual appearance in single static frames, there is no quantitative methodology for evaluating such static bias in the latent representation compared to bias toward dynamic information (e.g. motion). We tackle this challenge by proposing a novel approach for quantifying the static and dynamic biases of any spatiotemporal model. To show the efficacy of our approach, we analyse two widely studied tasks, action recognition and video object segmentation. Our key findings are threefold: (i) Most examined spatiotemporal models are biased toward static information; although, certain two-stream architectures with cross-connections show a better balance between the static and dynamic information captured. (ii) Some datasets that are commonly assumed to be biased toward dynamics are actually biased toward static information. (iii) Individual units (channels) in an architecture can be biased toward static, dynamic or a combination of the two.11Project page and code
Matthew Kowal, Mennatullah Siam, Md. Amirul Islam, Neil D. B. Bruce, Richard P. Wildes, Konstantinos G. Derpanis
CVPR4
2021 Simpler Does It: Generating Semantic Labels with Objectness Guidance
Md. Amirul Islam, Matthew Kowal, Sen Jia 0002, Konstantinos G. Derpanis, Neil D. B. Bruce
BMVC5
2021 Global Pooling, More than Meets the Eye: Position Information is Encoded Channel-Wise in CNNs
abstract
In this paper, we challenge the common assumption that collapsing the spatial dimensions of a 3D (spatial-channel) tensor in a convolutional neural network (CNN) into a vector via global pooling removes all spatial information. Specifically, we demonstrate that positional information is encoded based on the ordering of the channel dimensions, while semantic information is largely not. Following this demonstration, we show the real world impact of these findings by applying them to two applications. First, we propose a simple yet effective data augmentation strategy and loss function which improves the translation invariance of a CNN’s output. Second, we propose a method to efficiently determine which channels in the latent representation are responsible for (i) encoding overall position information or (ii) region-specific positions. We first show that semantic segmentation has a significant reliance on the overall position channels to make predictions. We then show for the first time that it is possible to perform a ‘region-specific’ attack, and degrade a network’s performance in a particular part of the input. We believe our findings and demonstrated applications will benefit research areas concerned with understanding the characteristics of CNNs. Code is available at: https://github.com/islamamirul/PermuteNet.
Md. Amirul Islam, Matthew Kowal, Sen Jia 0002, Konstantinos G. Derpanis, Neil D. B. Bruce
ICCV5
2021 Shape or Texture: Understanding Discriminative Features in CNNs
Md. Amirul Islam, Matthew Kowal, Patrick Esser, Sen Jia 0002, Björn Ommer, Konstantinos G. Derpanis, Neil D. B. Bruce
ICLR7
2021 Relative Saliency and Ranking: Models, Metrics, Data and Benchmarks
abstract
Salient object detection is a problem that has been considered in detail and many solutions have been proposed. In this paper, we argue that work to date has addressed a problem that is relatively ill-posed. Specifically, there is not universal agreement about what constitutes a salient object when multiple observers are queried. This implies that some objects are more likely to be judged salient than others, and implies a relative rank exists on salient objects. Initially, we present a novel deep learning solution based on a hierarchical representation of relative saliency and stage-wise refinement. Further to this, we present data, analysis and baseline benchmark results towards addressing the problem of salient object ranking. Methods for deriving suitable ranked salient object instances are presented, along with metrics suitable to measuring algorithm performance. In addition, we show how a derived dataset can be successively refined to provide cleaned results that correlate well with pristine ground truth in its characteristics and value for training and testing models. Finally, we provide a comparison among prevailing algorithms that address salient object ranking or detection to establish initial baselines providing a basis for comparison with future efforts addressing this problem. The source code and data are publicly available via our project page: ryersonvisionlab.github.io/cocosalrank.
Mahmoud Kalash, Md. Amirul Islam, Neil D. B. Bruce
IEEE Trans. Pattern Anal. Mach. Intell.3
2020 Feature Binding with Category-Dependant MixUp for Semantic Segmentation and Adversarial Robustness
Md. Amirul Islam, Matthew Kowal, Konstantinos G. Derpanis, Neil D. B. Bruce
BMVC4
2020 Revisiting Saliency Metrics: Farthest-Neighbor Area Under Curve
abstract
In this paper, we propose a new metric to address the long-standing problem of center bias in saliency evaluation. We first show that distribution-based metrics cannot measure saliency performance across datasets due to ambiguity in the choice of standard deviation, especially for Convolutional Neural Networks. Therefore, our proposed metric is AUC-based because ROC curves are relatively robust to the standard deviation problem. However, this requires sufficient unique values in the saliency prediction to compute AUC scores. Secondly, we propose a global smoothing function for the problem of few value degrees in predicted saliency output. Compared with random noise, our smoothing function can create unique values without losing the existing relative saliency relationship. Finally, we show our proposed AUC-based metric can generate a more directional negative set for evaluation, denoted as Farthest-Neighbor AUC (FN-AUC). Our experiments show FN-AUC can measure spatial biases, central and peripheral, more effectively than S-AUC without penalizing the fixation locations.
Sen Jia 0002, Neil D. B. Bruce
CVPR2
2020 How much Position Information Do Convolutional Neural Networks Encode?
Md. Amirul Islam, Sen Jia 0002, Neil D. B. Bruce
ICLR3
2020 Distributed Iterative Gating Networks for Semantic Segmentation
abstract
In this paper, we present a canonical structure for controlling information flow in neural networks with an efficient feedback routing mechanism based on a strategy of Distributed Iterative Gating (DIGNet). The structure of this mechanism derives from a strong conceptual foundation, and presents a light-weight mechanism for adaptive control of computation similar to recurrent convolutional neural networks by integrating feedback signals with a feed forward architecture. In contrast to other RNN formulations, DIGNet generates feedback signals in a cascaded manner that implicitly carries information from all the layers above. This cascaded feedback propagation by means of the propagator gates is found to be more effective compared to other feedback mechanisms that use feedback from output of either the corresponding stage or from the previous stage. Experiments reveal the high degree of capability that this recurrent approach with cascaded feedback presents over feed-forward baselines and other recurrent models for pixel-wise labeling problems on three challenging datasets, PASCAL VOC 2012, COCO-Stuff, and ADE20K.
Rezaul Karim, Md. Amirul Islam, Neil D. B. Bruce
WACV3
2020 EML-NET: An Expandable Multi-Layer NETwork for saliency prediction
Sen Jia 0002, Neil D. B. Bruce
Image Vis. Comput.2
2019 In-depth Evaluation and Experimental Analysis of a Weight Pruning Genetic Algorithm
abstract
The topic of optimizing neural networks parameters has garnered considerable attention over the years, more so with the continued succession of ever larger, denser networks. In this work, we present evaluation and analysis of our proposed Weight Pruning Genetic Algorithm, with pruning specific mutators. The results show that evolved weight pruning of MNIST trained multilayer perceptrons and convolutional networks can remove as much as 72.4% and 89.6% of fully connected layer parameters, and yield improvements in test set accuracy without the use of retraining.
Sasa Janjic, Parimala Thulasiraman, Neil D. B. Bruce
CEC3
2019 Lossless Image Compression Using List Update Algorithms
Arezoo Abdollahi, Neil D. B. Bruce, Shahin Kamali, Rezaul Karim
SPIRE2
2019 Recurrent Iterative Gating Networks for Semantic Segmentation
abstract
In this paper, we present an approach for Recurrent Iterative Gating called RIGNet. The core elements of RIGNet involve recurrent connections that control the flow of information in neural networks in a top-down manner, and different variants on the core structure are considered. The iterative nature of this mechanism allows for gating to spread in both spatial extent and feature space. This is revealed to be a powerful mechanism with broad compatibility with common existing networks. Analysis shows how gating interacts with different network characteristics, and we also show that more shallow networks with gating may be made to perform better than much deeper networks that do not include RIGNet modules.
Rezaul Karim, Md. Amirul Islam, Neil D. B. Bruce
WACV3
2018 Semantics Meet Saliency: Exploring Domain Affinity and Models for Dual-Task Prediction
Md. Amirul Islam, Mahmoud Kalash, Neil D. B. Bruce
BMVC3
2018 Revisiting Salient Object Detection: Simultaneous Detection, Ranking, and Subitizing of Multiple Salient Objects
abstract
Salient object detection is a problem that has been considered in detail and many solutions proposed. In this paper, we argue that work to date has addressed a problem that is relatively ill-posed. Specifically, there is not universal agreement about what constitutes a salient object when multiple observers are queried. This implies that some objects are more likely to be judged salient than others, and implies a relative rank exists on salient objects. The solution presented in this paper solves this more general problem that considers relative rank, and we propose data and metrics suitable to measuring success in a relative object saliency landscape. A novel deep learning solution is proposed based on a hierarchical representation of relative saliency and stage-wise refinement. We also show that the problem of salient object subitizing can be addressed with the same network, and our approach exceeds performance of any prior work across all metrics considered (both traditional and newly proposed).
Md. Amirul Islam, Mahmoud Kalash, Neil D. B. Bruce
CVPR3
2018 Capturing real-world gaze behaviour: live and unplugged
abstract
Understanding human gaze behaviour has benefits from scientific understanding to many application domains. Current practices constrain possible use cases, requiring experimentation restricted to a lab setting or controlled environment. In this paper, we demonstrate a flexible unconstrained end-to-end solution that allows for collection and analysis of gaze data in real-world settings. To achieve these objectives, rich 3D models of the real world are derived along with strategies for associating experimental eye-tracking data with these models. In particular, we demonstrate the strength of photogrammetry in allowing these capabilities to be realized, and demonstrate the first complete solution for 3D gaze analysis in large-scale outdoor environments using standard camera technology without fiducial markers. The paper also presents techniques for quantitative analysis and visualization of 3D gaze data. As a whole, the body of techniques presented provides a foundation for future research, with new opportunities for experimental studies and computational modeling efforts.
Karishma Singh, Mahmoud Kalash, Neil D. B. Bruce
ETRA3
2018 Redundancy in Convolutional Neural Networks: Insights on Model Compression and Structure
abstract
Much work has been done on making convolutional models larger and more robust, but recent works have shown there is significant redundancy in the models, suggesting that these models are vastly more complex than necessary. In this work we explore the degree to which the representational redundancy of common pretrained models can be reduced, through evolved weight pruning of the fully connected layers and energy-based filter pruning. From investigating the effects of such pruning on model structure and accuracy, our findings show that this manner of analysis can help influence model parameter selection and architecture.
Sasa Janjic, Parimala Thulasiraman, Neil D. B. Bruce
IJCNN3
2017 Salient Object Detection using a Context-Aware Refinement Network
Md. Amirul Islam, Mahmoud Kalash, Mrigank Rochan, Neil D. B. Bruce, Yang Wang 0003
BMVC4
2017 Gated Feedback Refinement Network for Dense Image Labeling
abstract
Effective integration of local and global contextual information is crucial for dense labeling problems. Most existing methods based on an encoder-decoder architecture simply concatenate features from earlier layers to obtain higher-frequency details in the refinement stages. However, there are limits to the quality of refinement possible if ambiguous information is passed forward. In this paper we propose Gated Feedback Refinement Network (G-FRNet), an end-to-end deep learning framework for dense labeling tasks that addresses this limitation of existing methods. Initially, G-FRNet makes a coarse prediction and then it progressively refines the details by efficiently integrating local and global contextual information during the refinement stages. We introduce gate units that control the information passed forward in order to filter out ambiguity. Experiments on three challenging dense labeling datasets (CamVid, PASCAL VOC 2012, and Horse-Cow Parsing) show the effectiveness of our method. Our proposed approach achieves state-of-the-art results on the CamVid and Horse-Cow Parsing datasets, and produces competitive results on the PASCAL VOC 2012 dataset.
Md. Amirul Islam, Mrigank Rochan, Neil D. B. Bruce, Yang Wang 0003
CVPR3
2017 Movers, Shakers, and Those Who Stand Still: Visual Attention-grabbing Techniques in Robot Teleoperation
abstract
We designed and evaluated a series of teleoperation interface techniques that aim to draw operator attention while mitigating negative effects of interruption. Monitoring live teleoperation video feeds, for example to search for survivors in search and rescue, can be cognitively taxing, particularly for operators driving multiple robots or monitoring multiple cameras. To reduce workload, emerging computer vision techniques can automatically identify and indicate (cue) salient points of potential interest for the operator. However, it is not clear how to cue such points to a preoccupied operator -- whether cues would be distracting and a hindrance to operators -- and how the design of the cue may impact operator cognitive load, attention drawn, and primary task performance. In this paper, we detail our iterative design process for creating a range of visual attention-grabbing cues that are grounded in psychological literature on human attention, and two formal evaluations that measure attention-grabbing capability and impact on operator performance. Our results show that visually cueing on-screen points of interest does not distract operators, that operators perform poorly without the cues, and detail how particular cue design parameters impact operator cognitive load and task performance. Specifically, full-screen cues can lower cognitive load, but can increase response time; animated cues may improve accuracy, but increase cognitive load. Finally, from this design process we provide tested, and theoretically grounded cues for attention drawing in teleoperation.
Daniel J. Rea, Stela Hanbyeol Seo, Neil D. B. Bruce, James Everett Young
HRI3
2017 Tortoise and the Hare Robot: Slow and steady almost wins the race, but finishes more safely
abstract
We investigated the effects of changing the tele-operation feel of operating a robot by modifying its speed and acceleration profiles, and found that reducing a robot's maximum speed by half can reduce collisions by 32%, while only increasing navigation task time by 10%. Teleoperated robots are increasingly popular for enabling people to remotely attend meetings, explore dangerous areas, or view tourist destinations. As these robots are being designed to work in crowded areas with people, obstacles, or even unpredictable debris, interfaces that support piloting them in a safe and controlled manner are important for successful teleoperation. We investigate modifying a teleoperated robot's speed and acceleration profiles on an operator remotely navigating through an obstacle course. Our results indicate that lower maximum speeds result in lower operator workload, fewer collisions, and are only slightly slower than other profiles with a higher maximum speed. Our results raise questions about how robot designers should think about physical robot capability design and default driving software settings, the robot control interface, and the relation of robot speed to control.
Daniel J. Rea, Mahdi Rahmani Hanzaki, Neil D. B. Bruce, James Everett Young
RO-MAN3
2016 A Deeper Look at Saliency: Feature Contrast, Semantics, and Beyond
abstract
In this paper we consider the problem of visual saliency modeling, including both human gaze prediction and salient object segmentation. The overarching goal of the paper is to identify high level considerations relevant to deriving more sophisticated visual saliency models. A deep learning model based on fully convolutional networks (FCNs) is presented, which shows very favorable performance across a wide variety of benchmarks relative to existing proposals. We also demonstrate that the manner in which training data is selected, and ground truth treated is critical to resulting model behaviour. Recent efforts have explored the relationship between human gaze and salient objects, and we also examine this point further in the context of FCNs. Close examination of the proposed and alternative models serves as a vehicle for identifying problems important to developing more comprehensive models going forward.
Neil D. B. Bruce, Christopher Catton, Sasa Janjic
CVPR1
2016 Factors underlying inter-observer agreement in gaze patterns: predictive modelling and analysis
abstract
In viewing an image or real-world scene, different observers may exhibit different viewing patterns. This is evidently due to a variety of different factors, involving both bottom-up and top-down processing. In the literature addressing prediction of visual saliency, agreement in gaze patterns across observers is often quantified according to a measure of inter-observer congruency (IOC). Intuitively, common viewership patterns may be expected to diagnose certain image qualities including the capacity for an image to draw attention, or perceptual qualities of an image relevant to applications in human computer interaction, visual design and other domains. Moreover, there is value in determining the extent to which different factors contribute to inter-observer variability, and corresponding dependence on the type of content being viewed. In this paper, we assess the extent to which different types of features contribute to variability in viewing patterns across observers. This is accomplished in considering correlation between image derived features and IOC values, and based on the capacity for more complex feature sets to predict IOC based on a regression model. Experimental results demonstrate the value of different feature types for predicting IOC. These results also establish the relative importance of top-down and bottom-up information in driving gaze and provide new insight into predictive analysis for gaze behavior associated with perceptual characteristics of images.
Shafin Rahman, Neil D. B. Bruce
ETRA2
2016 Predicting task from eye movements: On the importance of spatial distribution, dynamics, and image features
Jonathan F. G. Boisvert, Neil D. B. Bruce
Neurocomputing2
2016 Sparse coding in early visual representation: From specific properties to general principles
Neil D. B. Bruce, Shafin Rahman, Diana Carrier
Neurocomputing1
2016 Weakly supervised object localization and segmentation in videos
Mrigank Rochan, Shafin Rahman, Neil D. B. Bruce, Yang Wang 0003
Image Vis. Comput.3
2015 Saliency weighted quality assessment of tone-mapped images
abstract
Different Tone-Mapping operators (TMOs) produce different Low Dynamic Range (LDR) images based on a single High Dynamic Range (HDR) image. The Tone-Mapped image Quality Index (TMQI) algorithm provides a quantitative means of assessing the quality of resultant LDR images. In this paper we test the hypothesis that TMQI predictions of human image quality can be further aligned with human judgement of image quality in considering visual attention, or regions that humans are predicted to fixate within a scene. We propose a modified version of the TMQI algorithm, a Saliency weighted Tone-Mapped Quality Index (STMQI) which demonstrates higher correlation with subjective ranking scores than the standard TMQI metric.
Hamid Reza Nasrinpour, Neil D. B. Bruce
ICIP2
2015 Saliency, Scale and Information: Towards a Unifying Theory
abstract
In this paper we present a definition for visual saliency grounded in information theory. This proposal is shown to relate to a variety of classic research contributions in scale-space theory, interest point detection, bilateral filtering, and to existing models of visual saliency. Based on the proposed definition of visual saliency, we demonstrate results competitive with the state-of-the art for both prediction of human fixations, and segmentation of salient objects. We also characterize different properties of this model including robustness to image transformations, and extension to a wide range of other data types with 3D mesh models serving as an example. Finally, we relate this proposal more generally to the role of saliency computation in visual information processing and draw connections to putative mechanisms for saliency computation in human vision.
Shafin Rahman, Neil D. B. Bruce
NIPS2
2014 Towards fine-grained fixation analysis: distilling out context dependence
abstract
In this paper, we explore the problem of analyzing gaze patterns towards attributing greater meaning to observed fixations. In recent years, there have been a number of efforts that attempt to categorize fixations according to their properties. Given that there are a multitude of factors that may contribute to fixational behavior, including both bottom-up and top-down influences on neural mechanisms for visual representation and saccadic control, efforts to better understand factors that may contribute to any given fixation may play an important role in augmenting raw fixation data. A grand objective of this line of thinking is in explaining the reason for any observed fixation as a combination of various latent factors. In the current work, we do not seek to solve this problem in general, but rather to factor out the role of the holistic structure of a scene as one observable, and quantifiable factor that plays a role in determining fixational behavior. Statistical methods and approximations to achieve this are presented, and supported by experimental results demonstrating the efficacy of the proposed methods.
Neil D. B. Bruce
ETRA1
2014 More human than human?: a visual processing approach to exploring believability of android faces
abstract
The issue of believability is core to android science, the challenge of creating a robot that can pass as a near human. While researchers are making great strides in improving the quality of androids and their likeness to people, it is simultaneously important to develop theoretical foundations behind believability, and experimental methods for exploring believability. In this paper, we explore a visual processing approach to investigating the believability of android faces, and present results from a study comparing current-generation android faces to humans' faces. We show how android faces are still not quite as believable as humans, and provide some mechanisms that may be used to investigate and compare believability in future projects.
Masayuki Nakane, James Everett Young, Neil D. B. Bruce
HAI3
2014 Examining visual saliency prediction in naturalistic scenes
abstract
Given the significant number of potential applications, visual saliency has increasingly become an area of interest in image and vision research. Many different strategies for predicting visual saliency have been proposed, that differ in their composition or rationale, and with a significant focus on improving performance across standard benchmarks. Recent benchmarks considering a large number of algorithms have further provided an understanding of the behavior of different algorithms. Performance evaluation has primarily focused on indoor and outdoor images of urban environments, many of which are composed, and contain salient objects. In this work, we test the performance of a number of the better performing algorithms on data derived from naturalistic scenes. In addition, given the strong connection to human vision, we test a putative model for early visual processing in primates tied to spectral energy and normalization. Results demonstrate significant differences between common datasets, and natural images. Performance analysis of the second-order contrast model also provides additional insight concerning the role of spectral energy in determining saliency. Finally we include analysis that demonstrates statistical properties of images that tend to imply common gaze patterns across observers.
Shafin Rahman, Mrigank Rochan, Yang Wang 0003, Neil D. B. Bruce
ICIP4
2014 ExpoBlend: Information preserving exposure blending based on normalized log-domain entropy
Neil D. B. Bruce
Comput. Graph.1
2013 Non-linear normalized entropy based exposure blending
Neil D. B. Bruce
Graphics Interface1
2013 Surround-see: enabling peripheral vision on smartphones during active use
abstract
Mobile devices are endowed with significant sensing capabilities. However, their ability to 'see' their surroundings, during active use, is limited. We present Surround-See, a self-contained smartphone equipped with an omni-directional camera that enables peripheral vision around the device to augment daily mobile tasks. Surround-See provides mobile devices with a field-of-view collinear to the device screen. This capability facilitates novel mobile tasks such as, pointing at objects in the environment to interact with content, operating the mobile device at a physical distance and allowing the device to detect user activity, even when the user is not holding it. We describe Surround-See's architecture, and demonstrate applications that exploit peripheral 'seeing' capabilities during active use of a mobile device. Users confirm the value of embedding peripheral vision capabilities on mobile devices and offer insights for novel usage methods.
Xing-Dong Yang, Khalad Hasan, Neil D. B. Bruce, Pourang Irani
UIST3
2009 Harris corners in the real world: A principled selection criterion for interest points based on ecological statistics
abstract
In this paper, we consider whether statistical regularities in natural images might be exploited to provide an improved selection criterion for interest points. One approach that has been particularly influential in this domain, is the Harris corner detector. The impetus for the selection criterion for Harris corners, proposed in early work and which remains in use to this day, is based on an intuitive mathematical definition constrained by the need for computational parsimony. In this paper, we revisit this selection criterion free of the computational constraints that existed 20 years ago, and also importantly, taking advantage of the regularities observed in natural image statistics. Based on the motivating factors of stability and richness of structure, a selection threshold for Harris corners is proposed based on a definition of optimality with respect to the structure observed in natural images. As a whole, the paper affords considerable insight into why existing approaches for selecting interest points work, and also their shortcomings. We also demonstrate how a proposal that is inspired by the properties of natural image statistics might be applied to overcome these shortcomings.
Neil D. B. Bruce, Pierre Kornprobst
CVPR1
2009 On the role of context in probabilistic models of visual saliency
abstract
In recent years, many principled probabilistic definitions for the determination of visual saliency have been proposed. Moreover, there has been increased focus on the role of context in the determination of visual salience. Prior efforts have shed some light on how context may help in predicting the location of, or presence of features associated with an object in the context of detection or recognition. Nevertheless, there remains a variety of manners in which context may be exploited towards providing better judgements of salient content. In this light, we investigate the role of context in the probabilistic determination of salience while presenting a number of potential avenues for future research.
Neil D. B. Bruce, Pierre Kornprobst
ICIP1
2007 Visual Correlates of Fixation Selection: A Look at the Spatial Frequency Domain
abstract
A representation for observing local image content is proposed for the purpose of considering the distinguishing characteristics of visual content that tends to draw a human observers gaze. Within this representation, the spectral profile distinguishing fixated from non-fixated locations is considered. Finally, the possibility of designing saliency operators based on the proposed local magnitude spectrum representation is explored, revealing a promising domain for predicting human gaze patterns.
Neil D. B. Bruce, Daniel P. Loach, John K. Tsotsos
ICIP (3)1
2006 A statistical basis for visual field anisotropies
Neil D. B. Bruce, John K. Tsotsos
Neurocomputing1
2005 Saliency Based on Information Maximization
abstract
A model of bottom-up overt attention is proposed based on the principle of maximizing information sampled from a scene. The proposed opera(cid:173) tion is based on Shannon's self-information measure and is achieved in a neural circuit, which is demonstrated as having close ties with the cir(cid:173) cuitry existent in the primate visual cortex. It is further shown that the proposed saliency measure may be extended to address issues that cur(cid:173) rently elude explanation in the domain of saliency based models. Results on natural images are compared with experimental eye tracking data re(cid:173) vealing the efficacy of the model in predicting the deployment of overt attention as compared with existing efforts.
Neil D. B. Bruce, John K. Tsotsos
NIPS1
2005 Features that draw visual attention: an information theoretic perspective
Neil D. B. Bruce
Neurocomputing1
2003 Evolutionary design of context-free attentional operators
abstract
A framework for simulating the visual attention system in primates is presented. Each stage of the attentional hierarchy is chosen with consideration for both psychophysics and mathematical optimality. A set of attentional operators are derived that act on basic image channels of intensity, hue, and orientation to produce maps representing perceptual importance of each image pixel. The development of such operators is realized within the context of a genetic optimization. The model includes the notion of an information domain where feature maps are transformed to a domain that more closely represents the response one might expect from the human visual system. The model is applied to a number of natural images to assess its efficacy in predicting guidance of attention in arbitrary natural scenes.
Neil D. B. Bruce, Ed Jernigan
ICIP (1)1