James J. Clark

dblp:60/4425 · also James Clark 0005 · DBLP profile ↗
← Back
69ranked-venue papers
17as first author
21since 2021 · last 2026
0000-0002-4512-6171ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 46 · 15 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 6 first-author · 9 since 2021Systems, architecture and hardware · 15 · 6 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 Ego4OOD: Rethinking Egocentric Video Domain Generalization via Covariate Shift Scoring
Zahra Vaseqi, James J. Clark
ICPR (11)2
2026 Latency-Aware Pruning and Quantization of Self-Supervised Speech Transformers for Edge Devices
abstract
The growing adoption of self-supervised learning transformers for speech (speech SSL) is constrained by their significant computational and memory demands, making deployment on resource-constrained edge devices challenging. We propose a latency-aware compression framework that integrates structured pruning and quantization to address these challenges. Guided by a latency model that considers the combined effects of pruning and quantization, our method dynamically identifies and removes less critical blocks while maintaining task performance, avoiding the inefficiencies of over-pruning and under-pruning seen in prior approaches. Unlike prior methods specialized in either post-training compression without fine-tuning data or in cases where fine-tuning data is available, our method is effective in both settings. Experimental results show that, in task-agnostic compression, our method achieves a 4.2× speedup on the Hikey970 edge development platform, outperforming previous task-agnostic pruning methods in most tasks, while requiring only 21–24 GPU hours—a 3× reduction compared to prior methods. Additionally, our method achieves a lower word error rate of 7.8% using task-specific pruning, while reducing computational overhead by approximately 19.4% in terms of GFLOPs compared to previous task-specific methods. Finally, our method consistently achieves higher accuracy than the state-of-the-art post-training compression approach across various latency speedup constraints, even without fine-tuning data.
Seyed Milad Ebrahimipour, Seyyed Hasan Mozafari, James J. Clark, Warren J. Gross, Brett H. Meyer
ACM Trans. Embed. Comput. Syst.3
2025 Decoupling Training-Free Guided Diffusion by ADMM
abstract
In this paper, we consider the conditional generation problem by guiding off-the-shelf unconditional diffusion models with differentiable loss functions in a plug-and-play fashion. While previous research has primarily focused on balancing the unconditional diffusion model and the guided loss through a tuned weight hyperparameter, we propose a novel framework that distinctly decouples these two components. Specifically, we introduce two variables x and z, to represent the generated samples governed by the unconditional generation model and the guidance function, respectively. This decoupling reformulates conditional generation into two manageable subproblems, unified by the constraint x = z. Leveraging this setup, we develop a new algorithm based on the Alternating Direction Method of Multipliers (ADMM) to adaptively balance these components. Additionally, we establish the equivalence between the diffusion reverse step and the proximal operator of ADMM and provide a detailed convergence analysis of our algorithm under certain mild assumptions. Our experiments demonstrate that our proposed method ADMMDiff consistently generates high-quality samples while ensuring strong adherence to the conditioning criteria. It outperforms existing methods across a range of conditional generation tasks, including image generation with various guidance and controllable motion synthesis.
Youyuan Zhang, Zehua Liu, Zenan Li, James J. Clark, Xujie Si
CVPR5
2025 Selective Unlearning via Representation Erasure Using Domain Adversarial Training
abstract
When deploying machine learning models in the real world, we often face the challenge of “unlearning” specific data points or subsets after training. Inspired by Domain-Adversarial Training of Neural Networks (DANN), we propose a novel algorithm,SURE, for targeted unlearning.SURE treats the process as a domain adaptation problem, where the “forget set” (data to be removed) and a validation set from the same distribution form two distinct domains. We train a domain classifier to discriminate between representations from the forget and validation sets.Using a gradient reversal strategy similar to DANN, we perform gradient updates to the representations to “fool” the domain classifier and thus obfuscate representations belonging to the forget set. Simultaneously, gradient descent is applied to the retain set (original training data minus the forget set) to preserve its classification performance. Unlike other unlearning approaches whose training objectives are built based on model outputs, SURE directly manipulates the representations.This is key to ensure robustness against a set of more powerful attacks than currently considered in the literature, that aim to detect which examples were unlearned through access to learned embeddings. Our thorough experiments reveal that SURE has a better unlearning quality to utility trade-off compared to other standard unlearning techniques for deep neural networks.
Nazanin Mohammadi Sepahvand, Eleni Triantafillou, Hugo Larochelle, Doina Precup, James J. Clark, Daniel M. Roy 0001, Gintare Karolina Dziugaite
ICLR5
2025 SEMU-Net: A Segmentation-Based Corrector for Fabrication Process Variations of Nanophotonics with Microscopic Images
abstract
Integrated silicon photonic devices, which manipulate light to transmit and process information on a silicon-on-insulator chip, are highly sensitive to structural variations. Minor deviations during nanofabrication-the precise process of building structures at the nanometer scale-such as over-or under-etching, corner rounding, and unintended defects, can significantly impact performance. To address these challenges, we introduce SEMU-Net, a comprehensive set of methods that automatically segments scanning electron microscope (SEM) images and uses them to train two deep neural network models based on U-Net and its variants. The predictor model anticipates fabrication-induced variations, while the corrector model adjusts the design to address these issues, ensuring that the final fabricated structures closely align with the intended specifications. Experimental results show that the segmentation U-Net reaches an average IoU score of 99.30%, while the corrector attention U-Net in a tandem architecture achieves an average IoU score of 98.67%.
Rambod Azimi, Yijian Kong, Dusan Gostimirovic, James J. Clark, Odile Liboiron-Ladouceur
WACV4
2025 FastVideoEdit: Leveraging Consistency Models for Efficient Text-to-Video Editing
abstract
Diffusion models have demonstrated remarkable capabilities in text-to-image and text-to-video generation, opening up possibilities for video editing based on textual input. However, the computational cost associated with sequential sampling in diffusion models poses challenges for efficient video editing. Existing approaches relying on image generation models for video editing suffer from time-consuming one-shot fine-tuning, additional condition extraction, or DDIM inversion, making real-time applications impractical. In this work, we propose FastVideoEdit, an efficient zero-shot video editing approach inspired by Consistency Models (CMs). By leveraging the self-consistency property of CMs, we eliminate the need for time-consuming inversion or additional condition extraction, reducing editing time. Our method enables direct mapping from source video to target video with strong preservation ability through attention control. This results in improved speed advantages, as fewer sampling steps can be used while maintaining comparable generation quality. Experimental results validate the state-of-the-art performance and speed advantages of FastVideoEdit across evaluation metrics encompassing editing speed, temporal consistency, and text-video alignment. The source code is available at github.com/youyuan-zhang/Fast/VideoEdit.
Youyuan Zhang, Xuan Ju, James J. Clark
WACV3
2024 An egocentric video and eye-tracking dataset for visual search in convenience stores
abstract
We introduce an egocentric video and eye-tracking dataset, comprised of 108 first-person videos of 36 shoppers searching for three different products (orange juice, KitKat chocolate bars, and canned tuna) in a convenience store, along with the frame-centered eye fixation locations for each video frame. The dataset also includes demographic information about each participant in the form of an 11-question survey. The paper describes two applications using the dataset — an analysis of eye fixations during search in the store, and a training of a clustered saliency model for predicting saliency of viewers engaged in product search in the store. The fixation analysis shows that fixation duration statistics are very similar to those found in image and video viewing, suggesting that similar visual processing is employed during search in 3D environments and during viewing of imagery on computer screens. A clustering technique was applied to the questionnaire data, which resulted in two clusters being detected. Based on these clusters, personalized saliency prediction models were trained on the store fixation data, which provided improved performance in prediction saliency on the store video data compared to state-of-the art universal saliency prediction methods.
Sansitha Panchadsaram, Rezvan Sherkati, James J. Clark
Comput. Vis. Image Underst.4
2023 Efficient 1D Grouped Convolution for PyTorch a Case Study: Fast On-Device Fine-Tuning for SqueezeBERT
abstract
Grouped convolution has been observed to be an effective approximation for convolution in many DNN applications. For example, SqueezeBERT, which is a light and fast BERT language processing model, utilizes 1D grouped convolutions. Though SqueezeBERT is well-optimized for inference on edge devices, it suffers from poor memory management during fine-tuning (training). This results in longer fine-tuning time on resource-limited GPUs compared to the original BERT model, BERT-base, despite being specifically designed for edge devices. We study this behavior and show that this poor memory management originates from the use of 1D grouped convolutions in SqueezeBERT. We re-implement 1D grouped convolutions using fully-connected layers, addressing the poor memory allocation and data locality of 1D grouped convolutions. We show that our method is well-suited for edge devices with limited memory; further, it has a negligible effect on inference speed. When utilizing our method, we observe a 42 % reduction in fine-tuning time for SqueezeBERT on edge devices.
Seyyed Hasan Mozafari, James J. Clark, Warren J. Gross, Brett H. Meyer
ASAP2
2023 Clustered Saliency Prediction
Rezvan Sherkati, James J. Clark
BMVC2
2023 High-Throughput Edge Inference for BERT Models via Neural Architecture Search and Pipeline
abstract
There has been growing interest in improving the BERT inference throughput on resource-constrained edge devices for a satisfactory user experience. One methodology is to employ heterogeneous computing, which utilizes multiple processing elements to accelerate inference. Another methodology is to deploy Neural Architecture Search (NAS) to find optimal solutions in accuracy-throughput design space. In this paper, for the first time, we incorporate NAS with pipelining for BERT models. We show that performing NAS with pipelining achieves on average 53% higher throughput, compared to NAS with a homogeneous system.
Hung-Yang Chang, Seyyed Hasan Mozafari, James J. Clark, Brett H. Meyer, Warren J. Gross
ACM Great Lakes Symposium on VLSI3
2023 Training Acceleration of Frequency Domain CNNs Using Activation Compression
abstract
Reducing the complexity of training convolutional neural networks results in lower energy consumption expended during training, or higher accuracy by admitting a greater number of training epochs within a training time budget. During backpropagation, a considerable amount of temporary data is offloaded from GPU memory to CPU memory, increasing training time. In this paper, we address this training time overhead by introducing an activation compression technique for frequency domain convolutional neural networks. Applying this compression technique on frequency domain AlexNet results in activation compression of 57.7%, and a reduction of training time by 23%, with a negligible effect on classification accuracy.
Seyyed Hasan Mozafari, James J. Clark, Warren J. Gross, Brett H. Meyer
ISCAS2
2023 GliTr: Glimpse Transformers with Spatiotemporal Consistency for Online Action Prediction
abstract
Many online action prediction models observe complete frames to locate and attend to informative subregions in the frames called glimpses and recognize an ongoing action based on global and local information. However, in applications with constrained resources, an agent may not be able to observe the complete frame, yet must still locate useful glimpses to predict an incomplete action based on local information only. In this paper, we develop Glimpse Transformers (GliTr), which observe only narrow glimpses at all times, thus predicting an ongoing action and the following most informative glimpse location based on the partial spatiotemporal information collected so far. In the absence of a ground truth for the optimal glimpse locations for action recognition, we train GliTr using a novel spatiotemporal consistency objective: We require GliTr to attend to the glimpses with features similar to the corresponding complete frames (i.e. spatial consistency) and the resultant class logits at time t equivalent to the ones predicted using whole frames up to t (i.e. temporal consistency). Inclusion of our proposed consistency objective yields ∼ 10% higher accuracy on the Something-Something-v2 (SSv2) dataset than the baseline cross-entropy objective. Overall, despite observing only ∼ 33% of the total area per frame, GliTr achieves 53.02% and 93.91% accuracy on the SSv2 and Jester datasets, respectively.
Samrudhdhi B. Rangrej, Kevin J. Liang, Tal Hassner, James J. Clark
WACV4
2023 Grow-push-prune: Aligning deep discriminants for effective structural network compression
Qing Tian 0003, Tal Arbel, James J. Clark
Comput. Vis. Image Underst.3
2022 Fast Heterogeneous Task Mapping for Reducing Edge DNN Latency
abstract
To meet DNN inference latency constraints on resource-constrained edge devices, we employ heterogeneous computing, utilizing multiple processing elements (e.g. CPU + GPU) accelerate inference. This leads to the challenge of efficiently mapping DNN operations to heterogeneous processing elements. For this task, we introduce a novel genetic algorithm (GA) optimizer. Through intelligent initialization and a customized mutation operation, we are able to evaluate 20x fewer generations while finding superior configurations compared with a baseline GA. Using our mapping optimizer, we find device placement configurations that achieve 15%, 24%, and 31% inference speed-up for BERT, SqueezeBERT, and InceptionV3,respectively.
Murray L. Kornelsen, Seyyed Hasan Mozafari, James J. Clark, Brett H. Meyer, Warren J. Gross
ASAP3
2022 Work-in-Progress: Utilizing latency and accuracy predictors for efficient hardware-aware NAS
abstract
With the increased size and complexity of state-of-the-art language models such as BERT, deploying them on resource-constrained devices has become challenging. Latency-aware Neural Architecture Search (NAS) is an effective solution for finding an efficient implementation of complex models that satisfy hardware limitations. However, collecting on-device accuracy and latency feedback would significantly slow down the search process, making NAS impractical. To address this, we propose a low-cost method that models both accuracy and latency of BERT-based models on the target device, NVIDIA Jetson TX2, and removes the hardware-related delays from the search loop. Using a Random Forest regressor, our predictors outperform the state-of-the-art and achieve up to 57x speedup while finding a set of near-optimal models.
Negin Firouzian, Seyyed Hasan Mozafari, James J. Clark, Warren J. Gross, Brett H. Meyer
CODES+ISSS3
2022 Consistency driven Sequential Transformers Attention Model for Partially Observable Scenes
abstract
Most hard attention models initially observe a complete scene to locate and sense informative glimpses, and predict class-label of a scene based on glimpses. However, in many applications (e.g., aerial imaging), observing an entire scene is not always feasible due to the limited time and resources available for acquisition. In this paper, we develop a Sequential Transformers Attention Model (STAM) that only partially observes a complete image and predicts informative glimpse locations solely based on past glimpses. We design our agent using DeiT-distilled [44] and train it with a one-step actorcritic algorithm. Furthermore, to improve classification performance, we introduce a novel training objective, which enforces consistency between the class distribution predicted by a teacher model from a complete image and the class distribution predicted by our agent using glimpses. When the agent senses only 4% of the total image area, the inclusion of the proposed consistency loss in our training objective yields 3% and 8% higher accuracy on ImageNet and fMoW datasets, respectively. Moreover, our agent outperforms previous state-of-the-art by observing nearly 27% and 42% fewer pixels in glimpses on ImageNet and fMoW.
Samrudhdhi B. Rangrej, Chetan L. Srinidhi, James J. Clark
CVPR3
2022 A Framework for Video-Text Retrieval with Noisy Supervision
abstract
A key challenge in extending vision-linguistic models to new video domains is curating large annotated datasets. We propose a framework that leverages videos with noisy linguistic descriptions, such as sports broadcasts, to train a model using an uncurated dataset. We introduce an unsupervised model that uses the corpus membership between a target and an auxiliary corpus to assign a relevance probability to the linguistic description of examples in the target domain. We examine these probabilities to evaluate the effect of noisy data in the video-text retrieval task. Our framework provides a domain-invariant recipe for enhancing multi-modal datasets by reducing the noise without requiring the costly manual curation effort. We show that our unsupervised model improves the performance of the video-text retrieval model using readily available hockey broadcast videos with closed-captioning. Furthermore, we propose a multi-modal cross-correlation objective function to obtain additional performance gains. We showcase our proposed framework in the context of a new multi-modal dataset of temporally labeled hockey videos with noisy textual descriptions.
Zahra Vaseqi, Pengnan Fan, James J. Clark, Martin D. Levine
ICMI3
2022 CES-KD: Curriculum-based Expert Selection for Guided Knowledge Distillation
abstract
Knowledge distillation (KD) is an effective tool for compressing deep classification models for edge devices. However, the performance of KD is affected by the large capacity gap between the teacher and student networks. Recent methods have resorted to a multiple teacher assistant (TA) setting for KD, which sequentially decreases the size of the teacher model to relatively bridge the size gap between these models. This paper proposes a new technique called Curriculum Expert Selection for Knowledge Distillation (CES-KD) to efficiently enhance the learning of a compact student under the capacity gap problem. This technique is built upon the hypothesis that a student network should be guided gradually using stratified teaching curriculum as it learns easy (hard) data samples better and faster from a lower (higher) capacity teacher network. Specifically, our method is a gradual TA-based KD technique that selects a single teacher per input image based on a curriculum driven by the difficulty in classifying the image. In this work, we empirically verify our hypothesis and rigorously experiment with CIFAR-10, CIFAR-100, CINIC-10, and ImageNet datasets and show improved accuracy on VGG-like models, ResNets, and WideResNets architectures.
Ibtihel Amara, Maryam Ziaeefard, Brett H. Meyer, Warren J. Gross, James J. Clark
ICPR5
2022 Efficient Fine-Tuning of BERT Models on the Edge
abstract
Resource-constrained devices are increasingly the deployment targets of machine learning applications. Static models, however, do not always suffice for dynamic environments. On-device training of models allows for quick adaptability to new scenarios. With the increasing size of deep neural networks, as noted with the likes of BERT and other natural language processing models, comes increased resource requirements, namely memory, computation, energy, and time. Furthermore, training is far more resource intensive than inference. Resource-constrained on-device learning is thus doubly difficult, especially with large BERT-like models. By reducing the memory usage of fine-tuning, pre-trained BERT models can become efficient enough to fine-tune on resource-constrained devices. We propose Freeze And ReconFigure (FAR), a memory-efficient training regime for BERT-like models that reduces the memory usage of activation maps during fine-tuning by avoiding unnecessary parameter updates. FAR reduces fine-tuning time on the DistilBERT model and CoLA dataset by 30 %, and time spent on memory operations by 47%. More broadly, reductions in metric performance on the GLUE and SQuAD datasets are around 1% on average.
Danilo Vucetic, Mohammadreza Tayaranian, Maryam Ziaeefard, James J. Clark, Brett H. Meyer, Warren J. Gross
ISCAS4
2021 A Probabilistic Hard Attention Model For Sequentially Observed Scenes
Samrudhdhi B. Rangrej, James J. Clark
BMVC2
2021 Task dependent deep LDA pruning of neural networks
Qing Tian 0003, Tal Arbel, James J. Clark
Comput. Vis. Image Underst.3
2019 Decoupled Hybrid 360° Panoramic Stereo Video
abstract
Although virtual reality (VR) and 360° panoramic videos amplify the viewer's immersion, they bring unique challenges from a content creator's perspective. As the audience gazes all around, the intended region of interest may not be seen at the right moment. In this paper, we present a novel hybrid 360° panoramic video camera rig that has both stereoscopic and monoscopic regions. Such non-uniform distribution of disparity serves as a new cinematographic tool. We also propose a custom camera rig capable of recording the hybrid format and is meant to be used similarly to a conventional camera: track the subject of interest and keep it inside the stereoscopic region. Camera tracking reduces motion blur of the subject. The rotational motion caused by tracking are decoupled from the stitched output video to reduce cybersickness. The stitched hybrid format was acceptable when tested on three different settings where the camera rig went through: no motion, yaw rotation, and any motion.
Kevin Kyunghwan Ra, James J. Clark
3DV2
2018 Going From Image to Video Saliency: Augmenting Image Salience With Dynamic Attentional Push
abstract
We present a novel method to incorporate the recent advent in static saliency models to predict the saliency in videos. Our model augments the static saliency models with the Attentional Push effect of the photographer and the scene actors in a shared attention setting. We demonstrate that not only it is imperative to use static Attentional Push cues, noticeable performance improvement is achievable by learning the time-varying nature of Attentional Push. We propose a multi-stream Convolutional Long Short-Term Memory network (ConvLSTM) structure which augments state-of-the-art in static saliency models with dynamic Attentional Push. Our network contains four pathways, a saliency pathway and three Attentional Push pathways. The multi-pathway structure is followed by an augmenting convnet that learns to combine the complementary and time-varying outputs of the ConvLSTMs by minimizing the relative entropy between the augmented saliency and viewers fixation patterns on videos. We evaluate our model by comparing the performance of several augmented static saliency models with state-of-the-art in spatiotemporal saliency on three largest dynamic eye tracking datasets, HOLLYWOOD2, UCF-Sport and DIEM. Experimental results illustrates that solid performance gain is achievable using the proposed methodology.
Siavash Gorji, James J. Clark
CVPR2
2018 Structured deep Fisher pruning for efficient facial trait classification
Qing Tian 0003, Tal Arbel, James J. Clark
Image Vis. Comput.3
2017 Attentional Push: A Deep Convolutional Network for Augmenting Image Salience with Shared Attention Modeling in Social Scenes
abstract
We present a novel visual attention tracking technique based on Shared Attention modeling. By considering the viewer as a participant in the activity occurring in the scene, our model learns the loci of attention of the scene actors and use it to augment image salience. We go beyond image salience and instead of only computing the power of image regions to pull attention, we also consider the strength with which the scene actors push attention to the region in question, thus the term Attentional Push. We present a convolutional neural network (CNN) which augments standard saliency models with Attentional Push. Our model contains two pathways: an Attentional Push pathway which learns the gaze location of the scene actors and a saliency pathway. These are followed by a shallow augmented saliency CNN which combines them and generates the augmented saliency. For training, we use transfer learning to initialize and train the Attentional Push CNN by minimizing the classification error of following the actors gaze location on a 2-D grid using a large-scale gaze-following dataset. The Attentional Push CNN is then fine-tuned along with the augmented saliency CNN to minimize the Euclidean distance between the augmented saliency and ground truth fixations using an eye-tracking dataset, annotated with the head and the gaze location of the scene actors. We evaluate our model on three challenging eye fixation datasets, SALICON, iSUN and CAT2000, and illustrate significant improvements in predicting viewers fixations in social scenes.
Siavash Gorji, James J. Clark
CVPR2
2016 Shannon information based adaptive sampling for action recognition
abstract
This paper investigates the effects of sampling on action recognition performance. Currently, dense (regular grid) sampling and uniform random sampling are popular strategies that achieve state-of-the-art performance. However, they are data-blind and pay equal attention to locations of different informativeness. In this paper, a Shannon information based adaptive sampling approach is proposed for action recognition. Results of different sampling approaches are compared on three benchmark datasets: the basic KTH and the challenging HMDB51 and UCF101 datasets. The method is shown to improve recognition accuracy as well as computational efficiency over the current state-of-the-art using less than one percent of the total pixels.
Qing Tian 0003, Tal Arbel, James J. Clark
ICPR3
2016 Hierarchical Spatio-Temporal Probabilistic Graphical Model with Multiple Feature Fusion for Binary Facial Attribute Classification in Real-World Face Videos
abstract
Recent literature shows that facial attributes, i.e., contextual facial information, can be beneficial for improving the performance of real-world applications, such as face verification, face recognition, and image search. Examples of face attributes include gender, skin color, facial hair, etc. How to robustly obtain these facial attributes (traits) is still an open problem, especially in the presence of the challenges of real-world environments: non-uniform illumination conditions, arbitrary occlusions, motion blur and background clutter. What makes this problem even more difficult is the enormous variability presented by the same subject, due to arbitrary face scales, head poses, and facial expressions. In this paper, we focus on the problem of facial trait classification in real-world face videos. We have developed a fully automatic hierarchical and probabilistic framework that models the collective set of frame class distributions and feature spatial information over a video sequence. The experiments are conducted on a large real-world face video database that we have collected, labelled and made publicly available. The proposed method is flexible enough to be applied to any facial classification problem. Experiments on a large, real-world video database McGillFaces [1] of 18,000 video frames reveal that the proposed framework outperforms alternative approaches, by up to 16.96 and 10.13%, for the facial attributes of gender and facial hair, respectively.
Meltem Demirkus, Doina Precup, James J. Clark, Tal Arbel
IEEE Trans. Pattern Anal. Mach. Intell.3
2015 Half-Occluded Regions and Detection of Pseudoscopy
abstract
We propose that left- and right-half-occlusion regions contain information that can distinguish between stereoscopic and pseudoscopic display conditions. A machine vision method is presented based on this idea, which detects pseudo copy using only the histograms of left and right half-occlusion pixel locations. Two psychophysical experiments are described which study the ability of human viewers to detect, or be influenced by, pseudoscopic display. The results of this study show that, during free viewing of HD 3D video imagery, humans judged pseudoscopic imagery to be of lower quality than stereoscopic imagery. Subjects performed at a 72% rate in deciding whether a short 5-second video clip was presented stereoscopically or pseudoscopically. Subjects were observed to fixate on half-occlusion regions with a frequency of 15.8%, as opposed to a frequency of 7.8% indicated by random chance. Viewers more frequently (17.4%) fixated on half-occlusion regions when making correct decisions then when they were incorrect (11.9%).
Jonathan Bouchard, Yasin Nazzar, James J. Clark
3DV3
2015 Hierarchical temporal graphical model for head pose estimation and subsequent attribute classification in real-world videos
Meltem Demirkus, Doina Precup, James J. Clark, Tal Arbel
Comput. Vis. Image Underst.3
2014 Probabilistic Temporal Head Pose Estimation Using a Hierarchical Graphical Model
Meltem Demirkus, Doina Precup, James J. Clark, Tal Arbel
ECCV (1)3
2014 Multi-layer temporal graphical model for head pose estimation in real-world videos
abstract
Head pose estimation has been receiving a lot of attention due to its wide range of possible applications. However, most approaches in the literature have focused on head pose estimation in controlled environments. Head pose estimation has recently begun to be applied to real-world environments. However, the focus has been on estimation from single images or video frames. Furthermore, most approaches frame the problem as classification into a set of coarse pose bins, rather than performing continuous pose estimation. The proposed multi-layer probabilistic temporal graphical model robustly estimates continuous head pose angle while leveraging the strengths of multiple features into account. Experiments performed on a large, real-world video database show that our approach not only significantly outperforms alternative head pose approaches, but also provides a pose probability assigned at each video frame, which permits robust temporal, probabilistic fusion of pose information over the entire video sequence.
Meltem Demirkus, Doina Precup, James J. Clark, Tal Arbel
ICIP3
2014 Robust semi-automatic head pose labeling for real-world face video sequences
Meltem Demirkus, James J. Clark, Tal Arbel
Multim. Tools Appl.2
2012 Enhancing visuospatial map learning through action on cellphones
abstract
The visuospatial learning of a map on cellphone displays was examined. The spatial knowledge of human participants was assessed after they had learned the relative positions of London Underground stations on a map via passive, marginally active, or active exploration. Following learning, the participants were required to answer questions in relation to the spatial representation and distribution of the stations on the map. Performances were compared between conditions involving (1) without auditory cues versus continuous auditory cues; (2) without auditory cues versus noncontinuous auditory cues; and (3) continuous auditory cues versus noncontinuous auditory cues. Results showed that the participants perfomed better following active and marginally-active explorations, as compared to purely passive learning. These results also suggest that under specific conditions (i.e., continuous sound with extremely fast tempo) there is no benefit to spatial abilities from active exploration over passive observation; while continuous sound with moderate to fast tempo is effective for simple actions (i.e., key press).
Mounia Ziat, Carmen E. Au, Amin Haji Abolhassani, James J. Clark
ACM Trans. Appl. Percept.4
2011 Spatial and probabilistic codebook template based head pose estimation from unconstrained environments
abstract
In unconstrained environments, head pose detection can be very challenging due to the joint and arbitrary occurrence of facial expressions, background clutter, partial occlusions and illumination conditions. Despite the wide range of head pose literature, most current methods can address this problem only up to a certain degree, and mostly for restricted scenarios. In this paper, we address the problem of head pose classification from real world images with large appearance variation. We represent each pose with a probabilistic and spatial template learned from facial codewords. The inference of the best template representing a test image is achieved probabilistically and spatially at the codebook. The experimental results are obtained from 5500 video frames collected under different illumination and background conditions. Our probabilistic framework is shown to outperform the current state-of-the-art in head pose classification.
Meltem Demirkus, Boris N. Oreshkin, James J. Clark, Tal Arbel
ICIP3
2011 Visual Task Inference Using Hidden Markov Models
Amin Haji Abolhassani, James J. Clark
IJCAI2
2011 Mirrormap: augmenting 2d mobile maps with virtual mirrors
abstract
In this paper, we describe our MirrorMap system, a 2D mobile map system that is augmented with live videos. In most major modern cities, traffic cameras and publicly accessible webcams are in abundance. These cameras provide live coverage of given city and are often positioned such that they have superior views of the scene. As such, it would be beneficial if a person who is wayfinding could have access to these video feeds, to acquire greater information about the area as they are route planning. Moreover, it would be beneficial if the system could provide the video feeds in a natural and familiar way that maintains the spatial relationships between the position of the corresponding camera and the person. We adopt a method called Virtual Mirroring. For each camera source, we place a virtual mirror in its stead. The mirror reflects the feed from the camera, and what results is the appearance that there are mirrors located in the position of the cameras. Akin to the well placed mirrors in convenience stores that provide shopkeepers with views of aisles he or she would not normally be able to see from the cash register, the user can point her (or his) mobile device in the direction of the source camera and see the virtual mirrors that are displaying the video feed. By so doing, the user can see additional views of the environment she would not normally be able to see from where she is standing. This additional information can inform her route planning.
Carmen E. Au, Victor Ng, James J. Clark
Mobile HCI3
2011 Integrating multiple views with virtual mirrors to facilitate scene understanding
abstract
In this article, an image integration technique called Virtual Mirroring (VM) is evaluated. VM is a technique that combines multiple 2D views of a 3D scene into a single composite image by overlaying views onto virtual mirrors. Given multiple views of a scene, one view is augmented with the remaining views by placing virtual mirrors on the first view and overlaying onto them the corresponding remaining views. Unlike a standard array presentation, where 2D views are not integrated and simply placed adjacent to one another, the VM presentation preserves the relative location, orientation, and scale between views. As such, it is our contention that humans will fare better at performing certain visual tasks, such as scene identification, when viewing a 3D scene via a VM presentation than when viewing an array presentation. We performed an experiment on 12 participants, where participants were required to identify 96 scenes both with a VM and an array presentation and we compared their % correctness and response times. Moreover, we studied the effects of adding an auditory attentional load on performance. We found that regardless of load, participants were able to identify scenes using VM presentation with greater accuracy and at greater speeds.
Carmen E. Au, James J. Clark
ACM Trans. Appl. Percept.2
2010 Photometric stereo using LCD displays
James J. Clark
Image Vis. Comput.1
2009 A sequential Bayesian approach to color constancy using non-uniform filters
Sandra Skaff, Tal Arbel, James J. Clark
Comput. Vis. Image Underst.3
2008 Estimating Surface Reflectance Spectra for Underwater Color Vision
abstract
This paper introduces a novel mathematical approach to surface spectral reflectance estimation in unknown underwater environments using uncalibrated color cameras. The approach derives surface spectral estimates without explicitly modeling the underwater medium characteristics such as light scattering and absorption. The latter two phenomena are dependent upon two parameters, which are the distance of the object from the camera and the depth of the object in water. The proposed approach does not require these parameters to be specified in advance. Spectral models are useful for underwater applications, where subtle differences in color need to be distinguished. Such models are also useful for fusing information from multiple images. We show that the proposed approach yields promising results. 1
Sandra Skaff, James J. Clark, Ioannis M. Rekleitis
BMVC2
2008 Video game design using an eye-movement-dependent model of visual attention
abstract
Eye movements can be used to infer the allocation of covert attention. In this article, we propose to model the allocation of attention in a task-dependent manner based on different eye movement conditions, specifically fixation and pursuit. We show that the image complexity at eye fixation points during fixation, and the pursuit direction during pursuit are significant factors in attention allocation. Results of the study are applied to the design of an interactive computer game. Real-time eye movement information is taken as one of inputs for the game. The utility of such eye information for controlling game difficulty is shown.
Li Jie, James J. Clark
ACM Trans. Multim. Comput. Commun. Appl.2
2007 Game Design Guided by Visual Attention
Li Jie, James J. Clark
ICEC2
2004 An Unsupervised, Online Learning Framework for Moving Object Detection
Vinod Nair, James J. Clark
CVPR (2)2
2004 A Temporal Stability Approach to Position and Attention-Shift-Invariant Recognition
abstract
Incorporation of visual-related self-action signals can help neural networks learn invariance. We describe a method that can produce a network with invariance to changes in visual input caused by eye movements and covert attention shifts. Training of the network is controlled by signals associated with eye movements and covert attention shifting. A temporal perceptual stability constraint is used to drive the output of the network toward remaining constant across temporal sequences of saccadic motions and covert attention shifts. We use a four-layer neural network model to perform the position-invariant extraction of local features and temporal integration of invariant presentations of local features in a bottom-up structure. We present results on both simulated data and real images to demonstrate that our network can acquire both position and attention shift invariance.
Muhua Li, James J. Clark
Neural Comput.2
2001 Guest Editor's Introduction
James J. Clark
Int. J. Comput. Vis.1
2001 Active Shape-from-Shadows with Controlled Illuminant Trajectories
James J. Clark, Lei Wang 0032
Int. J. Comput. Vis.1
2001 View-based route-learning with self-organizing neural networks
F. Hamze, James J. Clark
Image Vis. Comput.2
2000 A Temporal-Difference Model of Perceptual Stability in Color Vision
abstract
The authors consider the problem of how humans can maintain a stable perception of object color across saccades in spite of the changes in sensory input caused by the spatially nonhomogeneous receptor spectral sensitivities. The authors propose a method, based on a temporal-difference reinforcement learning scheme, for constructing associations between pre- and post-motor stimuli and which yields a constant color perception across saccades.
James J. Clark, J. Kevin O'Regan
ICPR1
1999 An Integral Formulation for Differential Photometric Stereo
abstract
In this paper we present an integral formulation of the active differential photometric stereo algorithm proposed by Clark (1992) and by Iwahori et al. [1992, 1994). The algorithm presented in this paper does not require measurement of derivatives of image quantities, but requires instead the computation of integrals of image quantities. Thus the algorithm is more robust to sensor noise and light source position errors than the Clark-Iwahori algorithm. We show that the algorithm presented in the paper can be efficiently implemented in practice with a planar distributed light source, and present experimental results demonstrating the efficacy of the algorithm.
James J. Clark, Holly Pekau
CVPR1
1998 Spatial Attention and Saccadic Camera Motion
abstract
An important aspect of computer-controlled camera motion systems is that of the generation of saccadic movements, which shift the camera gaze quickly from one fixation position to another. Recent psychophysical experiments suggest that there exists a causal connection between spatial shifts in visual attention and the production of saccadic eye movements in humans. Motivated by this experimental evidence, we propose a winner-take-all based model of exogenous spatial attention, involving both sustained and transient feature detection channels, and link it to the targetting and triggering of saccadic eye movements. We show that this model accounts for a range of oculomotor phenomena observed in human subjects. We describe the application of this model to a robotic camera gaze control system.
James J. Clark
ICRA1
1997 Trajectories for optimal temporal integration in active vision systems
abstract
We describe a general technique for specifying trajectories of controllable imaging parameters in an active vision system so that temporal integration processes are optimized. The technique assumes that a Kalman filter is used to perform the temporal integration of measurements and is based on determining, at each point in time, the set of imaging parameter values that minimizes the trace of the state estimate error covariance matrix. We present the application of this technique to the active vision task of extracting the location and orientation of a plane from shadows cast on it with a position controlled light source.
James J. Clark, Lei Wang 0032
ICRA1
1995 Parameterized Surface Fitting via MAP Estimation for Binocular Stereo
abstract
We present a novel method for reconstructing three dimensional surfaces from stereo intensity data. We employ a set of competing surface hypotheses based on parameterized models. We use maximum a posteriori (MAP) estimation and demonstrate a connection to the Hough transform. Experimental results are given showing the effectiveness of the algorithm.
Michael J. Weisman, Alan L. Yuille, James J. Clark
ICRA3
1994 MIMD Image Analysis with Local Agents
abstract
Complex systems that are typically thought of as exhibiting "life-like" qualities are observed to have many aspects in common with MIMD parallel computing systems. We present a computational model of a class of MIMD systems based on this observation. We describe the application of the model to the parallelization of image analysis tasks. We describe an example of our approach applied to the problem of extracting the circuit design from the layout image of a CMOS integrated circuit.>
James J. Clark, Robert P. Hewes
ICIP (3)1
1994 Active Shape and Depth Extraction from Shadow Images
abstract
We present a recursive estimation scheme for determining surface shape and depth. This technique relies on the control over the way in which shadows are cast in scenes, through variation of the position of illuminant. The estimation process is based on the iterated extended Kalman filter, with a near-optimal control of the movement of the light source to reduce the sensitivity of the state estimate to the sensory noise.>
Lei Wang 0032, James J. Clark
ICIP (1)2
1993 The Harvard Binocular Head
abstract
The ability to dynamically control the imaging parameters such as camera position, focus, and aperture is satisfied through the use of special purpose hardware such as a robotic head. This paper presents the design and control aspects of the Harvard Head, a binocular image acquisition system. We present three applications of the head in vision tasks concentrating on the computation of depth from controlled camera motion.
Nicola J. Ferrier, James J. Clark
Int. J. Pattern Recognit. Artif. Intell.2
1992 Active photometric stereo
abstract
An active vision technique for determining the absolute depth of surfaces is described. The algorithm assumes a very general model for the reflectance properties of the surface, and is valid for most of the shading models commonly used in computer vision work. The algorithm relies on the controlled motion of a point light source, which is not at infinity but relatively close to the surface and to the camera. The sensitivity of the computed depth values to errors in the measured quantities is derived, allowing a confidence measure for the depth to be determined. The confidence measure can aid in the estimation of accurate depth values from multiple image measurements taken over time. A method based on robust estimation that permits an unbiased estimate of the depth values to be obtained is presented. The results of experiments on synthetic and real-world imagery are reported, illustrating the efficacy of the active photometric stereo algorithm.>
James J. Clark
CVPR1
1991 VLSI sensori-motor systems
abstract
The authors describe their efforts in developing sensorimotor chips which contain arrays of sensing elements, circuitry for processing the raw sensor data into forms relevant to motion-related tasks, circuitry for generating motion signals based on the processed sensor data and the goals of the system, and an operating system which selects a unique motor command from a set of usually conflicting motion signals. Sensors of this kind are intended for use in robots. They would greatly reduce the computational burden from that imposed with standard sensing techniques that use high bandwidth video cameras or even tactile sensing arrays, since they would generate motor signals directly rather than sensor signals from which another computer would have to generate motor signals.>
James J. Clark, Daniel J. Friedman
ICRA1
1990 On the sampling and reconstruction of time-warped bandlimited signals
abstract
The sampling and reconstruction problem for time-warped bandlimited signals is examined, and its relationship to the more general sampling problem for nonbandlimited signals is discussed. In particular, time-warped bandlimited signals in L/sup 2/ are considered, covering approximation of L/sup 2/ signals, representational ambiguity, and approximation with nonaffine warpings.>
Douglas Cochran, James J. Clark
ICASSP2
1990 Shape from shading via the fusion of specular and Lambertian image components
abstract
A closed-form analytical solution to the problem of obtaining surface-normal information from the specular and Lambertian components of the image of a surface is presented. It is shown that this algebraic approach to the fusion of data is extremely sensitive to noise, and therefore two alternative approaches based on the minimization of energy functionals are provided. The first approach weights the specular and Lambertian information uniformly over the image with respect to a smoothness constraint. The second approach weights the specular and Lambertian component adaptively according to a measure of the sensitivity of the algebraically derived surface normals to noise in the measured image components. This results in a greater dependence on the smoothness constraint in the parts of the image where the surface reconstruction process is most sensitive to noise and provides a more accurate reconstruction of the surface than the uniform weighting technique.>
James J. Clark, Alan L. Yuille
ICPR (1)1
1990 On local detection of moving edges
abstract
The authors propose a detection framework with multiple velocity channels for moving edges based on a generalization of J. Canny's edge detector (1986). Finite state machines (FSMs) are set up at discrete lattice points in the image plane and operate based on the outputs of all velocity channels. The outputs of the FSMs denote whether there are edges at their corresponding positions, and their states record the edge velocities. In the temporal dimension, statistics are attached to the edges to aid in removing phantom edges.>
Ten-lee Hwang, James J. Clark
ICPR (1)2
1990 A spatio-temporal generalization of Canny's edge detector
abstract
Moving step edges are modeled as the product of a deterministic function in space and a stochastic function in time which captures the edge shapes and the temporal uncertainties, respectively. Under J. Canny's (IEEE Trans. on Pattern Analysis and Machine Intelligence, vol.PAMI-8, p.679-98, Nov. 1986) original optimality criteria, a set of optimal edge detectors is derived. They are in a product form, i.e., a product of a spatial function and a temporal function. The spatial function is Canny's edge detector in one dimension and the temporal function can be well approximated by the exponential function. Generalizing Canny's edge detector to the temporal domain is not only theoretically interesting, but also practically useful. The generalization of Canny's edge detectors provides better immunity to noise and can serve as one of the tools in understanding the temporal behavior of moving edges. They have been used in a data-fusion framework to detect moving edges and their normal velocities simultaneously. For completeness, the authors derive some properties of the optimal edge detectors and compare them with Gabor filters.>
Ten-lee Hwang, James J. Clark
ICPR (1)2
1990 Management of sensory-motor activity in mobile robots
abstract
Consideration is given to the management of conflicting demands on the use of the motor units of the robot from the active sensors and manipulatory processes. An operating-system-like facility is introduced for managing the conflicting motor requests produced by multiple active sensing and manipulation tasks. A proposed implementation of this sensory-motor management facility based on the input-output-timed-automata (IOTA) abstraction is introduced. The IOTA abstraction is general enough to allow the enforcement of temporal constraints and to permit the specification and modification of task priorities based on the goals and current state of the robot. The IOTA-based sensory-motor operating system acts as a scheduler for various motor requests submitted by the active sensor and manipulation systems. The sensory-motor system is more general than the subsumption architecture of R. Brooks (1986) and can implement in a straightforward fashion alteration of priorities.>
Azer Bestavros, James J. Clark, Nicola J. Ferrier
ICRA2
1989 A depth recovery algorithm using defocus information
abstract
Pentland (1987) proposed an algorithm for sensing scene depth by measuring the amount of defocus of an image. The algorithm is interesting in the sense that no correspondence problems are involved. The authors discuss the difference between shape from defocus and shape from focus and then present a two-phase algorithm where the defocus process is modeled as a two-dimensional Gaussian point-spread function. During the first phase (calibration phase), a camera system parameter is determined off-line. In the next phase (depth-recovery phase), this parameter is used to recover on-line the scene depth by taking two images of the same scene, but with a different amount of defocus. Some implementation issues are addressed, and test results on real images are provided.>
Ten-lee Hwang, James J. Clark, Alan L. Yuille
CVPR2
1989 Control of visual attention in mobile robots
abstract
The authors describe a control system for a binocular image acquisition mechanism, for use in mobile robotic systems, which allows shifts in focus of attenuation to be made in a natural, device-independent manner. The control method is based on the modal control technique proposed by R.W. Brockett (Proc. IEEE Robotics Autom. Conf., 1988). The shifts are accomplished by altering the feedback gains applied to the visual feedback paths in the position and velocity control loops of the binocular camera system. By altering these gains, a feature-selection operation can be performed by which the saliency of a given feature is enhanced, while the saliency of other features is reduced. Two experiments performed with the system to demonstrate modal control of attention are discussed.>
James J. Clark, Nicola J. Ferrier
ICRA1
1989 Authenticating Edges Produced by Zero-Crossing Algorithms
abstract
It is shown that zero-crossing edge detection algorithms can produce edges that do not correspond to significant image intensity changes. Such edges are called phantom or spurious. A method for classifying zero crossings as corresponding to authentic or phantom edges is presented. The contrast of an authentic edge is shown to increase and the contrast of phantom edges to decrease with a decrease in the filter scale. Thus, a phantom edge is truly a phantom in that the closer one examines it, the weaker it becomes. The results of applying the classification schemes described to synthetic and authentic signals in one and two dimensions are given. The significance of the phantom edges is examined with respect to their frequency and strength relative to the authentic edges, and it is seen that authentic edges are denser and stronger, on the average, than phantom edges.>
James J. Clark
IEEE Trans. Pattern Anal. Mach. Intell.1
1988 Modal Control Of An Attentive Vision System
abstract
Cambridge, MA A vision system for use in a mobile robot system, or in n fixed multi-tasking industrial robot requires attentive control. Attentive control refers to the process by which the direction of gaze of the visual sensors are determined, nlong with the determination of what processing is required to be applied to the sensed images based on the goals of the robot and the tasks it is performing. This paper describes the implementation of a. mnt,ion control system which allows the attentive control of a binocular vision system. Attentive inputs to the system specify the type of visual feedback that the oculo-motor control system will use. The MDL language developed by Brockett [7] i8 used to communicate between the attentive planner and the motion controller.
James J. Clark, Nicola J. Ferrier
ICCV1
1988 A magnetic field based compliance matching sensor for high resolution, high compliance tactile sensing
abstract
A description is given of a general approach for producing high-resolution tactile sensors that are highly compliant. The approach is based on the idea of compliance matching, wherein the high compliance of the contacting element is matched to the low compliance of the sensing element. A prototype tactile sensor is proposed which uses magnetic fields as the matching medium. The author details the design of an important component of this device, 64*64 element array of magnetic field sensors.>
James J. Clark
ICRA1
1988 Singularity Theory and Phantom Edges in Scale Space
abstract
The process of detecting edges in a one-dimensional signal by finding the zeros of the second derivative of the signal can be interpreted as the process of detecting the critical points of a general class of contrast functions that are applied to the signal. It is shown that the second derivative of the contrast function at the critical point is related to the classification of the associated edge as being phantom or authentic. The contrast of authentic edges decreases with filter scale, while the contrast of phantom edges are shown to increase with scale. As the filter scale increases, an authentic edge must either turn into a phantom edge or join with a phantom edge and vanish. The points in the scale space at which these events occur are seen to be singular points of the contrast function. Using ideas from singularity, or catastrophy theory, the scale map contours near these singular points are found to be either vertical or parabolic.>
James J. Clark
IEEE Trans. Pattern Anal. Mach. Intell.1
1985 A systolic parallel processor for the rapid computation of multiresolution edge images using the ▿2G operator
James J. Clark, Peter D. Lawrence
J. Parallel Distributed Comput.1