Bertram E. Shi

dblp:37/5620 · also Bertram Emil Shi · DBLP profile ↗
← Back
92ranked-venue papers
7as first author
15since 2021 · last 2025
0000-0001-9167-7495ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 62 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 2 first-author · 3 since 2021Systems, architecture and hardware · 17 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 10 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021
YearPublicationVenuePosition
2025 Dynamic Prompting Improves Turn-taking in Embodied Spoken Dialogue Systems
abstract
The ability to coordinate turn taking during spoken dialogue is crucial for an embodied spoken dialogue system (SDS), e.g., in a humanoid robot. The SDS needs to model transitions in the conversational floor, which describes each party’s stance (either speaking or listening). Further, the SDS needs to signal its perception of the floor to the human, so that they can coordinate floor transitions and resolve conflicts. Conventional SDS employ standalone modules to control floor transitions but do not produce timely and appropriate responses. Recent end-to-end audio LLMs generate responses quickly, but do not coordinate floor transitions as accurately. In this work, we propose an SDS architecture that dynamically adjusts its prompts to an end-to-end audio LLM based upon its perception of the conversational floor state. The LLM output determines not only the audio output, but also the perceived floor state. This enables the system to signal its stance to the human, both when listening and when speaking. We conducted an experiment where a humanoid robot administered a semi-structured interview with human subjects. Results show that, compared with baseline systems using static prompts, dynamic prompting enables the LLM to model floor transitions more accurately, to generate more appropriate signalling, and to interrupt less, leading to smoother turn-taking in dialogue.
Dingdong Liu, Xiaoyu Mo, Fugee Tsung, Xiaojuan Ma, Bertram E. Shi
RO-MAN6
2024 OAT: Object-Level Attention Transformer for Gaze Scanpath Prediction
Yini Fang, Jingling Yu, Haozheng Zhang, Ralf van der Lans, Bertram E. Shi
ECCV (55)5
2024 Merging Multiple Datasets for Improved Appearance-Based Gaze Estimation
Liang Wu 0009, Bertram E. Shi
ICPR (14)2
2024 A Humanoid Robot Dialogue System Architecture Targeting Patient Interview Tasks
abstract
Humanoid robots are promising approach to automating patient interviews routinely conducted by medical staff. Their human-like appearance enables them to use the full gamut of verbal and behavioral cues that are critical to a successful interview. On the other hand, anthropomorphism can induce expectations of human-level performance by the robot. Not meeting such expectations degrades the quality of interaction. Specifically, humans expect rich real-time interactions during speech exchange, such as backchanneling and barge-ins. The nature of the patient interview task differs from most other scenarios where task oriented dialogue systems have been used, as there is increased potential of engagement breakdown during interaction. We describe a dialogue system architecture that improves the performance of humanoid robots on the patient interview task. Our architecture adds a nested inner real-time control loop to improve the timeliness of the robot’s responses based on the notion of "stance", an elaboration of the concept of a "turn", common in most existing dialogue systems. It also expands the dialogue state to monitor not only task progress, but also human engagement. Experiments using a humanoid robot running our proposed architecture reveal improved performance on interview tasks in terms of the perceived timeliness of responses and users’ impressions of the system.
Dingdong Liu, Yejin Bang, Ho Shu Chan, Rita Frieske, Hoo Choun Chung, Jay Nieles, Tianjia Zhang, Kien T. Pham 0001, Wai Yi Rosita Cheng, Yini Fang, Qifeng Chen 0001, Pascale Fung, Xiaojuan Ma, Bertram E. Shi
RO-MAN15
2024 User Engagement Correlates Better with Behavioral than Physiological Measures in a Virtual Reality Robotic Rehabilitation System
abstract
Robotic systems to assist with movement rehabil-itation are transitioning from providing fixed pre-programmed assistance towards adaptive challenge-oriented strategies that present patients with tasks that are demanding yet achiev-able. This promotes active engagement, which is crucial for stimulating neural plasticity and promoting recovery. While it has been well established that varying the challenge level can affect user engagement, measuring engagement during task performance has received less attention. To investigate this issue, we developed a virtual reality (VR) robotic system for upper limb rehabilitation using a line-tracing task that measures physiological and behavioral signals. Challenge level can be modulated by introducing force noise disturbance. We con-ducted a preliminary study on 12 participants, measuring user engagement and physiological/behavioral signals at different noise (challenge) levels. Our findings align with the predictions of flow channel theory. Engagement peaks at an intermediate challenge level. While past work considered only physiological measures, our results reveal that behavioral measures are better correlated with user engagement. Physiological measures correlate better with arousal. This work takes a step toward systems that dynamically adapt task parameters to optimize user engagement.
Haofei Wang 0001, Bertram E. Shi
SMC3
2023 RMES: Real-Time Micro-Expression Spotting Using Phase From Riesz Pyramid
abstract
Micro-expressions (MEs) are involuntary and subtle facial expressions that are thought to reveal feelings people are trying to hide. ME spotting detects the temporal intervals containing MEs in videos. Detecting such quick and subtle motions from long videos is difficult. Recent works leverage detailed facial motion representations, such as the optical flow, and deep learning models, leading to high computational complexity. To reduce computational complexity and achieve real-time operation, we propose RMES, a real-time ME spotting framework. We represent motion using phase computed by Riesz Pyramid, and feed this motion representation into a three-stream shallow CNN, which predicts the likelihood of each frame belonging to an ME. In comparison to optical flow, phase provides more localized motion estimates, which are essential for ME spotting, resulting in higher performance. Using phase also reduces the required computation of the ME spotting pipeline by 77.8%. Despite its relative simplicity and low computational complexity, our framework achieves state-of-the-art performance on two public datasets: CAS(ME)2and SAMM Long Videos.
Yini Fang, Didan Deng, Liang Wu 0009, Frederic Jumelle, Bertram E. Shi
ICME5
2023 Towards High Performance Low Complexity Calibration in Appearance Based Gaze Estimation
abstract
Appearance-based gaze estimation from RGB images provides relatively unconstrained gaze tracking from commonly available hardware. The accuracy of subject-independent models is limited partly by small intra-subject and large inter-subject variations in appearance, and partly by a latent subject-dependent bias. To improve estimation accuracy, we have previously proposed a gaze decomposition method that decomposes the gaze angle into the sum of a subject-independent gaze estimate from the image and a subject-dependent bias. Estimating the bias from images outperforms previously proposed calibration algorithms, unless the amount of calibration data is prohibitively large. This paper extends that work with a more complete characterization of the interplay between the complexity of the calibration dataset and estimation accuracy. In particular, we analyze the effect of the number of gaze targets, the number of images used per gaze target and the number of head positions in calibration data using a new NISLGaze dataset, which is well suited for analyzing these effects as it includes more diversity in head positions and orientations for each subject than other datasets. A better understanding of these factors enables low complexity high performance calibration. Our results indicate that using only a single gaze target and single head position is sufficient to achieve high quality calibration. However, it is useful to include variability in head orientation as the subject is gazing at the target. Our proposed estimator based on these studies (GEDDNet) outperforms state-of-the-art methods by more than 6.3%. One of the surprising findings of our work is that the same estimator yields the best performance both with and without calibration. This is convenient, as the estimator works well "straight out of the box," but can be improved if needed by calibration. However, this seems to violate the conventional wisdom that train and test conditions must be matched. To better understand the reasons, we provide a new theoretical analysis that specifies the conditions under which this can be expected. The dataset is available at http://nislgaze.ust.hk. Source code is available at https://github.com/HKUST-NISL/GEDDnet.
Zhaokang Chen, Bertram E. Shi
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 HGCN-GJS: Hierarchical Graph Convolutional Network with Groupwise Joint Sampling for Trajectory Prediction
abstract
Pedestrian trajectory prediction is of great importance for downstream tasks, such as autonomous driving and mobile robot navigation. Realistic models of the social interactions within the crowd is crucial for accurate pedestrian trajectory prediction. However, most existing methods do not capture group level interactions well, focusing only on pairwise interactions and neglecting group-wise interactions. In this work, we propose a hierarchical graph convolutional network, HGCN-GJS, for trajectory prediction which well leverages group level interactions within the crowd. Furthermore, we introduce a joint sampling scheme that captures co-dependencies between pedestrian trajectories during trajectory generation. Based on group information, this scheme ensures that generated trajectories within each group are consistent with each other, but enables different groups to act more independently. We demonstrate that our proposed network achieves state of the art performance on all datasets we have considered.
Yuying Chen, Xiaodong Mei 0001, Bertram E. Shi, Ming Liu 0001
IROS4
2022 CI-AVSR: A Cantonese Audio-Visual Speech Datasetfor In-car Command Recognition
abstract
With the rise of deep learning and intelligent vehicles, the smart assistant has become an essential in-car component to facilitate driving and provide extra functionalities. In-car smart assistants should be able to process general as well as car-related commands and perform corresponding actions, which eases driving and improves safety. However, there is a data scarcity issue for low resource languages, hindering the development of research and applications. In this paper, we introduce a new dataset, Cantonese In-car Audio-Visual Speech Recognition (CI-AVSR), for in-car command recognition in the Cantonese language with both video and audio data. It consists of 4,984 samples (8.3 hours) of 200 in-car commands recorded by 30 native Cantonese speakers. Furthermore, we augment our dataset using common in-car background noises to simulate real environments, producing a dataset 10 times larger than the collected one. We provide detailed statistics of both the clean and the augmented versions of our dataset. Moreover, we implement two multimodal baselines to demonstrate the validity of CI-AVSR. Experiment results show that leveraging the visual signal improves the overall performance of the model. Although our best model can achieve a considerable quality on the clean test set, the speech recognition quality on the noisy data is still inferior and remains an extremely challenging task for real in-car speech recognition systems. The dataset and code will be released at https://github.com/HLTCHKUST/CI-AVSR.
Wenliang Dai, Samuel Cahyawijaya, Tiezheng Yu, Elham J. Barezi, Peng Xu 0008, Cheuk Tung Yiu, Rita Frieske, Holy Lovenia, Genta Indra Winata, Qifeng Chen 0001, Xiaojuan Ma, Bertram E. Shi, Pascale Fung
LREC12
2022 ASCEND: A Spontaneous Chinese-English Dataset for Code-switching in Multi-turn Conversation
abstract
Code-switching is a speech phenomenon occurring when a speaker switches language during a conversation. Despite the spontaneous nature of code-switching in conversational spoken language, most existing works collect code-switching data from read speech instead of spontaneous speech. ASCEND (A Spontaneous Chinese-English Dataset) is a high-quality Mandarin Chinese-English code-switching corpus built on spontaneous multi-turn conversational dialogue sources collected in Hong Kong. We report ASCEND’s design and procedure for collecting the speech data, including annotations. ASCEND consists of 10.62 hours of clean speech, collected from 23 bilingual speakers of Chinese and English. Furthermore, we conduct baseline experiments using pre-trained wav2vec 2.0 models, achieving a best performance of 22.69% character error rate and 27.05% mixed error rate.
Holy Lovenia, Samuel Cahyawijaya, Genta Indra Winata, Peng Xu 0008, Yan Xu 0012, Zihan Liu 0001, Rita Frieske, Tiezheng Yu, Wenliang Dai, Elham J. Barezi, Qifeng Chen 0001, Xiaojuan Ma, Bertram E. Shi, Pascale Fung
LREC13
2022 Automatic Speech Recognition Datasets in Cantonese: A Survey and New Dataset
abstract
Automatic speech recognition (ASR) on low resource languages improves the access of linguistic minorities to technological advantages provided by artificial intelligence (AI). In this paper, we address the problem of data scarcity for the Hong Kong Cantonese language by creating a new Cantonese dataset. Our dataset, Multi-Domain Cantonese Corpus (MDCC), consists of 73.6 hours of clean read speech paired with transcripts, collected from Cantonese audiobooks from Hong Kong. It comprises philosophy, politics, education, culture, lifestyle and family domains, covering a wide range of topics. We also review all existing Cantonese datasets and analyze them according to their speech type, data source, total size and availability. We further conduct experiments with Fairseq S2T Transformer, a state-of-the-art ASR model, on the biggest existing dataset, Common Voice zh-HK, and our proposed MDCC, and the results show the effectiveness of our dataset. In addition, we create a powerful and robust Cantonese ASR model by applying multi-dataset learning on MDCC and Common Voice zh-HK.
Tiezheng Yu, Rita Frieske, Peng Xu 0008, Samuel Cahyawijaya, Cheuk Tung Shadow Yiu, Holy Lovenia, Wenliang Dai, Elham J. Barezi, Qifeng Chen 0001, Xiaojuan Ma, Bertram E. Shi, Pascale Fung
LREC11
2021 Ensembling With a Fixed Parameter Budget: When Does It Help and Why?
abstract
Given a fixed parameter budget, one can build a single large neural network or create a memory-split ensemble: a pool of several smaller networks with the same total parameter count as the single network. A memory-split ensemble can outperform its single model counterpart (Lobacheva et al., 2020): a phenomenon known as the memory-split advantage (MSA). The reasons for MSA are still not yet fully understood. In particular, it is difficult in practice to predict when it will exist. This paper sheds light on the reasons underlying MSA using random feature theory. We study the dependence of the MSA on several factors: the parameter budget, the training set size, the L2 regularization and the Stochastic Gradient Descent (SGD) hyper-parameters. Using the bias-variance decomposition, we show that MSA exists when the reduction in variance due to the ensemble (\ie, \textit{ensemble gain}) exceeds the increase in squared bias due to the smaller size of the individual networks (\ie, \textit{shrinkage cost}). Taken together, our theoretical analysis demonstrates that the MSA mainly exists for the small parameter budgets relative to the training set size, and that memory-splitting can be understood as a type of regularization. Adding other forms of regularization, \eg L2 regularization, reduces the MSA. Thus, the potential benefit of memory-splitting lies primarily in the possibility of speed-up via parallel computation. Our empirical experiments with deep neural networks and large image datasets show that MSA is not a general phenomenon, but mainly exists when the number of training iterations is small.
Didan Deng, Bertram E. Shi
ACML2
2021 AVGCN: Trajectory Prediction using Graph Convolutional Networks Guided by Human Attention
abstract
Pedestrian trajectory prediction is a critical yet challenging task especially for crowded scenes. We suggest that introducing an attention mechanism to infer the importance of different neighbors is critical for accurate trajectory prediction in scenes with varying crowd size. In this work, we propose a novel method, AVGCN, for trajectory prediction utilizing graph convolutional networks (GCN) based on human attention (A denotes attention, V denotes visual field constraints). First, we train an attention network that estimates the importance of neighboring pedestrians, using gaze data collected as subjects perform a bird’s eye view crowd navigation task. Then, we incorporate the learned attention weights modulated by constraints on the pedestrian’s visual field into a trajectory prediction network that uses a GCN to aggregate information from neighbors efficiently. AVGCN also considers the stochastic nature of pedestrian trajectories by taking advantage of variational trajectory prediction. Our approach achieves state-of-the-art performance on several trajectory prediction benchmarks, and the lowest average prediction error over all considered benchmarks.
Yuying Chen, Ming Liu 0002, Bertram E. Shi
ICRA4
2021 Active head rolls enhance sonar-based auditory localization performance
abstract
Animals utilize a variety of active sensing mechanisms to perceive the world around them. Echolocating bats are an excellent model for the study of active auditory localization. The big brown bat (Eptesicus fuscus), for instance, employs active head roll movements during sonar prey tracking. The function of head rolls in sound source localization is not well understood. Here, we propose an echolocation model with multi-axis head rotation to investigate the effect of active head roll movements on sound localization performance. The model autonomously learns to align the bat's head direction towards the target. We show that a model with active head roll movements better localizes targets than a model without head rolls. Furthermore, we demonstrate that active head rolls also reduce the time required for localization in elevation. Finally, our model offers key insights to sound localization cues used by echolocating bats employing active head movements during echolocation.
Lakshitha P. Wijesinghe, Melville Wohlgemuth, Richard H. Y. So, Jochen Triesch, Cynthia F. Moss, Bertram E. Shi
PLoS Comput. Biol.6
2021 Using Eye Gaze to Enhance Generalization of Imitation Networks to Unseen Environments
abstract
Vision-based autonomous driving through imitation learning mimics the behavior of human drivers by mapping driver view images to driving actions. This article shows that performance can be enhanced via the use of eye gaze. Previous research has shown that observing an expert's gaze patterns can be beneficial for novice human learners. We show here that neural networks can also benefit. We trained a conditional generative adversarial network to estimate human gaze maps accurately from driver-view images. We describe two approaches to integrating gaze information into imitation networks: eye gaze as an additional input and gaze modulated dropout. Both significantly enhance generalization to unseen environments in comparison with a baseline vanilla network without gaze, but gaze-modulated dropout performs better. We evaluated performance quantitatively on both single images and in closed-loop tests, showing that gaze modulated dropout yields the lowest prediction error, the highest success rate in overtaking cars, the longest distance between infractions, lowest epistemic uncertainty, and improved data efficiency. Using Grad-CAM, we show that gaze modulated dropout enables the network to concentrate on task-relevant areas of the image.
Yuying Chen, Ming Liu 0001, Bertram E. Shi
IEEE Trans. Neural Networks Learn. Syst.4
2020 MIMAMO Net: Integrating Micro- and Macro-Motion for Video Emotion Recognition
abstract
Spatial-temporal feature learning is of vital importance for video emotion recognition. Previous deep network structures often focused on macro-motion which extends over long time scales, e.g., on the order of seconds. We believe integrating structures capturing information about both micro- and macro-motion will benefit emotion prediction, because human perceive both micro- and macro-expressions. In this paper, we propose to combine micro- and macro-motion features to improve video emotion recognition with a two-stream recurrent network, named MIMAMO (Micro-Macro-Motion) Net. Specifically, smaller and shorter micro-motions are analyzed by a two-stream network, while larger and more sustained macro-motions can be well captured by a subsequent recurrent network. Assigning specific interpretations to the roles of different parts of the network enables us to make choice of parameters based on prior knowledge: choices that turn out to be optimal. One of the important innovations in our model is the use of interframe phase differences rather than optical flow as input to the temporal stream. Compared with the optical flow, phase differences require less computation and are more robust to illumination changes. Our proposed network achieves state of the art performance on two video emotion datasets, the OMG emotion dataset and the Aff-Wild dataset. The most significant gains are for arousal prediction, for which motion information is intuitively more informative. Source code is available at https://github.com/wtomin/MIMAMO-Net.
Didan Deng, Zhaokang Chen, Yuqian Zhou, Bertram E. Shi
AAAI4
2020 CoMoGCN: Coherent Motion Aware Trajectory Prediction with Graph Representation
Yuying Chen, Bertram E. Shi, Ming Liu 0001
BMVC3
2020 Multitask Emotion Recognition with Incomplete Labels
abstract
We train a unified model to perform three tasks: facial action unit detection, expression classification, and valence-arousal estimation. We address two main challenges of learning the three tasks. First, most existing datasets are highly imbalanced. Second, most existing datasets do not contain labels for all three tasks. To tackle the first challenge, we apply data balancing techniques to experimental datasets. To tackle the second challenge, we propose an algorithm for the multitask model to learn from missing (incomplete) labels. This algorithm has two steps. We first train a teacher model to perform all three tasks, where each instance is trained by the ground truth label of its corresponding task. Secondly, we refer to the outputs of the teacher model as the soft labels. We use the soft labels and the ground truth to train the student model. We find that most of the student models outperform their teacher model on all the three tasks. Finally, we use model ensembling to boost performance further on the three tasks. Our code is publicly available1.1https://github.com/wtomin/multitask-Emotion-Recognition-withIncomplete-Labels
Didan Deng, Zhaokang Chen, Bertram E. Shi
FG3
2020 Offset Calibration for Appearance-Based Gaze Estimation via Gaze Decomposition
abstract
Appearance-based gaze estimation provides relatively unconstrained gaze tracking. However, subject-independent models achieve limited accuracy partly due to individual variations. To improve estimation, we propose a gaze decomposition method that enables low complexity calibration, i.e., using calibration data collected when subjects view only one or a few gaze targets and the number of images per gaze target is small. Lowering the complexity of calibration makes it more convenient and less timeconsuming for the user, and more widely applicable. Motivated by our finding that the inter-subject squared bias exceeds the intra-subject variance for a subject-independent estimator, we decompose the gaze estimate into the sum of a subject-independent term estimated from the input image by a deep convolutional network and a subject-dependent bias term. During training, both the weights of the deep network and the bias terms are estimated. During testing, if no calibration data is available, we can set the bias term to zero. Otherwise, the bias term can be estimated from images of the subject gazing at known gaze targets. Experimental results on three datasets show that without calibration, our method outperforms state-of-the-art by at least 6.3%. For low complexity calibration sets, our method outperforms other calibration methods. More complex calibration algorithms do not outperform our method until the size of the calibration set is excessively large. Even then, the gains obtained by alternatives are small, e.g., only 0.1° lower error for 64 gaze targets. Source code is available at https://github.com/czk32611/Gaze-Decomposition.
Zhaokang Chen, Bertram E. Shi
WACV2
2019 A gaze model improves autonomous driving
abstract
End-to-end behavioral cloning trained by human demonstration is now a popular approach for vision-based autonomous driving. A deep neural network maps drive-view images directly to steering commands. However, the images contain much task-irrelevant data. Humans attend to behaviorally relevant information using saccades that direct gaze towards important areas. We demonstrate that behavioral cloning also benefits from active control of gaze. We trained a conditional generative adversarial network (GAN) that accurately predicts human gaze maps while driving in both familiar and unseen environments. We incorporated the predicted gaze maps into end-to-end networks for two behaviors: following and overtaking. Incorporating gaze information significantly improves generalization to unseen environments. We hypothesize that incorporating gaze information enables the network to focus on task critical objects, which vary little between environments, and ignore irrelevant elements in the background, which vary greatly.
Yuying Chen, Lei Tai, Haoyang Ye, Ming Liu 0001, Bertram E. Shi
ETRA6
2019 Task-embedded online eye-tracker calibration for improving robustness to head motion
abstract
Remote eye trackers are widely used for screen-based interactions. They are less intrusive than head mounted eye trackers, but are generally quite sensitive to head movement. This leads to the requirement for frequent recalibration, especially in applications requiring accurate eye tracking. We propose here an online calibration method to compensate for head movements if estimates of the gaze targets are available. For example, in dwell-time based gaze typing it is reasonable to assume that for correct selections, the user's gaze target during the dwell-time was at the key center. We use this assumption to derive an eye-position dependent linear transformation matrix for correcting the measured gaze. Our experiments show that the proposed method significantly reduces errors over a large range of head movements.
Jimin Pi, Bertram E. Shi
ETRA2
2019 Gaze awareness improves collaboration efficiency in a collaborative assembly task
abstract
In building human robot interaction systems, it would be helpful to understand how humans collaborate, and in particular, how humans use others' gaze behavior to estimate their intent. Here we studied the use of gaze in a collaborative assembly task, where a human user assembled an object with the assistance of a human helper. We found that the being aware of the partner's gaze significantly improved collaboration efficiency. Task completion times were much shorter when gaze communication was available, than when it was blocked. In addition, we found that the user's gaze was more likely to lie on the object of interest in the gaze-aware case than the gaze-blocked case. In the context of human-robot collaboration systems, our results suggest that gaze data in the period surrounding verbal requests will be more informative and can be used to predict the target object.
Haofei Wang 0001, Bertram E. Shi
ETRA2
2019 Gaze Training by Modulated Dropout Improves Imitation Learning
abstract
Imitation learning by behavioral cloning is a prevalent method that has achieved some success in vision-based autonomous driving. The basic idea behind behavioral cloning is to have the neural network learn from observing a human expert's behavior. Typically, a convolutional neural network learns to predict the steering commands from raw driver-view images by mimicking the behaviors of human drivers. However, there are other cues, such as gaze behavior, available from human drivers that have yet to be exploited. Previous researches have shown that novice human learners can benefit from observing experts' gaze patterns. We present here that deep neural networks can also profit from this. We propose a method, gaze-modulated dropout, for integrating this gaze information into a deep driving network implicitly rather than as an additional input. Our experimental results demonstrate that gaze-modulated dropout enhances the generalization capability of the network to unseen scenes. Prediction error in steering commands is reduced by 23.5% compared to uniform dropout. Running closed loop in the simulator, the gaze-modulated dropout net increased the average distance travelled between infractions by 58.5%. Consistent with these results, the gazemodulated dropout net shows lower model uncertainty.
Yuying Chen, Lei Tai, Ming Liu 0001, Bertram E. Shi
IROS5
2019 Using Variable Dwell Time to Accelerate Gaze-Based Web Browsing with Two-Step Selection
abstract
In order to avoid the “Midas Touch” problem, gaze-based interfaces for selection often introduce a dwell time: a fixed amount of time the user must fixate upon an object before it is selected. Past interfaces have used a uniform dwell time across all objects. Here, we propose a gaze-based browser using a two-step selection policy with variable dwell time. In the first step, a command (e.g., “back” or “select”) is chosen from a menu using a dwell time that is constant across the different commands. In the second step, if the “select” command is chosen, the user selects a hyperlink using a dwell time that varies between different hyperlinks. We assign shorter dwell times to more likely hyperlinks and longer dwell times to less likely hyperlinks. In order to infer the likelihood each hyperlink will be selected, we have developed a probabilistic model of natural gaze behavior while surfing the web. We have evaluated a number of heuristic and probabilistic methods for varying the dwell times using both simulation and experiment. Our results demonstrate that varying dwell time improves the user experience in comparison with fixed dwell time, resulting in fewer errors and increased speed. While all of the methods for varying dwell time resulted in improved performance, the probabilistic models yielded much greater gains than the simple heuristics. The best performing model reduces error rate by 50% compared to 100ms uniform dwell time while maintaining a similar response time. It reduces response time by 60% compared to 300ms uniform dwell time while maintaining a similar error rate.
Zhaokang Chen, Bertram E. Shi
Int. J. Hum. Comput. Interact.2
2018 Appearance-Based Gaze Estimation Using Dilated-Convolutions
Zhaokang Chen, Bertram E. Shi
ACCV (6)2
2018 SLAM-based localization of 3D gaze using a mobile eye tracker
abstract
Past work in eye tracking has focused on estimating gaze targets in two dimensions (2D), e.g. on a computer screen or scene camera image. Three-dimensional (3D) gaze estimates would be extremely useful when humans are mobile and interacting with the real 3D environment. We describe a system for estimating the 3D locations of gaze using a mobile eye tracker. The system integrates estimates of the user's gaze vector from a mobile eye tracker, estimates of the eye tracker pose from a visual-inertial simultaneous localization and mapping (SLAM) algorithm, a 3D point cloud map of the environment from a RGB-D sensor. Experimental results indicate that our system produces accurate estimates of 3D gaze over a much larger range than remote eye trackers. Our system will enable applications, such as the analysis of 3D human attention and more anticipative human robot interfaces.
Haofei Wang 0001, Jimin Pi, Tong Qin 0001, Shaojie Shen, Bertram E. Shi
ETRA5
2017 Photorealistic facial expression synthesis by the conditional difference adversarial autoencoder
abstract
Photorealistic facial expression synthesis from single face image can be widely applied to face recognition, data augmentation for emotion recognition or entertainment. This problem is challenging, in part due to a paucity of labeled facial expression data, making it difficult for algorithms to disambiguate changes due to identity and changes due to expression. In this paper, we propose the conditional difference adversarial autoencoder (CDAAE) for facial expression synthesis. The CDAAE takes a facial image of a previously unseen person and generates an image of that person's face with a target emotion or facial action unit (AU) label. The CDAAE adds a feedforward path to an autoencoder structure connecting low level features at the encoder to features at the corresponding level at the decoder. It handles the problem of disambiguating changes due to identity and changes due to facial expression by learning to generate the difference between low-level features of images of the same person but with different facial expressions. The CDAAE structure can be used to generate novel expressions by combining and interpolating between facial expressions/action units within the training set. Our experimental results demonstrate that the CDAAE can preserve identity information when generating facial expression for unseen subjects more faithfully than previous approaches. This is especially advantageous when training with small databases.
Yuqian Zhou, Bertram E. Shi
ACII2
2017 Feedback Networks
abstract
Urrently, the most successful learning models in computer vision are based on learning successive representations followed by a decision layer. This is usually actualized through feedforward multilayer neural networks, e.g. ConvNets, where each layer forms one of such successive representations. However, an alternative that can achieve the same goal is a feedback based approach in which the representation is formed in an iterative manner based on a feedback received from previous iterations output. We establish that a feedback based approach has several core advantages over feedforward: it enables making early predictions at the query time, its output naturally conforms to a hierarchical structure in the label space (e.g. a taxonomy), and it provides a new basis for Curriculum Learning. We observe that feedback develops a considerably different representation compared to feedforward counterparts, in line with the aforementioned advantages. We provide a general feedback based learning architecture, instantiated using existing RNNs, with the endpoint results on par or better than existing feedforward networks and the addition of the above advantages.
Amir Zamir, Te-Lin Wu, Lin Sun 0004, Bokui Shen, Bertram E. Shi, Jitendra Malik, Silvio Savarese
CVPR5
2017 Pose-Independent Facial Action Unit Intensity Regression Based on Multi-Task Deep Transfer Learning
abstract
Facial expression recognition plays an increasingly important role in human behavior analysis and human computer interaction. Facial action units (AUs) coded by the Facial Action Coding System (FACS) provide rich cues for the interpretation of facial expressions. Much past work on AU analysis used only frontal view images, but natural images contain a much wider variety of poses. The FG 2017 Facial Expression Recognition and Analysis challenge (FERA 2017) requires participants to estimate the AU occurrence and intensity under nine different pose angles. This paper proposes a multi-task deep network addressing the AU intensity estimation sub-challenge of FERA 2017. The network performs the tasks of pose estimation and pose-dependent AU intensity estimation simultaneously. It merges the pose-dependent AU intensity estimates into a single estimate using the estimated pose. The two tasks share transferred bottom layers of a deep convolutional neural network (CNN) pre-trained on ImageNet. Our model outperforms the baseline results, and achieves a balanced performance among nine pose angles for most AUs.
Yuqian Zhou, Jimin Pi, Bertram E. Shi
FG3
2017 Probabilistic adjustment of dwell time for eye typing
abstract
Requiring a dwell time before selection is a common way to solve “Midas-touch problem” in gaze-based interaction. Choosing the dwell time involves a tradeoff between unintentional selection for short dwell times and slow text entry for long dwell times. We propose a probabilistic model for gaze based selection, which adjusts the dwell time based on the probability of each letter based on the past letters selected. By reformulating the entire problem of gaze-based selection probabilistically, we can naturally integrate the probability of each character naturally and with very few prior assumptions and very few free parameters. It automatically assigns shorter dwell times to more likely characters and longer dwell times to less likely characters. Our experimental results demonstrate that the proposed technique speeds up typing without loss in accuracy. The concept of this can be generalized to other dwell-based applications, leading to more efficient gaze system interaction.
Jimin Pi, Bertram E. Shi
HSI2
2017 Lattice Long Short-Term Memory for Human Action Recognition
abstract
Human actions captured in video sequences are threedimensional signals characterizing visual appearance and motion dynamics. To learn action patterns, existing methods adopt Convolutional and/or Recurrent Neural Networks (CNNs and RNNs). CNN based methods are effective in learning spatial appearances, but are limited in modeling long-term motion dynamics. RNNs, especially Long Short- Term Memory (LSTM), are able to learn temporal motion dynamics. However, naively applying RNNs to video sequences in a convolutional manner implicitly assumes that motions in videos are stationary across different spatial locations. This assumption is valid for short-term motions but invalid when the duration of the motion is long.,,In this work, we propose Lattice-LSTM (L2STM), which extends LSTM by learning independent hidden state transitions of memory cells for individual spatial locations. This method effectively enhances the ability to model dynamics across time and addresses the non-stationary issue of long-term motion dynamics without significantly increasing the model complexity. Additionally, we introduce a novel multi-modal training procedure for training our network. Unlike traditional two-stream architectures which use RGB and optical flow information as input, our two-stream model leverages both modalities to jointly train both input gates and both forget gates in the network rather than treating the two streams as separate entities with no information about the other. We apply this end-to-end system to benchmark datasets (UCF-101 and HMDB-51) of human action recognition. Experiments show that on both datasets, our proposed method outperforms all existing ones that are based on LSTM and/or CNNs of similar model complexities.
Lin Sun 0004, Kui Jia, Kevin Chen 0001, Dit-Yan Yeung, Bertram E. Shi, Silvio Savarese
ICCV5
2017 Learning multisensory cue integration on mobile robots
abstract
Developmental robotics seeks to build robots that learn to interact with the environment largely autonomously. These robots can calibrate their sensorimotor competencies on their own, much like developing children. In this paper, we build a developmental model of image stabilization based on the active efficient coding (AEC) framework and apply the model to a real robotic platform. In the visual system of primates, the optokinetic response (OKR) and the vestibulo-ocular reflex (VOR) cooperate to ensure image stabilization during relative motion between the observer and the environment. Inspired by these biological findings, our model integrates visual, inertial and motor encoder sensory cues. The sensory processing and the motor policy co-develop. The visual processing is based on a sparse coding algorithm. Motor behavior is learned using reinforcement learning. Our results show that the stabilization performance is improved by integrating visual and inertial inputs. Importantly, the weighting between the two inputs is learned automatically as the robot interacts with the environment.
Chong Zhang 0002, Jochen Triesch, Bertram E. Shi
ICRA3
2017 Learning multisensory neural controllers for robot arm tracking
abstract
Humans learn multisensory eye-hand coordination starting from infancy without supervision. For an example, they learn to track their hands by exploiting various sensory modalities, such as vision and proprioception. This integration occurs as they learn to perceive the world around them and their relationship to it. Most prior work has focused on the role of vision, as it is a primary sensory source for humans. However, it is interesting to study how vision and proprioception interact. We propose a system which combines visual and proprioceptive information to learn the eye-hand coordination skills that enable a robot to fixate its camera gaze on the end effector of its arm. In our model, visual cues are part of the feedback control loop, whereas proprioceptive cues are part of a feedforward control loop. Both controllers, as well as the sensory transform from raw visual information to a neural sensory representation are learned as the robot performs motor babbling movements. Visual information is encoded by sparse coding. The basis functions that emerge are similar to the receptive fields in the human visual cortex. An actor-critic reinforcement learning algorithm is used to drive eye motor neurons fusing visual and proprioceptive cues. We model and test the system in the iCub simulation environment. Our results suggest that these sensory modalities are capable of jointly learning model parameters to perform the tracking task. The evolved policy has characteristics that are qualitatively similar to the human oculomotor plant.
Lakshitha P. Wijesinghe, Marco Antonelli, Jochen Triesch, Bertram E. Shi
IJCNN4
2017 Action unit selective feature maps in deep networks for facial expression recognition
abstract
Facial expression recognizers based on handcrafted features have achieved satisfactory performance on many databases. Recently, deep neural networks, e. g. deep convolutional neural networks (CNNs) have been shown to boost performance on vision tasks. However, the mechanisms exploited by CNNs are not well established. In this paper, we establish the existence and utility of feature maps selective to action units in a deep CNN trained by transfer learning. We transfer a network pre-trained on the Image-Net dataset to the facial expression recognition task using the Karolinska Directed Emotional Faces (KDEF), Radboud Faces Database(RaFD) and extended Cohn-Kanade (CK+) database. We demonstrate that higher convolutional layers of the deep CNN trained on generic images are selective to facial action units. We also show that feature selection is critical in achieving robustness, with action unit selective feature maps being more critical in the facial expression recognition task. These results support the hypothesis that both human and deeply learned CNNs use similar mechanisms for recognizing facial expressions.
Yuqian Zhou, Bertram E. Shi
IJCNN2
2017 HOTS: A Hierarchy of Event-Based Time-Surfaces for Pattern Recognition
abstract
This paper describes novel event-based spatio-temporal features called time-surfaces and how they can be used to create a hierarchical event-based pattern recognition architecture. Unlike existing hierarchical architectures for pattern recognition, the presented model relies on a time oriented approach to extract spatio-temporal features from the asynchronously acquired dynamics of a visual scene. These dynamics are acquired using biologically inspired frameless asynchronous event-driven vision sensors. Similarly to cortical structures, subsequent layers in our hierarchy extract increasingly abstract features using increasingly large spatio-temporal windows. The central concept is to use the rich temporal information provided by events to create contexts in the form of time-surfaces which represent the recent temporal activity within a local spatial neighborhood. We demonstrate that this concept can robustly be used at all stages of an event-based hierarchical model. First layer feature units operate on groups of pixels, while subsequent layer feature units operate on the output of lower level feature units. We report results on a previously published 36 class character recognition task and a four class canonical dynamic card pip task, achieving near 100 percent accuracy on each. We introduce a new seven class moving face recognition task, achieving 79 percent accuracy.This paper describes novel event-based spatio-temporal features called time-surfaces and how they can be used to create a hierarchical event-based pattern recognition architecture. Unlike existing hierarchical architectures for pattern recognition, the presented model relies on a time oriented approach to extract spatio-temporal features from the asynchronously acquired dynamics of a visual scene. These dynamics are acquired using biologically inspired frameless asynchronous event-driven vision sensors. Similarly to cortical structures, subsequent layers in our hierarchy extract increasingly abstract features using increasingly large spatio-temporal windows. The central concept is to use the rich temporal information provided by events to create contexts in the form of time-surfaces which represent the recent temporal activity within a local spatial neighborhood. We demonstrate that this concept can robustly be used at all stages of an event-based hierarchical model. First layer feature units operate on groups of pixels, while subsequent layer feature units operate on the output of lower level feature units. We report results on a previously published 36 class character recognition task and a four class canonical dynamic card pip task, achieving near 100 percent accuracy on each. We introduce a new seven class moving face recognition task, achieving 79 percent accuracy.
Xavier Lagorce, Garrick Orchard, Francesco Galluppi, Bertram E. Shi, Ryad Benosman
IEEE Trans. Pattern Anal. Mach. Intell.4
2016 A two layer disparity selective simple cell model
abstract
The responses of disparity selective complex cells in the mammalian visual system are often modeled by the disparity energy model. This model linearly combines inputs from binocular simple cell units, whose responses are computed by the combination of left and right eye inputs through linear binocular receptive fields, followed by half wave rectification and an expansive nonlinearity. While many complex cells have responses that can be modeled by the standard disparity energy model, there are many that cannot be. While the standard disparity energy model has been extended to account for some specific types of tuning (e.g. cells that are “monocularly responsive” in that they only respond to input from one eye, yet are disparity tuned indicating that they receive input from both eyes), actual neurons display a wider range of ocular dominance indices and disparity selectivities than can be fully accounted for by these models. Here, we describe a two layer simple cell model that can be used to construct complex cells that fully cover this range. The model combines the responses from a first layer of model monocular simple cells. By adjusting the weights from the monocular to a second binocular layer, the model can exhibit more diverse tuning properties than previously proposed models. We also show that these weights can be learned by sparse coding, and that if so, there is a strong relationship between distribution of tuning properties in the population and the input disparity statistics.
Sijing Cheng, Qiuyan Peng, Bertram E. Shi
IJCNN3
2016 Simultaneous learning of the structure and kinematic model of an articulated body from point clouds
abstract
We present an algorithm that simultaneously identifies the structure and learns the kinematic model of an articulated body from point cloud data and joint angle data with minimal prior assumptions. The robot arm is represented by a point cloud. The structure is identified by segmenting the point cloud to subsets corresponding to different links according to a minimum matching error criterion that includes spatial continuity. The kinematic model is represented as a set of rigid transformations, which are learned by a stochastic gradient decent-based iterative closest point (ICP) algorithm. The learned transformations are then used to infer the structural relationship between segments. We validate the experiment on both synthetic data and data acquired by a Kinect depth sensor viewing a robot arm moving three of its joints. Our results demonstrate that point cloud can be correctly segmented into different rigid parts; the rigid transformations associated with each part can be learned and the kinematic structure can be correctly inferred. Our proposed approach makes a step toward a unified framework for robot self-discovery and learning forward kinematic models purely with data from depth sensors and joint angle encoders.
Bertram E. Shi
IJCNN2
2016 Unsupervised learning of depth during coordinated head/eye movements
abstract
Autonomous robots and humans need to create a coherent 3D representation of their peripersonal space in order to interact with nearby objects. Recent studies in visual neuroscience suggest that the small coordinated head/eye movements that humans continually perform during fixation provides useful depth information. In this work, we mimic such a behavior on a humanoid robot and propose a computational model that extracts depth information without requiring the kinematic model of the robot. First, we show that, during fixational head/eye movements, proprioceptive cues and optic flow lie on a low dimensional subspace that is a function of the depth of the target. Then, we use the generative adaptive subspace self-organizing map (GASSOM) to learn these depth-dependent subspaces. The depth of the target is eventually decoded using a winner-take-all strategy. The proposed model is validated on a simulated model of the iCub robot.
Marco Antonelli, Michele Rucci, Bertram E. Shi
IROS3
2015 Human Action Recognition Using Factorized Spatio-Temporal Convolutional Networks
abstract
Human actions in video sequences are three-dimensional (3D) spatio-temporal signals characterizing both the visual appearance and motion dynamics of the involved humans and objects. Inspired by the success of convolutional neural networks (CNN) for image classification, recent attempts have been made to learn 3D CNNs for recognizing human actions in videos. However, partly due to the high complexity of training 3D convolution kernels and the need for large quantities of training videos, only limited success has been reported. This has triggered us to investigate in this paper a new deep architecture which can handle 3D signals more effectively. Specifically, we propose factorized spatio-temporal convolutional networks (FstCN) that factorize the original 3D convolution kernel learning as a sequential process of learning 2D spatial kernels in the lower layers (called spatial convolutional layers), followed by learning 1D temporal kernels in the upper layers (called temporal convolutional layers). We introduce a novel transformation and permutation operator to make factorization in FstCN possible. Moreover, to address the issue of sequence alignment, we propose an effective training and inference strategy based on sampling multiple video clips from a given action video sequence. We have tested FstCN on two commonly used benchmark datasets (UCF-101 and HMDB-51). Without using auxiliary training videos to boost the performance, FstCN outperforms existing CNN based methods and achieves comparable performance with a recent method that benefits from using auxiliary training videos.
Lin Sun 0004, Kui Jia, Dit-Yan Yeung, Bertram E. Shi
ICCV4
2015 On the utility of sparse neural representations in adaptive behaving agents
abstract
A number of unsupervised learning algorithms seeking to account for the receptive field properties of simple cells in the mammalian primary visual cortex have been proposed. Among these are principal component analysis and sparse coding. While it appears that the receptive field properties learned by sparse coding match those measured in cortical cells better than those learned by principal component analysis, it is still not clear why biological neural systems might prefer to use sparse codes. In this paper we explore another reason why sparse representations might be preferred over principal component analysis by studying the utility of different coding schemes in an adaptive behaving agent. We suggest that the qualitative properties of representations based on sparse coding are more stable in the presence of changes in the input statistics than those of representations based on principal component analysis. We demonstrate this by examining representations learned on binocular visual input with different disparity distributions. Our results show that in encoding retinal disparity, the properties of sparse codes are more stable, and that this has important implications in adaptive agents, where the statistics change over time. In particular, in an agent who jointly learns a representation for binocular visual inputs along with a vergence control policy, the learned behavior is unstable when actions are driven by PCA based representations, but stable and self-calibrating when driven by sparse coding based representations.
Thusitha N. Chandrapala, Bertram E. Shi, Jochen Triesch
IJCNN2
2015 Laplacian Auto-Encoders: An explicit learning of nonlinear data manifold
Kui Jia, Lin Sun 0004, Shenghua Gao, Zhan Song, Bertram E. Shi
Neurocomputing5
2015 Learning Slowness in a Sparse Model of Invariant Feature Detection
abstract
Primary visual cortical complex cells are thought to serve as invariant feature detectors and to provide input to higher cortical areas. We propose a single model for learning the connectivity required by complex cells that integrates two factors that have been hypothesized to play a role in the development of invariant feature detectors: temporal slowness and sparsity. This model, the generative adaptive subspace self-organizing map (GASSOM), extends Kohonen's adaptive subspace self-organizing map (ASSOM) with a generative model of the input. Each observation is assumed to be generated by one among many nodes in the network, each being associated with a different subspace in the space of all observations. The generating nodes evolve according to a first-order Markov chain and generate inputs that lie close to the associated subspace. This model differs from prior approaches in that temporal slowness is not an externally imposed criterion to be maximized during learning but, rather, an emergent property of the model structure as it seeks a good model of the input statistics. Unlike the ASSOM, the GASSOM does not require an explicit segmentation of the input training vectors into separate episodes. This enables us to apply this model to an unlabeled naturalistic image sequence generated by a realistic eye movement model. We show that the emergence of temporal slowness within the model improves the invariance of feature detectors trained on this input.
Thusitha N. Chandrapala, Bertram E. Shi
Neural Comput.2
2014 Intrinsically motivated learning of visual motion perception and smooth pursuit
abstract
Developmental robots require cognitive structures that can learn perception-action cycles via interactions with the environment. Here, we extend the efficient coding hypothesis, which has been used to model the development of sensory processing in isolation, to model the development of the perception-action cycle. Our extension combines sparse coding and reinforcement learning so that sensory processing and behavior co-develop to optimize a shared intrinsic motivational signal: the fidelity of the neural encoding of the sensory input under resource constraints. Applying this framework to a model of a robot actively observing a time-varying environment leads to the simultaneous development of visual smooth pursuit behavior and model neurons similar to cortical neurons selective to visual motion. We suggest that this general principle may form the basis for a unified and integrated approach to learning many other perception/action loops.
Chong Zhang 0002, Yu Zhao 0031, Jochen Triesch, Bertram E. Shi
ICRA4
2014 The generative Adaptive Subspace Self-Organizing Map
abstract
The Adaptive Subspace Self Organized Map (ASSOM) is a model that incorporates sparsity, nonlinear pooling, topological organization and temporal continuity to learn invariant feature detectors, each corresponding to one node of the network. Temporal continuity is implemented by grouping inputs into "training episodes". Each episode contains samples from one invariance class and is mapped to a particular node during training. However, this explicit grouping makes application of this algorithm for natural image sequences difficult, since the grouping is generally not known a priori. This work proposes a probabilistic generative model of the ASSOM that addresses this problem. Each node of the ASSOM generates input vectors from one invariance class. Training sequences are generated by nodes that are chosen according to a Markov process. We demonstrate that this model can learn invariant feature detectors similar to those found in the primary visual cortex from an unlabeled sequence of input images generated by a realistic model of eye movements. Performance is comparable to the original ASSOM algorithm, but without the need for explicit grouping into training episodes.
Thusitha N. Chandrapala, Bertram E. Shi
IJCNN2
2013 Combining texture and stereo disparity cues for real-time face detection
Feijun Jiang, Mika Fischer, Hazim Kemal Ekenel, Bertram E. Shi
Signal Process. Image Commun.4
2012 Efficient and robust integration of face detection and head pose estimation
Feijun Jiang, Hazim Kemal Ekenel, Bertram E. Shi
ICPR3
2012 Active Vision During Coordinated Head/Eye Movements in a Humanoid Robot
abstract
While looking at a point in the scene, humans continually perform smooth eye movements to compensate for involuntary head rotations. Since the optical nodal points of the eyes do not lie on the head rotation axes, this behavior yields useful 3-D information in the form of visual parallax. Here, we describe the replication of this behavior in a humanoid robot. We have developed a method for egocentric distance estimation based on the parallax that emerges during compensatory head/eye movements. This method was tested in a robotic platform equipped with an anthropomorphic neck and two binocular pan-tilt units specifically designed to reproduce the visual input signals experienced by humans. We show that this approach yields accurate and robust estimation of egocentric distance within the space nearby the agent. These results provide a further demonstration of how behavior facilitates the solution of complex perceptual problems.
Xutao Kuang, Mark Gibson, Bertram E. Shi, Michele Rucci
IEEE Trans. Robotics3
2011 Effective discretization of Gabor features for real-time face detection
abstract
We describe a real-time face detector based on Gabor features. While Gabor features often lead to improved performance, they are often avoided as they are perceived as being computationally expensive. We address this in two ways. First, we propose an efficient discrete encoding method for the Gabor feature vector. This enables us to use a computationally efficient multi-stage classifier based on boosting and winnowing. Second, we accelerate computationally complex computations using the parallelization provided by graphics processing units (GPUs). With these innovations, the resulting detector runs at 16.8 fps for 640 × 480 images on a PC equipped with an i5 CPU and a GTX 465 graphic card.
Feijun Jiang, Bertram E. Shi, Mika Fischer, Hazim Kemal Ekenel
ICIP2
2011 The role of orientation diversity in binocular vergence control
abstract
Neurons tuned to binocular disparity in area V1 are hypothesized to be responsible for short latency binocular vergence movements, which align the two eyes on the same object as it moves in depth. Disparity selective neurons in V1 are not only selective to disparity, but also to other visual stimulus dimensions, in particular orientation. In this work, we explore the role of neurons tuned to different orientations in binocular vergence control. We trained an artificial binocular vision system to execute corrective vergence movements based on the outputs of disparity selective neurons tuned to different orientations and scales. As might be expected, we find that neurons tuned to vertical orientations have the strongest effect on the vergence eye movements. The effect of neurons tuned to other orientations decreases as the tuned orientation approaches horizontal. Although adding neurons tuned to non-vertical orientations does not appear to improve vergence tracking accuracy, we find that neurons tuned to non-vertical orientations still play critical roles in binocular vergence control. First, they decrease the time required to learn the vergence control strategy. Second, they also increase the effective range of vergence control.
Chao Qu, Bertram E. Shi
IJCNN2
2011 Self-Organizing Neural Population Coding for improving robotic visuomotor coordination
abstract
We present an extension of Kohonen's Self Organizing Map (SOM) algorithm called the Self Organizing Neural Population Coding (SONPC) algorithm. The algorithm adapts online the neural population encoding of sensory and motor coordinates of a robot according to the underlying data distribution. By allocating more neurons towards area of sensory or motor space which are more frequently visited, this representation improves the accuracy of a robot system on a visually guided reaching task. We also suggest a Mean Reflection method to solve the notorious border effect problem encountered with SOMs for the special case where the latent space and the data space dimensions are the same.
Piotr Dudek, Bertram E. Shi
IJCNN3
2010 GPU implemention of fast Gabor filters
abstract
With their parallel multi-core architecture, Programmable Graphics Processing Units (GPUs) are well suited for implementing biologically-inspired visual processing algorithms, such as Gabor filtering. We compare several GPU implementations of Gabor filtering. On the same graphics card (an NVIDIA GeForce 9800 GTX+) and for convolution kernel radii from 8 to 48 pixels, an algorithm that decomposes Gabor filtering into a number of simpler steps results in an algorithm that is 2.2 to 33 times faster than direct 2D convolution and 2.8 to 6.6 times faster than a FFT based approach. Surprisingly, in comparison with an optimized algorithm for Gabor filtering running on a PC (Core2 Duo 3.16GHz), it is only 4-10 times faster. The PC can efficiently implement a recursive 1D filter, which requires far fewer arithmetic operations than convolution. However, due to data dependencies, this recursive filter typically runs slower than 1D convolution on the GPU. This highlights the importance of simultaneously considering both arithmetic and memory operations in porting algorithms to GPUs.
Bertram E. Shi
ISCAS2
2010 Autonomous Development of Vergence Control Driven by Disparity Energy Neuron Populations
abstract
We present a simple optimization criterion that leads to autonomous development of a sensorimotor feedback loop driven by the neural representation of the depth in the mammalian visual cortex. Our test bed is an active stereo vision system where the vergence angle between the two eyes is controlled by the output of a population of disparity-selective neurons. By finding a policy that maximizes the total response across the neuron population, the system eventually tracks a target as it moves in depth. We characterized the tracking performance of the resulting policy using objects moving both sinusoidally and randomly in depth. Surprisingly, the system can even learn how to track based on stimuli it cannot track: even though the closed loop 3 dB tracking bandwidth of the system is 0.3 Hz, correct tracking policies are learned for input stimuli moving as fast as 0.75 Hz.
Bertram E. Shi
Neural Comput.2
2009 Maximizing neural responses leads to sensori-motor coordination of binocular vergence
abstract
We present a simple optimization criterion based on the neural representation of the depth in the mammalian visual cortex that leads to self organization of a sensori-motor feedback loop. We study an active stereo vision system where the vergence angle between the two eyes is controlled by the output of a population of disparity selective neurons. We show that finding a control policy that maximizes the total response across an entire population of disparity selective neurons results in a system that automatically learns to track a target as it moves in depth. We characterized the tracking performance of the resulting policy using objects moving both sinusoidally and randomly in depth. The closed loop 3 dB tracking bandwidth of the system was 0.3 Hz.
Bertram E. Shi
CIMSIVP2
2009 Normalized phase shift motion energy neuron populations for image velocity estimation
abstract
Motion energy neurons are commonly used in biologically motivated algorithms for image velocity estimation. These algorithms typically use a large population of neurons tuned to different locations in the spatial-temporal frequency domain, with each neuron requiring a complex-valued separate spatio-temporal image filter. Here, we show that it is possible to construct a large population of motion energy neurons by combining the outputs of a much fewer number of filters with differing phase shifts. With spatial pooling, the velocity estimation using this phase-tuned population is more reliable than estimation using a more conventional frequency-tuned population. In addition, we show that by normalizing the population response, we can obtain a confidence measure for the resulting velocity estimates.
Yicong Meng, Bertram E. Shi
IJCNN2
2009 Extending Phase Mechanism to Differential Motion Opponency for Motion Pop-out
abstract
We extend the concept of phase tuning, a ubiquitous mechanism in sensory neurons including motion and disparity detection neurons, to the motion contrast detection. We demonstrate that motion contrast can be detected by phase shifts between motion neuronal responses in different spatial regions. By constructing the differential motion opponency in response to motions in two different spatial regions, varying motion contrasts can be detected, where similar motion is detected by zero phase shifts and differences in motion by non-zero phase shifts. The model can exhibit either enhancement or suppression of responses by either different or similar motion in the surrounding. A primary advantage of the model is that the responses are selective to relative motion instead of absolute motion, which could model neurons found in neurophysiological experiments responsible for motion pop-out detection.
Yicong Meng, Bertram E. Shi
NIPS2
2009 Integrating contrast invariance into a model for cortical orientation map formation
Laura Y. Zhao, Bertram E. Shi
Neurocomputing2
2009 Disparity Estimation by Pooling Evidence From Energy Neurons
abstract
In this paper, we propose an algorithm for disparity estimation from disparity energy neurons that seeks to maintain simplicity and biological plausibility, while also being based upon a formulation that enables us to interpret the model outputs probabilistically. We use the Bayes factor from statistical hypothesis testing to show that, in contradiction to the implicit assumption of many previously proposed biologically plausible models, a larger response from a disparity energy neuron does not imply more evidence for the hypothesis that the input disparity is close to the preferred disparity of the neuron. However, we find that the normalized response can be interpreted as evidence, and that information from different orientation channels can be combined by pooling the normalized responses. Based on this insight, we propose an algorithm for disparity estimation constructed out of biologically plausible operations. Our experimental results on real stereograms show that the algorithm outperforms a previously proposed coarse-to-fine model. In addition, because its outputs can be interpreted probabilistically, the model also enables us to identify occluded pixels or pixels with incorrect disparity estimates.
Eric K. C. Tsang, Bertram E. Shi
IEEE Trans. Neural Networks2
2008 Improved illumination invariance using a color edge representation based on Double Opponent neurons
abstract
We describe an evaluation framework that provides a quantitative measure on the performance of a neural network color constancy model. In this framework, the responses of three models of color constancy to a set of color edges under varying illuminating conditions are computed. We study a model based on double opponent cells, as well as two variants of the Retinex model. Evaluation metrics on the modelspsila capabilities to discriminate among different color edges and resist illuminant induced changes are measured using this framework, we confirm the advantage of incorporating spectral opponency into the color constancy model.
Javy H. Y. Lau, Bertram E. Shi
IJCNN2
2008 Neuromorphic implementation of active gaze and vergence control
abstract
We present an active stereo system with gaze and vergence control driven by model of the disparity selective neurons in the mammalian visual cortex. The hardware consists of a mobile stereo camera and three Multimap boards. The Multimap boards compute multiple cortical maps responding to target locations, orientations and disparities, and generate movement commands to track target in space. Each board can compute more than 10 cortical maps at 320*240 pixel resolution and 25 frames per second, and consumes 3.5W.
Eric K. C. Tsang, Stanley Y. M. Lam, Yicong Meng, Bertram E. Shi
ISCAS4
2008 A Two Stage Energy Model Exhibiting Selectivity to Changing Disparity
Xiaojiang Guo, Bertram E. Shi
ISNN (1)2
2008 Normalization Enables Robust Validation of Disparity Estimates from Neural Populations
abstract
Binocular fusion takes place over a limited region smaller than one degree of visual angle (Panum's fusional area), which is on the order of the range of preferred disparities measured in populations of disparity-tuned neurons in the visual cortex. However, the actual range of binocular disparities encountered in natural scenes extends over tens of degrees. This discrepancy suggests that there must be a mechanism for detecting whether the stimulus disparity is inside or outside the range of the preferred disparities in the population. Here, we compare the efficacy of several features derived from the population responses of phase-tuned disparity energy neurons in differentiating between in-range and out-of-range disparities. Interestingly, some features that might be appealing at first glance, such as the average activation across the population and the difference between the peak and average responses, actually perform poorly. On the other hand, normalizing the difference between the peak and average responses results in a reliable indicator. Using a probabilistic model of the population responses, we improve classification accuracy by combining multiple features. A decision rule that combines the normalized peak to average difference and the peak location significantly improves performance over decision rules based on either measure in isolation. In addition, classifiers using normalized difference are also robust to mismatch between the image statistics assumed by the model and the actual image statistics.
Eric K. C. Tsang, Bertram E. Shi
Neural Comput.2
2008 Adaptive Gain Control for Spike-Based Map Communication in a Neuromorphic Vision System
abstract
To support large numbers of model neurons, neuromorphic vision systems are increasingly adopting a distributed architecture, where different arrays of neurons are located on different chips or processors. Spike-based protocols are used to communicate activity between processors. The spike activity in the arrays depends on the input statistics as well as internal parameters such as time constants and gains. In this paper, we investigate strategies for automatically adapting these parameters to maintain a constant firing rate in response to changes in the input statistics. We find that under the constraint of maintaining a fixed firing rate, a strategy based upon updating the gain alone performs as well as an optimal strategy where both the gain and the time constant are allowed to vary. We discuss how to choose the time constant and propose an adaptive gain control mechanism whose operation is robust to changes in the input statistics. Our experimental results on a mobile robotic platform validate the analysis and efficacy of the proposed strategy.
Yicong Meng, Bertram E. Shi
IEEE Trans. Neural Networks2
2007 Active Visual Tracking of Heading Direction By Combining Motion Energy Neurons
abstract
We describe a robotic vision system that aligns a camera's optical axis with its direction of translation by estimating the focus of expansion. Visual processing is based on functional models of populations of neurons in cortical areas VI through MST. Populations of motion energy neurons tuned to different orientations, positions and directions of motion are successively transformed into a population of neurons that collectively encode the focus of expansion at 25 frames per second. We characterize the performance of the system translating through a cluttered environment, and show that the performance is robust to variations in system parameters.
Stanley Y. M. Lam, Bertram E. Shi
ISCAS2
2007 Sensor Integration in Autonomous Systems
abstract
In this survey introduction to the special session of the same title at the IEEE International Symposium on Circuits and Systems, we review past work in the development of bio-inspired circuits and platforms for sensing and actuation, as well as recent work that has been done in integrating these systems into applications.
Bertram E. Shi, Csaba Rekeczky
ISCAS1
2007 Probabilistic Modelling of Phase-tuned Disparity Energy Neuron Populations
abstract
We present a low dimensional Bayes probabilistic model for the population of binocular disparity energy neurons centered at the same retinal location, but selective to different disparities by phase shifts between the left and right monocular receptive fields. The model accurately predicts response distributions of simulated binocular disparity energy neurons. It provides a probabilistic explanation for the decrease in the reliability of the population responses for large disparities. Applied to vergence control, it generates more reliable responses than using the preferred disparity of the most responsive neuron as the control signal.
Eric K. C. Tsang, Bertram E. Shi
ISCAS2
2007 Extending position/phase-shift tuning to motion energy neurons improves velocity discrimination
abstract
We extend position and phase-shift tuning, concepts already well established in the disparity energy neuron literature, to motion energy neurons. We show that Reichardt-like detectors can be considered examples of position tuning, and that motion energy filters whose complex valued spatio-temporal receptive fields are space-time separable can be considered examples of phase tuning. By combining these two types of detectors, we obtain an architecture for constructing motion energy neurons whose center frequencies can be adjusted by both phase and posi- tion shifts. Similar to recently described neurons in the primary visual cortex, these new motion energy neurons exhibit tuning that is between purely space- time separable and purely speed tuned. We propose a functional role for this intermediate level of tuning by demonstrating that comparisons between pairs of these motion energy neurons can reliably discriminate between inputs whose velocities lie above or below a given reference velocity.
Yiu Man Lam, Bertram E. Shi
NIPS2
2007 Estimating disparity with confidence from energy neurons
abstract
Binocular fusion takes place over a limited region smaller than one degree of visual angle (Panum's fusional area), which is on the order of the range of preferred disparities measured in populations of disparity-tuned neurons in the visual cortex. However, the actual range of binocular disparities encountered in natural scenes ranges over tens of degrees. This discrepancy suggests that there must be a mechanism for detecting whether the stimulus disparity is either inside or outside of the range of the preferred disparities in the population. Here, we present a statistical framework to derive feature in a population of V1 disparity neuron to determine the stimulus disparity within the preferred disparity range of the neural population. When optimized for natural images, it yields a feature that can be explained by the normalization which is a common model in V1 neurons. We further makes use of the feature to estimate the disparity in natural images. Our proposed model generates more correct estimates than coarse-to-fine multiple scales approaches and it can also identify regions with occlusion. The approach suggests another critical role for normalization in robust disparity estimation.
Eric K. C. Tsang, Bertram E. Shi
NIPS2
2007 Recursive Anisotropic 2-D Gaussian Filtering Based on a Triple-Axis Decomposition
abstract
We describe a recursive algorithm for anisotropic 2-D Gaussian filtering, based on separating the filter into the cascade of three, rather two, 1-D filters. The filters operate along axes obtained by integer horizontal and/or vertical pixel shifts. This eliminates interpolation, which removes spatial inhomogeneity in the filter, and produces more elliptically shaped kernels. It also results in a more regular filter structure, which facilitates implementation in DSP chips. Finally, it improves matching between filters with the same eccentricity and width, but different orientations. Our analysis and experiments indicate that the computational complexity is similar to an algorithm that operates along two axes (<11 ms for a 512 x 512 image using a 3.2-GHz Pentium 4 PC). On the other hand, given a limited set of basis filter axes, there is an orientation dependent lower bound on the achievable aspect ratios.
Stanley Y. M. Lam, Bertram E. Shi
IEEE Trans. Image Process.2
2006 A Scalable FPGA Implementation of Cellular Neural Networks for Gabor-type Filtering
abstract
We describe an implementation of Gabor-type filters on field programmable gate arrays using the cellular neural network (CNN) architecture. The CNN template depends upon the parameters (e.g., orientation, bandwidth) of the Gabor-type filter and can be modified at runtime so that the functionality of Gabor-type filter can be changed dynamically. Our implementation uses the Euler method to solve the ordinary differential equation describing the CNN. The design is scalable to allow for different pixel array sizes, as well as simultaneous computation of multiple filter outputs tuned to different orientations and bandwidths. For 1024 pixel frames, an implementation on a Xilinx Virtex XC2V1000-4 device uses 1842 slices, operates at 120 MHz and achieves 23,000 Euler iterations over one frame per second.
Ocean Y. H. Cheung, Philip H. W. Leong, Eric K. C. Tsang, Bertram E. Shi
IJCNN4
2006 Neuromorphic Translational Ego-motion Estimation using Log-Polar Motion Energy
abstract
We describe a neuromorphic hardware system for detecting patterns of expanding visual motion due to forward translation, which can be used for self-motion estimation. The system consists of a set of motion-energy filters, modeling the functional properties of direction tuned neurons in primary visual cortex. We apply the filters to log-polar sampled images to facilitate the construction of filters tuned to radial and angular directions of motion. We then describe how their outputs can be used to estimate direction of heading. We implemented this algorithm in a hardware system, which captures images from a camera and performs the required computations at 30 frames per second for images with 352 by 288 resolution.
Stanley Y. M. Lam, Eric K. C. Tsang, Yicong Meng, Bertram E. Shi
IJCNN4
2006 An Efficient Spike-Based Communication Protocol for Neurally Inspired Feature Maps
abstract
We describe the implementation of an efficient and flexible protocol for communicating biologically inspired feature maps between processors in a multi-processor parallel hardware system. A feature map is a retinotopic array of neurons sharing the same feature selectivity, e.g. to spatial contrast changes. The retina and visual cortex are thought to compute many feature maps in interpreting the environment. Our spike-based encoding method exploits the sparsity of these maps. We also use contrast normalization to amplify responses in areas of small contrast, while maintaining selectivity in areas with large contrast. This communication architecture will enable us to easily expand our system to handle more complex models.
Yicong Meng, Stanley Y. M. Lam, Eric K. C. Tsang, Bertram E. Shi
IJCNN4
2006 Active Binocular Gaze Control Inspired by Superior Colliculus
abstract
We describe the construction of visual maps used to guide binocular saccade-vergence movements in an active vision system. The model neurons in the maps are inspired by binocular disparity selective neurons found in the superior colliculus (SC). We modify the standard disparity energy model so that it is non-orientation selective (isotropic), like the neurons in SC. We analyze the differences between the standard and isotropic models, arguing that the isotropic disparity energy is well suited for gaze control. Finally, we describe the hardware implementation of this model, which we use to demonstrate that these model neurons can be effective in guiding binocular gaze control of a robotic vision system.
Eric K. C. Tsang, Bertram E. Shi
IJCNN2
2006 Expandable hardware for computing cortical feature maps
abstract
We describe expandable hardware architecture for the rapid simulation of feature maps inspired by the visual cortex. Feature maps are retinotopically organized arrays of neurons selective to different combinations of visual features. The responses of these maps are believed to be important for the brain to merge information from different visual cues. This architecture is based around a custom designed board containing DSP and FPGA chips. It is modular in the sense that additional boards can be integrated into the system to accommodate cortical models of increasing complexity.
Bertram E. Shi, Eric K. C. Tsang, Stanley Y. M. Lam, Yicong Meng
ISCAS1
2005 Implementation of Gabor-Type Filters on Field Programmable Gate Arrays
Ocean Y. H. Cheung, Philip H. W. Leong, Eric K. C. Tsang, Bertram E. Shi
FPT4
2004 A CNN model of multi-dimensional stimulus selectivity in primary visual cortex
abstract
We describe a neuromorphic approach to implementing model visual cortical neurons using a four-layer cellular neural network (CNN) chips. A key challenge is that visual cortical neurons are simultaneously selective along many stimulus dimensions, including retinal position, spatial frequency, orientation, temporal frequency, direction of motion, and binocular disparity. The ubiquity of intra-cortical feedback interconnections also implies that the neurons should operate in parallel and in continuous time. We discuss the modeling and implementation considerations that lead naturally to four layer networks, and describe the current status of our work in building silicon networks of tens of thousands of neurons.
Bertram E. Shi
IJCNN1
2004 A Preference for Phase-Based Disparity in a Neuromorphic Implementation of the Binocular Energy Model
abstract
The relative depth of objects causes small shifts in the left and right retinal positions of these objects, called binocular disparity. This letter describes an electronic implementation of a single binocularly tuned complex cell based on the binocular energy model, which has been proposed to model disparity-tuned complex cells in the mammalian primary visual cortex. Our system consists of two silicon retinas representing the left and right eyes, two silicon chips containing retinotopic arrays of spiking neurons with monocular Gabor-type spatial receptive fields, and logic circuits that combine the spike outputs to compute a disparity-selective complex cell response. The tuned disparity can be adjusted electronically by introducing either position or phase shifts between the monocular receptive field profiles. Mismatch between the monocular receptive field profiles caused by transistor mismatch can degrade the relative responses of neurons tuned to different disparities. In our system, the relative responses between neurons tuned by phase encoding are better matched than neurons tuned by position encoding. Our numerical sensitivity analysis indicates that the relative responses of phase-encoded neurons that are least sensitive to the receptive field parameters vary the most in our system. We conjecture that this robustness may be one reason for the existence of phase-encoded disparity-tuned neurons in biological neural systems.
Eric K. C. Tsang, Bertram E. Shi
Neural Comput.2
2003 Cortically-inspired visual processing with a four layer cellular neural network
abstract
This paper describes a four layer cellular neural network architecture implementing image processing inspired by the functionality of neurons in the visual cortex: linear orientation selective filtering and half wave rectification. The network implements both even and odd symmetric Gabor-like filters simultaneously, with pairs of layers representing the positive and negative components of the filter outputs. Each layer is an array of analog nonlinear continuous time processing elements ("cells" or "neurons"), each corresponding to one pixel in the image. Because neurons are feedback interconnected only with neurons from nearest neighbor pixels, we can easily implement this network in VLSI. For example, a recent implementation filters a 32 x 64 pixel image in parallel within a few milliseconds while dissipating only a few milliwatts. This paper analyzes the dynamics of this network mathematically, deriving the spatial transfer functions of the orientation selective filters and proving stability.
Bertram E. Shi
IJCNN1
2003 Flooring the observation probability for robust ASR in impulsive noise
abstract
Impulsive noise usually introduces sudden mismatches between the observation features and the acoustic models trained with clean speech, which drastically degrades the performance of automatic speech recognition (ASR) systems. This paper presents a novel method to directly suppress the adverse effect of impulsive noise on recognition. In this method, according to the noise sensitivity of each feature dimension, the observation vector is divided into several subvectors, each of which is assigned to a suitable flooring threshold. In recognition stage, observation probability of each feature sub-vector is floored at the Gaussian mixture level. Thus, the unreliable relative probability difference caused by impulsive noise is eliminated, and the expected correct state sequence recovers the priority of being chosen in decoding. Experimental evaluations on Aurora2 database show that the proposed method achieves the average error rate reduction (ERR) of 61.62% and 84.32% in simulated impulsive noise and machinegun noise environment, respectively, while maintaining high performance for clean speech recognition.
Pei Ding, Bertram E. Shi, Pascale Fung
INTERSPEECH2
2003 A Neuromorphic Multi-chip Model of a Disparity Selective Complex Cell
abstract
The relative depth of objects causes small shifts in the left and right ret- inal positions of these objects, called binocular disparity. Here, we describe a neuromorphic implementation of a disparity selective com- plex cell using the binocular energy model, which has been proposed to model the response of disparity selective cells in the visual cortex. Our system consists of two silicon chips containing spiking neurons with monocular Gabor-type spatial receptive fields (RF) and circuits that combine the spike outputs to compute a disparity selective complex cell response. The disparity selectivity of the cell can be adjusted by both position and phase shifts between the monocular RF profiles, which are both used in biology. Our neuromorphic system performs better with phase encoding, because the relative responses of neurons tuned to dif- ferent disparities by phase shifts are better matched than the responses of neurons tuned by position shifts.
Eric K. C. Tsang, Bertram E. Shi
NIPS2
2001 A one-pass strategy for keyword spotting and verification
abstract
One common method for keyword spotting in unconstrained speech is based upon a two pass strategy consisting of Viterbi-decoding to detect and segment possible keyword hits, followed by the computation of a confidence measure to verify those hits. In this paper, we propose a simple one-pass strategy where computation of the confidence measure is computed simultaneously with a Viterbi-like decoding stage. However, backtracking is not required, which when coupled with the need for only a single pass through the utterance significantly reduces the memory requirements of this algorithm. This feature makes it well suited for devices where processing power and memory are limited. Experimental results on a connected digits task show that performance of the decoding is comparable to that using a Viterbi search with backtracking. Experimental results on spotting days of the week in continuous speech indicate that the confidence measure calculated is effective in reducing the number of false alarms.
Chak Shun Lai, Bertram E. Shi
ICASSP2
2000 Soft GPD for minimum classification error rate training
abstract
Minimum classification error (MCE) rate training is a discriminative training method which seeks to minimize an empirical estimate of the error probability derived over a training set. The segmental generalized probabilistic descent (GPD) algorithm for MCE uses the log likelihood of the best path as a discriminant function to estimate the error probability. This paper shows that by using a discriminant function similar to the auxiliary function used in EM, we can obtain a "soft" version of GPD in the sense that information about all possible paths is retained. Complexity is similar to segmental GPD. For certain parameter values, the algorithm is equivalent to segmental GPD. By modifying the misclassification measure usually used, we can obtain an algorithm for embedded MCE training for continuous speech which does not require a separate N-best search to determine competing classes. Experimental results show error rate reduction of 20% compared with maximum likelihood training.
Bertram E. Shi, Kaisheng Yao
ICASSP1
2000 Residual noise compensation for robust speech recognition in nonstationary noise
abstract
We present a model-based noise compensation algorithm for robust speech recognition in nonstationary noisy environments. The effect of noise is split into a stationary part, compensated by parallel model combination, and a time varying residual. The evolution of residual noise parameters is represented by a set of state space models. The state space models are updated by Kalman prediction and the sequential maximum likelihood algorithm. Prediction of residual noise parameters from different mixtures are fused, and the fused noise parameters are used to modify the linearized likelihood score of each mixture. Noise compensation proceeds in parallel with recognition. Experimental results demonstrate that the proposed algorithm improves recognition performance in highly nonstationary environments, compared with parallel model combination alone.
Kaisheng Yao, Bertram E. Shi, Pascale Fung
ICASSP2
2000 Visual Tracking with Subpixel Resolution using an Analog VLSI Computational Sensor
abstract
This paper describes the application of a computational vision sensor to active binocular tracking. The sensor outputs are used to control the vergence angles of the two cameras and the tilt angle of the head so that the center pixels of the sensor arrays image the same point in the environment. One distinguishing feature of the sensor used here is the possibility to resolve target motions with subpixel resolution. This is due to the use of a phase based algorithm which integrates information over multiple pixels.
Ziyi Lu, Bertram E. Shi
ICRA2
2000 Residual noise compensation by a sequential EM algorithm for robust speech recognition in nonstationary noise
abstract
We model noise as a stationary component plus a time varying residual. The stationary part is estimated off-line and compensated using Log-Add noise compensation. The time varying residual is estimated and compensated using a sequential EM algorithm. The residual noise compensation proceeds in parallel with the recognition process. Experimental results demonstrate that the proposed algorithm improves the recognition performance not only in highly nonstationary noise but also in slow-varying noise, compared with Log-Add noise compensation alone.
Kaisheng Yao, Bertram E. Shi, Satoshi Nakamura 0001
INTERSPEECH2
2000 A 1-D local image velocity sensor using Gabor filtering
abstract
We propose gradient-based algorithm for local velocity estimation that uses a Gabor-type filter as a spatial pre-filtering stage. An integrated differentiation and division circuit extracts the local image velocity in continuous time. We describe the design, implementation and the simulation results of a 1-D velocity sensor consisting of 27 pixels. Each pixel generates two outputs: the local image velocity and the local motion direction.
Kwok Kit Lau, Bertram E. Shi
ISCAS2
2000 Binocular visual feedback with CNN sensors
abstract
This paper describes the use of a CNN sensor to provide visual feedback signals in an active binocular vision system. The sensor outputs are used to control the vergence angles of the two cameras and the tilt angle of the head so that the center pixels of the sensors image the same point in the environment. The sensors contain 2D arrays of phototransistors for image input and CNN-based processing circuits that convolve the input image with filters similar to even and odd Gabor filters. We use the odd and even filter outputs in a phase based algorithm, which enables target motions to be detected with sub-pixel resolution. Experimental results demonstrating the ability of the system to reconstruct target motions in 3D are described.
Ziyi Lu, Bertram E. Shi
ISCAS2
1999 Real-Time Gabor-Type Filtering Using Analog Focal Plane Image Processors
abstract
We describe focal plane image processors which filter images in space by filters which are similar to even and odd Gabor filters. All processing is done using analog circuits fabricated on the same die as the photosensors allowing combined sensing and processing at rates as high as 5000 frames per second. Both 1D and 2D sensors have been designed, fabricated and tested. In the 2D sensor the tuned orientation can be steered and the filter response scaled by adjusting external bias voltages controlling parameters such as the gains of the analog processing circuits. Because both even and odd Gabor-type filter outputs are calculated, sensors can be used in Gabor phase-based algorithms. As a simple example, we have embedded a 1D sensor on a mobile robot platform to perform image fixation.
Bertram E. Shi
CVPR1
1999 Channel and noise adaptation via HMM mixture mean transform and stochastic matching
abstract
We present a non-linear model transformation for adapting Gaussian mixture HMMs using both static and dynamic MFCC observation vectors to additive noise and constant system tilt. This transformation depends upon a few compensation coefficients which can be estimated from channel distorted speech via maximum-likelihood stochastic matching. Experimental results validate the effectiveness of the adaptation. We also provide an adaptation strategy which can result in improved performance at reduced computational cost compared with a straightforward implementation of stochastic matching.
Shuen Kong Wong, Bertram E. Shi
ICASSP2
1999 Liftered forward masking procedure for robust digits recognition
abstract
Using TI digits recognition experiments, we show that a combination of two dynamic speech features, Liftered Forward Masked (LFM) MFCC and 2-D cepstrum, can improve system robustness to additive Volvo noise while maintaining system performance comparable to standard MFCC features in clean conditions. Through experiments, we show that the information extracted by forward masking and by the 2D cepstrum are in some sense orthogonal. By combining the LFM MFCC and the 2-D cepstrum plus \\Delta 2-D cepstrum, we achieve a recognition rate above 90% on the TI connected digits task, even in additive Volvo noise condition with SNR as low as 0dB. This corresponds to a SNR gain over 30dB compared with standard MFCC plus dynamic and acceleration coefficients.
Kaisheng Yao, Bertram E. Shi, Pascale Fung
EUROSPEECH2
1998 A non-linear model transformation for ML stochastic matching in additive noise
abstract
We present a non-linear model transformation for adapting Gaussian mixture HMMs using both static and dynamic MFCC observation vectors to the presence of additive noise. This transformation depends upon a few compensation coefficients which can be estimated from a short training token of noise. Alternatively, one can also apply maximum-likelihood stochastic matching to estimate the compensation coefficients from speech embedded in noise. This can eliminate the need for segmentation of pure noise from speech for the estimation and can also compensate for inaccuracies in the estimation of the compensation coefficients as well as those due to the approximations used in deriving the transformation.
Shuen Kong Wong, Bertram E. Shi
MMSP2
1997 An Analog VLSI Neural Network for Phase-based Machine Vision
Bertram E. Shi, Kwok Fai Hui
NIPS1
1988 End effector actuation with a solid state motor
abstract
An actuator technology program is in progress to develop miniature motors, joints, and links suitable for modular placement in structures approaching the human hand in size, dexterity and weight. The end-effector device described is a small gripper with tweezer action. This is a first step in a program which will evolve from basic precision miniature grippers to more sophisticated and dextrous multi-jointed end effectors. The gripper can hold objects up to 0.5 in. wide, and it is intended for a circuit-board-level scale of operation. The jaw motor is piezoelectric and works by resonance operation in the longitudinal vibration mode to move a slide plate in a linear fashion. The slide plate is coupled to a mechanism in a 0.5 in./sup 3/-volume housing which converts the linear slide motion to lateral jaw action. The motion will eventually be controlled by linear encoders for position control and one or more strain gages to develop a force servo controller. Motor performance, modeling, and the smaller gripper are described.>
Jeffery S. Schoenwald, Paul M. Beckham, P. M. Rattner, B. Vanderlip, Bertram E. Shi
ICRA5