Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jie Yang 0001

dblp:12/1198-1 · DBLP profile ↗
← Back
91ranked-venue papers
10as first author
0since 2021 · last 2017
0000-0003-2896-1498ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 56 · 6 first-authorArtificial intelligence and machine learning · 35 · 2 first-authorHuman-computer interaction and ubiquitous computing · 18 · 2 first-authorSystems, architecture and hardware · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-authorComputer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
18 papers
Image recognition and object detection · 37% Representation and self-supervised learning · 14% Video understanding and tracking · 12%
Human-computer interaction and pervasive computing
11 papers
Collaborative and social computing · 44% Ubiquitous computing and smart environments · 22% Learning and educational technologies · 16%
Computer graphics and multimedia
8 papers
Multimedia analysis and retrieval · 56% Image and video processing · 24% Computational photography and imaging · 7%
Databases, data mining, and information retrieval
3 papers
Data mining · 54% Information retrieval · 41% Machine learning and data management · 5%

Topics — the 30 heaviest of 74, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Collaborative and social computing
remote collaboration
0.232008
Predicting Visual Focus of Attention From Intention in Remote Collaborative Tasks · IEEE Trans. Multim. 2008
Effects of task properties, partner actions, and message content on eye gaze patterns in a collaborative task · CHI 2005
DOVE: drawing over video environment · ACM Multimedia 2003
Computer vision › Image recognition and object detection
scene recognition
0.212013
Categorization of Multiple Objects in a Scene Using a Biased Sampling Strategy · Int. J. Comput. Vis. 2013
Computer vision › Image recognition and object detection › object labeling
semi-automatic object labeling
0.222009
Semi-Automatically Labeling Objects in Images · IEEE Trans. Image Process. 2009
SmartLabel: an object labeling tool using iterated harmonic energy minimization · ACM Multimedia 2006
Learning and educational technologies
intelligent tutoring systems
0.112012
A paradigm for handwriting-based intelligent tutors · Int. J. Hum. Comput. Stud. 2012
Machine learning › Representation and self-supervised learning
discriminative learning with pairwise constraints
0.122006
A Discriminative Learning Framework with Pairwise Constraints for Video Object Classification · IEEE Trans. Pattern Anal. Mach. Intell. 2006
A Discriminative Learning Framework with Pairwise Constraints for Video Object Classification · CVPR (2) 2004
Computer vision › Video understanding and tracking › video classification
video object classification
0.122006
A Discriminative Learning Framework with Pairwise Constraints for Video Object Classification · IEEE Trans. Pattern Anal. Mach. Intell. 2006
A Discriminative Learning Framework with Pairwise Constraints for Video Object Classification · CVPR (2) 2004
Collaborative and social computing › computer-supported cooperative work › collaborative applications
collaborative physical tasks
0.122005
Effects of task properties, partner actions, and message content on eye gaze patterns in a collaborative task · CHI 2005
DOVE: drawing over video environment · ACM Multimedia 2003
Ubiquitous computing and smart environments
context-aware computing
0.122005
Predicting human interruptibility with sensors · ACM Trans. Comput. Hum. Interact. 2005
Predicting human interruptibility with sensors: a Wizard of Oz feasibility study · CHI 2003
Ubiquitous computing and smart environments › interruption management
interruptibility prediction
0.122005
Predicting human interruptibility with sensors · ACM Trans. Comput. Hum. Interact. 2005
Predicting human interruptibility with sensors: a Wizard of Oz feasibility study · CHI 2003
Machine learning › Representation and self-supervised learning › representation learning › feature extraction
bag-of-features
0.112009
A biased sampling strategy for object categorization · ICCV 2009
Machine learning › Probabilistic and Bayesian machine learning › sampling
feature sampling
0.112009
A biased sampling strategy for object categorization · ICCV 2009
Computer vision › Image recognition and object detection
image annotation
0.112009
Semi-Automatically Labeling Objects in Images · IEEE Trans. Image Process. 2009
Machine learning › Learning paradigms › semi-supervised learning › graph-based semi-supervised learning
label propagation
0.112009
Sparsity induced similarity measure for label propagation · ICCV 2009
Computer vision › Image recognition and object detection › image classification
object classification
0.112009
A biased sampling strategy for object categorization · ICCV 2009
Machine learning › Learning paradigms
semi-supervised learning
0.112009
Sparsity induced similarity measure for label propagation · ICCV 2009
Machine learning › Representation and self-supervised learning
similarity measure
0.112009
Sparsity induced similarity measure for label propagation · ICCV 2009
Information retrieval
relevance feedback
0.112009
Semi-Automatically Labeling Objects in Images · IEEE Trans. Image Process. 2009
Computer vision › 3D vision
local feature descriptor
0.112008
A deformable local image descriptor · CVPR 2008
Multimedia analysis and retrieval › video analysis
compressed domain analysis
0.112008
Mining Appearance Models Directly From Compressed Video · IEEE Trans. Multim. 2008
Multimedia analysis and retrieval › image retrieval
content-based image retrieval
0.112008
Object fingerprints for content analysis with applications to street landmark localization · ACM Multimedia 2008
Image and video processing › feature detection
landmark detection
0.112008
Object fingerprints for content analysis with applications to street landmark localization · ACM Multimedia 2008
Multimedia analysis and retrieval
video content analysis
0.112008
Mining Appearance Models Directly From Compressed Video · IEEE Trans. Multim. 2008
Computer vision › Video understanding and tracking › object tracking
appearance modeling
0.112007
Robust Object Tracking Via Online Dynamic Spatial Bias Appearance Models · IEEE Trans. Pattern Anal. Mach. Intell. 2007
Computer vision › Video understanding and tracking
object tracking
0.112007
Robust Object Tracking Via Online Dynamic Spatial Bias Appearance Models · IEEE Trans. Pattern Anal. Mach. Intell. 2007
Collaborative and social computing › remote collaboration
remote assistance
0.112007
Sharing a single expert among multiple partners · CHI 2007
Computer vision › Face, body and person analysis › face alignment
active appearance model fitting
0.112006
Robust AAM Fitting by Fusion of Images and Disparity Data · CVPR (2) 2006
Computer vision › Image recognition and object detection
object labeling
0.112006
SmartLabel: an object labeling tool using iterated harmonic energy minimization · ACM Multimedia 2006
Computer vision › 3D vision › stereo vision
stereo matching
0.112006
Robust AAM Fitting by Fusion of Images and Disparity Data · CVPR (2) 2006
Data mining
clustering
0.112006
SmartLabel: an object labeling tool using iterated harmonic energy minimization · ACM Multimedia 2006
Data mining
semi-supervised learning
0.112006
SmartLabel: an object labeling tool using iterated harmonic energy minimization · ACM Multimedia 2006

Methods — techniques the papers use, named apart from their topics

biased sampling · 0.4superpixel over-segmentation · 0.2saliency detection · 0.2quadtree partitioning · 0.2harmonic solution · 0.2gaussian fields · 0.2machine learning · 0.1relevance feedback · 0.1iterated harmonic energy minimization · 0.1gaussian random field · 0.1l1 sparse decomposition · 0.1bag-of-words · 0.1segmentation · 0.1salient region detection · 0.1multi-size support regions · 0.1motion vector · 0.1mixture of gaussians · 0.1local-to-global similarity · 0.1
YearPublicationVenuePosition
2017 Gazing point dependent eye gaze estimation
Hong Cheng 0002, Yanli Ji, Lu Yang 0002, Yang Zhao 0024, Jie Yang 0001
Pattern Recognit.7
2016 Algorithm and VLSI Architecture of Edge-Directed Image Upscaling for 4k Display System
abstract
High-quality and cost-efficient image upscaling design is very important for many real-time video processing applications, especially when the display panel resolution reaches ultrahigh definition. Compared with New Edge-Directed Interpolation (NEDI) based implicit edge directional upscaling, explicit methods require less computational resource and more easily reach real-time performance, especially when the required image definition and upscaling ratio are very high. Nevertheless, the investigation of applications of explicit methods in video processing systems remains largely missing arguably because it is commonly believed that explicit edge-directed interpolation tends to introduce unexpected artifacts because of inaccurate detection and hence its image quality is relatively poor. This paper proposes an explicit edge-directed adaptive interpolation method that leverages more sophisticated edge detection and orientation estimation algorithms to avoid misinterpolation, thereby providing similar or even better image quality than those with implicit methods. Targeting the real-time 4K video display system, the proposed edge-directed image upscaling algorithm is further implemented with an efficient very-large-scale integration (VLSI) architecture. The experimental results demonstrate that the proposed interpolation algorithm outperforms previous explicit and implicit edge-directed methods in both objective and subjective tests. The presented VLSI implementation further demonstrates that the maximum output video sequence of the proposed interpolation method can reach 4k × 2k@60 Hz with a reasonable hardware cost.
Qiubo Chen, Hongbin Sun 0001, Xuchong Zhang, Huibin Tao, Jie Yang 0001, Jizhong Zhao, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.5
2016 Sparsity-Induced Similarity Measure and Its Applications
abstract
The structures of feature vectors-based semisupervised/supervised learning have gained considerable interest in recent years due to their effectiveness for better object modeling and classification. In many machine learning and computer vision tasks, a critical issue is the similarity between two feature vectors. In this paper, we present a novel technique to measure similarities among feature vectors by decomposing each feature vector as an ℓ1sparse linear combination of the rest of the feature vectors. The main idea is that the coefficients in such sparse decomposition reflect the features' neighborhood structure, thus providing better similarity measures among the decomposed feature vector and the rest of the feature vectors. The proposed approach is applied to label propagation and action recognition, and is evaluated on several commonly used datasets. The experimental results show that the proposed sparsity-induced similarity measure significantly improves the performance of both label propagation and action recognition.
Hong Cheng 0002, Zicheng Liu 0001, Lei Hou 0019, Jie Yang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2013 Categorization of Multiple Objects in a Scene Using a Biased Sampling Strategy
Lei Yang 0063, Nanning Zheng 0001, Yang Yang 0025, Jie Yang 0001
Int. J. Comput. Vis.5
2012 A paradigm for handwriting-based intelligent tutors
Lisa Anthony, Jie Yang 0001, Kenneth R. Koedinger
Int. J. Hum. Comput. Stud.2
2012 Multi-support-region image descriptors and its application to street landmark localization
Hong Cheng 0002, Zicheng Liu 0001, Jie Yang 0001
Mach. Vis. Appl.3
2011 A unified context assessing model for object categorization
Lei Yang 0063, Nanning Zheng 0001, Jie Yang 0001
Comput. Vis. Image Underst.3
2010 Modeling Urban Scenes in the Spatial-Temporal Space
Jiong Xu, Jie Yang 0001
ACCV (2)3
2010 Learning feature transforms for object detection from panoramic images
abstract
We present a novel technique to detect objects from panoramic images using existing object detectors trained from perspective images. By leveraging existing object detectors, we save the cost of training a new detector which requires tedious and time consuming training data collection and labeling. The core of our technique is learning a feature transform which is represented by Gaussian Process Regression (GPR). Feature vectors computed directly from panoramic images are transformed into new feature vectors in such a way that the existing classifier has much better detection rate on the transformed feature vectors. Our feature transform has the interesting property that it not only corrects for the geometric distortions resulted from panoramic imaging process, but also corrects for the pose mismatches between the objects on the panoramic images and those on the training images. Our experiments show that we are able to successfully apply an existing car detector trained on perspective images to panoramic images which have both geometric distortions and larger pose variations.
Hong Cheng 0002, Zicheng Liu 0001, Jie Yang 0001
ICME3
2010 Locality preserving and global discriminant projection with prior information
Honggang Zhang 0002, Weihong Deng, Jun Guo 0002, Jie Yang 0001
Mach. Vis. Appl.4
2009 Categorization of Multiple Objects in a Scene without Semantic Segmentation
Lei Yang 0063, Nanning Zheng 0001, Yang Yang 0066, Jie Yang 0001
ACCV (1)5
2009 Sparsity induced similarity measure for label propagation
abstract
Graph-based semi-supervised learning has gained considerable interests in the past several years thanks to its effectiveness in combining labeled and unlabeled data through label propagation for better object modeling and classification. A critical issue in constructing a graph is the weight assignment where the weight of an edge specifies the similarity between two data points. In this paper, we present a novel technique to measure the similarities among data points by decomposing each data point as an L1sparse linear combination of the rest of the data points. The main idea is that the coefficients in such a sparse decomposition reflect the point's neighborhood structure thus providing better similarity measures among the decomposed data point and the rest of the data points. The proposed approach is evaluated on four commonly-used data sets and the experimental results show that the proposed Sparsity Induced Similarity (SIS) measure significantly improves label propagation performance. As an application of the SIS-based label propagation, we show that the SIS measure can be used to improve the Bag-of-Words approach for scene classification.
Hong Cheng 0002, Zicheng Liu 0001, Jie Yang 0001
ICCV3
2009 A biased sampling strategy for object categorization
abstract
In this paper, we present a biased sampling strategy for object class modeling, which can effectively circumvent the scene matching problem commonly encountered in statistical image-based object categorization. The method optimally combines the bottom-up, biologically inspired saliency information with loose, top-down class prior information to form a probabilistic distribution for feature sampling. When sampling over different positions and scales of patches, the weak spatial coherency is preserved by a segment-based analysis. We evaluate the proposed sampling strategy within the bag-of-features (BoF) object categorization framework on three public data sets. Our technique outperforms other state-of-the-art sampling technologies, and leads to a better performance in object categorization on VOC2008 dataset.
Lei Yang 0063, Nanning Zheng 0001, Jie Yang 0001
ICCV3
2009 PFID: Pittsburgh fast-food image dataset
abstract
We introduce the first visual dataset of fast foods with a total of 4,545 still images, 606 stereo pairs, 303 3600videos for structure from motion, and 27 privacy-preserving videos of eating events of volunteers. This work was motivated by research on fast food recognition for dietary assessment. The data was collected by obtaining three instances of 101 foods from 11 popular fast food chains, and capturing images and videos in both restaurant conditions and a controlled lab setting. We benchmark the dataset using two standard approaches, color histogram and bag of SIFT features in conjunction with a discriminative classifier. Our dataset and the benchmarks are designed to stimulate research in this area and will be released freely to the research community.
Kapil Dhingra, Lei Yang 0063, Rahul Sukthankar, Jie Yang 0001
ICIP6
2009 A new tracking method for small infrared targets
abstract
We report on a new approach to tracking small infrared targets. The method improves on existing target trackers by combining mean-shift tracker with Kalman filtering and by updating the tracking parameters through the measurement of the complexity of the target region. We have further developed a nonlinear algorithm to improve the robustness of the traditional mean-shift tracker for small infrared targets. Experimental results demonstrate a superior performance of our method compared to existing target trackers, particularly in the environment of strong measurement noise and large variation of illumination.
Lei Yang 0063, Weiping Lu, Jie Yang 0001
ICIP3
2009 Towards virtually cooking Chinese food
abstract
Chinese food is delicious but cooking Chinese food is a very complex process which involves in combing ingredients at the right time and temperature. In this paper, we present a multimedia technology for virtually cooking Chinese food. We focus on a popular Chinese food, shredded potato. We propose to use the theory from mechanics of materials to model shredded potato in its cooking process. The shredded potato is initially simplified with changed deformable beams by adjusting model parameters during the cooking process when shredded potato gets dasiasoftpsila deformations. We use superposition principle to cope with the multi-load problem when potato shreds pile up and contact with each other. We describe the modeling techniques and implementation issues in detail. We show the result from the proposed modeling technique by comparing it with the real image.
Jinghao Fei, Jie Yang 0001, Jianping Fan 0002
ICME2
2009 Fast food recognition from videos of eating for calorie estimation
abstract
Accurate and passive acquisition of dietary data from patients is essential for a better understanding of the etiology of obesity and development of effective weight management programs. Self-reporting is currently the main method for such data acquisition. However, studies have shown that data obtained by self-reporting seriously underestimate food intake and thus do not accurately reflect the real habitual behavior of individuals. Computer food recognition programs have not yet been developed. In this paper, we present a study for recognizing foods from videos of eating, which are directly recorded in restaurants by a web camera. From recognition results, our method then estimates food calories of intake. We have evaluated our method on a database of 101 foods from 9 food restaurants in USA and obtained promising results.
Jie Yang 0001
ICME2
2009 An image-based automatic Arabic translation system
Yi Chang 0001, Datong Chen, Ying Zhang 0048, Jie Yang 0001
Pattern Recognit.4
2009 Optimization of a training set for more robust face detection
Jie Chen 0001, Xilin Chen 0001, Jie Yang 0001, Shiguang Shan, Ruiping Wang 0001, Wen Gao 0001
Pattern Recognit.3
2009 Semi-Automatically Labeling Objects in Images
abstract
Labeling objects in images plays a crucial role in many visual learning and recognition applications that need training data, such as image retrieval, object detection and recognition. Manually creating object labels in images is time consuming and, thus, becomes impossible for labeling a large image dataset. In this paper, we present a family of semi-automatic methods based on a graph-based semi-supervised learning algorithm for labeling objects in images. We first present SmartLabel that proposes to label images with reduced human input by iteratively computing the harmonic solutions to minimize a quadratic energy function on the Gaussian fields. SmartLabel tackles the problem of lacking negative data in the learning by embedding relevance feedback after the first iteration, which also leads to one limitation of SmartLabel-needing additional human supervision. To overcome the limitation and enhance SmartLabel, we propose SmartLabel-2 that utilizes a novel scheme to sample negative examples automatically, replace regular patch partitioning in SmartLabel by quadtree partitioning and applies image over-segmentation (superpixels) to extract smooth object contours. Evaluation on six diverse object categories have indicated that SmartLabel-2 can achieve promising results with a small amount of labeled data (e.g., 1%-5% of image size) and obtain close-to-fine extraction of object contours on different kinds of objects.
Jie Yang 0001
IEEE Trans. Image Process.2
2008 A deformable local image descriptor
abstract
This paper presents a novel local image descriptor that is robust to general image deformations. A limitation with traditional image descriptors is that they use a single support region for each interest point. For general image deformations, the amount of deformation for each location varies and is unpredictable such that it is difficult to choose the best scale of the support region. To overcome this difficulty, we propose to use multiple support regions of different sizes surrounding an interest point. A feature vector is computed for each support region, and the concatenation of these feature vectors forms the descriptor for this interest point. Furthermore, we propose a new similarity measure model, Local-to-Global Similarity (LGS) model, for point matching that takes advantage of the multi-size support regions. Each support region acts as a ‘weak’ classifier and the weights of these classifiers are learned in an unsupervised manner. The proposed approach is evaluated on a number of images with real and synthetic deformations. The experiment results show that our method outperforms existing techniques under different deformations.
Hong Cheng 0002, Zicheng Liu 0001, Nanning Zheng 0001, Jie Yang 0001
CVPR4
2008 Layered object categorization
abstract
In this paper, we propose a novel framework of object categorization, namely layered object categorization, which takes advantage of hierarchical category information and performs object categorization at different levels. The proposed hierarchical structure of object categories is built bottom-up and top-down simultaneously accordingly to cognitive rules. First, part-based models are learnt to evaluate structure similarities at the basic level and objects are divided into basic categories. Then the decision cues for object categorization at different layers are optimally selected. Prior knowledge about inter-category relationships is utilized to infer objectspsila higher inclusive concept labels, while the most discriminative visual details of each category at the lower specific levels are selected automatically. We evaluate the proposed method with a hierarchical database and show promising results. The layered object categorization provides an efficient way for dynamically adapting the object categorization results to different applications.
Lei Yang 0063, Jie Yang 0001, Nanning Zheng 0001, Hong Cheng 0002
ICPR2
2008 Handwritten Chinese character recognition using Local Discriminant Projection with Prior Information
abstract
In this paper, we propose a new method to model the manifold of handwritten Chinese characters using the local discriminant projection. We utilize a cascade framework that combines global similarity with local discriminative cues to recognize Chinese characters. We find the similarity of different characters using a nearest-neighbor (NN) classifier, and followed by the Local Discriminant Projection with Prior Information (LDPPI) to map similar characters within a cluster to a low-dimensional space. We evaluate the proposed method on two large public datasets, ETL9B which contains 607,200 handwritten characters from 200 people, and HCL2000 which contains 3,755,000 characters written by 1,000 people. The experimental results demonstrate that the proposed method achieves 0.74% error rate on ETL9B database and 1.88% on HCL2000 database.
Honggang Zhang 0002, Jie Yang 0001, Weihong Deng, Jun Guo 0002
ICPR2
2008 A sparse representation of physical activity video in the study of obesity
abstract
Wearable visual devices have many emerging applications in human health monitoring, such as the study of food intake and physical activities. However, these devices produce large amounts of data which must be processed efficiently and effectively. In this paper, we present a video- based health monitoring system for physical activity studies. We utilize an efficient signal extraction method to process and reduce the effects of noises on field-acquired image data. We further develop a maximum likelihood estimator to estimate the walking speed and classify activity patterns (walking or running) in real-time under both indoor and outdoor environments. We demonstrate the feasibility of the proposed method using three different test video sequences.
Robert J. Sclabassi, Qiang Liu 0033, John D. Fernstrom, Madelyn H. Fernstrom, Jie Yang 0001, Mingui Sun
ISCAS6
2008 Object fingerprints for content analysis with applications to street landmark localization
abstract
An object can be a basic unit for multimedia content analysis. Besides similarity among common objects, each object has its own unique characteristics which we cannot find in other surrounding objects in multimedia data. We call such unique characteristics object fingerprints. In this paper, we propose a novel approach to extract and match object fingerprints for multimedia content analysis. In particular, we focus on the problem of street landmark localization from images. Instead of modeling and matching a street landmark as a whole, our proposed approach extracts the landmark’s object fingerprints in a given image and match to a new image or video in order to localize the landmark. We formulate matching the landmark’s object fingerprints as a classification problem solved by a cascade of 1NN classifiers. We develop a street landmark localization system that combines salient region detection, segmentation, and object fingerprint extraction techniques for the purpose. To evaluate, we have compiled a novel dataset which consists of 15 U.S. street landmarks ’ images and videos. Our experiments on this dataset show superior performance to state-of-the-art recognition algorithms [20, 33]. The proposed approach can also be well generalized to other objects of interest and content analysis tasks. We demonstrate the feasibility through the application of our approach to refine web image search results and obtained encouraging results.
Jie Yang 0001
ACM Multimedia2
2008 Modeling Background and Segmenting Moving Objects from Compressed Video
abstract
Modeling background and segmenting moving objects are significant techniques for video surveillance and other video processing applications. Most existing methods of modeling background and segmenting moving objects mainly operate in the spatial domain at pixel level. In this paper, we present three new algorithms (running average, median, mixture of Gaussians) modeling background directly from compressed video, and a two-stage segmentation approach based on the proposed background models. The proposed methods utilize discrete cosine transform (DCT) coefficients (including ac coefficients) at block level to represent background, and adapt the background by updating DCT coefficients. The proposed segmentation approach can extract foreground objects with pixel accuracy through a two-stage process. First a new background subtraction technique in the DCT domain is exploited to identify the block regions fully or partially occupied by foreground objects, and then pixels from these foreground blocks are further classified in the spatial domain. The experimental results show the proposed background modeling algorithms can achieve comparable accuracy to their counterparts in the spatial domain, and the associated segmentation scheme can visually generate good segmentation results with efficient computation. For instance, the computational cost of the proposed median and MoG algorithms are only 40.4% and 20.6% of their counterparts in the spatial domain for background construction.
Weiqiang Wang 0001, Jie Yang 0001, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2008 Mining Appearance Models Directly From Compressed Video
abstract
In this paper, we propose an approach for learning appearance models of moving objects directly from compressed video. The appearance of a moving object changes dynamically in video due to varying object poses, lighting conditions, and partial occlusions. Efficiently mining the appearance models of objects is a crucial and challenging technology to support content-based video coding, clustering, indexing, and retrieval at the object level. The proposed approach learns the appearance models of moving objects in the spatial-temporal dimension of video data by taking advantage of the MPEG video compression format. It detects a moving object and recovers the trajectory of each macroblock covered by the object using the motion vector present in the compressed stream. The appearances are then reconstructed in the DCT domain along the object's trajectory, and modeled as a mixture of Gaussians (MoG) using DCT coefficients. We prove that, under certain assumptions, the MoG model learned from the DCT domain can achieve pixel-level accuracy when transformed back to the spatial domain, and has a better band-selectivity compared to the MoG model learned in the spatial domain. We finally cluster the MoG models to merge the appearance models of the same object together for object-level content analysis.
Datong Chen, Qiang Liu 0033, Mingui Sun, Jie Yang 0001
IEEE Trans. Multim.4
2008 Predicting Visual Focus of Attention From Intention in Remote Collaborative Tasks
abstract
While shared visual space plays a very important role in remote collaboration on physical tasks, it is challenging and expensive to track users' focus of attention (FOA) during these tasks. In this paper, we propose to identify a user's FOA from his/her intention based on task properties, people's actions in the workspace, and conversational content. We employ a conditional Markov model to characterize a subject's FOA. We demonstrate the feasibility of the proposed method using a collaborative laboratory task in which one partner (the helper) instructs another (the worker) on how to assemble online puzzles. We model a helper's FOA using task properties, workers' actions, and conversational content. The accuracy of the model ranged from 65.40% for puzzles with easy-to-name pieces to 74.25% for puzzles with more difficult-to-name pieces. The proposed model can be used to predicate a user's FOA in a remote collaborative task without tracking the user's eye gaze.
Jiazhi Ou, Lui Min Oh, Susan R. Fussell, Tal Blum, Jie Yang 0001
IEEE Trans. Multim.5
2007 Benefits of Handwritten Input for Students Learning Algebra Equation Solving
Lisa Anthony, Jie Yang 0001, Kenneth R. Koedinger
AIED2
2007 Sharing a single expert among multiple partners
abstract
Expertise to assist people on complex tasks is often in short supply. One solution to this problem is to design systems that allow remote experts to help multiple people in simultaneously. As a first step towards building such a system, we studied experts' attention and communication as they assisted two novices at the same time in a co-located setting. We compared simultaneous instruction when the novices are being instructed to do the same task or different tasks. Using machine learning, we attempted to identify speech markers of upcoming attention shifts that could serve as input to a remote assistance system.
Jeffrey Wong, Lui Min Oh, Jiazhi Ou, Carolyn P. Rosé, Jie Yang 0001, Susan R. Fussell
CHI5
2007 Enhancing a Driver's Situation Awareness using a Global View Map
abstract
This paper proposes a novel method to enhance a driver's situation awareness by dynamically providing a global view of surroundings for the driver. The surroundings of a vehicle are captured by an omni-directional vision system mounted on the top of the vehicle. The video stream from the camera is processed to detect nearby vehicles. Positions of these detected objects are overlaid on a global view of a local map (e.g., an aerial imagery or satellite imagery map). We establish the relationship between the omni-directional vision system and the global view map. The global view map dynamically provides a realistic perspective view of the driving environment. This map can be projected onto an HUD on the windshield. By looking at the display, a driver can have a global picture of the situation and potentially produce a good driving strategy. We illustrate the proposed method by dynamically mapping a video stream onto Google Earth map.
Hong Cheng 0002, Zicheng Liu 0001, Nanning Zheng 0001, Jie Yang 0001
ICME4
2007 A Video Processing Approach to the Study of Obesity
abstract
Currently, a significant obstacle in identifying the most important factors that contribute to obesity is the lack of appropriate tools to evaluate food consumption and physical activity on a daily basis. This paper presents a novel video-based approach to estimating the energy intake and expenditure of individuals who are over-weight or obese. We propose to utilize a miniature video camera to chronically record data surrounding an individual's daily activity. From the recorded data, energy intake and expenditure are extracted objectively using signal processing algorithms. We discuss two algorithms which are both featured with low computational complexity, suitable for processing large sets of field-acquired video data. These algorithms estimate: (i) energy intake (caloric content) by determining the physical dimensions of the foods consumed; and (ii) energy expenditure by evaluating the extent of the subject's physical activity. We also report initial experimental results to demonstrate the effectiveness of the new approach.
Robert J. Sclabassi, Qiang Liu 0033, Jie Yang 0001, John D. Fernstrom, Madelyn H. Fernstrom, Mingui Sun
ICME4
2007 Robust Object Tracking Via Online Dynamic Spatial Bias Appearance Models
abstract
This paper presents a robust object tracking method via a spatial bias appearance model learned dynamically in video. Motivated by the attention shifting among local regions of a human vision system during object tracking, we propose to partition an object into regions with different confidences and track the object using a dynamic spatial bias appearance model (DSBAM) estimated from region confidences. The confidence of a region is estimated to re ect the discriminative power of the region in a feature space, and the probability of occlusion. We propose a novel hierarchical Monte Carlo (HAMC) algorithm to learn region confidences dynamically in every frame. The algorithm consists of two levels of Monte Carlo processes implemented using two particle filtering procedures at each level and can efficiently extract high confidence regions through video frames by exploiting the temporal consistency of region confidences. A dynamic spatial bias map is then generated from the high confidence regions, and is employed to adapt the appearance model of the object and to guide a tracking algorithm in searching for correspondences in adjacent frames of video images. We demonstrate feasibility of the proposed method in video surveillance applications. The proposed method can be combined with many other existing tracking systems to enhance the robustness of these systems.
Datong Chen, Jie Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 Detecting social interactions of the elderly in a nursing home environment
abstract
Social interaction plays an important role in our daily lives. It is one of the most important indicators of physical or mental changes in aging patients. In this article, we investigate the problem of detecting social interaction patterns of patients in a skilled nursing facility using audio/visual records. Our studies consist of both a “Wizard of Oz” style study and an experimental study of various sensors and detection models for detecting and summarizing social interactions among aging patients and caregivers. We first simulate plausible sensors using human labeling on top of audio and visual data collected from a skilled nursing facility. The most useful sensors and robust detection models are determined using the simulated sensors. We then present the implementation of some real sensors based on video and audio analysis techniques and evaluate the performance of these implementations in detecting interactions. We conclude the article with discussions and future work.
Datong Chen, Jie Yang 0001, Robert G. Malkin, Howard D. Wactlar
ACM Trans. Multim. Comput. Commun. Appl.2
2006 Robust AAM Fitting by Fusion of Images and Disparity Data
abstract
Active Appearance Models (AAMs) have been popularly used to represent the appearance and shape variations of human faces. Fitting an AAM to images recovers the face pose as well as its deformable shape and varying appearance. Successful fitting requires that the AAM is sufficiently generic such that it covers all possible facial appearances and shapes in the images. Such a generic AAM is often difficult to be obtained in practice, especially when the image quality is low or when occlusion occurs. To achieve robust AAM fitting under such circumstances, this paper proposes to incorporate the disparity data obtained from a stereo camera with the image fitting process. We develop an iterative multi-level algorithm that combines efficient AAM fitting to 2D images and robust 3D shape alignment to disparity data. Experiments on tracking faces in low-resolution images captured from meeting scenarios show that the proposed method achieves better performance than the original 2D AAM fitting algorithm. We also demonstrate an application of the proposed method to a facial expression recognition task.
Joerg Liebelt, Jing Xiao 0006, Jie Yang 0001
CVPR (2)3
2006 Towards the Application of a Handwriting Interface for Mathematics Learning
abstract
We believe handwriting input may be able to provide significant advantages over typing, especially in the mathematics learning domain. The use of handwriting may result in decreased extraneous cognitive load on students, and it may provide better support for the two-dimensional spatial components of mathematics when compared to existing typing-based tools. Here we report progress towards the application of a handwriting interface for mathematics learning. We introduce a prototype system that allows students to use handwriting input to solve algebraic equations in an intelligent tutor. We discuss strategies to improve the existing handwriting system and apply it to math learning. Although the recognition accuracy of current handwriting engines may not be at a level suitable for use by students, we hypothesize that this may be realistically improved via advance training of the engine on a large corpus, as well as via techniques similar to co-training
Lisa Anthony, Jie Yang 0001, Kenneth R. Koedinger
ICME2
2006 People Identification with Limited Labels in Privacy-Protected Video
abstract
People identification is an essential task for video content analysis in a surveillance system. To construct a good classifier requires a large amount of training data, which may not be obtained in some scenario. In this paper, we propose an approach to augment insufficient training data by labeling identical video images that have removed people's identities by masking faces. We show user study results that human subjects can perform reasonably well in labeling pairwise constraints from face obscured images. We also present a new discriminative learning algorithm WPKLR to handle uncertainties in pairwise constraints. The effectiveness of the proposed approach is demonstrated using video captured in a nursing home environment. The experiments show that the WPKLR approach can obtain a high accuracy of people identification using limited labeled data and noisy pairwise constraints, and meanwhile minimize the risk of exposing people's identities
Yi Chang 0001, Datong Chen, Jie Yang 0001
ICME4
2006 Directing Attention in Online Aggregate Sensor Streams via Auditory Blind Value Assignment
abstract
Multiparty collaborative applications in which groups of people act in concert to achieve some real-world goal abound. In these situations, it is useful for a central planning agent to receive online audio-visual information from all participants. However, as the size of the group grows, it becomes difficult to process all the sensory streams; cognitive overload prevents direct analysis of sensory streams for situational awareness. To avoid this situation, an automatic method is needed to assign value to each stream and direct the attention of the planning agent to those streams which are most valuable. We present an audio-based blind value assignment (BVA) method to address this problem, and experiments demonstrating the method's efficacy. We demonstrate that use of audio BVA techniques results in automatic value judgments which are broadly similar to human value judgments and superior to automatic judgments based on video information
Robert G. Malkin, Datong Chen, Jie Yang 0001, Alex Waibel
ICME3
2006 A Multimedia System for Route Sharing and Video-Based Navigation
abstract
Trip planning and in-vehicle navigation are crucial tasks for easier and safer driving. The existing navigation systems are based on machine intelligence without allowing human knowledge incorporation. These systems give turn guidance with abstract visual instruction and have not reached the potential of minimizing driver's cognitive load, which is the amount of mental processing power required. In this paper, we describe the development of a multimedia system that makes driving and navigation safer and easier by offering tools for route sharing in trip planning and video-based route guidance during driving. The system provides a multimodal interface for a user to share his/her route with others by drawing on a digital map, naturally incorporating human knowledge into the trip planning process. The system gives driving instructions by overlaying navigational arrows onto live video and providing synthesized voice to reduce the driver's cognitive load, in addition to presenting landmark images for key maneuvers. We describe our observations which had motivated the development of the system, detailed architecture and user interfaces, and finally discusses our initial test findings in the real-road driving context
Jie Yang 0001, Jing Zhang 0011
ICME2
2006 Webdove: a Web-Based Collaboration System for Physical Tasks
abstract
While many systems are available for audio-visual people collaboration and data collaboration, systems for collaboration on physical objects are few. In this paper, we present WebDOVE, a system designed to address the needs of collaborative physical tasks. WebDOVE supports both live video streams and pen-based gesture recognition in multi-party bidirectional communication via inexpensive Web cameras. WebDOVE allows distributed collaborators to draw over video streams to produce and interpret pointing and representational gestures as readily as they do in face-to-face settings. To accommodate potential diverse platform requirements from different participants, WebDOVE is designed to be a Web-based platform-independent and browser-independent collaboration solution. We show via experiments that despite WebDOVE's platform independency, it requires moderate network bandwidth and CPU load, which make WebDOVE a practical solution for real-world applications
Jiazhi Ou, Yong Rut, Jie Yang 0001
ICME4
2006 Multimodal estimation of user interruptibility for smart mobile telephones
abstract
Context-aware computer systems are characterized by the ability to consider user state information in their decision logic. One example application of context-aware computing is the smart mobile telephone. Ideally, a smart mobile telephone should be able to consider both social factors (i.e., known relationships between contactor and contactee) and environmental factors (i.e., the contactee's current locale and activity) when deciding how to handle an incoming request for communication.Toward providing this kind of user state information and improving the ability of the mobile phone to handle calls intelligently, we present work on inferring environmental factors from sensory data and using this information to predict user interruptibility. Specifically, we learn the structure and parameters of a user state model from continuous ambient audio and visual information from periodic still images, and attempt to associate the learned states with user-reported interruptibility levels. We report experimental results using this technique on real data, and show how such an approach can allow for adaptation to specific user preferences.
Robert G. Malkin, Datong Chen, Jie Yang 0001, Alex Waibel
ICMI3
2006 Combining audio and video to predict helpers' focus of attention in multiparty remote collaboration on physical tasks
abstract
The increasing interest in supporting multiparty remote collaboration has created both opportunities and challenges for the research community. The research reported here aims to develop tools to support multiparty remote collaborations and to study human behaviors using these tools. In this paper we first introduce an experimental multimedia (video and audio) system with which an expert can collaborate with several novices. We then use this system to study helpers' focus of attention (FOA) during a collaborative circuit assembly task. We investigate the relationship between FOA and language as well as activities using multimodal (audio and video) data, and use learning methods to predict helpers' FOA. We process different modalities separately and fusion the results to make a final decision. We employ a sliding window-based delayed labeling method to automatically predict changes in FOA in real time using only the dialogue among the helper and workers. We apply an adaptive background subtraction method and support vector machine to recognize the worker's activities from the video. To predict the helper's FOA, we make decisions using the information of joint project boundaries and workers' recent activities. The overall prediction accuracies are 79.52% using audio only and 81.79% using audio and video combined.
Jiazhi Ou, Yanxin Shi, Jeffrey Wong, Susan R. Fussell, Jie Yang 0001
ICMI5
2006 SmartLabel: an object labeling tool using iterated harmonic energy minimization
abstract
Labeling objects in images is an essential prerequisite for many visual learning and recognition applications that depend on training data, such as image retrieval, object detection and recognition. Manually creating labels in images is not only time-consuming but also subject to human labeling errors, and eventually, becomes impossible for a large scale image database. Semi-supervised learning (SSL)algorithms such as Gaussian random field (GRF)can be applied to labeling objects in images since they have the ability to include a large amount of unlabeled data while requiring only a small amount of labeled data. However, the one-shot property of GRF prevents it from achieving good labeling performance. In this paper, we presents a novel object labeling tool, SmartLabel, to semi-automatically label objects in images. The algorithm of SmartLabel has four innovations over GRF:1)soft labeling,2)graph construction with spatial constraints, 3)iterated harmonic energy minimization, and 4)using relevance feedback to incorporate human interaction in the loop. As demonstrated in datasets of six object categories, the proposed SmartLabel not only works effectively even with a very small amount of user input (e.g., 1 .5%of image size)but also achieves significant improvement over GRF.
Jie Yang 0001
ACM Multimedia2
2006 A Discriminative Learning Framework with Pairwise Constraints for Video Object Classification
abstract
To deal with the problem of insufficient labeled data in video object classification, one solution is to utilize additional pairwise constraints that indicate the relationship between two examples, i.e., whether these examples belong to the same class or not. In this paper, we propose a discriminative learning approach which can incorporate pairwise constraints into a conventional margin-based learning framework. Different from previous work that usually attempts to learn better distance metrics or estimate the underlying data distribution, the proposed approach can directly model the decision boundary and, thus, require fewer model assumptions. Moreover, the proposed approach can handle both labeled data and pairwise constraints in a unified framework. In this work, we investigate two families of pairwise loss functions, namely, convex and nonconvex pairwise loss functions, and then derive three pairwise learning algorithms by plugging in the hinge loss and the logistic loss functions. The proposed learning algorithms were evaluated using a people identification task on two surveillance video data sets. The experiments demonstrated that the proposed pairwise learning algorithms considerably outperform the baseline classifiers using only labeled data and two other pairwise learning algorithms with the same amount of pairwise constraints.
Jian Zhang 0003, Jie Yang 0001, Alex Hauptmann 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2005 Effects of task properties, partner actions, and message content on eye gaze patterns in a collaborative task
abstract
Helpers providing guidance for collaborative physical tasks shift their gaze between the workspace, supply area, and instructions. Understanding when and why helpers gaze at each area is important both for a theoretical understanding of collaboration on physical tasks and for the design of automated video systems for remote collaboration. In a laboratory experiment using a collaborative puzzle task, we recorded helpers' gaze while manipulating task complexity and piece differentiability. Helpers gazed toward the pieces bay more frequently when pieces were difficult to differentiate and less frequently over repeated trials. Preliminary analyses of message content show that helpers tend to look at the pieces bay when describing the next piece and at the workspace when describing where it goes. The results are consistent with a grounding model of communication, in which helpers seek visual evidence of understanding unless they are confident that they have been understood. The results also suggest the feasibility of building automated video systems based on remote helpers' shifting visual requirements.
Jiazhi Ou, Lui Min Oh, Jie Yang 0001, Susan R. Fussell
CHI3
2005 Integrating co-training and recognition for text detection
abstract
Training a good text detector requires a large amount of labeled data, which can be very expensive to obtain. Co-training has been shown to be a powerful semi-supervised learning tool for solving many problems using a large amount of unlabeled data. However, augmented data from a co-training process could potentially degrade the performance of classifiers due to added noises from unlabeled data. This paper makes two contributions by proposing a modified co-training scheme for text detection. First, to get cleaner augmented data, the new algorithm integrates some authority knowledge of unlabeled data into co-training. Text recognition output of each selected unlabeled image patch is used as the authority that is combined with classifier prediction to decide if the sample will be added to the augmented set. Second, instead of evenly combining predictions of two co-training classifiers, a weighted combination is learned and used to produce the final prediction. Contributions of the new algorithm have been evaluated on a standard text detection dataset.
Datong Chen, Jie Yang 0001
ICME3
2005 Analyzing and predicting focus of attention in remote collaborative tasks
abstract
To overcome the limitations of current technologies for remote collaboration, we propose a system that changes a video feed based on task properties, people’s actions, and message properties. First, we examined how participants manage different visual resources in a laboratory experiment using a collaborative task in which one partner (the helper) instructs another (the worker) how to assemble online puzzles. We analyzed helpers’ eye gaze as a function of the aforementioned parameters. Helpers gazed at the set of alternative pieces more frequently when it was harder for workers to differentiate these pieces, and less frequently over repeated trials. The results further suggest that a helper’s desired focus of attention can be predicted based on task properties, his/her partner’s actions, and message properties. We propose a conditional Markov model classifier to explore the feasibility of predicting gaze based on these properties. The accuracy of the model ranged from 65.40 % for puzzles with easyto-name pieces to 74.25 % for puzzles with more difficult to name pieces. The results suggest that we can use our model to automatically manipulate video feeds to show what helpers want to see when they want to see it. Categories and Subject Descriptors H5.3. Information interfaces and presentation (e.g., HCI): Group and organizational interfaces – collaborative computing, computer-supported collaborative work
Jiazhi Ou, Lui Min Oh, Susan R. Fussell, Tal Blum, Jie Yang 0001
ICMI5
2005 Detection of text on road signs from video
abstract
A fast and robust framework for incrementally detecting text on road signs from video is presented in this paper. This new framework makes two main contributions. 1) The framework applies a divide-and-conquer strategy to decompose the original task into two subtasks, that is, the localization of road signs and the detection of text on the signs. The algorithms for the two subtasks are naturally incorporated into a unified framework through a feature-based tracking algorithm. 2) The framework provides a novel way to detect text from video by integrating two-dimensional (2-D) image features in each video frame (e.g., color, edges, texture) with the three-dimensional (3-D) geometric structure information of objects extracted from video sequence (such as the vertical plane property of road signs). The feasibility of the proposed framework has been evaluated using 22 video sequences captured from a moving vehicle. This new framework gives an overall text detection rate of 88.9% and a false hit rate of 9.2%. It can easily be applied to other tasks of text detection from video and potentially be embedded in a driver assistance system.
Xilin Chen 0001, Jie Yang 0001
IEEE Trans. Intell. Transp. Syst.3
2005 Predicting human interruptibility with sensors
abstract
A person seeking another person's attention is normally able to quickly assess how interruptible the other person currently is. Such assessments allow behavior that we consider natural, socially appropriate, or simply polite. This is in sharp contrast to current computer and communication systems, which are largely unaware of the social situations surrounding their usage and the impact that their actions have on these situations. If systems could model human interruptibility, they could use this information to negotiate interruptions at appropriate times, thus improving human computer interaction.This article presents a series of studies that quantitatively demonstrate that simple sensors can support the construction of models that estimate human interruptibility as well as people do. These models can be constructed without using complex sensors, such as vision-based techniques, and therefore their use in everyday office environments is both practical and affordable. Although currently based on a demographically limited sample, our results indicate a substantial opportunity for future research to validate these results over larger groups of office workers. Our results also motivate the development of systems that use these models to negotiate interruptions at socially appropriate times.
James Fogarty, Scott E. Hudson, Christopher G. Atkeson, Daniel Avrahami, Jodi Forlizzi, Sara B. Kiesler, Johnny C. Lee, Jie Yang 0001
ACM Trans. Comput. Hum. Interact.8
2004 A Discriminative Learning Framework with Pairwise Constraints for Video Object Classification
Jian Zhang 0003, Jie Yang 0001, Alex Hauptmann 0001
CVPR (2)3
2004 Multimodal detection of human interaction events in a nursing home environment
abstract
In this paper, we propose a multimodal system for detecting human activity and interaction patterns in a nursing home. Activities of groups of people are firstly treated as interaction patterns between any pair of partners and are then further broken into individual activities and behavior events using a multi-level context hierarchy graph. The graph is implemented using a dynamic Bayesian network to statistically model the multi-level concepts. We have developed a coarse-to-fine prototype system to illustrate the proposed concept. Experimental results have demonstrated the feasibility of the proposed approaches. The objective of this research is to automatically create concise and comprehensive reports of activities and behaviors of patients to support physicians and caregivers in a nursing facility.
Datong Chen, Robert G. Malkin, Jie Yang 0001
ICMI3
2004 Incremental detection of text on road signs from video with application to a driving assistant system
abstract
This paper proposes a fast and robust framework for incrementally detecting text on road signs from natural scene video. The new framework makes two main contributions. First, the framework applies a Divide-and-Conquer strategy to decompose the original task into two sub-tasks, that is, localization of road signs and detection of text. The algorithms for the two sub-tasks are smoothly incorporated into a unified framework through a real time tracking algorithm. Second, the framework provides a novel way for text detection from video by integrating 2D features in each video frame (e.g., color, edges, texture) with 3D information available in a video sequence (e.g., object structure). The feasibility of the proposed framework has been evaluated on the video sequences captured from a moving vehicle. The new framework can be applied to a driving assistant system and other tasks of text detection from video.
Xilin Chen 0001, Jie Yang 0001
ACM Multimedia3
2004 Gestures Over Video Streams to Support Remote Collaboration on Physical Tasks
abstract
This article considers tools to support remote gesture in video systems being used to complete collaborative physical tasks-tasks in which two or more individuals work together manipulating three-dimensional objects in the real world. We first discuss the process of conversational grounding during collaborative physical tasks, particularly the role of two types of gestures in the grounding process: pointing gestures, which are used to refer to task objects and locations, and representational gestures, which are used to represent the form of task objects and the nature of actions to be used with those objects. We then consider ways in which both pointing and representational gestures can be instantiated in systems for remote collaboration on physical tasks. We present the results of two studies that use a "surrogate" approach to remote gesture, in which images are intended to express the meaning of gestures through visible embodiments, rather than direct views of the hands. In Study 1, we compare performance with a cursor-based pointing device that allows remote partners to point to objects in a video feed of the work area to performance side-by-side or with the video system alone. In Study 2, we compare performance with two variations of a pen-based drawing tool that allows for both pointing and representational gestures to performance with video alone. The results suggest that simple surrogate gesture tools can be used to convey gestures from remote sites, but that the tools need to be able to convey representational as well as pointing gestures to be effective. The results further suggest that an automatic erasure function, in which drawings disappear a few seconds after they were created, is more beneficial for collaboration than tools requiring manual erasure. We conclude with a discussion of the theoretical and practical implications of the results, as well as several areas for future research.
Susan R. Fussell, Leslie D. Setlock, Jie Yang 0001, Jiazhi Ou, Elizabeth Mauer, Adam D. I. Kramer
Hum. Comput. Interact.3
2004 Automatic detection and recognition of signs from natural scenes
abstract
In this paper, we present an approach to automatic detection and recognition of signs from natural scenes, and its application to a sign translation task. The proposed approach embeds multiresolution and multiscale edge detection, adaptive searching, color analysis, and affine rectification in a hierarchical framework for sign detection, with different emphases at each phase to handle the text in different sizes, orientations, color distributions and backgrounds. We use affine rectification to recover deformation of the text regions caused by an inappropriate camera view angle. The procedure can significantly improve text detection rate and optical character recognition (OCR) accuracy. Instead of using binary information for OCR, we extract features from an intensity image directly. We propose a local intensity normalization method to effectively handle lighting variations, followed by a Gabor transform to obtain local features, and finally a linear discriminant analysis (LDA) method for feature selection. We have applied the approach in developing a Chinese sign translation system, which can automatically detect and recognize Chinese signs as input from a camera, and translate the recognized text into English.
Xilin Chen 0001, Jie Yang 0001, Jing Zhang 0011, Alex Waibel
IEEE Trans. Image Process.2
2003 Predicting human interruptibility with sensors: a Wizard of Oz feasibility study
abstract
A person seeking someone else's attention is normally able to quickly assess how interruptible they are. This assessment allows for behavior we perceive as natural, socially appropriate, or simply polite. On the other hand, today's computer systems are almost entirely oblivious to the human world they operate in, and typically have no way to take into account the interruptibility of the user. This paper presents a Wizard of Oz study exploring whether, and how, robust sensor-based predictions of interruptibility might be constructed, which sensors might be most useful to such predictions, and how simple such sensors might be.The study simulates a range of possible sensors through human coding of audio and video recordings. Experience sampling is used to simultaneously collect randomly distributed self-reports of interruptibility. Based on these simulated sensors, we construct statistical models predicting human interruptibility and compare their predictions with the collected self-report data. The results of these models, although covering a demographically limited sample, are very promising, with the overall accuracy of several models reaching about 78%. Additionally, a model tuned to avoiding unwanted interruptions does so for 90% of its predictions, while retaining 75% overall accuracy.
Scott E. Hudson, James Fogarty, Christopher G. Atkeson, Daniel Avrahami, Jodi Forlizzi, Sara B. Kiesler, Johnny C. Lee, Jie Yang 0001
CHI8
2003 SMaRT: the Smart Meeting Room Task at ISL
abstract
As computational and communications systems become increasingly smaller, faster, more powerful, and more integrated, the goal of interactive, integrated meeting support rooms is slowly becoming reality. It is already possible, for instance, to rapidly locate task-related information during a meeting, filter it, and share it with remote users. Unfortunately, the technologies that provide such capabilities are as obstructive as they are useful - they force humans to focus on the tool rather than the task. Thus the veneer of utility often hides the true costs of use, which are longer, less focused human interactions. To address this issue, we present our current research efforts towards SMaRT: the Smart Meeting Room Task. The goal of SMaRT is to provide meeting support services that do not require explicit human-computer interaction. Instead, by monitoring the activities in the meeting room using both video and audio analysis, the room is able to react appropriately to users' needs and allow the users to focus on their own goals.
Alex Waibel, Tanja Schultz, Michael Bett, Matthias Denecke, Robert G. Malkin, Ivica Rogina, Rainer Stiefelhagen, Jie Yang 0001
ICASSP (4)8
2003 Calibration of a Hybrid Camera Network
abstract
Visual surveillance using a camera network has imposed new challenges to camera calibration. An essential problem is that a large number of cameras may not have a common field of view or even be synchronized well. We propose to use a hybrid camera network that consists of catadioptric and perspective cameras for a visual surveillance task. The relations between multiple views of a scene captured from different cameras can be then calibrated under the catadioptric camera's coordinate system. We address the important issue of how to calibrate the hybrid camera network. We calibrate the hybrid camera network in three steps. First, we calibrate the catadioptric camera using only the vanishing points. In order to reduce computational complexity, we calibrate the camera without the mirror first and then calibrate the catadioptric camera system. Second, we determine 3D positions of some points using as few as two spatial parallel lines and some equidistance points. Finally, we calibrate other perspective cameras based on these known spatial points.
Xilin Chen 0001, Jie Yang 0001, Alex Waibel
ICCV2
2003 Automatically Labeling Video Data Using Multi-class Active Learning
abstract
Labeling video data is an essential prerequisite for many vision applications that depend on training data, such as visual information retrieval, object recognition, and human activity modelling. However, manually creating labels is not only time-consuming but also subject to human errors, and eventually, becomes impossible for a very large amount of data (e.g. 24/7 surveillance video). To minimize the human effort in labeling, we propose a unified multiclass active learning approach for automatically labeling video data. We include extending active learning from binary classes to multiple classes and evaluating several practical sample selection strategies. The experimental results show that the proposed approach works effectively even with a significantly reduced amount of labeled data. The best sample selection strategy can achieve more than a 50% error reduction over random sample selection.
Jie Yang 0001, Alex Hauptmann 0001
ICCV2
2003 Gestural communication over video stream: supporting multimodal interaction for remote collaborative physical tasks
abstract
We present a system integrating gesture and live video to support collaboration on physical tasks. The architecture combines network IP cameras, desktop PCs, and tablet PCs to allow a remote helper to draw on a video feed of a workspace as he/she provides task instructions. A gesture recognition component enables the system both to normalize freehand drawings to facilitate communication with remote partners and to use pen-based input as a camera control device. Results of a preliminary user study suggest that our gesture over video communication system enhances task performance over traditional video-only systems. Implications for the design of multimodal systems to support collaborative physical tasks are also discussed.
Jiazhi Ou, Susan R. Fussell, Xilin Chen 0001, Leslie D. Setlock, Jie Yang 0001
ICMI5
2003 DOVE: drawing over video environment
abstract
We demonstrate a multimedia system that integrates pen-based gesture and live video to support collaboration on physical tasks. The system combines network IP cameras, desktop PCs, and tablet PCs (or PDAs) to allow a remote helper to draw on a video feed of a workspace as he/she provides task instructions. A gesture recognition component enables the system both to normalize freehand drawings to facilitate communication with remote partners and to use pen-based input as a camera control device. The system also embeds some tools, such as controlled video delay, gesture delay, and remote camera pan-tilt-zoom control. The system provides a software environment for studying multimodal/multimedia communication for remote collaborative physical tasks.
Jiazhi Ou, Xilin Chen 0001, Susan R. Fussell, Jie Yang 0001
ACM Multimedia4
2002 Towards non-cooperative iris recognition systems
abstract
Iris Technology has been successfully applied to person verification and identification. However, all commercial products require user cooperation for iris image capture. This paper examines the new challenges of iris recognition when extended to less cooperative situations. With the current stress on security and surveillance, this has been an important consideration. First, a summary of research findings of the past decade on iris recognition is described Then we identified new challenges that will be encountered when extending these methods to less cooperative situations. The difficulties are great and this paper describes some initial work into this area. One difficulty studied is the loss of iris details captured. We propose a modified Kolmogorov complexity measure based on maximum Shannon entropy of wavelet packet reconstruction to quantify the iris information. Real-time eye-corner tracking, iris segmentation and feature extraction algorithms are implemented. Video images of the iris are captured by an ordinary CCD camera with a zoom lens. Experiments are performed and the performances and analysis of iris code method and correlation method are described. Several useful findings were reached albeit from a small database. The iris codes were found to contain almost all the discriminating information. Our correlation approach coupled with nearest neighbour classification outperforms the conventional thresholding method for iris recognition with degraded images.
Eric Sung, Xilin Chen 0001, Jie Yang 0001
ICARCV4
2002 Automatic detection and translation of text from natural scenes
abstract
Large amounts of information are embedded in natural scenes. Signs are good examples of natural objects with high information content. In this paper, we discuss problems in automatic detection and translation of text from natural scenes. We describe the chal1enges of automatic text detection and propose methods to address these chal1enges. We extend example based machine translation technology for sign translation and present a prototype system for Chinese sign translation. This system is capable of capturing images, automatically detecting and recognizing text, and translating the text into English. The translation can be displayed on a palm size PDA, or synthesized as a voice output message over the earphones.
Jie Yang 0001, Xilin Chen 0001, Jing Zhang 0011, Ying Zhang 0048, Alex Waibel
ICASSP1
2002 Towards Monitoring Human Activities Using an Omnidirectional Camera
abstract
We propose an approach for monitoring human activities in an indoor environment using an omnidirectional camera. Robustly tracking people is prerequisite for modeling and recognizing human activities. An omnidirectional camera mounted on the ceiling is less prone to problems of occlusion. We use the Markov Random Field (MRF) to present both background and foreground, and adapt models effectively against environment changes. We employ a deformable model to adapt the foreground models to optimally match objects in different position within a pattern of view of the omnidirectional camera. In order to monitor human activity, we represent positions of people as spatial points and analyze moving trajectories within a time-spatial window. The method provides an efficient way to monitoring high-level human activities without exploring identities.
Xilin Chen 0001, Jie Yang 0001
ICMI2
2002 Flexi-Modal and Multi-Machine User Interfaces
abstract
We describe our system which facilitates collaboration using multiple modalities, including speech, handwriting, gestures, gaze tracking, direct manipulation, large projected touch-sensitive displays, laser pointer tracking, regular monitors with a mouse and keyboard, and wireless networked handhelds. Our system allows multiple, geographically dispersed participants to simultaneously and flexibly mix different modalities using the right interface at the right time on one or more machines. We discuss each of the modalities provided, how they were integrated in the system architecture, and how the user interface enabled one or more people to flexibly use one or more devices.
Brad A. Myers, Robert G. Malkin, Michael Bett, Alex Waibel, Ben Bostwick, Rob Miller 0001, Jie Yang 0001, Matthias Denecke, Edgar Seemann, Choon Hong Peck, Dave Kong, Jeffrey Nichols 0001, William L. Scherlis
ICMI7
2002 A PDA-Based Sign Translator
abstract
We propose an effective approach for a PDA-based sign system and present the sign translator. Its main functions include three parts: detection, recognition and translation. Automatic detection and recognition of text in natural scenes is a prerequisite for the automatic sign translator. In order to make the system robust for text detection in various natural scenes, the detection approach efficiently embeds multi-resolution, adaptive search in a hierarchical framework with different emphases at each layer. We also introduce an intensity-based OCR method to recognize characters in various fonts and lighting conditions, where we employ the Gabor transform to obtain local features, and LDA for selection and classification of features. The recognition rate is 92.4% for the testing set obtained from the natural sign. A sign is different from the normal used sentence. It is brief with a lot of abbreviations and place nouns. We only briefly introduce a rule-based place name translation. We have integrated all these functions in a PDA, which can capture sign images, auto segment and recognize the Chinese sign, and translate it into English.
Jing Zhang 0011, Xilin Chen 0001, Jie Yang 0001, Alex Waibel
ICMI3
2002 Automatic sign translation
abstract
Large amounts of information is embedded in the natural scenes. Signs are good examples of objects in natural environments which have rich information content. In this paper, we present our efforts in the automatic sign translation. We describe the challenges in the automatic sign translation and introduce the architecture of our current system for automatic detection and translation of Chinese signs. Two data-driven machine translation methods: Example Based Machine Translation (EBMT) and Statistical Machine Translation (SMT) are compared for the task of translating Chinese signs into English. We report the experimental results of both methods that are trained from a small bilingual sign corpus combined with a bilingual glossary. The experiment results indicate that EBMT generates more correct translations while SMT is better at inferring unseen patterns. We are currently working on developing a multi-engine machine translation system that can incrementally learn from the data and combine the results from EBMT and SMT.
Ying Zhang 0048, Bing Zhao 0005, Jie Yang 0001, Alex Waibel
INTERSPEECH3
2002 Automatic Detection of Signs with Affine Transformation
abstract
In this paper, we propose an approach for detecting signs from natural scenes. The approach efficiently embeds multiresolution, adaptive search, and affine rectification algorithms in a hierarchical framework, with different emphases at each layer. We combine in multi-resolution and multi-scale edge detection techniques to effectively detect text in different sizes. By using the cites from text inside the image, we introduce affine rectification transformation to recover deformation of the text region caused by air inappropriate camera view angle. This procedure can significantly improve text detection rate and OCR (Optical Character Recognition) accuracy. Experimental results have demonstrated feasibility of the proposed algorithms. We have applied the proposed approach to a Chinese sign translation system, which can automatically detect Chinese text input from a camera, recognize the text, and translate the recognized text into English or voice stream.
Xilin Chen 0001, Jie Yang 0001, Jing Zhang 0011, Alex Waibel
WACV2
2002 A PDA-based Face Recognition System
abstract
In this paper, we present a PDA-based face recognition system as well as some of the associated challenges of developing a PDA-based face recognition system. We describe a prototype system built from an off the shelf PDA, and introduce algorithms for image preprocessing to enhance the quality of the image by sharpening focus, and normalizing both lighting condition and head rotation. We use a unified LDA/PCA algorithm for face recognition. The algorithm maximizes the LDA criterion directly without a separate PCA step, which eliminates the possibility of losing discriminative information due to a separate PCA step. We demonstrate effectiveness of these algorithms and the feasibility of this system by experimental results. This system has many applications including information retrieval and law enforcement.
Jie Yang 0001, Xilin Chen 0001, William Kunz
WACV1
2002 Modeling focus of attention for meeting indexing based on multiple cues
abstract
A user's focus of attention plays an important role in human-computer interaction applications, such as a ubiquitous computing environment and intelligent space, where the user's goal and intent have to be continuously monitored. We are interested in modeling people's focus of attention in a meeting situation. We propose to model participants' focus of attention from multiple cues. We have developed a system to estimate participants' focus of attention from gaze directions and sound sources. We employ an omnidirectional camera to simultaneously track participants' faces around a meeting table and use neural networks to estimate their head poses. In addition, we use microphones to detect who is speaking. The system predicts participants' focus of attention from acoustic and visual information separately. The system then combines the output of the audio- and video-based focus of attention predictors. We have evaluated the system using the data from three recorded meetings. The acoustic information has provided 8% relative error reduction on average compared to only using one modality. The focus of attention model can be used as an index for a multimedia meeting record. It can also be used for analyzing a meeting.
Rainer Stiefelhagen, Jie Yang 0001, Alex Waibel
IEEE Trans. Neural Networks2
2001 An Adaptive Algorithm for Text Detection from Natural Scenes
abstract
We present a new adaptive algorithm for automatic detection of text from a natural scene. The initial cues of text regions are first detected from the captured image/video. An adaptive color modeling and searching algorithm is then utilized near the initial text cues, to discriminate text/non-text regions. EM optimization algorithm is used for color modeling, under the constraint of text layout relations for a specific language. The proposed algorithm combines the advantages of several previous approaches for text detection, and utilizes a focus-of-attention approach for text finding. The whole algorithm is applied in a prototype system that can automatically detect and recognize sign input from a video camera, and translate the signs into English text or voice streams. We present evaluation results of our algorithm on this system.
Jiang Gao, Jie Yang 0001
CVPR (2)2
2001 A direct LDA algorithm for high-dimensional data - with application to face recognition
Hua Yu 0008, Jie Yang 0001
Pattern Recognit.2
2000 Face Recognition in a Meeting Room
abstract
We investigate the recognition of human faces in a meeting room. The major challenges of identifying human faces in this environment include low quality of input images, poor illumination, unrestricted head poses and continuously changing facial expressions and occlusion. In order to address these problems we propose a novel algorithm, dynamic space warping (DSW). The basic idea of the algorithm is to combine local features under certain spatial constraints. We compare DSW with the eigenface approach on data collected from various meetings. We have tested both front and profile face images and images with two stages of occlusion. The experimental results indicate that the DSW approach outperforms the eigenface approach in both cases.
Ralph Gross, Jie Yang 0001, Alex Waibel
FG2
2000 Segmenting Hands of Arbitrary Color
abstract
Hand segmentation is a prerequisite for many gesture recognition tasks. Color has been widely used for hand segmentation. However, many approaches rely on predefined skin color models. It is very difficult to predefine a color model in a mobile application where the light condition may change dramatically over time. We propose a novel statistical approach to hand segmentation based on Bayes decision theory. The proposed method requires no predefined skin color model. Instead it generates a hand color model and a background color model for a given image, and uses these models to classify each pixel in the image as either a hand pixel or a background pixel. Models are generated using a Gaussian mixture model with the restricted EM algorithm. Our method is capable of segmenting hands of arbitrary color in a complex scene. It performs well even when there is a significant overlap between hand and background colors, or when the user wears gloves. We show that the Bayes decision method is superior to a commonly used method by comparing their upper bound performance. Experimental results demonstrate the feasibility of the proposed method.
Xiaojin Zhu 0001, Jie Yang 0001, Alex Waibel
FG2
2000 Partial Information in Multimodal Dialogue
Matthias Denecke, Jie Yang 0001
ICMI2
2000 Growing Gaussian Mixture Models for Pose Invariant Face Recognition
abstract
A major challenge for face recognition algorithms lies in the variance faces undergo while changing pose. This problem is typically addressed by building view dependent models based on face images taken from predefined head poses. However, it is impossible to determine all head poses beforehand in an unrestricted setting such as a meeting room, where people can move and interact freely. We present an approach to pose invariant face recognition. We employ Gaussian mixture models to characterize human faces and model pose variance with different numbers of mixture components. The optimal number of mixture components for each person is automatically learned from training data by growing the mixture models. The proposed algorithm is tested on real data recorded in a meeting room. The experimental results indicate that the new method outperforms standard eigenface and Gaussian mixture model approaches. Our algorithm achieved as much as 42% error reduction compared to the standard eigenface approach on the same test data.
Ralph Gross, Jie Yang 0001, Alex Waibel
ICPR2
2000 Simultaneous Tracking of Head Poses in a Panoramic View
abstract
In this paper we present an approach to simultaneously estimate gaze directions of multiple people in the view of a panoramic camera. Human faces are located and tracked using a probabilistic skin-color model and motion detection. Neural networks are used to estimate head poses of the detected faces. With this approach, it is possible to simultaneously track the locations of multiple people around a meeting table and estimate their gaze directions using only a panoramic camera. We have achieved an accuracy of 9 degrees for head pan estimation and 6 degrees for tilt estimation for a multi-user system.
Rainer Stiefelhagen, Jie Yang 0001, Alex Waibel
ICPR2
2000 Dialogue management for multimodal user registration
abstract
... information with a user. It is widely used in hospitals, hotels and conferences. In this paper, we propose an approach to interactive user registration by combining face recognition, speech recognition and speech synthesis technologies together through an efficient dialogue manager. In order to minimize a user's effort, we employ a new dialogue management model based on a finite state automaton (FSA), which uses a Baysian network to fuse the user's information from multiple channels (e.g., face image, speech, records stored in a pre-constructed database) to reliably estimate the confidence about user identity. Instead of fixing weights, the FSA adjusts its weights dynamically by integrating partial information from multiple information sources. This is achieved by maximizing an objective function to determine an optimal action at each succeeding state according to current confidence and information cues. Thus the transition between states can be done along the shortest path from the initial state to the goal state. We have developed a multimodal user registration system to demonstrate the feasibility of the proposed approach.
Fei Huang 0002, Jie Yang 0001, Alex Waibel
INTERSPEECH2
2000 Towards Unrestricted Lip Reading
abstract
Lip reading provides useful information in speech perception and language understanding, especially when the auditory speech is degraded. However, many current automatic lip reading systems impose some restrictions on users. In this paper, we present our research efforts in the Interactive System Laboratory, towards unrestricted lip reading. We first introduce a top–down approach to automatically track and extract lip regions. This technique makes it possible to acquire visual information in real-time without limiting the user's freedom of movement. We then discuss normalization algorithms to preprocess images for different lightning conditions (global illumination and side illumination). We also compare different visual preprocessing methods such as raw image, Linear Discriminant Analysis (LDA), and Principle Component Analysis (PCA). We demonstrate the feasibility of the proposed methods by the development of a modular system for flexible human–computer interaction via both visual and acoustic speech. The system is based on an extension of the existing state-of-the-art speech recognition system, a modular Multiple State–Time Delayed Neural Network (MS–TDNN) system. We have developed adaptive combination methods at several different levels of the recognition network. The system can automatically track a speaker and extract his/her lip region in real-time. The system has been evaluated under different noisy conditions such as white noise, music, and mechanical noise. The experimental results indicate that the system can achieve up to 55% error reduction using visual information in addition to the acoustic signal.
Uwe Meier, Rainer Stiefelhagen, Jie Yang 0001, Alex Waibel
Int. J. Pattern Recognit. Artif. Intell.3
1999 Modeling focus of attention for meeting indexing
abstract
Visual cues, such as gesturing, looking at each other or monitoring each others facial expressions, play an important role in meetings.Such information can be used for indexing of multimedia meeting recordings.In this paper, we present an approach to detect who is looking at whom during a meeting.Our proposal is to employ Hidden Markov Models to characterize participants' focus of attention by using gaze information as well as knowledge about the number and positions of people present in a meeting.The number and positions of the participants faces are detected in the field of view of a panoramic camera.We use neural networks to estimate the directions of participants' gaze from camera images.We discuss the implementation of the approach in detail including system architecture, data collection, and evaluation.The system has achieved an accuracy rate of up to 93 % in detecting focus of attention on test sequences taken from meetings.We have used focus of attention as an index in a multimedia meeting browser.
Rainer Stiefelhagen, Jie Yang 0001, Alex Waibel
ACM Multimedia (1)2
1999 Multimodal people ID for a multimedia meeting browser
abstract
A meeting browser is a system that allows users to review a multimedia meeting record from a variety of indexing methods. Identification of meeting participants is essential for creating such a multimedia meeting record. Moreover, knowing who is speaking can enhance the performance of speech recognition and indexing meeting transcription. In this paper, we present an approach that identifies meeting participants by fusing multimodal inputs. We use face ID, speaker ID, color appearance ID, and sound source directional ID to identify and track meeting. After describing the different modules in detail, we will discuss a framework for combining the information sources. Integration of the multimodal people ID into the multimedia meeting browser is in its preliminary stage.
Jie Yang 0001, Xiaojin Zhu 0001, Ralph Gross, John Kominek, Alex Waibel
ACM Multimedia (1)1
1998 Skin-Color Modeling and Adaptation
Jie Yang 0001, Weier Lu, Alex Waibel
ACCV (2)1
1998 Visual Tracking for Multimodal Human Computer Interaction
abstract
In this paper, we present visual tracking techniques for multimodal human computer interaction.First, we discuss techniques for tracking human faces in which human skin-color is used as a major feature.An adaptive stochastic model has been developed to characterize the skin-color distributions.Based on the maximum likelihood method, the model parameters can be adapted for different people and different lighting conditions.The feasibility of the model has been demonstrated by the development of a real-time face tracker.The system has achieved a rate of 30-t-frames/second using a low-end workstation with a framegrabber and a camera.We also present a top-down approach for tracking facial features such as eyes, nostrils, and lip comers.These real-time visual tracking techniques have been successfully applied to many applications such as gaze tracking, and lipreading.The face tracker has been combined with a microphone array for extracting speech signal from a specific person.The gaze tracker has been combined with a speech recognizer in a multimodal interface for controlling a panoramic image viewer.
Jie Yang 0001, Rainer Stiefelhagen, Uwe Meier, Alex Waibel
CHI1
1997 Gaze tracking for multimodal human-computer interaction
abstract
This paper discusses the problem of gaze tracking and its applications to multimodal human-computer interaction. The function of a gaze tracking system can be either passive or active. For example, a system can identify user's message target by monitoring the user's gaze, or the user could use his gaze to directly control an application or launch actions. We have developed a real-time gaze tracking system that estimates the 3D position and rotation (pose) of a user's head. We demonstrate the applications of the gaze tracker to human-computer interaction by two examples. The first example shows that gaze tracker can help speech recognition systems by switching language model and grammar based on user's gaze information. The second example illustrates the combination of the gaze tracker and a speech recognizer to view a panorama image.
Rainer Stiefelhagen, Jie Yang 0001
ICASSP2
1997 Multimodal interfaces for multimedia information agents
abstract
When humans communicate they take advantage of a rich spectrum of cues. Some are verbal and acoustic. Some are non-verbal and non-acoustic. Signal processing technology has devoted much attention to the recognition of speech, as a single human communication signal. Most other complementary communication cues, however, remain unexplored and unused in human-computer interaction. In this paper we show that the addition of non-acoustic or non-verbal cues can significantly enhance robustness, flexibility, naturalness and performance of human-computer interaction. We demonstrate computer agents that use speech, gesture, handwriting, pointing, spelling jointly for more robust, natural and flexible human-computer interaction in the various tasks of an information worker: information creation, access, manipulation or dissemination.
Alex Waibel, Bernhard Suhm, Minh Tue Vo, Jie Yang 0001
ICASSP4
1997 Real-time lip-tracking for lipreading
abstract
This paper presents a new approach to lip tracking for lipreading.Instead of only tracking features on lips, we propose to track lips along with other facial features such as pupils and nostril.In the new approach, the face is rst located in an image using a stochastic skin-color model, the eyes, lip-corners and nostrils are then located and tracked inside the facial region.The new approach can eectively improve the robustness of lip-tracking and simplify automatic detection and recovery of tracking failure.The feasibility of the proposed approach has been demonstrated by implementation of a lip tracking system.The system has been tested by a database that contains 900 image sequences of dierent speakers spelling words.The system has successfully extract lip regions from the image sequences to obtain training data for the audio-visual speech recognition system.The system has been also applied to extract the lip region in real-time from live video images to obtain the visual input for an audio-visual speech recognition system.On test sequences we have achieved a reduction of the number of frames with tracking failures by a factor of two using detection and prediction of outliers in the set of found features.
Rainer Stiefelhagen, Uwe Meier, Jie Yang 0001
EUROSPEECH3
1997 Human action learning via hidden Markov model
abstract
To successfully interact with and learn from humans in cooperative modes, robots need a mechanism for recognizing, characterizing, and emulating human skills. In particular, it is our interest to develop the mechanism for recognizing and emulating simple human actions, i.e., a simple activity in a manual operation where no sensory feedback is available. To this end, we have developed a method to model such actions using a hidden Markov model (HMM) representation. We proposed an approach to address two critical problems in action modeling: classifying human action-intent, and learning human skill, for which we elaborated on the method, procedure, and implementation issues in this paper. This work provides a framework for modeling and learning human actions from observations. The approach can be applied to intelligent recognition of manual actions and high-level programming of control input within a supervisory control paradigm, as well as automatic transfer of human skills to robotic systems.
Jie Yang 0001, Yangsheng Xu
IEEE Trans. Syst. Man Cybern. Part A1
1996 Focus of attention: Towards low bitrate video tele-conferencing
abstract
Low bitrate video tele-conferencing requires adapting algorithms that may work perfectly well in a high-bitrate situation. When a slow transmission rate is unacceptable, compromise must be reached among the demands of speed, bandwidth limits and image quality. In this paper we present an approach to low bitrate video tele-conferencing by focusing attention on important information. We show that by selectively degrading the quality of less important regions, more important regions can be sent without loss of quality but with greatly reduced bandwidth requirements. A prototype system has been developed to demonstrate the concept. The experimental results show significant savings of required bandwidth for video subjected to the changes.
Jie Yang 0001, Leejay Wu, Alex Waibel
ICIP (2)1
1996 A real-time face tracker
abstract
The authors present a real-time face tracker. The system has achieved a rate of 30+ frames/second using an HP-9000 workstation with a frame grabber and a Canon VC-Cl camera. It can track a person's face while the person moves freely (e.g., walks, jumps, sits down and stands up) in a room. Three types of models have been employed in developing the system. First, they present a stochastic model to characterize skin color distributions of human faces. The information provided by the model is sufficient for tracking a human face in various poses and views. This model is adaptable to different people and different lighting conditions in real-time. Second, a motion model is used to estimate image motion and to predict the search window. Third, a camera model is used to predict and compensate for camera motion. The system can be applied to teleconferencing and many HCI applications including lip reading and gaze tracking. The principle in developing this system can be extended to other tracking problems such as tracking the human hand.
Jie Yang 0001, Alex Waibel
WACV1
1995 Towards Human-Robot Coordination: Skill Modeling and Transferring via Hidden Markov Model
abstract
Automatic modeling and transferring human skill to a robot is an important step towards creating an intelligent robot in a cooperative environment where humans and robots can complement and enhance each other's performance in a reciprocal manner. We first classify two distinct categories of skills, i.e., action skill and reaction skill. Then we address how the hidden Markov model can be used for modeling these two types of human skill. We focus on the reaction learning scheme and discuss the issues and problems associated with the hidden Markov model approach.
Sangheng Xu, Jie Yang 0001
ICRA2
1994 Gesture Interface: Modeling and Learning
abstract
This paper presents a method for developing a gesture-based system using a multidimensional hidden Markov model (HMM). Instead of using geometric features, gestures are converted into sequential symbols. HMMs are employed to represent the gestures and their parameters are learned from the training data. Based on "the most likely performance" criterion, the gestures can be recognized by evaluating the trained HMMs. We have developed a prototype to demonstrate the feasibility of the proposed method. The system achieved 99.78% accuracy for a 9 gesture isolated recognition task. Encouraging results were also obtained from experiments of continuous gesture recognition. The proposed method is applicable to any multidimensional signal representation gesture, and will be a valuable tool in telerobotics and human computer interfacing.>
Jie Yang 0001, Yangsheng Xu
ICRA1
1994 Hidden Markov model approach to skill learning and its application to telerobotics
abstract
In this paper, we discuss the problem of how human skill can be represented as a parametric model using a hidden Markov model (HMM), and how an HMM-based skill model can be used to learn human skill. HMM is feasible to characterize a doubly stochastic process--measurable action and immeasurable mental states--that is involved in the skill learning. We formulated the learning problem as a multidimensional HMM and developed a testbed for a variety of skill learning applications. Based on "the most likely performance" criterion, the best action sequence can be selected from all previously measured action data by modeling the skill as an HMM. The proposed method has been implemented in the teleoperation control of space station robot system, and some important implementation issues have been discussed. The method allows a robot to learn human skill in certain tasks and to improve motion performance.
Jie Yang 0001, Yangsheng Xu
IEEE Trans. Robotics Autom.1