Sy Bor Wang

dblp:55/4859 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
0since 2021 · last 2007
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Probabilistic and Bayesian machine learning · 84% Video understanding and tracking · 16%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › structured prediction
conditional random field
0.112007
Hidden Conditional Random Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2007
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.112007
Hidden Conditional Random Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2007
Machine learning › Probabilistic and Bayesian machine learning
structured prediction
0.112007
Hidden Conditional Random Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2007
Computer vision › Video understanding and tracking
gesture recognition
0.112006
Hidden Conditional Random Fields for Gesture Recognition · CVPR (2) 2006
Machine learning › Probabilistic and Bayesian machine learning › structured prediction
hidden conditional random field
0.112006
Hidden Conditional Random Fields for Gesture Recognition · CVPR (2) 2006
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
discriminative latent variable model
0.012007
Hidden Conditional Random Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2007

Methods — techniques the papers use, named apart from their topics

latent variable model · 0.1conditional random field · 0.1hidden markov model · 0.1discriminative sequence modeling · 0.1
YearPublicationVenuePosition
2007 Detecting communication errors from visual cues during the system's conversational turn
abstract
Automatic detection of communication errors in conversational systems has been explored extensively in the speech community. However, most previous studies have used only acoustic cues. Visual information has also been used by the speech community to improve speech recognition in dialogue systems, but this visual information is only used when the speaker is communicating vocally. A recent perceptual study indicated that human observers can detect communication problems when they see the visual footage of the speaker during the system's reply. In this paper, we present work in progress towards the development of a communication error detector that exploits this visual cue. In datasets we collected or acquired, facial motion features and head poses were estimated while users were listening to the system response and passed to a classifier for detecting a communication error. Preliminary experiments have demonstrated that the speaker's visual information during the system's reply is potentially useful and accuracy of automatic detection is close to human performance.
Sy Bor Wang, David Demirdjian, Trevor Darrell
ICMI1
2007 Hidden Conditional Random Fields
abstract
We present a discriminative latent variable model for classification problems in structured domains where inputs can be represented by a graph of local observations. A hidden-state Conditional Random Field framework learns a set of latent variables conditioned on local features. Observations need not be independent and may overlap in space and time.
Ariadna Quattoni, Sy Bor Wang, Louis-Philippe Morency, Michael Collins 0001, Trevor Darrell
IEEE Trans. Pattern Anal. Mach. Intell.2
2006 Hidden Conditional Random Fields for Gesture Recognition
abstract
We introduce a discriminative hidden-state approach for the recognition of human gestures. Gesture sequences often have a complex underlying structure, and models that can incorporate hidden structures have proven to be advantageous for recognition tasks. Most existing approaches to gesture recognition with hidden states employ a Hidden Markov Model or suitable variant (e.g., a factored or coupled state model) to model gesture streams; a significant limitation of these models is the requirement of conditional independence of observations. In addition, hidden states in a generative model are selected to maximize the likelihood of generating all the examples of a given gesture class, which is not necessarily optimal for discriminating the gesture class against other gestures. Previous discriminative approaches to gesture sequence recognition have shown promising results, but have not incorporated hidden states nor addressed the problem of predicting the label of an entire sequence. In this paper, we derive a discriminative sequence model with a hidden state structure, and demonstrate its utility both in a detection and in a multi-way classification formulation. We evaluate our method on the task of recognizing human arm and head gestures, and compare the performance of our method to both generative hidden state and discriminative fully-observable models.
Sy Bor Wang, Ariadna Quattoni, Louis-Philippe Morency, David Demirdjian, Trevor Darrell
CVPR (2)1
2005 Inferring body pose using speech content
abstract
Untethered multimodal interfaces are more attractive than tethered ones because they are more natural and expressive for interaction. Such interfaces usually require robust vision-based body pose estimation and gesture recognition. In interfaces where a user is interacting with a computer using speech and arm gestures, the user's spoken keywords can be recognized in conjuction with a hypothesis of body poses. This co-occurence can reduce the number of body pose hypothesis for the vision based tracker. In this paper we show that incorporating speech-based body pose constraints can increase the robustness and accuracy of vision-based tracking systems.Next, we describe an approach for gesture recognition. We show how Linear Discriminant Analysis (LDA), can be employed to estimate 'good features' that can be used in a standard HMM-based gesture recognition system. We show that, by applying our LDA scheme, recognition errors can be significantly reduced over a standard HMM-based technique.We applied both techniques in a Virtual Home Desktop scenario. Experiments where the users controlled a desktop system using gestures and speech were conducted and the results show that the speech recognised in conjunction with body poses has increased the accuracy of the vision-based tracking system.
Sy Bor Wang, David Demirdjian
ICMI1