Edward C. Kaiser

dblp:87/279 · DBLP profile ↗
← Back
24ranked-venue papers
11as first author
0since 2021 · last 2013
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 11 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-authorArtificial intelligence and machine learning · 6 · 3 first-authorComputer networks · 3 · 1 first-authorSecurity and privacy · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
3 papers
Network security · 52% Cryptographic protocols and secure computation · 27% Blockchain and cryptocurrency security · 21%
Computer networks
4 papers
Internet of things and sensor networks · 62% Internet architecture and protocols · 28% Wireless networking · 9%
Human-computer interaction and pervasive computing
1 paper
Collaborative and social computing · 100%
Artificial intelligence
1 paper
Information extraction and text analysis · 67% Speech recognition and synthesis · 33%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Distributed systems · 54% Energy-efficient computing · 46%
Computer graphics and multimedia
1 paper
Audio and music processing · 100%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Network security › attack resilience › attack mitigation
denial-of-service defense
0.122007
mod_kaPoW: mitigating DoS with transparent proof-of-work · CoNEXT 2007
Design and implementation of network puzzles · INFOCOM 2005
Cryptographic protocols and secure computation › secret sharing
cheating detection
0.112009
Fides: remote anomaly-based cheat detection using client emulation · CCS 2009
Internet of things and sensor networks › camera sensor networks
wireless video sensor networks
0.122003
Panoptes: scalable low-power video sensor networking technologies · ACM Multimedia 2003
Panoptes: scalable low-power video sensor networking technologies · ACM Multimedia 2003
Collaborative and social computing
computer-mediated communication
0.112007
Multimodal redundancy across handwriting and speech during computer mediated human-human interactions · CHI 2007
Blockchain and cryptocurrency security › consensus protocol
proof-of-work
0.112007
mod_kaPoW: mitigating DoS with transparent proof-of-work · CoNEXT 2007
Network security › attack resilience › attack mitigation › denial-of-service defense
client puzzles
0.112005
Design and implementation of network puzzles · INFOCOM 2005
Distributed systems › distributed system architecture
client-server systems
0.012009
Fides: remote anomaly-based cheat detection using client emulation · CCS 2009
Energy-efficient computing
power management
0.022003
Panoptes: scalable low-power video sensor networking technologies · ACM Multimedia 2003
Panoptes: scalable low-power video sensor networking technologies · ACM Multimedia 2003
Natural language and speech › Information extraction and text analysis › syntactic parsing › grammar-based parsing
finite-state parsing
0.011999
Robust, Finite-State Parsing for Spoken Language Understanding · ACL 1999
Natural language and speech › Speech recognition and synthesis
spoken language understanding
0.011999
Robust, Finite-State Parsing for Spoken Language Understanding · ACL 1999
Natural language and speech › Information extraction and text analysis
syntactic parsing
0.011999
Robust, Finite-State Parsing for Spoken Language Understanding · ACL 1999
Audio and music processing
speech recognition
0.012007
Multimodal redundancy across handwriting and speech during computer mediated human-human interactions · CHI 2007
Wireless networking › WLAN
IEEE 802.11
0.012003
Panoptes: scalable low-power video sensor networking technologies · ACM Multimedia 2003

Methods — techniques the papers use, named apart from their topics

client emulation · 0.2anomaly detection · 0.2transparent proof-of-work · 0.1empirical study · 0.1iptables · 0.1hash-reversal puzzle · 0.1streaming · 0.1prioritization · 0.1event triggering · 0.1application-specific filtering · 0.1bitmapping · 0.0bit mapping · 0.0semantic parsing · 0.0finite state machine · 0.0
YearPublicationVenuePosition
2013 Demonstration of sketch-thru-plan: a multimodal interface for command and control
abstract
This paper demonstrates a multimodal system called Sketch-Thru-Plan (STP) that enables users to speak and draw doctrinal language and symbols in order to create courses of action. We argue that STP can meet many of the challenges inherent in building user interfaces for operations planning and command-and-control. The system is being transitioned to military organizations for use in planning courses of action.
Phil Cohen 0001, M. Cecelia Buchanan, Edward C. Kaiser, Michael J. Corrigan, Scott Lind, Matt Wesson
ICMI3
2009 Fides: remote anomaly-based cheat detection using client emulation
abstract
As a result of physically owning the client machine, cheaters in online games currently have the upper-hand when it comes to avoiding detection. To address this problem and turn the table on cheaters, this paper presents Fides, an anomaly-based cheat detection approach that remotely validates game execution. With Fides, a server-side Controller specifies how and when a client-side Auditor measures the game. To accurately validate measurements, the Controller partially emulates the client and collaborates with the server. This paper examines a range of cheat methods and initial measurements that counter them, showing that a Fides prototype is able to efficiently detect several existing cheats, including one state-of-the-art cheat that is advertised as "undetectable".
Edward C. Kaiser, Wu-chang Feng, Travis Schluessler
CCS1
2007 Multimodal redundancy across handwriting and speech during computer mediated human-human interactions
abstract
Lecturers, presenters and meeting participants often say what they publicly handwrite. In this paper, we report on three empirical explorations of such multimodal redundancy -- during whiteboard presentations, during a spontaneous brainstorming meeting, and during the informal annotation and discussion of photographs. We show that redundantly presented words, compared to other words used during a presentation or meeting, tend to be topic specific and thus are likely to be out-of-vocabulary. We also show that they have significantly higher tf-idf (term frequency-inverse document frequency) weights than other words, which we argue supports the hypothesis that they are dialogue-critical words. We frame the import of these empirical findings by describing SHACER, our recently introduced Speech and HAndwriting reCognizER, which can combine information from instances of redundant handwriting and speech to dynamically learn new vocabulary.
Edward C. Kaiser, Paulo Barthelmess, Candice Erdmann, Phil Cohen 0001
CHI1
2007 mod_kaPoW: mitigating DoS with transparent proof-of-work
abstract
Unwanted traffic remains a fundamental problem for networked systems. Proof-of-work (PoW) is a defense mechanism that adds a client-specific challenge at the start of a networked protocol. The challenge acts as a filter for clients based on their willingness to solve a computational task of varying difficulty. The difficulty is tailored to the individual client and is set proportional to its relative load on the server.
Edward C. Kaiser, Wu-chang Feng
CoNEXT1
2007 Workshop on tagging, mining and retrieval of human related activity information
abstract
Inexpensive and user friendly cameras, microphones, and other devices such as digital pens are making it increasingly easy to capture, store and process large amounts of data over a variety of media. Even though the barriers for data acquisition have been lowered, making use of these data remains challenging. The focus of the present workshop is on issues related to theory, methods and techniques for facilitating the organization, retrieval and reuse of multimodal information. The emphasis is on organization and retrieval of information related to human activity, i.e. that is generated and consumed by individuals and groups as they go about their work, learning and leisure.
Paulo Barthelmess, Edward C. Kaiser
ICMI2
2007 Toward content-aware multimodal tagging of personal photo collections
abstract
A growing number of tools is becoming available, that make use ofexisting tags to help organize and retrieve photos, facilitating the management and use of photo sets. The tagging on which these techniques rely remains a time consuming, labor intensive task that discourages many users. To address this problem, we aim to leverage the multimodal content of naturally occurring photo discussions among friends and families to automatically extract tags from a combination of conversational speech, handwriting, and photo content analysis. While naturally occurring discussions are rich sources of informationabout photos, methods need to be developed to reliably extract a set of discriminative tags from this noisy, unconstrained group discourse. To this end, this paper contributes ananalysis of pilot data identifying robust multimodal features examining the interplay between photo content and other modalities such as speech and handwriting. Our analysis is motivated by a search for design implications leading to the effective incorporation of automated location and person identification(e.g. based on GPS and facial recognition technologies) into a system able to extract tags from natural multimodal conversations.
Paulo Barthelmess, Edward C. Kaiser, David McGee
ICMI2
2006 Collaborative multimodal photo annotation over digital paper
abstract
The availability of metadata annotations over media content such as photos is known to enhance retrieval and organization, particularly for large data sets. The greatest challenge for obtaining annotations remains getting users to perform the large amount of tedious manual work that is required.In this paper we introduce an approach for semi-automated labeling based on extraction of metadata from naturally occurring conversations of groups of people discussing pictures among themselves.As the burden for structuring and extracting metadata is shifted from users to the system, new recognition challenges arise. We explore how multimodal language can help in 1) detecting a concise set of meaningful labels to be associated with each photo, 2) achieving robust recognition of these key semantic terms, and 3) facilitating label propagation via multimodal shortcuts. Analysis of the data of a preliminary pilot collection suggests that handwritten labels may be highly indicative of the semantics of each photo, as indicated by the correlation of handwritten terms with high frequency spoken ones. We point to initial directions exploring a multimodal fusion technique to recover robust spelling and pronunciation of these high-value terms from redundant speech and handwriting.
Paulo Barthelmess, Edward C. Kaiser, David McGee, Phil Cohen 0001
ICMI2
2006 Collaborative multimodal photo annotation over digital paper
abstract
The availability of metadata annotations over media content such as photos is known to enhance retrieval and organization, particularly for large data sets. The greatest challenge for obtaining annotations remains getting users to perform the large amount of tedious manual work that is required. In this demo we show a system for semi-automated labeling based on extraction of metadata from naturally occurring conversations of groups of people discussing pictures among themselves. The system supports a variety of collaborative label elicitation scenarios mixing co-located and distributed participants, operating primarily via speech, handwriting and sketching over tangible digital paper photo printouts. We demonstrate the real-time capabilities of the system by providing hands-on annotation experience for conference participants. Demo annotations are performed over public domain pictures portraying mainstream themes (e.g. from famous movies).
Paulo Barthelmess, Edward C. Kaiser, David McGee, Phil Cohen 0001
ICMI2
2006 Using redundant speech and handwriting for learning new vocabulary and understanding abbreviations
abstract
New language constantly emerges from complex, collaborative human-human interactions like meetings -- such as, for instance, when a presenter handwrites a new term on a whiteboard while saying it. Fixed vocabulary recognizers fail on such new terms, which often are critical to dialogue understanding. We present a proof-of-concept multimodal system that combines information from handwriting and speech recognition to learn the spelling, pronunciation and semantics of out-of-vocabulary terms from single instances of redundant multimodal presentation (e.g. saying a term while handwriting it). For the task of recognizing the spelling and semantics of abbreviated Gantt chart labels across a held-out test series of five scheduling meetings we show a significant relative error rate reduction of 37% when our learning methods are used and allowed to persist across the meeting series, as opposed to when they are not used.
Edward C. Kaiser
ICMI1
2006 Edge-splitting in a cumulative multimodal system, for a no-wait temporal threshold on information fusion, combined with an under-specified display
abstract
Predicting the end of user input turns in a multimodal system can be complex. User interactions vary across a spectrum from single, unimodal inputs to multimodal combinations delivered either simultaneously or sequentially. Early multimodal systems used a fixed duration temporal threshold to determine how long to wait for the next input before processing and integration. Several recent studies have proposed using dynamic or adaptive temporal thresholds to predict turn segmentation and thus achieve faster system response times. We introduce an approach that requires no temporal threshold. First we contrast current multimodal command interfaces to a new class of cumulative-observant multimodal systems that we introduce. Within that new system class we show how our technique of edge-splitting combined with our strategy for under-specified, no-wait, visual feedback resolves parsing problems that underlie turn segmentation errors. Test results show a 46.2% significant reduction in multimodal recognition errors, compared to not using these techniques.
Edward C. Kaiser, Paulo Barthelmess
INTERSPEECH1
2005 Distributed pointing for multimodal collaboration over sketched diagrams
abstract
A problem faced by groups that are not co-located but need to collaborate on a common task is the reduced access to the rich multimodal communicative context that they would have access to if they were collaborating face-to-face. Collaboration support tools aim to reduce the adverse effects of this restricted access to the fluid intermixing of speech, gesturing, writing and sketching by providing mechanisms to enhance the awareness of distributed participants of each others' actions.In this work we explore novel ways to leverage the capabilities of multimodal context-aware systems to bridge co-located and distributed collaboration contexts. We describe a system that allows participants at remote sites to collaborate in building a project schedule via sketching on multiple distributed whiteboards, and show how participants can be made aware of naturally occurring pointing gestures that reference diagram constituents as they are performed by remote participants.The system explores the multimodal fusion of pen, speech and 3D gestures, coupled to the dynamic construction of a semantic representation of the interaction, anchored on the sketched diagram, to provide feedback that overcomes some of the intrinsic ambiguities of pointing gestures.
Paulo Barthelmess, Edward C. Kaiser, David Demirdjian
ICMI2
2005 Design and implementation of network puzzles
abstract
Client puzzles have been proposed in a number of protocols as a mechanism for mitigating the effects of distributed denial of service (DDoS) attacks. In order to provide protection against simultaneous attacks across a wide range of applications and protocols, however, such puzzles must be placed at a layer common to all of them; the network layer. Placing puzzles at the IP layer fundamentally changes the service paradigm of the Internet, allowing any device within the network to push load back onto those it is servicing. An advantage of network layer puzzles over previous puzzle mechanisms is that they can be applied to all traffic from malicious clients, making it possible to defend against arbitrary attacks as well as making previously voluntary mechanisms mandatory. In this paper, we outline goals which must be met for puzzles to be deployed effectively at the network layer. We then describe the design, implementation, and evaluation of a system that meets these goals by supporting efficient, fine-grained control of puzzles at the network layer. In particular, we describe modifications to existing puzzle protocols that allow them to work at the network layer, a hint-based hash-reversal puzzle that allows for the generation and verification of fine-grained puzzles at line speed in the fast path of high-speed routers, and an iptables implementation that supports transparent deployment at arbitrary locations in the network.
Wu-chi Feng, Edward C. Kaiser, A. Luu
INFOCOM2
2005 Multimodal new vocabulary recognition through speech and handwriting in a whiteboard scheduling application
abstract
Our goal is to automatically recognize and enroll new vocabulary in a multimodal interface. To accomplish this our technique aims to leverage the mutually disambiguating aspects of co-referenced, co-temporal handwriting and speech. The co-referenced semantics are spatially and temporally determined by our multimodal interface for schedule chart creation. This paper motivates and describes our technique for recognizing out-of-vocabulary (OOV) terms and enrolling them dynamically in the system. We report results for the detection and segmentation of OOV words within a small multimodal test set. On the same test set we also report utterance, word and pronunciation level error rates both over individual input modes and multimodally. We show that combining information from handwriting and speech yields significantly better results than achievable by either mode alone.
Edward C. Kaiser
IUI1
2005 Panoptes: scalable low-power video sensor networking technologies
abstract
Video-based sensor networks can provide important visual information in a number of applications including: environmental monitoring, health care, emergency response, and video security. This article describes the Panoptes video-based sensor networking architecture, including its design, implementation, and performance. We describe two video sensor platforms that can deliver high-quality video over 802.11 networks with a power requirement less than 5 watts. In addition, we describe the streaming and prioritization mechanisms that we have designed to allow it to survive long-periods of disconnected operation. Finally, we describe a sample application and bitmapping algorithm that we have implemented to show the usefulness of our platform. Our experiments include an in-depth analysis of the bottlenecks within the system as well as power measurements for the various components of the system.
Wu-chi Feng, Edward C. Kaiser, Wu-chang Feng, Mikael Le Baillif
ACM Trans. Multim. Comput. Commun. Appl.2
2004 A multimodal learning interface for sketch, speak and point creation of a schedule chart
abstract
We present a video demonstration of an agent-based test bed application for ongoing research into multi-user, multimodal, computer-assisted meetings. The system tracks a two person scheduling meeting: one person standing at a touch sensitive whiteboard creating a Gantt chart, while another person looks on in view of a calibrated stereo camera. The stereo camera performs real-time, untethered, vision-based tracking of the onlooker's head, torso and limb movements, which in turn are routed to a 3D-gesture recognition agent. Using speech, 3D deictic gesture and 2D object de-referencing the system is able to track the onlooker's suggestion to move a specific milestone. The system also has a speech recognition agent capable of recognizing out-of-vocabulary (OOV) words as phonetic sequences. Thus when a user at the whiteboard speaks an OOV label name for a chart constituent while also writing it, the OOV speech is combined with letter sequences hypothesized by the handwriting recognizer to yield an orthography, pronunciation and semantics for the new label. These are then learned dynamically by the system and become immediately available for future recognition.
Edward C. Kaiser, David Demirdjian, Alexander Gruenstein, John Niekrasz, Matt Wesson
ICMI1
2003 Mutual disambiguation of 3D multimodal interaction in augmented and virtual reality
abstract
We describe an approach to 3D multimodal interaction in immersive augmented and virtual reality environments that accounts for the uncertain nature of the information sources. The resulting multimodal system fuses symbolic and statistical information from a set of 3D gesture, spoken language, and referential agents. The referential agents employ visible or invisible volumes that can be attached to 3D trackers in the environment, and which use a time-stamped history of the objects that intersect them to derive statistics for ranking potential referents. We discuss the means by which the system supports mutual disambiguation of these modalities and information sources, and show through a user study how mutual disambiguation accounts for over 45% of the successful 3D multimodal interpretations. An accompanying video demonstrates the system in action.
Edward C. Kaiser, Alex Olwal, David McGee, Hrvoje Benko, Andrea Corradini 0002, Phil Cohen 0001, Steven K. Feiner
ICMI1
2003 Panoptes: scalable low-power video sensor networking technologies
abstract
This demonstration will show the video sensor networking technologies developed at the OGI School of Science and Engineering. The general purpose video sensors allow programmers to create application-specific filtering, power management, and event triggering mechanisms. The demo will show a handful of video sensors operating under a variety of conditions including intermittent network connectivity as one might see in an environmental observation application.
Wu-chi Feng, Brian Code, Edward C. Kaiser, Mike Shea, Wu-chang Feng
ACM Multimedia3
2003 Panoptes: scalable low-power video sensor networking technologies
abstract
Video-based sensor networks can provide important visual information in a number of applications including: environmental monitoring, health care, emergency response, and video security. This paper describes the Panoptes video-based sensor networking architecture, including its design, implementation, and performance. We describe a video sensor platform that can deliver high-quality video over 802.11 networks with a power requirement of approximately 5 watts. In addition, we describe the streaming and prioritization mechanisms that we have designed to allow it to survive long-periods of disconnected operation. Finally, we describe a sample application and bitmapping algorithm that we have implemented to show the usefulness of our platform. Our experiments include an in-depth analysis of the bottlenecks within the system as well as power measurements for the various components of the system.
Wu-chi Feng, Brian Code, Edward C. Kaiser, Mike Shea, Wu-chang Feng, Louis Bavoil
ACM Multimedia3
2002 Implementation testing of a hybrid symbolic/statistical multimodal architecture
abstract
The design and implementation of hybrid symbolic/statistical architectures is a major area of interest in current multimodal system development. Such an architecture attempts to improve multimodal recognition and disambiguation rates by using corpus-based statistics to weight the contributions from various input streams. This is in contrast to current architectures that assume independence between input streams, and combine un-weighted posterior probabilities simply by taking their cross product.
Edward C. Kaiser, Phil Cohen 0001
INTERSPEECH1
1999 Robust, Finite-State Parsing for Spoken Language Understanding
abstract
Human understanding of spoken language appears to integrate the use of contextual expectations with acoustic level perception in a tightly-coupled, sequential fashion. Yet computer speech understanding systems typically pass the transcript produced by a speech recognizer into a natural language parser with no integration of acoustic and grammatical constraints. One reason for this is the complexity of implementing that integration. To address this issue we have created a robust, semantic parser as a single finite-state machine (FSM). As such, its run-time action is less complex than other robust parsers that are based on either chart or generalized left-right (GLR) architectures. Therefore, we believe it is ultimately more amenable to direct integration with a speech decoder.
Edward C. Kaiser
ACL1
1999 PROFER: predictive, robust finite-state parsing for spoken language
abstract
The natural language processing component of a speech understanding system is commonly a robust, semantic parser, implemented as either a chart-based transition network, or as a generalized left-right (GLR) parser. In contrast, we are developing a robust, semantic parser that is a single, predictive finite-state machine. Our approach is motivated by our belief that such a finite-state parser can ultimately provide an efficient vehicle for tightly integrating higher-level linguistic knowledge into speech recognition. We report on our development of this parser, with an example of its use, and a description of how it compares to both finite-state predictors and chart-based semantic parsers, while combining the elements of both.
Edward C. Kaiser, Michael Johnston, Peter A. Heeman
ICASSP1
1998 Beyond structured dialogues: factoring out grounding
abstract
Structured dialogue models are currently the only tools for easily building spoken dialogue systems. This approach, however, requires the dialogue designer to completely specify all dialogue behavior between the user and system, including how information is grounded between the user and the system. In this paper, we advocate factoring out the grounding behavior from structured dialogue models by using a general purpose dialogue manager that accounts for this behavior. This not only simplifies the specification of dialogue, but also allows more powerful mechanisms of grounding to be employed, which cannot be implemented within the framework of structured dialogues. 1. INTRODUCTION As speech recognition and speech synthesis continue to improve, spoken dialogue systems have started to emerge. However, significant barriers remain in building effective spoken dialogue systems. There will always be errors in speech recognition, and unfortunately, natural language understanding and dialogue...
Peter A. Heeman, Michael Johnston, Justin Denney, Edward C. Kaiser
ICSLP4
1998 Universal speech tools: the CSLU toolkit
abstract
A set of freely available, universal speech tools is needed to accelerate progress in the speech technology. The CSLU Toolkit represents an effort to make the core technology and fundamental infrastructure accessible, affordable and easy to use. The CSLU Toolkit has been under development for five years. This paper describes recent improvements, additions and uses of the CSLU Toolkit. 1. INTRODUCTION Since 1993, the Center for Spoken Language Understanding (CSLU) has focused on incorporating state-of-the-art spokenlanguage technology into a portable, comprehensive and easyto -use software environment. The result of these efforts is the CSLU Toolkit. The toolkit integrates learning materials, authoring tools and core technologies such as speech recognition, text-to-speech synthesis, facial animation and speech reading. The toolkit is designed to support basic research, development and education activities related to spoken language systems and human-computer interfaces. What are our ...
Stephen Sutton, Ronald A. Cole, Jacques de Villiers, Johan Schalkwyk, Pieter J. E. Vermeulen, Michael W. Macon, Yonghong Yan 0002, Edward C. Kaiser, Brian Rundle, Khaldoun Shobaki, John-Paul Hosom, Alexander Kain, Johan Wouters, Dominic W. Massaro, Michael M. Cohen
ICSLP8
1997 Bringing spoken language systems to the classroom
abstract
Currently, there are few opportunities for people to learn about and experiment with the latest spoken language technology. Furthermore, most research and development activities are restricted to a handful of academic and industrial labs. In order to make the technology less exclusive, it must become more accessible to the general population. This is now feasible with the development of the CSLU Toolkit which combines easy-to-use authoring tools with state-of-the-art human language technology. In this paper, we focus on the educational role of the toolkit and describe how it is being used in several local schools. 1. INTRODUCTION Research and development of spoken language systems is currently limited to relatively few academic and industrial laboratories. Because building such systems requires multidisciplinary expertise, sophisticated system development tools, language resources (e.g., pronunciation dictionaries), substantial computer resources and advanced technologies such as spee...
Stephen Sutton, Edward C. Kaiser, A. Cronk, Ronald A. Cole
EUROSPEECH2