EDBT 2026 Demo / reviewers in the wild / expert
Edward C. Kaiser
dblp:87/279
· DBLP profile ↗
24ranked-venue papers
11as first author
0since 2021 · last 2013
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 11 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-authorArtificial intelligence and machine learning · 6 · 3 first-authorComputer networks · 3 · 1 first-authorSecurity and privacy · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Network and information security
3 papers |
Network security · 52% Cryptographic protocols and secure computation · 27% Blockchain and cryptocurrency security · 21% | |
| Computer networks
4 papers |
Internet of things and sensor networks · 62% Internet architecture and protocols · 28% Wireless networking · 9% | |
| Human-computer interaction and pervasive computing
1 paper |
Collaborative and social computing · 100% | |
| Artificial intelligence
1 paper |
Information extraction and text analysis · 67% Speech recognition and synthesis · 33% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Distributed systems · 54% Energy-efficient computing · 46% | |
| Computer graphics and multimedia
1 paper |
Audio and music processing · 100% |
Topics — the 13 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Network security › attack resilience › attack mitigation
denial-of-service defense |
0.1 | 2 | 2007 | mod_kaPoW: mitigating DoS with transparent proof-of-work · CoNEXT 2007 Design and implementation of network puzzles · INFOCOM 2005 |
Cryptographic protocols and secure computation › secret sharing
cheating detection |
0.1 | 1 | 2009 | Fides: remote anomaly-based cheat detection using client emulation · CCS 2009 |
Internet of things and sensor networks › camera sensor networks
wireless video sensor networks |
0.1 | 2 | 2003 | Panoptes: scalable low-power video sensor networking technologies · ACM Multimedia 2003 Panoptes: scalable low-power video sensor networking technologies · ACM Multimedia 2003 |
Collaborative and social computing
computer-mediated communication |
0.1 | 1 | 2007 | Multimodal redundancy across handwriting and speech during computer mediated human-human interactions · CHI 2007 |
Blockchain and cryptocurrency security › consensus protocol
proof-of-work |
0.1 | 1 | 2007 | mod_kaPoW: mitigating DoS with transparent proof-of-work · CoNEXT 2007 |
Network security › attack resilience › attack mitigation › denial-of-service defense
client puzzles |
0.1 | 1 | 2005 | Design and implementation of network puzzles · INFOCOM 2005 |
Distributed systems › distributed system architecture
client-server systems |
0.0 | 1 | 2009 | Fides: remote anomaly-based cheat detection using client emulation · CCS 2009 |
Energy-efficient computing
power management |
0.0 | 2 | 2003 | Panoptes: scalable low-power video sensor networking technologies · ACM Multimedia 2003 Panoptes: scalable low-power video sensor networking technologies · ACM Multimedia 2003 |
Natural language and speech › Information extraction and text analysis › syntactic parsing › grammar-based parsing
finite-state parsing |
0.0 | 1 | 1999 | Robust, Finite-State Parsing for Spoken Language Understanding · ACL 1999 |
Natural language and speech › Speech recognition and synthesis
spoken language understanding |
0.0 | 1 | 1999 | Robust, Finite-State Parsing for Spoken Language Understanding · ACL 1999 |
Natural language and speech › Information extraction and text analysis
syntactic parsing |
0.0 | 1 | 1999 | Robust, Finite-State Parsing for Spoken Language Understanding · ACL 1999 |
Audio and music processing
speech recognition |
0.0 | 1 | 2007 | Multimodal redundancy across handwriting and speech during computer mediated human-human interactions · CHI 2007 |
Wireless networking › WLAN
IEEE 802.11 |
0.0 | 1 | 2003 | Panoptes: scalable low-power video sensor networking technologies · ACM Multimedia 2003 |
Methods — techniques the papers use, named apart from their topics
client emulation · 0.2anomaly detection · 0.2transparent proof-of-work · 0.1empirical study · 0.1iptables · 0.1hash-reversal puzzle · 0.1streaming · 0.1prioritization · 0.1event triggering · 0.1application-specific filtering · 0.1bitmapping · 0.0bit mapping · 0.0semantic parsing · 0.0finite state machine · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2013 | Demonstration of sketch-thru-plan: a multimodal interface for command and controlabstractThis paper demonstrates a multimodal system called Sketch-Thru-Plan (STP) that enables users to speak and draw doctrinal language and symbols in order to create courses of action. We argue that STP can meet many of the challenges inherent in building user interfaces for operations planning and command-and-control. The system is being transitioned to military organizations for use in planning courses of action. Phil Cohen 0001, M. Cecelia Buchanan, Edward C. Kaiser, Michael J. Corrigan, Scott Lind, Matt Wesson |
ICMI | 3 |
| 2009 | Fides: remote anomaly-based cheat detection using client emulationabstractAs a result of physically owning the client machine, cheaters in online games currently have the upper-hand when it comes to avoiding detection. To address this problem and turn the table on cheaters, this paper presents Fides, an anomaly-based cheat detection approach that remotely validates game execution. With Fides, a server-side Controller specifies how and when a client-side Auditor measures the game. To accurately validate measurements, the Controller partially emulates the client and collaborates with the server. This paper examines a range of cheat methods and initial measurements that counter them, showing that a Fides prototype is able to efficiently detect several existing cheats, including one state-of-the-art cheat that is advertised as "undetectable". Edward C. Kaiser, Wu-chang Feng, Travis Schluessler |
CCS | 1 |
| 2007 | Multimodal redundancy across handwriting and speech during computer mediated human-human interactionsabstractLecturers, presenters and meeting participants often say what they publicly handwrite. In this paper, we report on three empirical explorations of such multimodal redundancy -- during whiteboard presentations, during a spontaneous brainstorming meeting, and during the informal annotation and discussion of photographs. We show that redundantly presented words, compared to other words used during a presentation or meeting, tend to be topic specific and thus are likely to be out-of-vocabulary. We also show that they have significantly higher tf-idf (term frequency-inverse document frequency) weights than other words, which we argue supports the hypothesis that they are dialogue-critical words. We frame the import of these empirical findings by describing SHACER, our recently introduced Speech and HAndwriting reCognizER, which can combine information from instances of redundant handwriting and speech to dynamically learn new vocabulary. Edward C. Kaiser, Paulo Barthelmess, Candice Erdmann, Phil Cohen 0001 |
CHI | 1 |
| 2007 | mod_kaPoW: mitigating DoS with transparent proof-of-workabstractUnwanted traffic remains a fundamental problem for networked systems. Proof-of-work (PoW) is a defense mechanism that adds a client-specific challenge at the start of a networked protocol. The challenge acts as a filter for clients based on their willingness to solve a computational task of varying difficulty. The difficulty is tailored to the individual client and is set proportional to its relative load on the server. Edward C. Kaiser, Wu-chang Feng |
CoNEXT | 1 |
| 2007 | Workshop on tagging, mining and retrieval of human related activity informationabstractInexpensive and user friendly cameras, microphones, and other devices such as digital pens are making it increasingly easy to capture, store and process large amounts of data over a variety of media. Even though the barriers for data acquisition have been lowered, making use of these data remains challenging. The focus of the present workshop is on issues related to theory, methods and techniques for facilitating the organization, retrieval and reuse of multimodal information. The emphasis is on organization and retrieval of information related to human activity, i.e. that is generated and consumed by individuals and groups as they go about their work, learning and leisure. Paulo Barthelmess, Edward C. Kaiser |
ICMI | 2 |
| 2007 | Toward content-aware multimodal tagging of personal photo collectionsabstractA growing number of tools is becoming available, that make use ofexisting tags to help organize and retrieve photos, facilitating the management and use of photo sets. The tagging on which these techniques rely remains a time consuming, labor intensive task that discourages many users. To address this problem, we aim to leverage the multimodal content of naturally occurring photo discussions among friends and families to automatically extract tags from a combination of conversational speech, handwriting, and photo content analysis. While naturally occurring discussions are rich sources of informationabout photos, methods need to be developed to reliably extract a set of discriminative tags from this noisy, unconstrained group discourse. To this end, this paper contributes ananalysis of pilot data identifying robust multimodal features examining the interplay between photo content and other modalities such as speech and handwriting. Our analysis is motivated by a search for design implications leading to the effective incorporation of automated location and person identification(e.g. based on GPS and facial recognition technologies) into a system able to extract tags from natural multimodal conversations. Paulo Barthelmess, Edward C. Kaiser, David McGee |
ICMI | 2 |
| 2006 | Collaborative multimodal photo annotation over digital paperabstractThe availability of metadata annotations over media content such as photos is known to enhance retrieval and organization, particularly for large data sets. The greatest challenge for obtaining annotations remains getting users to perform the large amount of tedious manual work that is required.In this paper we introduce an approach for semi-automated labeling based on extraction of metadata from naturally occurring conversations of groups of people discussing pictures among themselves.As the burden for structuring and extracting metadata is shifted from users to the system, new recognition challenges arise. We explore how multimodal language can help in 1) detecting a concise set of meaningful labels to be associated with each photo, 2) achieving robust recognition of these key semantic terms, and 3) facilitating label propagation via multimodal shortcuts. Analysis of the data of a preliminary pilot collection suggests that handwritten labels may be highly indicative of the semantics of each photo, as indicated by the correlation of handwritten terms with high frequency spoken ones. We point to initial directions exploring a multimodal fusion technique to recover robust spelling and pronunciation of these high-value terms from redundant speech and handwriting. Paulo Barthelmess, Edward C. Kaiser, David McGee, Phil Cohen 0001 |
ICMI | 2 |
| 2006 | Collaborative multimodal photo annotation over digital paperabstractThe availability of metadata annotations over media content such as photos is known to enhance retrieval and organization, particularly for large data sets. The greatest challenge for obtaining annotations remains getting users to perform the large amount of tedious manual work that is required. In this demo we show a system for semi-automated labeling based on extraction of metadata from naturally occurring conversations of groups of people discussing pictures among themselves. The system supports a variety of collaborative label elicitation scenarios mixing co-located and distributed participants, operating primarily via speech, handwriting and sketching over tangible digital paper photo printouts. We demonstrate the real-time capabilities of the system by providing hands-on annotation experience for conference participants. Demo annotations are performed over public domain pictures portraying mainstream themes (e.g. from famous movies). Paulo Barthelmess, Edward C. Kaiser, David McGee, Phil Cohen 0001 |
ICMI | 2 |
| 2006 | Using redundant speech and handwriting for learning new vocabulary and understanding abbreviationsabstractNew language constantly emerges from complex, collaborative human-human interactions like meetings -- such as, for instance, when a presenter handwrites a new term on a whiteboard while saying it. Fixed vocabulary recognizers fail on such new terms, which often are critical to dialogue understanding. We present a proof-of-concept multimodal system that combines information from handwriting and speech recognition to learn the spelling, pronunciation and semantics of out-of-vocabulary terms from single instances of redundant multimodal presentation (e.g. saying a term while handwriting it). For the task of recognizing the spelling and semantics of abbreviated Gantt chart labels across a held-out test series of five scheduling meetings we show a significant relative error rate reduction of 37% when our learning methods are used and allowed to persist across the meeting series, as opposed to when they are not used. Edward C. Kaiser |
ICMI | 1 |
| 2006 | Edge-splitting in a cumulative multimodal system, for a no-wait temporal threshold on information fusion, combined with an under-specified displayabstractPredicting the end of user input turns in a multimodal system can be complex. User interactions vary across a spectrum from single, unimodal inputs to multimodal combinations delivered either simultaneously or sequentially. Early multimodal systems used a fixed duration temporal threshold to determine how long to wait for the next input before processing and integration. Several recent studies have proposed using dynamic or adaptive temporal thresholds to predict turn segmentation and thus achieve faster system response times. We introduce an approach that requires no temporal threshold. First we contrast current multimodal command interfaces to a new class of cumulative-observant multimodal systems that we introduce. Within that new system class we show how our technique of edge-splitting combined with our strategy for under-specified, no-wait, visual feedback resolves parsing problems that underlie turn segmentation errors. Test results show a 46.2% significant reduction in multimodal recognition errors, compared to not using these techniques. Edward C. Kaiser, Paulo Barthelmess |
INTERSPEECH | 1 |
| 2005 | Distributed pointing for multimodal collaboration over sketched diagramsabstractA problem faced by groups that are not co-located but need to collaborate on a common task is the reduced access to the rich multimodal communicative context that they would have access to if they were collaborating face-to-face. Collaboration support tools aim to reduce the adverse effects of this restricted access to the fluid intermixing of speech, gesturing, writing and sketching by providing mechanisms to enhance the awareness of distributed participants of each others' actions.In this work we explore novel ways to leverage the capabilities of multimodal context-aware systems to bridge co-located and distributed collaboration contexts. We describe a system that allows participants at remote sites to collaborate in building a project schedule via sketching on multiple distributed whiteboards, and show how participants can be made aware of naturally occurring pointing gestures that reference diagram constituents as they are performed by remote participants.The system explores the multimodal fusion of pen, speech and 3D gestures, coupled to the dynamic construction of a semantic representation of the interaction, anchored on the sketched diagram, to provide feedback that overcomes some of the intrinsic ambiguities of pointing gestures. Paulo Barthelmess, Edward C. Kaiser, David Demirdjian |
ICMI | 2 |
| 2005 | Design and implementation of network puzzlesabstractClient puzzles have been proposed in a number of protocols as a mechanism for mitigating the effects of distributed denial of service (DDoS) attacks. In order to provide protection against simultaneous attacks across a wide range of applications and protocols, however, such puzzles must be placed at a layer common to all of them; the network layer. Placing puzzles at the IP layer fundamentally changes the service paradigm of the Internet, allowing any device within the network to push load back onto those it is servicing. An advantage of network layer puzzles over previous puzzle mechanisms is that they can be applied to all traffic from malicious clients, making it possible to defend against arbitrary attacks as well as making previously voluntary mechanisms mandatory. In this paper, we outline goals which must be met for puzzles to be deployed effectively at the network layer. We then describe the design, implementation, and evaluation of a system that meets these goals by supporting efficient, fine-grained control of puzzles at the network layer. In particular, we describe modifications to existing puzzle protocols that allow them to work at the network layer, a hint-based hash-reversal puzzle that allows for the generation and verification of fine-grained puzzles at line speed in the fast path of high-speed routers, and an iptables implementation that supports transparent deployment at arbitrary locations in the network. Wu-chi Feng, Edward C. Kaiser, A. Luu |
INFOCOM | 2 |
| 2005 | Multimodal new vocabulary recognition through speech and handwriting in a whiteboard scheduling applicationabstractOur goal is to automatically recognize and enroll new vocabulary in a multimodal interface. To accomplish this our technique aims to leverage the mutually disambiguating aspects of co-referenced, co-temporal handwriting and speech. The co-referenced semantics are spatially and temporally determined by our multimodal interface for schedule chart creation. This paper motivates and describes our technique for recognizing out-of-vocabulary (OOV) terms and enrolling them dynamically in the system. We report results for the detection and segmentation of OOV words within a small multimodal test set. On the same test set we also report utterance, word and pronunciation level error rates both over individual input modes and multimodally. We show that combining information from handwriting and speech yields significantly better results than achievable by either mode alone. Edward C. Kaiser |
IUI | 1 |
| 2005 | Panoptes: scalable low-power video sensor networking technologiesabstractVideo-based sensor networks can provide important visual information in a number of applications including: environmental monitoring, health care, emergency response, and video security. This article describes the Panoptes video-based sensor networking architecture, including its design, implementation, and performance. We describe two video sensor platforms that can deliver high-quality video over 802.11 networks with a power requirement less than 5 watts. In addition, we describe the streaming and prioritization mechanisms that we have designed to allow it to survive long-periods of disconnected operation. Finally, we describe a sample application and bitmapping algorithm that we have implemented to show the usefulness of our platform. Our experiments include an in-depth analysis of the bottlenecks within the system as well as power measurements for the various components of the system. Wu-chi Feng, Edward C. Kaiser, Wu-chang Feng, Mikael Le Baillif |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2004 | A multimodal learning interface for sketch, speak and point creation of a schedule chartabstractWe present a video demonstration of an agent-based test bed application for ongoing research into multi-user, multimodal, computer-assisted meetings. The system tracks a two person scheduling meeting: one person standing at a touch sensitive whiteboard creating a Gantt chart, while another person looks on in view of a calibrated stereo camera. The stereo camera performs real-time, untethered, vision-based tracking of the onlooker's head, torso and limb movements, which in turn are routed to a 3D-gesture recognition agent. Using speech, 3D deictic gesture and 2D object de-referencing the system is able to track the onlooker's suggestion to move a specific milestone. The system also has a speech recognition agent capable of recognizing out-of-vocabulary (OOV) words as phonetic sequences. Thus when a user at the whiteboard speaks an OOV label name for a chart constituent while also writing it, the OOV speech is combined with letter sequences hypothesized by the handwriting recognizer to yield an orthography, pronunciation and semantics for the new label. These are then learned dynamically by the system and become immediately available for future recognition. Edward C. Kaiser, David Demirdjian, Alexander Gruenstein, John Niekrasz, Matt Wesson |
ICMI | 1 |
| 2003 | Mutual disambiguation of 3D multimodal interaction in augmented and virtual realityabstractWe describe an approach to 3D multimodal interaction in immersive augmented and virtual reality environments that accounts for the uncertain nature of the information sources. The resulting multimodal system fuses symbolic and statistical information from a set of 3D gesture, spoken language, and referential agents. The referential agents employ visible or invisible volumes that can be attached to 3D trackers in the environment, and which use a time-stamped history of the objects that intersect them to derive statistics for ranking potential referents. We discuss the means by which the system supports mutual disambiguation of these modalities and information sources, and show through a user study how mutual disambiguation accounts for over 45% of the successful 3D multimodal interpretations. An accompanying video demonstrates the system in action. Edward C. Kaiser, Alex Olwal, David McGee, Hrvoje Benko, Andrea Corradini 0002, Phil Cohen 0001, Steven K. Feiner |
ICMI | 1 |
| 2003 | Panoptes: scalable low-power video sensor networking technologiesabstractThis demonstration will show the video sensor networking technologies developed at the OGI School of Science and Engineering. The general purpose video sensors allow programmers to create application-specific filtering, power management, and event triggering mechanisms. The demo will show a handful of video sensors operating under a variety of conditions including intermittent network connectivity as one might see in an environmental observation application. Wu-chi Feng, Brian Code, Edward C. Kaiser, Mike Shea, Wu-chang Feng |
ACM Multimedia | 3 |
| 2003 | Panoptes: scalable low-power video sensor networking technologiesabstractVideo-based sensor networks can provide important visual information in a number of applications including: environmental monitoring, health care, emergency response, and video security. This paper describes the Panoptes video-based sensor networking architecture, including its design, implementation, and performance. We describe a video sensor platform that can deliver high-quality video over 802.11 networks with a power requirement of approximately 5 watts. In addition, we describe the streaming and prioritization mechanisms that we have designed to allow it to survive long-periods of disconnected operation. Finally, we describe a sample application and bitmapping algorithm that we have implemented to show the usefulness of our platform. Our experiments include an in-depth analysis of the bottlenecks within the system as well as power measurements for the various components of the system. Wu-chi Feng, Brian Code, Edward C. Kaiser, Mike Shea, Wu-chang Feng, Louis Bavoil |
ACM Multimedia | 3 |
| 2002 | Implementation testing of a hybrid symbolic/statistical multimodal architectureabstractThe design and implementation of hybrid symbolic/statistical architectures is a major area of interest in current multimodal system development. Such an architecture attempts to improve multimodal recognition and disambiguation rates by using corpus-based statistics to weight the contributions from various input streams. This is in contrast to current architectures that assume independence between input streams, and combine un-weighted posterior probabilities simply by taking their cross product. Edward C. Kaiser, Phil Cohen 0001 |
INTERSPEECH | 1 |
| 1999 | Robust, Finite-State Parsing for Spoken Language UnderstandingabstractHuman understanding of spoken language appears to integrate the use of contextual expectations with acoustic level perception in a tightly-coupled, sequential fashion. Yet computer speech understanding systems typically pass the transcript produced by a speech recognizer into a natural language parser with no integration of acoustic and grammatical constraints. One reason for this is the complexity of implementing that integration. To address this issue we have created a robust, semantic parser as a single finite-state machine (FSM). As such, its run-time action is less complex than other robust parsers that are based on either chart or generalized left-right (GLR) architectures. Therefore, we believe it is ultimately more amenable to direct integration with a speech decoder. Edward C. Kaiser |
ACL | 1 |
| 1999 | PROFER: predictive, robust finite-state parsing for spoken languageabstractThe natural language processing component of a speech understanding system is commonly a robust, semantic parser, implemented as either a chart-based transition network, or as a generalized left-right (GLR) parser. In contrast, we are developing a robust, semantic parser that is a single, predictive finite-state machine. Our approach is motivated by our belief that such a finite-state parser can ultimately provide an efficient vehicle for tightly integrating higher-level linguistic knowledge into speech recognition. We report on our development of this parser, with an example of its use, and a description of how it compares to both finite-state predictors and chart-based semantic parsers, while combining the elements of both. Edward C. Kaiser, Michael Johnston, Peter A. Heeman |
ICASSP | 1 |
| 1998 | Beyond structured dialogues: factoring out groundingabstractStructured dialogue models are currently the only tools for easily building spoken dialogue systems. This approach, however, requires the dialogue designer to completely specify all dialogue behavior between the user and system, including how information is grounded between the user and the system. In this paper, we advocate factoring out the grounding behavior from structured dialogue models by using a general purpose dialogue manager that accounts for this behavior. This not only simplifies the specification of dialogue, but also allows more powerful mechanisms of grounding to be employed, which cannot be implemented within the framework of structured dialogues. 1. INTRODUCTION As speech recognition and speech synthesis continue to improve, spoken dialogue systems have started to emerge. However, significant barriers remain in building effective spoken dialogue systems. There will always be errors in speech recognition, and unfortunately, natural language understanding and dialogue... Peter A. Heeman, Michael Johnston, Justin Denney, Edward C. Kaiser |
ICSLP | 4 |
| 1998 | Universal speech tools: the CSLU toolkitabstractA set of freely available, universal speech tools is needed to accelerate progress in the speech technology. The CSLU Toolkit represents an effort to make the core technology and fundamental infrastructure accessible, affordable and easy to use. The CSLU Toolkit has been under development for five years. This paper describes recent improvements, additions and uses of the CSLU Toolkit. 1. INTRODUCTION Since 1993, the Center for Spoken Language Understanding (CSLU) has focused on incorporating state-of-the-art spokenlanguage technology into a portable, comprehensive and easyto -use software environment. The result of these efforts is the CSLU Toolkit. The toolkit integrates learning materials, authoring tools and core technologies such as speech recognition, text-to-speech synthesis, facial animation and speech reading. The toolkit is designed to support basic research, development and education activities related to spoken language systems and human-computer interfaces. What are our ... Stephen Sutton, Ronald A. Cole, Jacques de Villiers, Johan Schalkwyk, Pieter J. E. Vermeulen, Michael W. Macon, Yonghong Yan 0002, Edward C. Kaiser, Brian Rundle, Khaldoun Shobaki, John-Paul Hosom, Alexander Kain, Johan Wouters, Dominic W. Massaro, Michael M. Cohen |
ICSLP | 8 |
| 1997 | Bringing spoken language systems to the classroomabstractCurrently, there are few opportunities for people to learn about and experiment with the latest spoken language technology. Furthermore, most research and development activities are restricted to a handful of academic and industrial labs. In order to make the technology less exclusive, it must become more accessible to the general population. This is now feasible with the development of the CSLU Toolkit which combines easy-to-use authoring tools with state-of-the-art human language technology. In this paper, we focus on the educational role of the toolkit and describe how it is being used in several local schools. 1. INTRODUCTION Research and development of spoken language systems is currently limited to relatively few academic and industrial laboratories. Because building such systems requires multidisciplinary expertise, sophisticated system development tools, language resources (e.g., pronunciation dictionaries), substantial computer resources and advanced technologies such as spee... Stephen Sutton, Edward C. Kaiser, A. Cronk, Ronald A. Cole |
EUROSPEECH | 2 |