VLDB 2026 Research / reviewers in the wild / expert
Jan Kleindienst
dblp:79/6986
· DBLP profile ↗
21ranked-venue papers
4as first author
0since 2021 · last 2018
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 8Human-computer interaction and ubiquitous computing · 7 · 1 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Deep learning architectures and training · 40% Question answering and dialogue systems · 30% Language models and text generation · 30% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Performance modeling and evaluation · 96% Distributed systems · 4% |
Topics — the 4 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Performance modeling and evaluation
benchmarking |
0.3 | 1 | 2018 | A Boo(n) for Evaluating Architecture Performance · ICML 2018 |
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
cloze-style question answering |
0.2 | 1 | 2016 | Text Understanding with the Attention Sum Reader Network · ACL (1) 2016 |
Distributed systems
middleware |
0.0 | 1 | 1996 | Lessons Learned from Implementing the CORBA Persistent Object Service · OOPSLA 1996 |
Data models and query languages › object-oriented database
persistent object system |
0.0 | 1 | 1996 | Lessons Learned from Implementing the CORBA Persistent Object Service · OOPSLA 1996 |
Methods — techniques the papers use, named apart from their topics
order statistics · 0.7ensemble · 0.2attention mechanism · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | A Boo(n) for Evaluating Architecture PerformanceabstractWe point out important problems with the common practice of using the best single model performance for comparing deep learning architectures, and we propose a method that corrects these flaws. Each time a model is trained, one gets a different result due to random factors in the training process, which include random parameter initialization and random data shuffling. Reporting the best single model performance does not appropriately address this stochasticity. We propose a normalized expected best-out-of-$n$ performance ($\text{Boo}_n$) as a way to correct these problems. Ondrej Bajgar, Rudolf Kadlec, Jan Kleindienst |
ICML | 3 |
| 2016 | Text Understanding with the Attention Sum Reader NetworkabstractSeveral large cloze-style context-questionanswer datasets have been introduced recently: the CNN and Daily Mail news data and the Children's Book Test.Thanks to the size of these datasets, the associated text comprehension task is well suited for deep-learning techniques that currently seem to outperform all alternative approaches.We present a new, simple model that uses attention to directly pick the answer from the context as opposed to computing the answer using a blended representation of words in the document as is usual in similar models.This makes the model particularly suitable for questionanswering problems where the answer is a single word from the document.Ensemble of our models sets new state of the art on all evaluated datasets. Rudolf Kadlec, Ondrej Bajgar, Jan Kleindienst |
ACL (1) | 4 |
| 2014 | Interactive Car Owner's Manual User StudyabstractWe present results of a user study with a prototype of an interactive speech-enabled car owner's user manual assistant. Its purpose is to help the driver learn about various car features and related procedures. The study focused on two scenarios -- when parked and while driving. We also used the Leap Motion gesture recognizer as an alternative to buttons. During the experiment we collected both objective driving data and subjective feedback. Results indicate that the users preferred the electronic user manual to the paper form, although they proposed numerous improvements. One particular concern was discoverability of content. The acceptance of Leap Motion gestures was low when driving, possibly impacted by short time allowed for practicing. Driver's distraction caused by interacting with the multimodal user manual was similar to that of receiving and reading text messages. Tomás Macek, Martin Labský, Jan Vystrcil, David Luksch, Tereza Kasparová, Ladislav Kunc, Jan Kleindienst |
AutomotiveUI | 7 |
| 2014 | Knowledge-based Dialog State TrackingabstractThis paper presents two discriminative knowledgebased dialog state trackers and their results on the Dialog State Tracking Challenge (DSTC) 2 and 3 datasets. The first tracker was submitted to the DSTC3 competition and scored second in the joint accuracy. The second tracker developed after the DSTC3 submission deadline gives even better results on the DSTC2 and DSTC3 datasets. It performs on par with the state of the art machine learning-based trackers while offering better interpretability. We summarize recent directions in the dialog state tracking (DST) and also discuss possible decomposition of the DST problem. Based on the results of DSTC2 and DSTC3 we analyze suitability of different techniques for each of the DST subproblems. Results of the trackers highlight the importance of Spoken Language Understanding (SLU) for the last two DSTCs. Rudolf Kadlec, Miroslav Vodolán, Jindrich Libovický, Jan Macek, Jan Kleindienst |
SLT | 5 |
| 2013 | Mostly passive information delivery in a carabstractIn this study we present and analyze a mostly passive infotainment approach to presenting information in a car. The passive style is similar to radio listening but content is generated on the fly and it is based on a mixture of personal information (calendar, emails) and public data (news, POI, jokes). The spoken part of the audio is machine synthesized. We explore two modes of operation. The first one is passive only. The second one is more interactive and speech commands are used to personalize the information mix and to request particular information items. Usability and distraction tests were conducted with both systems implemented using the Wizard of Oz technique. Both systems were assessed using multiple objective and subjective metrics and the results indicate that driver distraction was low for both systems. The users differed in the amount of interaction they preferred. Some users preferred more command-driven styles while others were happy with passive presentation. Most of the users were satisfied with the quality of synthesized speech and found it sufficient for the given purpose. In addition, feedback was collected from the subjects on what kind of information they liked listening to and how they would have preferred to ask for specific types of information. Tomás Macek, Tereza Kasparová, Jan Kleindienst, Ladislav Kunc, Martin Labský, Jan Vystrcil |
AutomotiveUI | 3 |
| 2012 | Impact of word error rate on driving performance while dictating short textsabstractThis paper describes the impact of speech recognition word error rate (WER) on driver's distraction in the context of short message dictation. A multi-modal dictation and error correction system was used in a simulated driving environment (Lane Change Test, LCT) to dictate text messages with prescribed semantic content. Driving accuracy was measured using several objective statistics produced by the LCT simulator. We report results for three datasets: 28 LCT trips by native US-English speakers at 40km/h, 23 more trips at 60km/h which had noise added in order to artificially increase WER levels and 22 LCT trips at 60km/h performed by non-native accented speakers. For the two datasets that used 60km/h we observed a moderate correlation between the driver's WER and driving performance statistics such as the mean deviation from ideal track (MDev) and the standard deviation of lateral position (SDLP). This correlation reached statistical significance for all of these statistics in the native dataset, and was significant for the overall SDLP in the non-native dataset. Additionally, we observed that higher WER levels lead to significantly lower message throughput and to significantly lower quality of sent messages, esp. for non-native speakers. Martin Labský, Jan Curín, Tomás Macek, Jan Kleindienst, Ladislav Kunc, Hoi Young, Ann Thymé-Gobbel, Holger Quast |
AutomotiveUI | 4 |
| 2011 | Dictating and editing short texts while driving: distraction and task completionabstractThis paper presents a multi-modal automotive dictation editor (codenamed ECOR) used to compose and correct text messages while driving. The goals are to keep driver's distraction minimal while achieving good task completion rates and times as well as acceptance by users. We report test results for a set of 28 native US-English speakers using the system while driving a standard lane-change-test (LCT) car simulator. The dictation editor was tested (1) without any display, (2) with a display showing the full edited text, and (3) with just the "active" part of text being shown. In all cases, the system provided extensive text-to-speech feedback in order to prevent the driver from having to look at the display. In addition, cell phone messaging and GPS destination entry were evaluated as reference tasks. The test subjects were instructed to send text messages containing prescribed semantic information, and were given a list of destinations for the GPS task. The levels of driver distraction (evaluated by car's deviation from an ideal track, reaction times, number of missed lane change signs, eye gaze information etc.) were compared between the 3 ECOR and the 2 reference tasks, and also to undistracted driving. Task completion was measured by the number and quality of messages sent out during a 4 minute LCT ride, and subjective feedback was collected via questionnaires. Results indicate that the eyes-free version keeps the distraction level acceptable while achieving good task completion rate. Both multi-modal versions caused more distraction than the eyes-free version and were comparable to the GPS entry task. For native speakers, the missing display for the eyes-free version did not impact quality of dictated text. By far, the cell phone texting task was the most distracting one. Text composition speed using dictation was faster than cell phone typing. Jan Curín, Martin Labský, Tomás Macek, Jan Kleindienst, Hoi Young, Ann Thymé-Gobbel, Holger Quast, Lars König |
AutomotiveUI | 4 |
| 2011 | Exercise Support System for Elderly: Multi-sensor Physiological State Detection and Usability Testing
Jan Macek, Jan Kleindienst |
INTERACT (2) | 2 |
| 2008 | SitCom: Virtual Smart-Room Environment for Multi-modal Perceptual Systems
Jan Curín, Jan Kleindienst |
IEA/AIE | 2 |
| 2008 | Creation and Visualization of User Behavior in Ambient Intelligent EnvironmentabstractAmbient intelligence is a rapidly developing field producing a new generation of applications and services demanding new, more complex strategies for usability testing and evaluation. In this paper, we identify caveats of traditional methods of application usability testing when applied to an ambient intelligent environment and we introduce a method for unified analysis of user behavior using a virtual environment visualization. For the usability testing in ambient environments, we propose to transform data (both data recorded in real settings and data artificially created by experts) into a virtual environment to cope with the following issues: a) overwhelming amount of data, b) proprietary data formats, and c) ethical codices issues. We describe a tool for creation of expert evaluation data and for transformation of data from video recordings and illustrate its use on two cases. Ivo Malý, Jan Curín, Jan Kleindienst, Pavel Slavík |
IV | 3 |
| 2007 | Interaction framework for home environment using speech and vision
Jan Kleindienst, Tomás Macek, Ladislav Serédi, Jan Sedivý |
Image Vis. Comput. | 1 |
| 2002 | CATCH-2004 Multi-Modal Browser: Overview Description with Usability AnalysisabstractThis paper takes a closer look at the user interface issues in our research multi-modal browser architecture. The browser framework, also briefly introduced in this paper, reuses single-modal browser technologies available for VoiceXML, WML, and HTML browsing. User interface actions on a particular browser are captured, converted to events, and distributed to the other browsers participating (possibly on different hosts) in the multi-modal framework. We have defined a synchronization protocol, which distributes such events with the help of the central component called the Virtual Proxy. The choice of the architecture and the synchronization primitives have profound consequences on handling certain interesting UI use cases. We particularly address those specified by the W3C MultiModal Requirements, which are related to the design of possible strategies of dealing with simultaneous input, solving input inconsistencies, and defining synchronization points. The proposed approaches are illustrated by examples. Jan Kleindienst, Ladislav Serédi, Pekka Kapanen, Janne Bergman |
ICMI | 1 |
| 2002 | An Approach to lightweight deployment of web servicesabstractWeb Services is gradually becoming the most popular distributed computing paradigm for the Internet. Although several vendor and research efforts are in progress, fully-fledged deployment of Web Services in a wide scale has not been accomplished yet. The present contribution describes a framework for lightweight deployment of Web Services. This framework can be seen as a contribution to a smooth transition step towards a complete large-scale deployment. The paper starts with a description of data access techniques supporting multiple data providers, devised and used in the scope of a European research project. It is illustrated how schemes employed to provide remote transparent access to the data providers evolved to a lightweight Web Services framework. Jaroslav Gergic, Jan Kleindienst, Y. Despotopoulos, John Soldatos 0001, Yiorgos Patikis, A. Anagnostou, Lazaros Polymenakos |
SEKE | 2 |
| 2001 | A DOM-based MVC Multi-modal e-Business
Stéphane H. Maes, Rafah Hosn, Jan Kleindienst, Tomás Macek, T. V. Raman 0001, Ladislav Serédi |
ICME | 3 |
| 2001 | Aspects of Design and Implementation of a Multi-Channel and Multi-Modal Information SystemabstractThe paper describes an architecture for multi-channel and multi-modal applications. First the design problem is explored and a proposal for a system that can handle multi-modal interaction and delivery of Internet content is proposed. The focus is pertained in some development aspects and the way they are addressed by using state-of-the-art tools. The various components are defined and described in detail. Finally, conclusions and a view of future work on the evolution of such systems is given. Vasiliki Demesticha, Jaroslav Gergic, Jan Kleindienst, Marion Mast, Lazaros Polymenakos, Henrik Schulz 0002, Ladislav Serédi |
ICSM | 3 |
| 2000 | An instantiable speech biometrics module with natural language interface: implementation in the telephony environmentabstractThis paper describes an implementation of the concept of conversational speech biometrics approach to personal authentication in the telephony environment. An application-independent module including a natural language-enabled part for verbal verification and identification and an acoustic speaker recognition engine for voice-print analysis are combined using a special verification/identification protocol which allows the application to adapt the session according to the dialog development and the query/command security. The results validate the feasibility and advantages of the concept of integrating the speaker and speech recognition technology. Users familiar with the system can log into the system with 2.7% or 3.2% false rejection and ca. 3/spl times/10/sup -11/% or 10/sup -6/% false acceptance rates in about 40 sec or 20 sec respectively. This makes speaker recognition for the first time deployable for high security applications even with today's technology-a claim that can't be made with other speaker recognition technology. The system has a client-server architecture and is suitable for various applications and platforms. Jirí Navrátil 0001, Jan Kleindienst, Stéphane H. Maes |
ICASSP | 2 |
| 2000 | Hierarchical feature-based translation for scalable natural language understandingabstractFor complex natural language understanding systems with a large number of statistically confusable but semantically different for-mal commands, there are many difficulties in performing an accu-rate translation of a user input into a formal command in a single step. This paper addresses scalability issues in natural language understanding, and describes a method for performing the transla-tion in a hierarchical manner. The hierarchical method improves the system accuracy, reduces the computational complexity of the translation, provides additional numerical robustness during train-ing and decoding, and permits a more efficient packaging of the components of the natural language understanding system. 1. Ganesh N. Ramaswamy, Jan Kleindienst |
INTERSPEECH | 2 |
| 1999 | A pervasive conversational interface for information interaction
Ganesh N. Ramaswamy, Jan Kleindienst, Daniel M. Coffman, Ponani S. Gopalakrishnan, Chalapathy Neti |
EUROSPEECH | 2 |
| 1998 | Automatic identification of command boundaries in a conversational natural language user interface
Ganesh N. Ramaswamy, Jan Kleindienst |
ICSLP | 2 |
| 1996 | Lessons Learned from Implementing the CORBA Persistent Object ServiceabstractIn this paper, the authors share their experiences gathered during the design and implementation of the CORBA Persistent Object Service. There are two problems related to a design and implementation of the Persistence Service: first, OMG intentionally leaves the functionality core of the Persistence Service unspecified; second, OMG encourages reuse of other Object Services without being specific enough in this respect. The paper identifies the key design issues implied both by the intentional lack of OMG specification and the limits of the implementation environment characteristics. At the same time, the paper discusses the benefits and drawbacks of reusing other Object Services, particularly the Relationship and Externalization Services, to support the Persistence Service. Surprisingly, the key lesson learned is that a direct reuse of these Object Services is impossible. Jan Kleindienst, Frantisek Plásil, Petr Tuma 0001 |
OOPSLA | 1 |
| 1996 | CORBA and Object Services
Jan Kleindienst, Frantisek Plásil, Petr Tuma 0001 |
SOFSEM | 1 |