VLDB 2026 Research / reviewers in the wild / expert
Michael Feld
dblp:26/2807
· DBLP profile ↗
24ranked-venue papers
6as first author
8since 2021 · last 2026
0000-0001-6755-5287ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 10 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-authorSystems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Looking for an Even Better Fit? A Personalized Incremental Reinforcement Learning Multimodal Object Referencing FrameworkabstractThe rapid advancement of the automotive industry toward automated and semi-automated vehicles has rendered traditional methods of vehicle interaction, such as touch-based and voice command systems, inadequate for a widening range of non-driving related tasks, such as referencing objects outside of the vehicle. Consequently, research has shifted toward gestural input (e.g., hand, gaze, and head pose gestures) as a more suitable mode of interaction during driving. However, due to the dynamic nature of driving and individual variation, there are significant differences in drivers’ gestural input performance. While, in theory, this inherent variability could be moderated by substantial data-driven machine learning models, prevalent methodologies lean toward constrained, single-instance trained models for object referencing. These models show a limited capacity to continuously adapt to the divergent behaviors of individual drivers and the variety of driving scenarios. To address this, we proposed, in our previous work, an incremental learning approach, adapting to changing user behavior. Although this method proved superior over state-of-the-art approaches, it still lacked in performance and had some computational resource limitations, mainly the small-sized memory needed to store a subset of old data for incremental learning. Therefore, in this work, we propose an alternative solution using reinforcement learning that can be used to overcome memory storage limitations, achieving a zero data storage approach for incremental learning without performance loss. Our newly enhanced solution has been added to our previous framework at https://github.com/amrgomaaelhady/IcRegress . Amr Gomaa, Kiran Gani, Michael Feld, Antonio Krüger |
ACM Trans. Interact. Intell. Syst. | 3 |
| 2024 | Bridging the Gap to Natural Language-based Grasp Predictions through Semantic Information ExtractionabstractEnabling multi-fingered robots to choose an appropriate grasp on an object from natural language instructions poses great difficulties for such systems. The diversity, imprecision, and limited information contained in the language make this task particularly challenging. However, speech serves humans as a natural communication interface that can aid robots in adapting to the environment more easily. Therefore, providing robots with relevant data about the objects they interact with is essential for them to understand how to carry out object manipulation tasks. By leveraging Named Entity Recognition (NER) to automatically extract semantic data, our work introduces a novel approach to text-based grasp predictions. Our methodology involves a multistage learning approach using a semantic information extractor that provides significant features to a grasp prediction model. To assess the effectiveness of our approach, we conducted experiments on an existing corpus and two corpora generated by ChatGPT. Our results demonstrate superior performance compared to similar grasp prediction models while overcoming limitations in the literature. Additionally, we open-source our training data for reproducibility and future research advancement. Niko Kleer, Martin Feick, Amr Gomaa, Michael Feld, Antonio Krüger |
IROS | 4 |
| 2024 | Looking for a better fit? An Incremental Learning Multimodal Object Referencing Framework adapting to Individual DriversabstractThe rapid advancement of the automotive industry towards automated and semi-automated vehicles has rendered traditional methods of vehicle interaction, such as touch-based and voice command systems, inadequate for a widening range of non-driving related tasks, such as referencing objects outside of the vehicle. Consequently, research has shifted toward gestural input (e.g., hand, gaze, and head pose gestures) as a more suitable mode of interaction during driving. However, due to the dynamic nature of driving and individual variation, there are significant differences in drivers’ gestural input performance. While, in theory, this inherent variability could be moderated by substantial data-driven machine learning models, prevalent methodologies lean towards constrained, single-instance trained models for object referencing. These models show a limited capacity to continuously adapt to the divergent behaviors of individual drivers and the variety of driving scenarios. To address this, we propose IcRegress, a novel regression-based incremental learning approach that adapts to changing behavior and the unique characteristics of drivers engaged in the dual task of driving and referencing objects. We suggest a more personalized and adaptable solution for multimodal gestural interfaces, employing continuous lifelong learning to enhance driver experience, safety, and convenience. Our approach was evaluated using an outside-the-vehicle object referencing use case, highlighting the superiority of the incremental learning models adapted over a single trained model across various driver traits such as handedness, driving experience, and numerous driving conditions. Finally, to facilitate reproducibility, ease deployment, and promote further research, we offer our approach as an open-source framework at https://github.com/amrgomaaelhady/IcRegress. Amr Gomaa, Guillermo Reyes, Michael Feld, Antonio Krüger |
IUI | 3 |
| 2024 | Incorporation of the Intended Task into a Vision-based Grasp Type Predictor for Multi-fingered Robotic GraspingabstractRobots that make use of multi-fingered or fully anthropomorphic end-effectors can engage in highly complex manipulation tasks. However, the choice of a suitable grasp for manipulating an object is strongly influenced by factors such as the physical properties of an object and the intended task. This makes predicting an appropriate grasping pose for carrying out a concrete task notably challenging. At the same time, current grasp type predictors rarely consider the task as a part of the prediction process. This work proposes a learning model that considers the task in addition to an object’s visual features for predicting a suitable grasp type. Furthermore, we generate a synthetic dataset by simulating robotic grasps on 3D object models based on the BarrettHand end-effector. With an angular similarity of 0.9 and above, our model achieves competitive prediction results compared to grasp type predictors that do not consider the intended task for learning grasps. Finally, to foster research in the field, we make our synthesized dataset available to the research community. Niko Kleer, Ole Keil, Martin Feick, Amr Gomaa, Tim Schwartz, Michael Feld |
RO-MAN | 6 |
| 2023 | It's all about you: Personalized in-Vehicle Gesture Recognition with a Time-of-Flight CameraabstractDespite significant advances in gesture recognition technology, recognizing gestures in a driving environment remains challenging due to limited and costly data and its dynamic, ever-changing nature. In this work, we propose a model-adaptation approach to personalize the training of a CNNLSTM model and improve recognition accuracy while reducing data requirements. Our approach contributes to the field of dynamic hand gesture recognition while driving by providing a more efficient and accurate method that can be customized for individual users, ultimately enhancing the safety and convenience of in-vehicle interactions, as well as driver’s experience and system trust. We incorporate hardware enhancement using a time-of-flight camera and algorithmic enhancement through data augmentation, personalized adaptation, and incremental learning techniques. We evaluate the performance of our approach in terms of recognition accuracy, achieving up to 90%, and show the effectiveness of personalized adaptation and incremental learning for a user-centered design. Guillermo Reyes, Amr Gomaa, Michael Feld |
AutomotiveUI | 3 |
| 2022 | Leveraging Publicly Available Textual Object Descriptions for Anthropomorphic Robotic Grasp PredictionsabstractRobotic systems using anthropomorphic end-effectors face tremendous challenges choosing a suitable pose for grasping an object. The fact that the choice of a grasp is influenced by the physical properties of an object, the intended task, and the environment results in a considerable amount of variables. The majority of models targeted towards enabling such robots to determine a suitable grasping pose rely on computer vision techniques, sometimes complemented by textual data. This paper investigates the potential of publicly available textual descriptions to predict a suitable grasping pose for anthropomorphic end-effectors. To this end, we have retrieved textual descriptions from Wikipedia, Wiktionary, and WordNet as well as a number of well-known dictionaries for 100 everyday objects. We analyze and compare the prediction quality of multiple learning methods while showing that a support vector machine-based approach can utilize this data for achieving a prediction accuracy above 0.75. Finally, we make our collected data available to the research community. Niko Kleer, Martin Feick, Michael Feld |
IROS | 3 |
| 2021 | ML-PersRef: A Machine Learning-based Personalized Multimodal Fusion Approach for Referencing Outside Objects From a Moving VehicleabstractOver the past decades, the addition of hundreds of sensors to modern vehicles has led to an exponential increase in their capabilities. This allows for novel approaches to interaction with the vehicle that go beyond traditional touch-based and voice command approaches, such as emotion recognition, head rotation, eye gaze, and pointing gestures. Although gaze and pointing gestures have been used before for referencing objects inside and outside vehicles, the multimodal interaction and fusion of these gestures have so far not been extensively studied. We propose a novel learning-based multimodal fusion approach for referencing outside-the-vehicle objects while maintaining a long driving route in a simulated environment. The proposed multimodal approaches outperform single-modality approaches in multiple aspects and conditions. Moreover, we also demonstrate possible ways to exploit behavioral differences between users when completing the referencing task to realize an adaptable personalized system for each driver. We propose a personalization technique based on the transfer-of-learning concept for exceedingly small data sizes to enhance prediction and adapt to individualistic referencing behavior. Our code is publicly available at https://github.com/amr-gomaa/ML-PersRef. Amr Gomaa, Guillermo Reyes, Michael Feld |
ICMI | 3 |
| 2021 | Multimodal Fusion Using Deep Learning Applied to Driver's Referencing of Outside-Vehicle ObjectsabstractThere is a growing interest in more intelligent natural user interaction with the car. Hand gestures and speech are already being applied for driver-car interaction. Moreover, multimodal approaches are also showing promise in the automotive industry. In this paper, we utilize deep learning for a multimodal fusion network for referencing objects outside the vehicle. We use features from gaze, head pose and finger pointing simultaneously to precisely predict the referenced objects in different car poses. We demonstrate the practical limitations of each modality when used for a natural form of referencing, specifically inside the car. As evident from our results, we overcome the modality specific limitations, to a large extent, by the addition of other modalities. This work highlights the importance of multimodal sensing, especially when moving towards natural user interaction. Furthermore, our user based analysis shows noteworthy differences in recognition of user behavior depending upon the vehicle pose. Abdul Rafey Aftab, Michael von der Beeck, Steven Rohrhirsch, Benoit Diotte, Michael Feld |
IV | 5 |
| 2020 | You Have a Point There: Object Selection Inside an Automobile Using Gaze, Head Pose and Finger PointingabstractSophisticated user interaction in the automotive industry is a fast emerging topic. Mid-air gestures and speech already have numerous applications for driver-car interaction. Additionally, multimodal approaches are being developed to leverage the use of multiple sensors for added advantages. In this paper, we propose a fast and practical multimodal fusion method based on machine learning for the selection of various control modules in an automotive vehicle. The modalities taken into account are gaze, head pose and finger pointing gesture. Speech is used only as a trigger for fusion. Single modality has previously been used numerous times for recognition of the user's pointing direction. We, however, demonstrate how multiple inputs can be fused together to enhance the recognition performance. Furthermore, we compare different deep neural network architectures against conventional Machine Learning methods, namely Support Vector Regression and Random Forests, and show the enhancements in the pointing direction accuracy using deep learning. The results suggest a great potential for the use of multimodal inputs that can be applied to more use cases in the vehicle. Abdul Rafey Aftab, Michael von der Beeck, Michael Feld |
ICMI | 3 |
| 2020 | Studying Person-Specific Pointing and Gaze Behavior for Multimodal Referencing of Outside Objects from a Moving VehicleabstractHand pointing and eye gaze have been extensively investigated in automotive applications for object selection and referencing. Despite significant advances, existing outside-the-vehicle referencing methods consider these modalities separately. Moreover, existing multimodal referencing methods focus on a static situation, whereas the situation in a moving vehicle is highly dynamic and subject to safety-critical constraints. In this paper, we investigate the specific characteristics of each modality and the interaction between them when used in the task of referencing outside objects (e.g. buildings) from the vehicle. We furthermore explore person-specific differences in this interaction by analyzing individuals' performance for pointing and gaze patterns, along with their effect on the driving task. Our statistical analysis shows significant differences in individual behaviour based on object's location (i.e. driver's right side vs. left side), object's surroundings, driving mode (i.e. autonomous vs. normal driving) as well as pointing and gaze duration, laying the foundation for a user-adaptive approach. Amr Gomaa, Guillermo Reyes, Alexandra Alles, Lydia Rupp, Michael Feld |
ICMI | 5 |
| 2017 | MADMACS - Multiadaptive Dialogue Management in Cyber-Physical EnvironmentsabstractToday, cyber-physical environments (CPEs) are omnipresent-for instance as smart homes, cars, shopping environments, business facilities, Industrie 4.0 factories, and smart cities. Characterized by a large number of individual systems and devices with their sensors and actuators, the interaction paradigm from the user's perspective is shifting towards system-environment interaction. Following this principle, a single user or user groups can freely choose an interaction modality to address the environment, which responds in multiadaptive manner. This paper presents a car repair garage scenario and introduces a dialogue platform, a device platform, as well as several related dialogue management and group interaction technologies. Collectively, they represent a major result of the MADMACS project aimed at developing a framework for multiadaptive interaction in and with CPEs. Yannick Körber, Vanessa Hahn, Mohammad Mehdi Moniri, Tim Schwartz, Michael Feld |
Intelligent Environments | 5 |
| 2016 | Incorporating the Driver's Focus of Attention into Automotive Applications in Real Traffic and in Simulator SetupsabstractIn this research we present an application for analyzing a driver's visual focus of attention in both real-life traffic and simulator setups. For this purpose we use the EyeVIUS [1] system featuring stationary eye-tracking and head-tracking functionality. In our real-life experiments, three dimensional representations of both the vehicle's interior and the outside environment are used. A real-time evaluation concerning the object in the driver's visual focus in these environments is then performed. We describe the functionality and the accuracy of our system, which is integrated in a fully functional vehicle in a real traffic environment. Considering the simulator setup, a seamless integration of our system into the OpenDS [2] driving simulator allows to automatically determine whether a user is engaged with virtual content of the simulator or focused on objects in the real world. Mohammad Mehdi Moniri, Dieter Merkel, Michael Feld, Christian Müller 0014 |
Intelligent Environments | 3 |
| 2016 | Combining Speech, Gaze, and Micro-gestures for the Multimodal Control of In-Car FunctionsabstractModern cars are already incredibly smart environments today due to the sheer number of sensors and processors packed into a small space. Likewise, new technologies in human-computer interaction increasingly find their way inside, e.g. eye tracking, speech interaction and gesture recognition. The support of new modalities is promising a reduction of driver distraction and a better handling of an increasing number of functions offered by in-vehicle systems. With multiple modalities to choose from, which can be combined arbitrarily via multimodal fusion, drivers can make a free choice depending on the demands of the situation and their preferences. Our paper presents a prototype in-car system that allows car features (like turning lights and windows) to be controlled by combinations of speech, gaze, and micro-gestures. We propose an interaction concept, sketch our architecture based on a domain-independent multimodal dialogue platform, and draw some first conclusions on the outcome. Robert Neßelrath, Mohammad Mehdi Moniri, Michael Feld |
Intelligent Environments | 3 |
| 2016 | Hybrid Teams of Humans, Robots, and Virtual Agents in a Production SettingabstractThis video paper describes the practical outcome of the first milestone of a project aiming at setting up a so-called Hybrid Team that can accomplish a wide variety of different tasks. In general, the aim is to realize and examine the collaboration of augmented humans with autonomous robots, virtual characters and SoftBots (purely software based agents) working together in a Hybrid Team to accomplish common tasks. The accompanying video shows a customized packaging scenario and can be downloaded from http://hysociatea.dfki.de/?p=441. Tim Schwartz, Michael Feld, Christian Bürckert, Svilen Dimitrov, Joachim Folz, Dieter Hutter, Peter Hevesi, Bernd Kiefer, Hans-Ulrich Krieger, Christoph Lüth, Dennis Mronga, Gerald Pirkl, Thomas Röfer, Torsten Spieldenner, Malte Wirkus, Ingo Zinnikus, Sirko Straube |
Intelligent Environments | 2 |
| 2015 | SiAM - Situation-Adaptive Multimodal Interaction for Innovative Mobility Concepts of the FutureabstractWhat does situation-adaptive technology mean for car drivers and how can it improve their lives? Why is multimodal interaction in the cockpit a critical ingredient? This contribution summarizes several important technological results of the three-year research project SiAM, which investigated these questions. Motivated by the story of an urban commuter, we illustrate three use cases for situation adaptivity: multimodal control of car functions, cognitive load aware interaction with the environment, and a persuasive intermodal travel assistant. Monika Pepik, Mohammad Mehdi Moniri, Robert Neßelrath, Tim Schwartz, Michael Feld, Yannick Körber, Matthieu Deru, Christian Müller 0014 |
Intelligent Environments | 5 |
| 2014 | SiAM-dp: A Platform for the Model-Based Development of Context-Aware Multimodal Dialogue ApplicationsabstractIntelligent Environments are highly interactive by integrating information and communication technology into the physical space. One goal is to provide user interfaces that are adaptive to the user and the environmental context, including the communication modalities. We present a new development platform for multimodal dialogue systems. A development approach based on semantic-models supports the creation of situation aware dialogue applications in a declarative way. Robert Neßelrath, Michael Feld |
Intelligent Environments | 2 |
| 2013 | Generating a Personalized UI for the Car: A User-Adaptive Rendering Architecture
Michael Feld, Gerrit Meixner, Angela Castronovo, Marc Seissler, Balaji Kalyanasundaram |
UMAP | 1 |
| 2012 | Personalized In-Vehicle Information Systems: Building an Application Infrastructure for Smart Cars in Smart SpacesabstractAlthough intelligent components in modern cars help to contribute to safe mobility, they lack important Smart Environment characteristics like personalization. Moreover, today's in-vehicle infotainment systems do not offer any interaction possibility between the passenger and the visible environment around the car. In this paper we combine several technologies and propose an approach on personalized user interaction in urban environments. We present a showcase that points out the interplay of personalized in-vehicle infotainment system and interaction with visible outside environment. Mohammad Mehdi Moniri, Michael Feld, Christian Müller 0014 |
Intelligent Environments | 2 |
| 2012 | Mobile texting: can post-ASR correction solve the issues? an experimental study on gain vs. costsabstractThe next big step in embedded, mobile speech recognition will be to allow completely free input as it is needed for messaging like SMS or email. However, unconstrained dictation remains error-prone, especially when the environment is noisy. In this paper, we compare different methods for improving a given free-text dictation system used to enter textbased messages in embedded mobile scenarios, where distraction, interaction cost, and hardware limitations enforce strict constraints over traditional scenarios. We present a corpus-based evaluation, measuring the trade-off between improvement of the word error rate versus the interaction steps that are required under various parameters. Results show that by post-processing the output of a "black box" speech recognizer (e.g. a web-based speech recognition service), a reduction of word error rate by 55% (10.3% abs.) can be obtained. For further error reduction, however, a richer representation of the original hypotheses (e.g. lattice) is necessary. Michael Feld, Saeedeh Momtazi, Farina Freigang, Dietrich Klakow, Christian Müller 0014 |
IUI | 1 |
| 2011 | The automotive ontology: managing knowledge inside the vehicle and sharing it between carsabstractCars have been increasingly equipped with technology, meeting the demand of people for safety, connectivity, and comfort. Upcoming HMIs provide access to in-car systems and web services in a personalized manner that facilitates a large array of functionality even while driving, with other passengers also benefiting from an enhanced experience. Such intelligent applications however depend on a solid basis to be effective: Personalization, adaptive HMI, situation-aware intelligent systems -- either of these require semantic knowledge about the user, the vehicle, the current driving situation. Advanced functions coexist with sensors, other functions, and even other vehicles. In such an environment, collaboration can be highly beneficial. Obtaining a common understanding of knowledge and providing a platform to exchange it is essential in order to reach the next level of intelligent in-car systems. This work describes the Automotive Ontology, which is located at the core of such an open platform. We give an overview of design areas relevant to automotive applications, as well as meta aspects that facilitate inference and reasoning. Michael Feld, Christian Müller 0014 |
AutomotiveUI | 1 |
| 2010 | Combining regression and classification methods for improving automatic speaker age recognitionabstractWe present a novel approach to automatic speaker age classification, which combines regression and classification to achieve competitive classification accuracy on telephone speech. Support vector machine regression is used to generate finer age estimates, which are combined with the posterior probabilities of well-trained discriminative gender classifiers to predict both the age and gender of a speaker. We show that this combination performs better than direct 7-class classifiers. The regressors and classifiers are trained using longterm features such as pitch and formants, as well as short-term (frame-based) features derived from MAP adaptation of GMMs that were trained on MFCCs. Charl Johannes van Heerden, Etienne Barnard, Marelie H. Davel, Christiaan van der Walt, Ewald van Dyk, Michael Feld, Christian Müller 0014 |
ICASSP | 6 |
| 2010 | Automatic speaker age and gender recognition in the car for tailoring dialog and mobile servicesabstractCar manufacturers are faced with a new challenge. While a new generation of “digital natives ” becomes a new customer group, the problem of aging society is still increasing. This em-phasizes the need of providing flexible in-car dialog that take into account the specific needs and preferences of the respective user (group). Along the lines of this year’s Interspeech motto “Spoken Language Processing for All”, we address the ques-tion how we find out which group the current user belongs to. Michael Feld, Felix Burkhardt, Christian Müller 0014 |
INTERSPEECH | 1 |
| 2010 | 2nd multimodal interfaces for automotive applications (MIAA 2010)abstractThis paper summarizes the main objectives of the 2nd IUI workshop on multimodal interfaces for automotive applications (MIAA 2010). Michael Feld, Christian Müller 0014, Tim Schwartz |
IUI | 1 |
| 2009 | Multilingual speaker age recognition: Regression analyses on the Lwazi corpusabstractMultilinguality represents an area of significant opportunities for automatic speech-processing systems: whereas multilingual societies are commonplace, the majority of speech-processing systems are developed with a single language in mind. As a step towards improved understanding of multilingual speech processing, the current contribution investigates how an important para-linguistic aspect of speech, namely speaker age, depends on the language spoken. In particular, we study how certain speech features affect the performance of an age recognition system for different South African languages in the Lwazi corpus. By optimizing our feature set and performing language-specific tuning, we are working towards true multilingual classifiers. As they are closely related, ASR and dialog systems are likely to benefit from an improved classification of the speaker. In a comprehensive corpus analysis on long-term features, we have identified features that exhibit characteristic behaviors for particular languages. In a follow-up regression experiment, we confirm the suitability of our feature selection for age recognition and present cross-language error rates. The mean absolute error ranges between 7.7 and 12.8 years for same-language predictors and rises to 14.5 years for cross-language predictors. Michael Feld, Etienne Barnard, Charl Johannes van Heerden, Christian Müller 0014 |
ASRU | 1 |