Hannes Schulz

dblp:12/2966 · DBLP profile ↗
← Back
31ranked-venue papers
10as first author
5since 2021 · last 2022
0000-0001-6408-9794ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 10 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Question answering and dialogue systems · 44% Representation and self-supervised learning · 17% Information extraction and text analysis · 13%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%
Theoretical computer science
1 paper
Information theory · 100%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%
Human-computer interaction and pervasive computing
1 paper
Interaction techniques and input · 50% Human-robot interaction · 50%

Topics — the 21 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning
contrastive learning
0.512021
Decomposed Mutual Information Estimation for Contrastive Representation Learning · ICML 2021
Natural language and speech › Question answering and dialogue systems › dialogue generation
dialogue response generation
0.512021
GRTr: Generative-Retrieval Transformers for Data-Efficient Dialogue Domain Adaptation · IEEE ACM Trans. Audio Speech Lang. Process. 2021
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
dialogue state tracking
0.512021
Overview of the Eighth Dialog System Technology Challenge: DSTC8 · IEEE ACM Trans. Audio Speech Lang. Process. 2021
Machine learning › Representation and self-supervised learning › mutual information
mutual information estimation
0.512021
Decomposed Mutual Information Estimation for Contrastive Representation Learning · ICML 2021
Information theory › information measures
mutual information
0.512021
Decomposed Mutual Information Estimation for Contrastive Representation Learning · ICML 2021
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.412019
The KnowRef Coreference Corpus: Removing Gender and Number Cues for Difficult Pronominal Anaphora Resolution · ACL (1) 2019
Natural language and speech › Information extraction and text analysis
coreference resolution
0.412019
The KnowRef Coreference Corpus: Removing Gender and Number Cues for Difficult Pronominal Anaphora Resolution · ACL (1) 2019
Natural language and speech › Information extraction and text analysis › coreference resolution
pronoun resolution
0.412019
The KnowRef Coreference Corpus: Removing Gender and Number Cues for Difficult Pronominal Anaphora Resolution · ACL (1) 2019
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.412019
Tell, Draw, and Repeat: Generating and Modifying Images Based on Continual Linguistic Instruction · ICCV 2019
Visual content generation and editing › image generation › controllable image generation
interactive image generation
0.412019
Tell, Draw, and Repeat: Generating and Modifying Images Based on Continual Linguistic Instruction · ICCV 2019
Visual content generation and editing › image generation
text-to-image generation
0.412019
Tell, Draw, and Repeat: Generating and Modifying Images Based on Continual Linguistic Instruction · ICCV 2019
Natural language and speech › Question answering and dialogue systems
task-oriented dialogue
0.312018
Towards Deep Conversational Recommendations · NeurIPS 2018
Recommender systems › interactive recommendation
conversational recommendation
0.312018
Towards Deep Conversational Recommendations · NeurIPS 2018
Computer vision › 3D vision
object pose estimation
0.212015
RGB-D object recognition and pose estimation based on pre-trained convolutional neural network features · ICRA 2015
Computer vision › Image recognition and object detection
object recognition
0.212015
RGB-D object recognition and pose estimation based on pre-trained convolutional neural network features · ICRA 2015
Computer vision › Image recognition and object detection › object recognition › multimodal object recognition
RGB-D object recognition
0.212015
RGB-D object recognition and pose estimation based on pre-trained convolutional neural network features · ICRA 2015
Computer vision › 3D vision › pose estimation › multi-modal pose estimation
RGB-D pose estimation
0.212015
RGB-D object recognition and pose estimation based on pre-trained convolutional neural network features · ICRA 2015
Natural language and speech › Question answering and dialogue systems › multimodal dialogue system
audio-visual scene-aware dialog
0.112021
Overview of the Eighth Dialog System Technology Challenge: DSTC8 · IEEE ACM Trans. Audio Speech Lang. Process. 2021
Natural language and speech › Question answering and dialogue systems
dialogue generation
0.112021
Decomposed Mutual Information Estimation for Contrastive Representation Learning · ICML 2021
Human-robot interaction › teleoperation
gesture-based robot control
0.112012
Real time interaction with mobile robots using hand gestures · HRI 2012
Interaction techniques and input
gesture input
0.112012
Real time interaction with mobile robots using hand gestures · HRI 2012

Methods — techniques the papers use, named apart from their topics

contrastive lower bound · 1.0chain rule of mutual information · 1.0recurrent neural network · 0.8convolutional neural network · 0.5retrieval fallback · 0.5pre-training and fine-tuning · 0.5generative-retrieval transformer · 0.5end-to-end dialog modeling · 0.5data augmentation · 0.4antecedent switching · 0.4sentiment analysis · 0.3neural architecture · 0.3depth imaging · 0.1
YearPublicationVenuePosition
2022 Semi-automatic Integrated Safety and Security Analysis for Automotive Systems
abstract
147
Markus Fockel, David Schubert, Roman Trentinaglia, Hannes Schulz, Wolfgang Kirmair
MODELSWARD4
2021 Decomposed Mutual Information Estimation for Contrastive Representation Learning
abstract
Recent contrastive representation learning methods rely on estimating mutual information (MI) between multiple views of an underlying context. E.g., we can derive multiple views of a given image by applying data augmentation, or we can split a sequence into views comprising the past and future of some step in the sequence. Contrastive lower bounds on MI are easy to optimize, but have a strong underestimation bias when estimating large amounts of MI. We propose decomposing the full MI estimation problem into a sum of smaller estimation problems by splitting one of the views into progressively more informed subviews and by applying the chain rule on MI between the decomposed views. This expression contains a sum of unconditional and conditional MI terms, each measuring modest chunks of the total MI, which facilitates approximation via contrastive bounds. To maximize the sum, we formulate a contrastive lower bound on the conditional MI which can be approximated efficiently. We refer to our general approach as Decomposed Estimation of Mutual Information (DEMI). We show that DEMI can capture a larger amount of MI than standard non-decomposed contrastive bounds in a synthetic setting, and learns better representations in a vision domain and for dialogue generation.
Alessandro Sordoni, Nouha Dziri, Hannes Schulz, Geoffrey J. Gordon, Philip Bachman, Remi Tachet des Combes
ICML3
2021 Overview of the Eighth Dialog System Technology Challenge: DSTC8
abstract
This paper introduces the Eighth Dialog System Technology Challenge. In line with recent challenges, the eighth edition focuses on applying end-to-end dialog technologies in a pragmatic way for multi-domain task-completion, noetic response selection, audio visual scene-aware dialog, and schema-guided dialog state tracking tasks. This paper describes the task definition, provided datasets, baselines and evaluation set-up for each track. We also summarize the results of the submitted systems to highlight the overall trends of the state-of-the-art technologies for the tasks.
Seokhwan Kim, Michel Galley, R. Chulaka Gunasekara, Adam Atkinson, Baolin Peng, Hannes Schulz, Jianfeng Gao 0001, Jinchao Li, Mahmoud Adada, Minlie Huang, Luis A. Lastras, Jonathan K. Kummerfeld, Walter S. Lasecki, Chiori Hori, Anoop Cherian, Tim K. Marks, Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara
IEEE ACM Trans. Audio Speech Lang. Process.7
2021 Editorial: Special Issue on the Eighth Dialog System Technology Challenge
Seokhwan Kim, Hannes Schulz, R. Chulaka Gunasekara, Chiori Hori, Abhinav Rastogi, Luis Fernando D'Haro
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 GRTr: Generative-Retrieval Transformers for Data-Efficient Dialogue Domain Adaptation
abstract
Domain adaptation has recently become a key problem in dialogue systems research. Deep learning, while being the preferred technique for modeling such systems, works best given massive training data. However, in real-world scenarios, such resources are rarely available for new domains, and the ability to train with a few dialogue examples can be considered essential. Pre-training on large data sources and adapting to the target data has become the standard method for few-shot problems within the deep learning framework. In this paper, we presentgrtr, a hybrid generative-retrieval model based on the large-scale general-purpose language model GPT[2] fine-tuned to the multi-domainmetalwoz dataset. In addition to robust and diverse response generation provided by the GPT[2], our model is able to estimate generation confidence, and is equipped with retrieval logic as a fallback for the cases when the estimate is low.grtr is the winning entry at the fast domain adaptation task of DSTC-8 in human evaluation ($>$4% improvement over the 2nd place system). It also attains superior performance to a series of baselines on automated metrics onmetalwoz andmultiwoz, a multi-domain dataset of goal-oriented dialogues. In this paper, we also conduct a study ofgrtr's performance in the setup of limited adaptation data, evaluating the model's overall response prediction performance onmetalwoz and goal-oriented performance onmultiwoz.
Igor Shalyminov, Alessandro Sordoni, Adam Atkinson, Hannes Schulz
IEEE ACM Trans. Audio Speech Lang. Process.4
2020 Fast Domain Adaptation for Goal-Oriented Dialogue Using a Hybrid Generative-Retrieval Transformer
abstract
Goal-oriented dialogue systems are now widely adopted in industry, where practical aspects of using them becomes of key importance. As such, it is expected from such systems to fit into a rapid prototyping cycle for new products and domains. For data-driven dialogue systems (especially those based on deep learning) that amounts to maintaining production-level performance having been provided with a few `seed' dialogue examples, normally referred to as data efficiency.With extremely data-dependent deep learning methods, the most promising way to achieve practical data efficiency is transfer learning-i.e., leveraging a greater, highly represented data source for training a base model, then fine-tuning it to available in-domain data.In this paper, we present a hybrid generative-retrieval model that can be trained using transfer learning. By using GPT-2 as the base model and fine-tuning it to the multidomain MetaLWOz dataset, we obtain a robust dialogue model able to perform both response generation and ranking1. Combining both, it outperforms several competitive generative-only and retrieval-only baselines, measured by language modeling quality on MetaLWOz as well as in goal- oriented metrics (Intent/Slot Fl-scores) on the MultiWoz corpus.
Igor Shalyminov, Alessandro Sordoni, Adam Atkinson, Hannes Schulz
ICASSP4
2019 The KnowRef Coreference Corpus: Removing Gender and Number Cues for Difficult Pronominal Anaphora Resolution
abstract
We introduce a new benchmark for coreference resolution and NLI, KnowRef, that targets common-sense understanding and world knowledge. Previous coreference resolution tasks can largely be solved by exploiting the number and gender of the antecedents, or have been handcrafted and do not reflect the diversity of naturally occurring text. We present a corpus of over 8,000 annotated text passages with ambiguous pronominal anaphora. These instances are both challenging and realistic. We show that various coreference systems, whether rule-based, feature-rich, or neural, perform significantly worse on the task than humans, who display high inter-annotator agreement. To explain this performance gap, we show empirically that state-of-the art models often fail to capture context, instead relying on the gender or number of candidate antecedents to make a decision. We then use problem-specific insights to propose a data-augmentation trick called antecedent switching to alleviate this tendency in models. Finally, we show that antecedent switching yields promising results on other tasks as well: we use it to achieve state-of-the-art results on the GAP coreference task.
Ali Emami, Paul Trichelair, Adam Trischler, Kaheer Suleman, Hannes Schulz, Jackie Chi Kit Cheung
ACL (1)5
2019 Tell, Draw, and Repeat: Generating and Modifying Images Based on Continual Linguistic Instruction
abstract
Conditional text-to-image generation is an active area of research, with many possible applications. Existing research has primarily focused on generating a single image from available conditioning information in one step. One practical extension beyond one-step generation is a system that generates an image iteratively, conditioned on ongoing linguistic input or feedback. This is significantly more challenging than one-step generation tasks, as such a system must understand the contents of its generated images with respect to the feedback history, the current feedback, as well as the interactions among concepts present in the feedback history. In this work, we present a recurrent image generation model which takes into account both the generated output up to the current step as well as all past instructions for generation. We show that our model is able to generate the background, add new objects, and apply simple transformations to existing objects. We believe our approach is an important step toward interactive generation. Code and data is available at: https://www.microsoft.com/en-us/research/project/generative-neural-visual-artist-geneva/.
Alaaeldin El-Nouby, Shikhar Sharma 0001, Hannes Schulz, R. Devon Hjelm, Layla El Asri, Samira Ebrahimi Kahou, Yoshua Bengio, Graham W. Taylor
ICCV3
2018 Towards Deep Conversational Recommendations
abstract
There has been growing interest in using neural networks and deep learning techniques to create dialogue systems. Conversational recommendation is an interesting setting for the scientific exploration of dialogue with natural language as the associated discourse involves goal-driven dialogue that often transforms naturally into more free-form chat. This paper provides two contributions. First, until now there has been no publicly available large-scale data set consisting of real-world dialogues centered around recommendations. To address this issue and to facilitate our exploration here, we have collected ReDial, a data set consisting of over 10,000 conversations centered around the theme of providing movie recommendations. We make this data available to the community for further research. Second, we use this dataset to explore multiple facets of conversational recommendations. In particular we explore new neural architectures, mechanisms and methods suitable for composing conversational recommendation systems. Our dataset allows us to systematically probe model sub-components addressing different parts of the overall problem domain ranging from: sentiment analysis and cold-start recommendation generation to detailed aspects of how natural language is used in this setting in the real world. We combine such sub-components into a full-blown dialogue system and examine its behavior.
Raymond Li, Samira Ebrahimi Kahou, Hannes Schulz, Vincent Michalski, Laurent Charlin, Christopher Joseph Pal
NeurIPS3
2017 Frames: a corpus for adding memory to goal-oriented dialogue systems
abstract
This paper proposes a new dataset, Frames, composed of 1369 human-human dialogues with an average of 15 turns per dialogue.This corpus contains goal-oriented dialogues between users who are given some constraints to book a trip and assistants who search a database to find appropriate trips.The users exhibit complex decision-making behaviour which involve comparing trips, exploring different options, and selecting among the trips that were discussed during the dialogue.To drive research on dialogue systems towards handling such behaviour, we have annotated and released the dataset and we propose in this paper a task called frame tracking.This task consists of keeping track of different semantic frames throughout each dialogue.We propose a rule-based baseline and analyse the frame tracking task through this baseline.
Layla El Asri, Hannes Schulz, Shikhar Sharma 0001, Jeremie Zumer, Justin Harris, Emery Fine, Rahul Mehrotra, Kaheer Suleman
SIGDIAL Conference2
2017 Object class segmentation of RGB-D video using recurrent convolutional neural networks
Mircea Serban Pavel, Hannes Schulz, Sven Behnke
Neural Networks2
2016 Semantic segmentation priors for object discovery
abstract
Reliable object discovery in realistic indoor scenes is a necessity for many computer vision and service robot applications. In these scenes, semantic segmentation methods have made huge advances in recent years. Such methods can provide useful prior information for object discovery by removing false positives and by delineating object boundaries. We propose a novel method that combines bottom-up object discovery and semantic priors for producing generic object candidates in RGB-D images. We use a deep learning method for semantic segmentation to classify colour and depth superpixels into meaningful categories. Separately for each category, we use saliency to estimate the location and scale of objects, and superpixels to find their precise boundaries. Finally, object candidates of all categories are combined and ranked. We evaluate our approach on the NYU Depth V2 dataset and show that we outperform other state-of-the-art object discovery methods in terms of recall.
Germán Martín García, Farzad Husain, Hannes Schulz, Simone Frintrop, Carme Torras, Sven Behnke
ICPR3
2016 Policy Networks with Two-Stage Training for Dialogue Systems
abstract
In this paper, we propose to use deep policy networks which are trained with an advantage actor-critic method for statistically optimised dialogue systems. First, we show that, on summary state and action spaces, deep Reinforcement Learning (RL) outperforms Gaussian Processes methods. Summary state and action spaces lead to good performance but require pre-engineering effort, RL knowledge, and domain expertise. In order to remove the need to define such summary spaces, we show that deep RL can also be trained efficiently on the original state and action spaces. Dialogue systems based on partially observable Markov decision processes are known to require many dialogues to train, which makes them unappealing for practical deployment. We show that a deep RL method based on an actor-critic architecture can exploit a small amount of data very efficiently. Indeed, with only a few hundred dialogues collected with a handcrafted policy, the actor-critic deep learner is considerably bootstrapped from a combination of supervised and batch RL. In addition, convergence to an optimal policy is significantly sped up compared to other deep RL methods initialized on the data with batch RL. All experiments are performed on a restaurant domain derived from the Dialogue State Tracking Challenge 2 (DSTC2) dataset.
Mehdi Fatemi, Layla El Asri, Hannes Schulz, Jing He 0010, Kaheer Suleman
SIGDIAL Conference3
2015 Depth and height aware semantic RGB-D perception with convolutional neural networks
Hannes Schulz, Nico Höft, Sven Behnke
ESANN1
2015 RGB-D object recognition and pose estimation based on pre-trained convolutional neural network features
abstract
Object recognition and pose estimation from RGB-D images are important tasks for manipulation robots which can be learned from examples. Creating and annotating datasets for learning is expensive, however. We address this problem with transfer learning from deep convolutional neural networks (CNN) that are pre-trained for image categorization and provide a rich, semantically meaningful feature set. We incorporate depth information, which the CNN was not trained with, by rendering objects from a canonical perspective and colorizing the depth channel according to distance from the object center. We evaluate our approach on the Washington RGB-D Objects dataset, where we find that the generated feature set naturally separates classes and instances well and retains pose manifolds. We outperform state-of-the-art on a number of subtasks and show that our approach can yield superior results when only little training data is available.
Max Schwarz, Hannes Schulz, Sven Behnke
ICRA2
2015 Recurrent convolutional neural networks for object-class segmentation of RGB-D video
abstract
Object-class segmentation is a computer vision task which requires labeling each pixel of an image with the class of the object it belongs to. Deep convolutional neural networks (DNN) are able to learn and exploit local spatial correlations required for this task. They are, however, restricted by their small, fixed-sized filters, which limits their ability to learn long-range dependencies. Recurrent Neural Networks (RNN), on the other hand, do not suffer from this restriction. Their iterative interpretation allows them to model long-range dependencies by propagating activity. This property might be especially useful when labeling video sequences, where both spatial and temporal long-range dependencies occur. In this work, we propose novel RNN architectures for object-class segmentation. We investigate three ways to consider past and future context in the prediction process by comparing networks that process the frames one by one with networks that have access to the whole sequence. We evaluate our models on the challenging NYU Depth v2 dataset for object-class segmentation and obtain competitive results.
Mircea Serban Pavel, Hannes Schulz, Sven Behnke
IJCNN2
2015 Two-layer contractive encodings for learning stable nonlinear features
Hannes Schulz, Kyunghyun Cho, Tapani Raiko, Sven Behnke
Neural Networks1
2014 Structured Prediction for Object Detection in Deep Neural Networks
Hannes Schulz, Sven Behnke
ICANN1
2013 Two-Layer Contractive Encodings with Shortcuts for Semi-supervised Learning
Hannes Schulz, Kyunghyun Cho, Tapani Raiko, Sven Behnke
ICONIP (1)1
2012 Learning Object-Class Segmentation with Convolutional Neural Networks
Hannes Schulz, Sven Behnke
ESANN1
2012 Real time interaction with mobile robots using hand gestures
abstract
We developed a robust real time hand gesture based interaction system to effectively communicate with a mobile robot which can operate in an outdoor environment. The system enables the user to operate a mobile robot using hand gesture based commands. In particular the system offers direct on site interaction providing better perception of environment to the user. To overcome the illumination challenges in outdoors, the system operates on depth images. Processed depth images are given as input to a convolutional neural network which is trained to detect static hand gestures.
Kishore Reddy Konda, Achim Königs, Hannes Schulz, Dirk Schulz 0001
HRI3
2012 Learning Two-Layer Contractive Encodings
Hannes Schulz, Sven Behnke
ICANN (1)1
2011 RoboCup 2011 Humanoid League Winners
Daniel D. Lee, Seung-Joon Yi, Stephen G. McGill, Sven Behnke, Marcell Missura, Hannes Schulz, Dennis W. Hong, Jeakweon Han, Michael A. Hopkins
RoboCup7
2011 Exploiting local structure in Boltzmann machines
Hannes Schulz, Andreas C. Müller 0001, Sven Behnke
Neurocomputing1
2010 Exploiting local structure in stacked Boltzmann machines
Hannes Schulz, Andreas C. Müller 0001, Sven Behnke
ESANN1
2010 Accelerating Large-Scale Convolutional Neural Networks with Parallel Graphics Multiprocessors
Dominik Scherer, Hannes Schulz, Sven Behnke
ICANN (3)2
2010 Topological features in locally connected RBMs
abstract
Unsupervised learning algorithms find ways to model latent structure present in the data. These latent structures can then serve as a basis for supervised classification methods. A common choice for unsupervised feature discovery is the Restricted Boltzmann Machine (RBM). Since the RBM is a general purpose learning machine, it is not particularly tailored for image data. Representations found by RBMs are consequently not image-like. Since it is essential to exploit the known topological structure for image analysis, it is desirable not to discard the topology property when learning new representations. Then, the same learning methods can be applied to the latent representation in a hierarchical manner. In this work, we propose a modification to the learning rule of locally connected RBMs, which ensures that topological image structure is preserved in the latent representation. To this end, we use a Gaussian kernel to transfer topological properties of the image space to the feature space. The learned model is then used as an initialization for a neural network trained to classify the images. We evaluate our approach on the MNIST and Caltech 101 datasets and demonstrate that we are able to learn topological feature maps.
Andreas C. Müller 0001, Hannes Schulz, Sven Behnke
IJCNN2
2010 Utilizing the Structure of Field Lines for Efficient Soccer Robot Localization
Hannes Schulz, Weichao Liu, Jörg Stückler, Sven Behnke
RoboCup1
2009 ILP, the Blind, and the Elephant: Euclidean Embedding of Co-proven Queries
Hannes Schulz, Kristian Kersting, Andreas Karwath
ILP1
2009 Designing User Interfaces for Smart-Applications for Operating Rooms and Intensive Care Units
Martin Christof Kindsmüller, Maral Haar, Hannes Schulz, Michael Herczeg
INTERACT (2)3
2006 Impact of Coprocessors on a Multithreaded Processor Design Using Prioritized Threads
abstract
Recently, multithreading became a standard technique to improve the processor utilization and system performance. Hardware support is provided for coarse-grained as well as simultaneous multithreading. In particular, embedded devices combine processor cores and varying sets of coprocessors to fulfill the requirements of their dedicated application field. In this paper, a simultaneous multithreaded processor is investigated that applies dynamic priorities for each thread on the instruction level. By means of a synchronization coprocessor, priorities of threads are dynamically adapted when other threads have to wait for a given thread. Based on simulations of a network-processing workload, two strategies of dynamic priority adaptation are evaluated and compared with static prioritization. As a result, performance gain can be shown.
Carsten Albrecht, Andreas C. Döring, Frank Penczek, Torben Schneider, Hannes Schulz
PDP5