Stefan Riezler

dblp:33/2631 · DBLP profile ↗
← Back
62ranked-venue papers
11as first author
12since 2021 · last 2024
0000-0002-3392-0191ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 55 · 11 first-author · 9 since 2021Databases, data management, data science and information retrieval · 4Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2024 VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street View
abstract
Incremental decision making in real-world environments is one of the most challenging tasks in embodied artificial intelligence. One particularly demanding scenario is Vision and Language Navigation (VLN) which requires visual and natural language understanding as well as spatial and temporal reasoning capabilities. The embodied agent needs to ground its understanding of navigation instructions in observations of a real-world environment like Street View. Despite the impressive results of LLMs in other research areas, it is an ongoing problem of how to best connect them with an interactive visual environment. In this work, we propose VELMA, an embodied LLM agent that uses a verbalization of the trajectory and of visual environment observations as contextual prompt for the next action. Visual information is verbalized by a pipeline that extracts landmarks from the human written navigation instructions and uses CLIP to determine their visibility in the current panorama view. We show that VELMA is able to successfully follow navigation instructions in Street View with only two in-context examples. We further finetune the LLM agent on a few thousand examples and achieve around 25% relative improvement in task completion over the previous state-of-the-art for two datasets.
Raphael Schumann, Wanrong Zhu, Weixi Feng, Tsu-Jui Fu, Stefan Riezler, William Yang Wang
AAAI5
2024 Prompting Large Language Models with Human Error Markings for Self-Correcting Machine Translation
abstract
While large language models (LLMs) pre-trained on massive amounts of unpaired language data have reached the state-of-the-art in machine translation (MT) of general domain texts, post-editing (PE) is still required to correct errors and to enhance term translation quality in specialized domains. In this paper we present a pilot study of enhancing translation memories (TM) produced by PE (source segments, machine translations, and reference translations, henceforth called PE-TM) for the needs of correct and consistent term translation in technical domains. We investigate a light-weight two-step scenario where at inference time, a human translator marks errors in the first translation step, and in a second step a few similar examples are extracted from the PE-TM to prompt an LLM. Our experiment shows that the additional effort of augmenting translations with human error markings guides the LLM to focus on a correction of the marked errors, yielding consistent improvements over automatic PE (APE) and MT from scratch.
Nathaniel Berger, Stefan Riezler, Miriam Exel, Matthias Huck
EAMT (1)2
2024 Text-to-OverpassQL: A Natural Language Interface for Complex Geodata Querying of OpenStreetMap
abstract
Abstract We present Text-to-OverpassQL, a task designed to facilitate a natural language interface for querying geodata from OpenStreetMap (OSM). The Overpass Query Language (OverpassQL) allows users to formulate complex database queries and is widely adopted in the OSM ecosystem. Generating Overpass queries from natural language input serves multiple use-cases. It enables novice users to utilize OverpassQL without prior knowledge, assists experienced users with crafting advanced queries, and enables tool-augmented large language models to access information stored in the OSM database. In order to assess the performance of current sequence generation models on this task, we propose OverpassNL,1 a dataset of 8,352 queries with corresponding natural language inputs. We further introduce task specific evaluation metrics and ground the evaluation of the Text-to-OverpassQL task by executing the queries against the OSM database. We establish strong baselines by finetuning sequence-to-sequence models and adapting large language models with in-context examples. The detailed evaluation reveals strengths and weaknesses of the considered learning strategies, laying the foundations for further research into the Text-to-OverpassQL task.
Michael Staniek, Raphael Schumann, Maike Züfle, Stefan Riezler
Trans. Assoc. Comput. Linguistics4
2023 Early Prediction of Sepsis Using Time Series Forecasting
abstract
Sepsis is a serious complication of an infection. Without quick treatment it can lead to organ failure and death. Early detection and treatment of sepsis can thus improve patient outcomes. Yet, their effectiveness often relies on awareness and acceptance of said procedures. In this work, we implement sepsis check based on a widely accepted guideline for sepsis recognition (Sepsis-3). Our implementation achieved F-score as high as 0.874. In addition to implementing the ruled-based approach to early sepsis detection, we use an existing data-driven transformer-based STraTS model [1] for time-series forecasting to support sepsis check and directly predicting sepsis label using 24-hour patient data in a fully data-driven setup. The advantage of time series forecasting is improved handling of missing data and the potential of applying the Sepsis-3 definition to unobserved forecast data. Additionally, we attempt to improve STraTS model by integrating a clinical text embedding module to enable multimodal learning. Both the original STraTS model and our refined STraTS+Text model perform good in both forecasting (masked MSE, mean squared error at approximately 5.24) and classification task (ROC-AUC, area under receiver operating characteristic curve at approximately 0.89).
Jinghua Xu, Natalia Minakova, Pablo Ortega Sanchez, Stefan Riezler
e-Science4
2023 Enhancing Supervised Learning with Contrastive Markings in Neural Machine Translation Training
abstract
Supervised learning in Neural Machine Translation (NMT) standardly follows a teacher forcing paradigm where the conditioning context in the model’s prediction is constituted by reference tokens, instead of its own previous predictions. In order to alleviate this lack of exploration in the space of translations, we present a simple extension of standard maximum likelihood estimation by a contrastive marking objective. The additional training signals are extracted automatically from reference translations by comparing the system hypothesis against the reference, and used for up/down-weighting correct/incorrect tokens. The proposed new training procedure requires one additional translation pass over the training set, and does not alter the standard inference setup. We show that training with contrastive markings yields improvements on top of supervised learning, and is especially useful when learning from postedits where contrastive markings indicate human error corrections to the original hypotheses.
Nathaniel Berger, Miriam Exel, Matthias Huck, Stefan Riezler
EAMT4
2023 Make More of Your Data: Minimal Effort Data Augmentation for Automatic Speech Recognition and Translation
abstract
Data augmentation is a technique to generate new training data based on existing data. We evaluate the simple and cost-effective method of concatenating the original data examples to build new training instances. Continued training with such augmented data is able to improve off-the-shelf Transformer and Conformer models that were optimized on the original data only. We demonstrate considerable improvements on the LibriSpeech-960h test sets (WER 2.83 and 6.87 for test-clean and test-other), which carry over to models combined with shallow fusion (WER 2.55 and 6.27). Our method of continued training also leads to improvements of up to 0.9 WER on the ASR part of CoVoST-2 for four non-English languages, and we observe that the gains are highly dependent on the size of the original training data. We compare different concatenation strategies and found that our method does not need speaker information to achieve its improvements. Finally, we demonstrate on two datasets that our methods also works for speech translation tasks.
Tsz Kin Lam, Shigehiko Schamoni, Stefan Riezler
ICASSP3
2023 Towards Inferential Reproducibility of Machine Learning Research
Michael Hagmann, Philipp Meier, Stefan Riezler
ICLR3
2022 Analyzing Generalization of Vision and Language Navigation to Unseen Outdoor Areas
abstract
Vision and language navigation (VLN) is a challenging visually-grounded language understanding task.Given a natural language navigation instruction, a visual agent interacts with a graph-based environment equipped with panorama images and tries to follow the described route.Most prior work has been conducted in indoor scenarios where best results were obtained for navigation on routes that are similar to the training routes, with sharp drops in performance when testing on unseen environments.We focus on VLN in outdoor scenarios and find that in contrast to indoor VLN, most of the gain in outdoor VLN on unseen data is due to features like junction type embedding or heading delta that are specific to the respective environment graph, while image information plays a very minor role in generalizing VLN to unseen outdoor areas.These findings show a bias to specifics of graph representations of urban environments, demanding that VLN tasks grow in scale and diversity of geographical environments.1
Raphael Schumann, Stefan Riezler
ACL (1)2
2021 Generating Landmark Navigation Instructions from Maps as a Graph-to-Text Problem
abstract
Raphael Schumann, Stefan Riezler. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Raphael Schumann, Stefan Riezler
ACL/IJCNLP (1)2
2021 Don't Search for a Search Method - Simple Heuristics Suffice for Adversarial Text Attacks
abstract
Recently more attention has been given to adversarial attacks on neural networks for natural language processing (NLP).A central research topic has been the investigation of search algorithms and search constraints, accompanied by benchmark algorithms and tasks.We implement an algorithm inspired by zeroth order optimization-based attacks and compare with the benchmark results in the TextAttack framework.Surprisingly, we find that optimizationbased methods do not yield any improvement in a constrained setup and slightly benefit from approximate gradient information only in unconstrained setups where search spaces are larger.In contrast, simple heuristics exploiting nearest neighbors without querying the target function yield substantial success rates in constrained setups, and nearly full success rate in unconstrained setups, at an order of magnitude fewer queries.We conclude from these results that current TextAttack benchmark tasks are too easy and constraints are too strict, preventing meaningful research on black-box adversarial text attacks.
Nathaniel Berger, Stefan Riezler, Sebastian Ebert, Artem Sokolov 0001
EMNLP (1)2
2021 Cascaded Models with Cyclic Feedback for Direct Speech Translation
abstract
Direct speech translation describes a scenario where only speech inputs and corresponding translations are available. Such data are notoriously limited. We present a technique that allows cascades of automatic speech recognition (ASR) and machine translation (MT) to exploit in-domain direct speech translation data in addition to out-of-domain MT and ASR data. After pre-training MT and ASR, we use a feed-back cycle where the downstream performance of the MT system is used as a signal to improve the ASR system by self-training, and the MT component is fine-tuned on multiple ASR outputs, making it more tolerant towards spelling variations. A comparison to end-to-end speech translation using components of identical architecture and the same data shows gains of up to 3.8 BLEU points on LibriVoxDeEn and up to 5.1 BLEU points on CoVoST for German-to-English speech translation.
Tsz Kin Lam, Shigehiko Schamoni, Stefan Riezler
ICASSP3
2021 On-the-Fly Aligned Data Augmentation for Sequence-to-Sequence ASR
abstract
We propose an on-the-fly data augmentation method for automatic speech recognition (ASR) that uses alignment information to generate effective training samples. Our method, called Aligned Data Augmentation (ADA) for ASR, replaces transcribed tokens and the speech representations in an aligned manner to generate previously unseen training pairs. The speech representations are sampled from an audio dictionary that has been extracted from the training corpus and inject speaker variations into the training examples. The transcribed tokens are either predicted by a language model such that the augmented data pairs are semantically close to the original data, or randomly sampled. Both strategies result in training pairs that improve robustness in ASR training. Our experiments on a Seq-to-Seq architecture show that ADA can be applied on top of SpecAugment, and achieves about 9-23% and 4-15% relative improvements in WER over SpecAugment alone on LibriSpeech 100h and LibriSpeech 960h test datasets, respectively.
Tsz Kin Lam, Mayumi Ohta, Shigehiko Schamoni, Stefan Riezler
Interspeech4
2020 Embedding Meta-Textual Information for Improved Learning to Rank
abstract
Neural approaches to learning term embeddings have led to improved computation of similarity and ranking in information retrieval (IR).So far neural representation learning has not been extended to meta-textual information that is readily available for many IR tasks, for example, patent classes in prior-art retrieval, topical information in Wikipedia articles, or product categories in e-commerce data.We present a framework that learns embeddings for meta-textual categories, and optimizes a pairwise ranking objective for improved matching based on combined embeddings of textual and meta-textual information.We show considerable gains in an experimental evaluation on cross-lingual retrieval in the Wikipedia domain for three language pairs, and in the Patent domain for one language pair.Our results emphasize that the mode of combining different types of information is crucial for model improvement.
Toshitaka Kuwa, Shigehiko Schamoni, Stefan Riezler
COLING3
2020 Correct Me If You Can: Learning from Error Corrections and Markings
abstract
Sequence-to-sequence learning involves a trade-off between signal strength and annotation cost of training data. For example, machine translation data range from costly expert-generated translations that enable supervised learning, to weak quality-judgment feedback that facilitate reinforcement learning. We present the first user study on annotation cost and machine learnability for the less popular annotation mode of error markings. We show that error markings for translations of TED talks from English to German allow precise credit assignment while requiring significantly less human effort than correcting/post-editing, and that error-marked data can be used successfully to fine-tune neural machine translation models.
Julia Kreutzer, Nathaniel Berger, Stefan Riezler
EAMT3
2020 LibriVoxDeEn: A Corpus for German-to-English Speech Translation and German Speech Recognition
abstract
We present a corpus of sentence-aligned triples of German audio, German text, and English translation, based on German audio books. The speech translation data consist of 110 hours of audio material aligned to over 50k parallel sentences. An even larger dataset comprising 547 hours of German speech aligned to German text is available for speech recognition. The audio data is read speech and thus low in disfluencies. The quality of audio and sentence alignments has been checked by a manual evaluation, showing that speech alignment quality is in general very high. The sentence alignment quality is comparable to well-used parallel translation data and can be adjusted by cutoffs on the automatic alignment score. To our knowledge, this corpus is to date the largest resource for German speech recognition and for end-to-end German-to-English speech translation.
Benjamin Beilharz, Sariya Karimova, Stefan Riezler
LREC4
2019 Self-Regulated Interactive Sequence-to-Sequence Learning
abstract
Not all types of supervision signals are created equal: Different types of feedback have different costs and effects on learning. We show how self-regulation strategies that decide when to ask for which kind of feedback from a teacher (or from oneself) can be cast as a learning-to-learn problem leading to improved cost-aware sequence-to-sequence learning. In experiments on interactive neural machine translation, we find that the self-regulator discovers an $ε$-greedy strategy for the optimal cost-quality trade-off by mixing different feedback types including corrections, error markups, and self-supervision. Furthermore, we demonstrate its robustness under domain shift and identify it as a promising alternative to active learning.
Julia Kreutzer, Stefan Riezler
ACL (1)2
2019 Interactive-Predictive Neural Machine Translation through Reinforcement and Imitation
Tsz Kin Lam, Shigehiko Schamoni, Stefan Riezler
MTSummit (1)3
2019 Leveraging implicit expert knowledge for non-circular machine learning in sepsis prediction
Shigehiko Schamoni, Holger A. Lindner, Verena Schneider-Lindner, Manfred Thiel, Stefan Riezler
Artif. Intell. Medicine5
2019 Learning Neural Sequence-to-Sequence Models from Weak Feedback with Bipolar Ramp Loss
abstract
In many machine learning scenarios, supervision by gold labels is not available and conse quently neural models cannot be trained directly by maximum likelihood estimation. In a weak supervision scenario, metric-augmented objectives can be employed to assign feedback to model outputs, which can be used to extract a supervision signal for training. We present several objectives for two separate weakly supervised tasks, machine translation and semantic parsing. We show that objectives should actively discourage negative outputs in addition to promoting a surrogate gold structure. This notion of bipolarity is naturally present in ramp loss objectives, which we adapt to neural models. We show that bipolar ramp loss objectives outperform other non-bipolar ramp loss objectives and minimum risk training on both weakly supervised tasks, as well as on a supervised machine translation task. Additionally, we introduce a novel token-level ramp loss objective, which is able to outperform even the best sequence-level ramp loss on both weakly supervised tasks.
Laura Jehl, Carolin Lawrence, Stefan Riezler
Trans. Assoc. Comput. Linguistics3
2018 Reliability and Learnability of Human Bandit Feedback for Sequence-to-Sequence Reinforcement Learning
abstract
We present a study on reinforcement learning (RL) from human bandit feedback for sequence-to-sequence learning, exemplified by the task of bandit neural machine translation (NMT).We investigate the reliability of human bandit feedback, and analyze the influence of reliability on the learnability of a reward estimator, and the effect of the quality of reward estimates on the overall RL task.Our analysis of cardinal (5-point ratings) and ordinal (pairwise preferences) feedback shows that their intra-and inter-annotator αagreement is comparable.Best reliability is obtained for standardized cardinal feedback, and cardinal feedback is also easiest to learn and generalize from.Finally, improvements of over 1 BLEU can be obtained by integrating a regressionbased reward estimator trained on cardinal feedback for 800 translations into RL for NMT.This shows that RL is possible even from small amounts of fairly reliable human feedback, pointing to a great potential for applications at larger scale.
Julia Kreutzer, Joshua Uyheng, Stefan Riezler
ACL (1)3
2018 Improving a Neural Semantic Parser by Counterfactual Learning from Human Bandit Feedback
abstract
Counterfactual learning from human bandit feedback describes a scenario where user feedback on the quality of outputs of a historic system is logged and used to improve a target system.We show how to apply this learning framework to neural semantic parsing.From a machine learning perspective, the key challenge lies in a proper reweighting of the estimator so as to avoid known degeneracies in counterfactual learning, while still being applicable to stochastic gradient optimization.To conduct experiments with human users, we devise an easy-to-use interface to collect human feedback on semantic parses.Our work is the first to show that semantic parsers can be improved significantly by counterfactual learning from logged human feedback data.
Carolin Lawrence, Stefan Riezler
ACL (1)2
2018 A Reinforcement Learning Approach to Interactive-Predictive Neural Machine Translation
abstract
We present an approach to interactivepredictive neural machine translation that attempts to reduce human effort from three directions: Firstly, instead of requiring humans to select, correct, or delete segments, we employ the idea of learning from human reinforcements in form of judgments on the quality of partial translations. Secondly, human effort is further reduced by using the entropy of word predictions as uncertainty criterion to trigger feedback requests. Lastly, online updates of the model parameters after every interaction allow the model to adapt quickly. We show in simulation experiments that reward signals on partial translations significantly improve character F-score and BLEU compared to feedback on full translations only, while human effort can be reduced to an average number of 5 feedback requests for every input.
Tsz Kin Lam, Julia Kreutzer, Stefan Riezler
EAMT3
2018 A user-study on online adaptation of neural machine translation to human post-edits
Sariya Karimova, Patrick Simianer, Stefan Riezler
Mach. Transl.3
2017 Bandit Structured Prediction for Neural Sequence-to-Sequence Learning
abstract
Bandit structured prediction describes a stochastic optimization framework where learning is performed from partial feedback.This feedback is received in the form of a task loss evaluation to a predicted output structure, without having access to gold standard structures.We advance this framework by lifting linear bandit learning to neural sequence-to-sequence learning problems using attention-based recurrent neural networks.Furthermore, we show how to incorporate control variates into our learning algorithms for variance reduction and improved generalization.We present an evaluation on a neural machine translation task that shows improvements of up to 5.89 BLEU points for domain adaptation from simulated bandit feedback.
Julia Kreutzer, Artem Sokolov 0001, Stefan Riezler
ACL (1)3
2017 Counterfactual Learning from Bandit Feedback under Deterministic Logging : A Case Study in Statistical Machine Translation
abstract
The goal of counterfactual learning for statistical machine translation (SMT) is to optimize a target SMT system from logged data that consist of user feedback to translations that were predicted by another, historic SMT system.A challenge arises by the fact that riskaverse commercial SMT systems deterministically log the most probable translation.The lack of sufficient exploration of the SMT output space seemingly contradicts the theoretical requirements for counterfactual learning.We show that counterfactual learning from deterministic bandit logs is possible nevertheless by smoothing out deterministic components in learning.This can be achieved by additive and multiplicative control variates that avoid degenerate behavior in empirical risk minimization.Our simulation experiments show improvements of up to 2 BLEU points by counterfactual learning from deterministic bandit feedback.
Carolin Lawrence, Artem Sokolov 0001, Stefan Riezler
EMNLP3
2016 Multimodal Pivots for Image Caption Translation
abstract
We present an approach to improve statistical machine translation of image descriptions by multimodal pivots defined in visual space. The key idea is to perform image retrieval over a database of images that are captioned in the target language, and use the captions of the most similar images for crosslingual reranking of translation outputs. Our approach does not depend on the availability of large amounts of in-domain parallel data, but only relies on available large datasets of monolingually captioned images, and on state-of-the-art convolutional neural networks to compute image similarities. Our experimental evaluation shows improvements of 1 BLEU point over strong baselines.
Julian Hitschler, Shigehiko Schamoni, Stefan Riezler
ACL (1)3
2016 Learning Structured Predictors from Bandit Feedback for Interactive NLP
Artem Sokolov 0001, Julia Kreutzer, Christopher Lo, Stefan Riezler
ACL (1)4
2016 Learning to translate from graded and negative relevance information
abstract
We present an approach for learning to translate by exploiting cross-lingual link structure in multilingual document collections. We propose a new learning objective based on structured ramp loss, which learns from graded relevance, explicitly including negative relevance information. Our results on English German translation of Wikipedia entries show small, but significant, improvements of our method over an unadapted baseline, even when only a weak relevance signal is used. We also compare our method to monolingual language model adaptation and automatic pseudo-parallel data extraction and find small improvements even over these strong baselines.
Laura Jehl, Stefan Riezler
COLING2
2016 A Full-Text Learning to Rank Dataset for Medical Information Retrieval
Vera Boteva, Demian Gholipour Ghalandari, Artem Sokolov 0001, Stefan Riezler
ECIR4
2016 A Corpus and Semantic Parser for Multilingual Natural Language Querying of OpenStreetMap
abstract
We present a corpus of 2,380 natural language queries paired with machine readable formulae that can be executed against world wide geographic data of the OpenStreetMap (OSM) database.We use the corpus to learn an accurate semantic parser that builds the basis of a natural language interface to OSM.Furthermore, we use response-based learning on parser feedback to adapt a statistical machine translation system for multilingual database access to OSM.Our framework allows to map fuzzy natural language expressions such as "nearby", "north of", or "in walking distance" to spatial polygons on an interactive map.Furthermore, it combines syntactic complexity and compositionality with a reasonable lexical variability of queries, making it an interesting new publicly available dataset for research on semantic parsing.
Carolin Haas, Stefan Riezler
HLT-NAACL2
2016 Stochastic Structured Prediction under Bandit Feedback
abstract
Stochastic structured prediction under bandit feedback follows a learning protocol where on each of a sequence of iterations, the learner receives an input, predicts an output structure, and receives partial feedback in form of a task loss evaluation of the predicted structure. We present applications of this learning scenario to convex and non-convex objectives for structured prediction and analyze them as stochastic first-order methods. We present an experimental evaluation on problems of natural language processing over exponential output spaces, and compare convergence speed across different objectives under the practical criterion of optimal task performance on development data and the optimization-theoretic criterion of minimal squared gradient norm. Best results under both criteria are obtained for a non-convex objective for pairwise preference learning under bandit feedback.
Artem Sokolov 0001, Julia Kreutzer, Stefan Riezler, Christopher Lo
NIPS3
2015 A Coactive Learning View of Online Structured Prediction in Statistical Machine Translation
abstract
We present a theoretical analysis of online parameter tuning in statistical machine translation (SMT) from a coactive learning view.This perspective allows us to give regret and generalization bounds for latent perceptron algorithms that are common in SMT, but fall outside of the standard convex optimization scenario.Coactive learning also introduces the concept of weak feedback, which we apply in a proofof-concept experiment to SMT, showing that learning from feedback that consists of slight improvements over predictions leads to convergence in regret and translation error rate.This suggests that coactive learning might be a viable framework for interactive machine translation.Furthermore, we find that surrogate translations replacing references that are unreachable in the decoder search space can be interpreted as weak feedback and lead to convergence in learning, if they admit an underlying linear model.
Artem Sokolov 0001, Stefan Riezler, Shay B. Cohen
CoNLL2
2015 Integrating a Large, Monolingual Corpus as Translation Memory into Statistical Machine Translation
Katharina Wäschle, Stefan Riezler
EAMT2
2015 Bandit structured prediction for learning from partial feedback in statistical machine translation
Artem Sokolov 0001, Stefan Riezler, Tanguy Urvoy
MTSummit2
2015 Response-based Learning for Machine Translation of Open-domain Database Queries
abstract
Response-based learning allows to adapt a statistical machine translation (SMT) system to an extrinsic task by extracting supervision signals from task-specific feedback.In this paper, we elicit response signals for SMT adaptation by executing semantic parses of translated queries against the Freebase database.The challenge of our work lies in scaling semantic parsers to the lexical diversity of opendomain databases.We find that parser performance on incorrect English sentences, which is standardly ignored in parser evaluation, is key in model selection.In our experiments, the biggest improvements in F1-score for returning the correct answer from a semantic parse for a translated query are achieved by selecting a parser that is carefully enhanced by paraphrases and synonyms.
Carolin Haas, Stefan Riezler
HLT-NAACL2
2015 Bag-of-Words Forced Decoding for Cross-Lingual Information Retrieval
abstract
Current approaches to cross-lingual information retrieval (CLIR) rely on standard retrieval models into which query translations by statistical machine translation (SMT) are integrated at varying degree.In this paper, we present an attempt to turn this situation on its head: Instead of the retrieval aspect, we emphasize the translation component in CLIR.We perform search by using an SMT decoder in forced decoding mode to produce a bag-ofwords representation of the target documents to be ranked.The SMT model is extended by retrieval-specific features that are optimized jointly with standard translation features for a ranking objective.We find significant gains over the state-of-the-art in a large-scale evaluation on cross-lingual search in the domains patents and Wikipedia.
Felix Hieber, Stefan Riezler
HLT-NAACL2
2015 Combining Orthogonal Information in Large-Scale Cross-Language Information Retrieval
abstract
System combination is an effective strategy to boost retrieval performance, especially in complex applications such as cross-language information retrieval (CLIR) where the aspects of translation and retrieval have to be optimized jointly. We focus on machine learning-based approaches to CLIR that need large sets of relevance-ranked data to train high-dimensional models. We compare these models under various measures of orthogonality, and present an experimental evaluation on two different domains (patents, Wikipedia) and two different language pairs (Japanese-English, German-English). We show that gains of over 10 points in MAP/NDCG can be achieved over the best single model by a linear combination of the models that contribute the most orthogonal information, rather than by combining the models with the best standalone retrieval performance.
Shigehiko Schamoni, Stefan Riezler
SIGIR2
2014 Response-based Learning for Grounded Machine Translation
abstract
We propose a novel learning approach for statistical machine translation (SMT) that allows to extract supervision signals for structured learning from an extrinsic re-sponse to a translation input. We show how to generate responses by grounding SMT in the task of executing a seman-tic parse of a translated query against a database. Experiments on the GEO-QUERY database show an improvement of about 6 points in F1-score for response-based learning over learning from refer-ences only on returning the correct an-swer from a semantic parse of a translated query. In general, our approach alleviates the dependency on human reference trans-lations and solves the reachability problem in structured learning for SMT. 1
Stefan Riezler, Patrick Simianer, Carolin Haas
ACL (1)1
2014 Learning to translate queries for CLIR
abstract
The statistical machine translation (SMT) component of cross-lingual information retrieval (CLIR) systems is often regarded as black box that is optimized for translation quality independent from the retrieval task. In recent work [10], SMT has been tuned for retrieval by training a reranker on $k$-best translations ordered according to their retrieval performance. In this paper we propose a decomposable proxy for retrieval quality that obviates the need for costly intermediate retrieval. Furthermore, we explore the full search space of the SMT decoder by directly optimizing decoder parameters under a retrieval-based objective. Experimental results for patent retrieval show our approach to be a promising alternative to the standard pipeline approach.
Artem Sokolov 0001, Felix Hieber, Stefan Riezler
SIGIR3
2014 On the Problem of Theoretical Terms in Empirical Computational Linguistics
abstract
Philosophy of science has pointed out a problem of theoretical terms in empirical sciences. This problem arises if all known measuring procedures for a quantity of a theory presuppose the validity of this very theory, because then statements containing theoretical terms are circular. We argue that a similar circularity can happen in empirical computational linguistics, especially in cases where data are manually annotated by experts. We define a criterion of T-non-theoretical grounding as guidance to avoid such circularities, and exemplify how this criterion can be met by crowdsourcing, by task-related data annotation, or by data in the wild. We argue that this criterion should be considered as a necessary condition for an empirical science, in addition to measures for reliability of data annotation.
Stefan Riezler
Comput. Linguistics1
2014 Online adaptation to post-edits for phrase-based statistical machine translation
Nicola Bertoldi, Patrick Simianer, Mauro Cettolo, Katharina Wäschle, Marcello Federico, Stefan Riezler
Mach. Transl.6
2013 Boosting Cross-Language Retrieval by Learning Bilingual Phrase Associations from Relevance Rankings
abstract
We present an approach to learning bilingual n-gram correspondences from relevance rankings of English documents for Japanese queries.We show that directly optimizing cross-lingual rankings rivals and complements machine translation-based cross-language information retrieval (CLIR).We propose an efficient boosting algorithm that deals with very large cross-product spaces of word correspondences.We show in an experimental evaluation on patent prior art search that our approach, and in particular a consensus-based combination of boosting and translation-based approaches, yields substantial improvements in CLIR performance.Our training and test data are made publicly available.
Artem Sokolov 0001, Laura Jehl, Felix Hieber, Stefan Riezler
EMNLP4
2013 Generative and Discriminative Methods for Online Adaptation in SMT
Katharina Wäschle, Patrick Simianer, Nicola Bertoldi, Stefan Riezler, Marcello Federico
MTSummit4
2012 Joint Feature Selection in Distributed Stochastic Learning for Large-Scale Discriminative Training in SMT
Patrick Simianer, Stefan Riezler, Chris Dyer
ACL (1)2
2012 Structural and Topical Dimensions in Multi-Task Patent Translation
Katharina Wäschle, Stefan Riezler
EACL2
2010 Learning Dense Models of Query Similarity from User Click Logs
Fabio De Bona, Stefan Riezler, Keith B. Hall, Massimiliano Ciaramita, Amac Herdagdelen, Maria Holmqvist
HLT-NAACL2
2010 Generalized syntactic and semantic models of query reformulation
abstract
We present a novel approach to query reformulation which combines syntactic and semantic information by means of generalized Levenshtein distance algorithms where the substitution operation costs are based on probabilistic term rewrite functions. We investigate unsupervised, compact and efficient models, and provide empirical evidence of their effectiveness. We further explore a generative model of query reformulation and supervised combination methods providing improved performance at variable computational costs. Among other desirable properties, our similarity measures incorporate information-theoretic interpretations of taxonomic relations such as specification and generalization.
Amac Herdagdelen, Massimiliano Ciaramita, Daniel Mahler, Maria Holmqvist, Keith B. Hall, Stefan Riezler, Enrique Alfonseca
SIGIR6
2010 Query Rewriting Using Monolingual Statistical Machine Translation
abstract
Long queries often suffer from low recall in Web search due to conjunctive term matching. The chances of matching words in relevant documents can be increased by rewriting query terms into new terms with similar statistical properties. We present a comparison of approaches that deploy user query logs to learn rewrites of query terms into terms from the document space. We show that the best results are achieved by adopting the perspective of bridging the “lexical chasm” between queries and documents by translating from a source language of user queries into a target language of Web documents. We train a state-of-the-art statistical machine translation model on query-snippet pairs from user query logs, and extract expansion terms from the query rewrites produced by the monolingual translation system. We show in an extrinsic evaluation in a real-world Web search task that the combination of a query-to-snippet translation model with a query language model achieves improved contextual query expansion compared to a state-of-the-art query expansion model that is trained on the same query log data.
Stefan Riezler, Yi Liu 0054
Comput. Linguistics1
2008 Translating Queries into Snippets for Improved Query Expansion
Stefan Riezler, Yi Liu 0054, Alexander Vasserman
COLING1
2008 Wide-Coverage Deep Statistical Parsing Using Automatic Dependency Structure Annotation
abstract
A number of researchers have recently conducted experiments comparing “deep” hand-crafted wide-coverage with “shallow” treebank- and machine-learning-based parsers at the level of dependencies, using simple and automatic methods to convert tree output generated by the shallow parsers into dependencies. In this article, we revisit such experiments, this time using sophisticated automatic LFG f-structure annotation methodologies with surprising results. We compare various PCFG and history-based parsers to find a baseline parsing system that fits best into our automatic dependency structure annotation technique. This combined system of syntactic parser and dependency structure annotation is compared to two hand-crafted, deep constraint-based parsers, RASP and XLE. We evaluate using dependency-based gold standards and use the Approximate Randomization Test to test the statistical significance of the results. Our experiments show that machine-learning-based shallow grammars augmented with sophisticated automatic dependency annotation technology outperform hand-crafted, deep, wide-coverage constraint grammars. Currently our best system achieves an f-score of 82.73% against the PARC 700 Dependency Bank, a statistically significant improvement of 2.18% over the most recent results of 80.55% for the hand-crafted LFG grammar and XLE parsing system and an f-score of 80.23% against the CBS 500 Dependency Bank, a statistically significant 3.66% improvement over the 76.57% achieved by the hand-crafted RASP grammar and parsing system.
Aoife Cahill, Michael Burke, Ruth O'Donovan, Stefan Riezler, Josef van Genabith, Andy Way
Comput. Linguistics4
2007 Statistical Machine Translation for Query Expansion in Answer Retrieval
Stefan Riezler, Alexander Vasserman, Ioannis Tsochantaridis, Vibhu O. Mittal, Yi Liu 0054
ACL1
2006 Grammatical Machine Translation
Stefan Riezler, John T. Maxwell III
HLT-NAACL1
2006 New Developments in Parsing Technology
abstract
New Developments in Parsing Technology is a collection of papers based on contributions to the International Workshop on Parsing Technology in the years 2000 and 2001. The publication formatof acollection might raisethefollowing questions: Isthe whole ofthe collection more than the sum of its previously published parts by virtue of an inspired selection of the most seminal papers in the area? Or does the collection go beyond a mere reprint of revised versions of workshop papers by including insightful overview articles or other previously unpublished material? In case of New Developments in Pars- ing Technology the answers to these questions are yes concerning added value by the inclusion of a previously unpublished invited talk by Michael Collins, and no concern- ing exceeding the sum of its previously published parts. Table 1 lists the table of contents of the book. The book starts out with an introduc- tory chapter written by the editors. In this article, the editors motivate an interest in parsing technology by listing 12 application areas that make crucial use of parsing tech- niques. Given the limited pool of candidate papers from two workshops, unfortunately
Stefan Riezler
Comput. Linguistics1
2004 Incremental Feature Selection and l1 Regularization for Relaxed Maximum-Entropy Modeling
Stefan Riezler, Alexander Vasserman
EMNLP1
2004 Speed and Accuracy in Shallow and Deep Stochastic Parsing
Ronald M. Kaplan, Stefan Riezler, Tracy Holloway King, John T. Maxwell III, Alexander Vasserman, Richard S. Crouch
HLT-NAACL2
2003 Statistical Sentence Condensation using Ambiguity Packing and Stochastic Disambiguation Methods for Lexical-Functional Grammar
Stefan Riezler, Tracy Holloway King, Richard S. Crouch, Annie Zaenen
HLT-NAACL1
2002 Parsing the Wall Street Journal using a Lexical-Functional Grammar and Discriminative Estimation Techniques
abstract
We present a stochastic parsing system consisting of a Lexical-Functional Grammar (LFG), a constraint-based parser and a stochastic disambiguation model. We report on the results of applying this system to parsing the UPenn Wall Street Journal (WSJ) treebank. The model combines full and partial parsing techniques to reach full grammar coverage on unseen data. The treebank annotations are used to provide partially labeled data for discriminative statistical estimation using exponential models. Disambiguation performance is evaluated by measuring matches of predicate-argument relations on two distinct test sets. On a gold standard of manually annotated f-structures for a subset of the WSJ treebank, this evaluation reaches 79% F-score. An evaluation on a gold standard of dependency relations for Brown corpus data achieves 76% F-score.
Stefan Riezler, Tracy Holloway King, Ronald M. Kaplan, Richard S. Crouch, John T. Maxwell III, Mark Johnson 0001
ACL1
2000 Lexicalized Stochastic Modeling of Constraint-Based Grammars using Log-Linear Measures and EM Training
abstract
We present a new approach to stochastic modeling of constraint-based grammars that is based on loglinear models and uses EM for estimation from unannotated data. The techniques are applied to an LFG grammar for German. Evaluation on an exact match task yields 86% precision for an ambiguity rate of 5.4, and 90% precision on a subcat frame match for an ambiguity rate of 25. Experimental comparison to training from a parsebank shows a 10% gain from EM training. Also, a new class-based grammar lexicalization is presented, showing a 10% gain over unlexicalized models.
Stefan Riezler, Detlef Prescher, Jonas Kuhn, Mark Johnson 0001
ACL1
2000 Using a Probabilistic Class-Based Lexicon for Lexical Ambiguity Resolution
Detlef Prescher, Stefan Riezler, Mats Rooth
COLING2
1999 Inside-Outside Estimation of a Lexicalized PCFG for German
abstract
The paper describes an extensive experiment in inside-outside estimation of a lexicalized probabilistic context free grammar for German verb-final clauses. Grammar and formalism features which make the experiment feasible are described. Successive models are evaluated on precision and recall of phrase markup.
Franz Beil, Glenn Carroll, Detlef Prescher, Stefan Riezler, Mats Rooth
ACL4
1999 Estimators for Stochastic "Unification-Based" Grammars
abstract
Log-linear models provide a statistically sound framework for Stochastic "Unification-Based" Grammars (SUBGs) and stochastic versions of other kinds of grammars. We describe two computationally-tractable ways of estimating the parameters of such grammars from a training corpus of syntactic analyses, and apply these to estimate a stochastic version of Lexical-Functional Grammar.
Mark Johnson 0001, Stuart Geman, Stephen Canon, Zhiyi Chi, Stefan Riezler
ACL5
1999 Inducing a Semantically Annotated Lexicon via EM-Based Clustering
abstract
We present a technique for automatic induction of slot annotations for subcategorization frames, based on induction of hidden classes in the EM framework of statistical estimation.The models are empirically evalutated by a general decision test.Induction of slot labeling for subcategorization frames is accomplished by a further application of EM, and applied experimentally on frame observations derived from parsing large corpora.We outline an interpretation of the learned representations as theoretical-linguistic decompositional lexical entries.
Mats Rooth, Stefan Riezler, Detlef Prescher
ACL2