Sebastian Padó

dblp:p/SebastianPado · DBLP profile ↗
← Back
61ranked-venue papers
13as first author
14since 2021 · last 2026
0000-0002-7529-6825ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 59 · 13 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 3Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
YearPublicationVenuePosition
2026 One Persona, Many Cues, Different Results: How Sociodemographic Cues Impact LLM Personalization
abstract
Personalization of LLMs by sociodemographic subgroup often improves user experience, but can also introduce or amplify biases and unfair outcomes across groups.Prior work has employed so-called personas, sociodemographic user attributes conveyed to a model, to study bias in LLMs by relying on a single cue to prompt a persona, such as user names or explicit attribute mentions.This disregards LLM sensitivity to prompt variation and the rarity of some cues in real interactions (external validity).We compare six commonly used persona cues across seven open and proprietary LLMs on four writing and advice tasks.While cues are overall highly correlated, they produce substantial variance in responses across personas that can change findings on persona-induced differences and bias.We therefore caution against claims based on single persona cues, especially when they are overly explicit and have low external validity.
Franziska Weeber, Vera Neplenbroek, Jan Batzner, Sebastian Padó
ACL (1)4
2025 Interpretable Text Embeddings and Text Similarity Explanation: A Survey
abstract
Text embeddings are a fundamental component in many NLP tasks, including classification, regression, clustering, and semantic search.However, despite their ubiquitous application, challenges persist in interpreting embeddings and explaining similarities between them.In this work, we provide a structured overview of methods specializing in inherently interpretable text embeddings and text similarity explanation, an underexplored research area.We characterize the main ideas, approaches, and tradeoffs.We compare means of evaluation, discuss overarching lessons learned and finally identify opportunities and open challenges for future research.
Juri Opitz, Lucas Möller, Andrianos Michail, Sebastian Padó, Simon Clematide
EMNLP4
2025 Explaining Neural News Recommendation with Attributions onto Reading Histories
abstract
An important aspect of responsible recommendation systems is the transparency of the prediction mechanisms. This is a general challenge for deep-learning-based systems such as the currently predominant neural news recommender architectures, which are optimized to predict clicks by matching candidate news items against users’ reading histories. Such systems achieve state-of-the-art click-prediction performance, but the rationale for their decisions is difficult to assess. At the same time, the economic and societal impact of these systems makes such insights very much desirable. In this article, we ask the question to what extent the recommendations of current news recommender systems are actually based on content-related evidence from reading histories. We approach this question from an explainability perspective. Building on the concept of integrated gradients, we present a neural news recommender that can accurately attribute individual recommendations to news items and words in input reading histories while maintaining a top scoring click-prediction performance. Using our method as a diagnostic tool, we find that: (a), a substantial number of users’ clicks on news are not explainable from reading histories, and many history-explainable items are actually skipped; (b), while many recommendations are based on content-related evidence in histories, for others the model does not attend to reasonable evidence, and recommendations stem from a spurious bias in user representations. Our code is publicly available at https://github.com/lucasmllr/xnrs .
Lucas Möller, Sebastian Padó
ACM Trans. Intell. Syst. Technol.2
2024 Towards Understanding the Relationship between In-context Learning and Compositional Generalization
abstract
According to the principle of compositional generalization, the meaning of a complex expression can be understood as a function of the meaning of its parts and of how they are combined. This principle is crucial for human language processing and also, arguably, for NLP models in the face of out-of-distribution data. However, many neural network models, including Transformers, have been shown to struggle with compositional generalization. In this paper, we hypothesize that forcing models to in-context learn can provide an inductive bias to promote compositional generalization. To test this hypothesis, we train a causal Transformer in a setting that renders ‘ordinary’ learning very difficult: we present it with different orderings of the training instance and shuffle instance labels. This corresponds to training the model on all possible few-shot learning problems attainable from the dataset. The model can solve the task, however, by utilizing earlier examples to generalize to later ones – i.e., in-context learning. In evaluations on the datasets, SCAN, COGS, and GeoQuery, models trained in this manner indeed show improved compositional generalization. This indicates the usefulness of in-context learning problems as an inductive bias for generalization.
Sungjun Han, Sebastian Padó
LREC/COLING2
2024 Multi-Dimensional Machine Translation Evaluation: Model Evaluation and Resource for Korean
abstract
Almost all frameworks for the manual or automatic evaluation of machine translation characterize the quality of an MT output with a single number. An exception is the Multidimensional Quality Metrics (MQM) framework which offers a fine-grained ontology of quality dimensions for scoring (such as style, fluency, accuracy, and terminology). Previous studies have demonstrated the feasibility of MQM annotation but there are, to our knowledge, no computational models that predict MQM scores for novel texts, due to a lack of resources. In this paper, we address these shortcomings by (a) providing a 1200-sentence MQM evaluation benchmark for the language pair English-Korean and (b) reframing MT evaluation as the multi-task problem of simultaneously predicting several MQM scores using SOTA language models, both in a reference-based MT evaluation setup and a reference-free quality estimation (QE) setup. We find that reference-free setup outperforms its counterpart in the style dimension while reference-based models retain an edge regarding accuracy. Overall, RemBERT emerges as the most promising model. Through our evaluation, we offer an insight into the translation quality in a more fine-grained, interpretable manner.
Dojun Park, Sebastian Padó
LREC/COLING2
2024 Approximate Attributions for Off-the-Shelf Siamese Transformers
abstract
Siamese encoders such as sentence transformers are among the least understood deep models.Established attribution methods cannot tackle this model class since it compares two inputs rather than processing a single one.To address this gap, we have recently proposed an attribution method specifically for Siamese encoders (Möller et al., 2023).However, it requires models to be adjusted and fine-tuned and therefore cannot be directly applied to off-the-shelf models.In this work, we reassess these restrictions and propose (i) a model with exact attribution ability that retains the original model's predictive performance and (ii) a way to compute approximate attributions for off-the-shelf models.We extensively compare approximate and exact attributions and use them to analyze the models' attendance to different linguistic aspects.We gain insights into which syntactic roles Siamese transformers attend to, confirm that they mostly ignore negation, explore how they judge semantically opposite adjectives, and find that they exhibit lexical bias.
Lucas Möller, Dmitry Nikolaev 0002, Sebastian Padó
EACL (1)3
2024 Media Bias Detection Across Families of Language Models
abstract
Iffat Maab, Edison Marrese-Taylor, Sebastian Padó, Yutaka Matsuo. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Iffat Maab, Edison Marrese-Taylor, Sebastian Padó, Yutaka Matsuo
NAACL-HLT3
2024 Language Learning, Representation, and Processing in Humans and Machines: Introduction to the Special Issue
abstract
Abstract Large Language Models (LLMs) and humans acquire knowledge about language without direct supervision. LLMs do so by means of specific training objectives, while humans rely on sensory experience and social interaction. This parallelism has created a feeling in NLP and cognitive science that a systematic understanding of how LLMs acquire and use the encoded knowledge could provide useful insights for studying human cognition. Conversely, methods and findings from the field of cognitive science have occasionally inspired language model development. Yet, the differences in the way that language is processed by machines and humans—in terms of learning mechanisms, amounts of data used, grounding and access to different modalities—make a direct translation of insights challenging. The aim of this edited volume has been to create a forum of exchange and debate along this line of research, inviting contributions that further elucidate similarities and differences between humans and LLMs.
Marianna Apidianaki, Abdellah Fourtassi, Sebastian Padó
Comput. Linguistics3
2024 Beyond Prompt Brittleness: Evaluating the Reliability and Consistency of Political Worldviews in LLMs
abstract
Abstract Due to the widespread use of large language models (LLMs), we need to understand whether they embed a specific “worldview” and what these views reflect. Recent studies report that, prompted with political questionnaires, LLMs show left-liberal leanings (Feng et al., 2023; Motoki et al., 2024). However, it is as yet unclear whether these leanings are reliable (robust to prompt variations) and whether the leaning is consistent across policies and political leaning. We propose a series of tests which assess the reliability and consistency of LLMs’ stances on political statements based on a dataset of voting-advice questionnaires collected from seven EU countries and annotated for policy issues. We study LLMs ranging in size from 7B to 70B parameters and find that their reliability increases with parameter count. Larger models show overall stronger alignment with left-leaning parties but differ among policy programs: They show a (left-wing) positive stance towards environment protection, social welfare state, and liberal society but also (right-wing) law and order, with no consistent preferences in the areas of foreign policy and migration.
Tanise Ceron, Neele Falk, Ana Baric, Dmitry Nikolaev 0002, Sebastian Padó
Trans. Assoc. Comput. Linguistics5
2023 Representation biases in sentence transformers
abstract
Variants of the BERT architecture specialised for producing full-sentence representations often achieve better performance on downstream tasks than sentence embeddings extracted from vanilla BERT.However, there is still little understanding of what properties of inputs determine the properties of such representations.In this study, we construct several sets of sentences with pre-defined lexical and syntactic structures and show that SOTA sentence transformers have a strong nominal-participant-set bias: cosine similarities between pairs of sentences are more strongly determined by the overlap in the set of their noun participants than by having the same predicates, lengthy nominal modifiers, or adjuncts.At the same time, the precise syntactic-thematic functions of the participants are largely irrelevant.
Dmitry Nikolaev 0002, Sebastian Padó
EACL2
2023 Multilingual estimation of political-party positioning: From label aggregation to long-input Transformers
abstract
Scaling analysis is a technique in computational political science that assigns a political actor (e.g.politician or party) a score on a predefined scale based on a (typically long) body of text (e.g. a parliamentary speech or an election manifesto).For example, political scientists have often used the left-right scale to systematically analyse political landscapes of different countries.NLP methods for automatic scaling analysis can find broad application provided they (i) are able to deal with long texts and (ii) work robustly across domains and languages.In this work, we implement and compare two approaches to automatic scaling analysis of political-party manifestos: label aggregation, a pipeline strategy relying on annotations of individual statements from the manifestos, and long-input-Transformer-based models, which compute scaling values directly from raw text.We carry out the analysis of the Comparative Manifestos Project dataset across 41 countries and 27 languages and find that the task can be efficiently solved by state-of-the-art models, with label aggregation producing the best results.
Dmitry Nikolaev 0002, Tanise Ceron, Sebastian Padó
EMNLP3
2023 An Attribution Method for Siamese Encoders
abstract
Despite the success of Siamese encoder models such as sentence transformers (ST), little is known about the aspects of inputs they pay attention to.A barrier is that their predictions cannot be attributed to individual features, as they compare two inputs rather than processing a single one.This paper derives a local attribution method for Siamese encoders by generalizing the principle of integrated gradients to models with multiple inputs.The output takes the form of feature-pair attributions and in case of STs it can be reduced to a token-token matrix.Our method involves the introduction of integrated Jacobians and inherits the advantageous formal properties of integrated gradients: it accounts for the model's full computation graph and is guaranteed to converge to the actual prediction.A pilot study shows that in case of STs few token pairs can dominate predictions and that STs preferentially focus on nouns and verbs.For accurate predictions, however, they need to attend to the majority of tokens and parts of speech.
Lucas Möller, Dmitry Nikolaev 0002, Sebastian Padó
EMNLP3
2022 Optimizing text representations to capture (dis)similarity between political parties
abstract
Even though fine-tuned neural language models have been pivotal in enabling "deep" automatic text analysis, optimizing text representations for specific applications remains a crucial bottleneck.In this study, we look at this problem in the context of a task from computational social science, namely modeling pairwise similarities between political parties.Our research question is what level of structural information is necessary to create robust text representation, contrasting a strongly informed approach (which uses both claim span and claim category annotations) with approaches that forgo one or both types of annotation with document structure-based heuristics.Evaluating our models on the manifestos of German parties for the 2021 federal election.We find that heuristics that maximize within-party over between-party similarity along with a normalization step lead to reliable party similarity prediction, without the need for manual annotation.
Tanise Ceron, Nico Blokker, Sebastian Padó
CoNLL3
2022 Constraining Linear-chain CRFs to Regular Languages
Sean Papay, Roman Klinger, Sebastian Padó
ICLR3
2020 Masking Actor Information Leads to Fairer Political Claims Detection
abstract
A central concern in Computational Social Sciences (CSS) is fairness: where the role of NLP is to scale up text analysis to large corpora, the quality of automatic analyses should be as independent as possible of textual properties.We analyze the performance of a state-of-theart neural model on the task of political claims detection (i.e., the identification of forwardlooking statements made by political actors) and identify a strong frequency bias: claims made by frequent actors are recognized better.We propose two simple debiasing methods which mask proper names and pronouns during training of the model, thus removing personal information bias.We find that (a) these methods significantly decrease frequency bias while keeping the overall performance stable; and (b) the resulting models improve when evaluated in an out-of-domain setting.
Erenay Dayanik, Sebastian Padó
ACL2
2020 Lost in Back-Translation: Emotion Preservation in Neural Machine Translation
abstract
Machine translation provides powerful methods to convert text between languages, and is therefore a technology enabling a multilingual world.An important part of communication, however, takes place at the non-propositional level (e.g., politeness, formality, emotions), and it is far from clear whether current MT methods properly translate this information.This paper investigates the specific hypothesis that the non-propositional level of emotions is at least partially lost in MT.We carry out a number of experiments in a back-translation setup and establish that (1) emotions are indeed partially lost during translation; (2) this tendency can be reversed almost completely with a simple re-ranking approach informed by an emotion classifier, taking advantage of diversity in the n-best list; (3) the re-ranking approach can also be applied to change emotions, obtaining a model for emotion style transfer.An in-depth qualitative analysis reveals that there are recurring linguistic changes through which emotions are toned down or amplified, such as change of modality.
Enrica Troiano, Roman Klinger, Sebastian Padó
COLING3
2020 Dissecting Span Identification Tasks with Performance Prediction
abstract
Span identification (in short, span ID) tasks such as chunking, NER, or code-switching detection, ask models to identify and classify relevant spans in a text.Despite being a staple of NLP, and sharing a common structure, there is little insight on how these tasks' properties influence their difficulty, and thus little guidance on what model families work well on span ID tasks, and why.We analyze span ID tasks via performance prediction, estimating how well neural architectures do on different tasks.Our contributions are: (a) we identify key properties of span ID tasks that can inform performance prediction; (b) we carry out a large-scale experiment on English data, building a model to predict performance for unseen span ID tasks that can support architecture choices; (c), we investigate the parameters of the meta model, yielding new insights on how model and task properties interact to affect span ID performance.We find, e.g., that span frequency is especially important for LSTMs, and that CRFs help when spans are infrequent and boundaries non-distinctive.
Sean Papay, Roman Klinger, Sebastian Padó
EMNLP (1)3
2020 DEbateNet-mig15: Tracing the 2015 Immigration Debate in Germany Over Time
abstract
DEbateNet-migr15 is a manually annotated dataset for German which covers the public debate on immigration in 2015. The building block of our annotation is the political science notion of a claim, i.e., a statement made by a political actor (a politician, a party, or a group of citizens) that a specific action should be taken (e.g., vacant flats should be assigned to refugees). We identify claims in newspaper articles, assign them to actors and fine-grained categories and annotate their polarity and date. The aim of this paper is two-fold: first, we release the full DEbateNet-mig15 corpus and document it by means of a quantitative and qualitative analysis; second, we demonstrate its application in a discourse network analysis framework, which enables us to capture the temporal dynamics of the political debate
Gabriella Lapesa, André Blessing, Nico Blokker, Erenay Dayanik, Sebastian Haunss, Jonas Kuhn, Sebastian Padó
LREC7
2020 RiQuA: A Corpus of Rich Quotation Annotation for English Literary Text
abstract
We introduce RiQuA (RIch QUotation Annotations), a corpus that provides quotations, including their interpersonal structure (speakers and addressees) for English literary text. The corpus comprises 11 works of 19th-century literature that were manually doubly annotated for direct and indirect quotations. For each quotation, its span, speaker, addressee, and cue are identified (if present). This provides a rich view of dialogue structures not available from other available corpora. We detail the process of creating this dataset, discuss the annotation guidelines, and analyze the resulting corpus in terms of inter-annotator agreement and its properties. RiQuA, along with its annotations guidelines and associated scripts, are publicly available for use, modification, and experimentation.
Sean Papay, Sebastian Padó
LREC2
2019 Who Sides with Whom? Towards Computational Construction of Discourse Networks for Political Debates
abstract
Understanding the structures of political debates (which actors make what claims) is essential for understanding democratic political decision-making.The vision of computational construction of such discourse networks from newspaper reports brings together political science and natural language processing.This paper presents three contributions towards this goal: (a) a requirements analysis, linking the task to knowledge base population; (b) a first release of an annotated corpus of claims on the topic of migration, based on German newspaper reports; (c) initial modeling results.
Sebastian Padó, André Blessing, Nico Blokker, Erenay Dayanik, Sebastian Haunss, Jonas Kuhn
ACL (1)1
2019 Crowdsourcing and Validating Event-focused Emotion Corpora for German and English
abstract
Sentiment analysis has a range of corpora available across multiple languages.For emotion analysis, the situation is more limited, which hinders potential research on crosslingual modeling and the development of predictive models for other languages.In this paper, we fill this gap for German by constructing deISEAR, a corpus designed in analogy to the well-established English ISEAR emotion dataset.Motivated by Scherer's appraisal theory, we implement a crowdsourcing experiment which consists of two steps.In step 1, participants create descriptions of emotional events for a given emotion.In step 2, five annotators assess the emotion expressed by the texts.We show that transferring an emotion classification model from the original English ISEAR to the German crowdsourced deISEAR via machine translation does not, on average, cause a performance drop.
Enrica Troiano, Sebastian Padó, Roman Klinger
ACL (1)2
2018 Leveraging Lexical Substitutes for Unsupervised Word Sense Induction
abstract
Word sense induction is the most prominent unsupervised approach to lexical disambiguation. It clusters word instances, typically represented by their bag-of-words contexts. Therefore, uninformative and ambiguous contexts present a major challenge. In this paper, we investigate the use of an alternative instance representation based on lexical substitutes, i.e., contextually suitable, meaning-preserving replacements. Using lexical substitutes predicted by a state-of-the-art automatic system and a simple clustering algorithm, we outperform bag-of-words instance representations and compete with much more complex structured probabilistic models. Furthermore, we show that an oracle based on manually-labeled lexical substitutes yields yet substantially higher performance. Taken together, this provides evidence for a complementarity between word sense induction and lexical substitution that has not been given much consideration before.
Domagoj Alagic, Jan Snajder, Sebastian Padó
AAAI3
2017 "Show Me the Cup": Reference with Continuous Representations
Marco Baroni, Gemma Boleda, Sebastian Padó
CICLing (1)3
2016 Model Architectures for Quotation Detection
abstract
Quotation detection is the task of locating spans of quoted speech in text.The state of the art treats this problem as a sequence labeling task and employs linear-chain conditional random fields.We question the efficacy of this choice: The Markov assumption in the model prohibits it from making joint decisions about the begin, end, and internal context of a quotation.We perform an extensive analysis with two new model architectures.We find that (a), simple boundary classification combined with a greedy prediction strategy is competitive with the state of the art; (b), a semi-Markov model significantly outperforms all others, by relaxing the Markov assumption.
Christian Scheible, Roman Klinger, Sebastian Padó
ACL (1)3
2016 Predictability of Distributional Semantics in Derivational Word Formation
abstract
Compositional distributional semantic models (CDSMs) have successfully been applied to the task of predicting the meaning of a range of linguistic constructions. Their performance on semi-compositional word formation process of (morphological) derivation, however, has been extremely variable, with no large-scale empirical investigation to date. This paper fills that gap, performing an analysis of CDSM predictions on a large dataset (over 30,000 German derivationally related word pairs). We use linear regression models to analyze CDSM performance and obtain insights into the linguistic factors that influence how predictable the distributional context of a derived word is going to be. We identify various such factors, notably part of speech, argument structure, and semantic regularity.
Sebastian Padó, Aurélie Herbelot, Max Kisselew, Jan Snajder
COLING1
2015 Distributional vectors encode referential attributes
abstract
Distributional methods have proven to excel at capturing fuzzy, graded aspects of meaning (Italy is more similar to Spain than to Germany).In contrast, it is difficult to extract the values of more specific attributes of word referents from distributional representations, attributes of the kind typically found in structured knowledge bases (Italy has 60 million inhabitants).In this paper, we pursue the hypothesis that distributional vectors also implicitly encode referential attributes.We show that a standard supervised regression model is in fact sufficient to retrieve such attributes to a reasonable degree of accuracy: When evaluated on the prediction of both categorical and numeric attributes of countries and cities, the model consistently reduces baseline error by 30%, and is not far from the upper bound.Further analysis suggests that our model is able to "objectify" distributional representations for entities, anchoring them more firmly in the external world in measurable ways.
Abhijeet Gupta, Gemma Boleda, Marco Baroni, Sebastian Padó
EMNLP4
2015 Design and realization of a modular architecture for textual entailment
abstract
Abstract A key challenge at the core of many Natural Language Processing (NLP) tasks is the ability to determine which conclusions can be inferred from a given natural language text. This problem, called theRecognition of Textual Entailment (RTE), has initiated the development of a range of algorithms, methods, and technologies. Unfortunately, research on Textual Entailment (TE), like semantics research more generally, is fragmented into studies focussing on various aspects of semantics such as world knowledge, lexical and syntactic relations, or more specialized kinds of inference. This fragmentation has problematic practical consequences. Notably, interoperability among the existing RTE systems is poor, and reuse of resources and algorithms is mostly infeasible. This also makes systematic evaluations very difficult to carry out. Finally, textual entailment presents a wide array of approaches to potential end users with little guidance on which to pick. Our contribution to this situation is the novel EXCITEMENT architecture, which was developed to enable and encourage the consolidation of methods and resources in the textual entailment area. It decomposes RTE into components with strongly typed interfaces. We specify (a) a modular linguistic analysis pipeline and (b) a decomposition of the ‘core’ RTE methods into top-level algorithms and subcomponents. We identify four major subcomponent types, including knowledge bases and alignment methods. The architecture was developed with a focus on generality, supporting all major approaches to RTE and encouraging language independence. We illustrate the feasibility of the architecture by constructing mappings of major existing systems onto the architecture. The practical implementation of this architecture forms the EXCITEMENT open platform. It is a suite of textual entailment algorithms and components which contains the three systems named above, including linguistic-analysis pipelines for three languages (English, German, and Italian), and comprises a number of linguistic resources. By addressing the problems outlined above, the platform provides a comprehensive and flexible basis for research and experimentation in textual entailment and is available as open source software under the GNU General Public License.
Sebastian Padó, Tae-Gil Noh, Asher Stern 0001, Rui Wang 0005, Roberto Zanoli
Nat. Lang. Eng.1
2014 Type and Thematic Fit in Logical Metonymy
Alessandra Zarcone, Sebastian Padó, Alessandro Lenci
CogSci2
2014 Towards Semantic Validation of a Derivational Lexicon
Britta D. Zeller, Sebastian Padó, Jan Snajder
COLING2
2014 Crowdsourcing Annotation of Non-Local Semantic Roles
abstract
This paper reports on a study of crowdsourcing the annotation of non-local (or implicit) frame-semantic roles, i.e., roles that are realized in the previous discourse context. We describe two annotation setups (marking and gap filling) and find that gap filling works considerably better, attaining an acceptable quality relatively cheaply. The produced data is available for research purposes.
Parvin Sadat Feizabadi, Sebastian Padó
EACL2
2014 What Substitutes Tell Us - Analysis of an "All-Words" Lexical Substitution Corpus
abstract
We present the first large-scale English "allwords lexical substitution" corpus.The size of the corpus provides a rich resource for investigations into word meaning.We investigate the nature of lexical substitute sets, comparing them to WordNet synsets.We find them to be consistent with, but more fine-grained than, synsets.We also identify significant differences to results for paraphrase ranking in context reported for the SEMEVAL lexical substitution data.This highlights the influence of corpus construction approaches on evaluation results.
Gerhard Kremer, Katrin Erk, Sebastian Padó, Stefan Thater
EACL3
2014 Polysemy Index for Nouns: an Experiment on Italian using the PAROLE SIMPLE CLIPS Lexical Database
Francesca Frontini, Valeria Quochi, Sebastian Padó, Monica Monachini, Jason Utt
LREC3
2014 Crosslingual and Multilingual Construction of Syntax-Based Vector Space Models
abstract
Syntax-based distributional models of lexical semantics provide a flexible and linguistically adequate representation of co-occurrence information. However, their construction requires large, accurately parsed corpora, which are unavailable for most languages. In this paper, we develop a number of methods to overcome this obstacle. We describe (a) a crosslingual approach that constructs a syntax-based model for a new language requiring only an English resource and a translation lexicon; and (b) multilingual approaches that combine crosslingual with monolingual information, subject to availability. We evaluate on two lexical semantic benchmarks in German and Croatian. We find that the models exhibit complementary profiles: crosslingual models yield higher accuracies while monolingual models provide better coverage. In addition, we show that simple multilingual models can successfully combine their strengths.
Jason Utt, Sebastian Padó
Trans. Assoc. Comput. Linguistics2
2013 DErivBase: Inducing and Evaluating a Derivational Morphology Resource for German
Britta D. Zeller, Jan Snajder, Sebastian Padó
ACL (1)3
2012 Inferring Covert Events in Logical Metonymies: a Probe Recognition Experiment
Alessandra Zarcone, Sebastian Padó, Alessandro Lenci
CogSci2
2012 Towards a model of formal and informal address in English
Manaal Faruqui, Sebastian Padó
EACL2
2012 LODifier: Generating Linked Data from Unstructured Text
Isabelle Augenstein, Sebastian Padó, Sebastian Rudolph
ESWC2
2012 French and German Corpora for Audience-based Text Type Classification
Amalia Todirascu, Sebastian Padó, Jennifer Krisch, Max Kisselew, Ulrich Heid
LREC2
2011 Generalized Event Knowledge in Logical Metonymy Resolution
Alessandra Zarcone, Sebastian Padó
CogSci2
2010 Assessing the Role of Discourse References in Entailment Inference
Shachar Mirkin, Ido Dagan, Sebastian Padó
ACL3
2010 Cross-lingual Induction of Selectional Preferences with Bilingual Vector Spaces
Yves Peirsman, Sebastian Padó
HLT-NAACL2
2010 A Flexible, Corpus-Driven Model of Regular and Inverse Selectional Preferences
abstract
We present a vector space–based model for selectional preferences that predicts plausibility scores for argument headwords. It does not require any lexical resources (such as WordNet). It can be trained either on one corpus with syntactic annotation, or on a combination of a small semantically annotated primary corpus and a large, syntactically analyzed generalization corpus. Our model is able to predict inverse selectional preferences, that is, plausibility scores for predicates given argument heads. We evaluate our model on one NLP task (pseudo-disambiguation) and one cognitive task (prediction of human plausibility judgments), gauging the influence of different parameters and comparing our model against other model classes. We obtain consistent benefits from using the disambiguation and semantic role information provided by a semantically tagged primary corpus. As for parameters, we identify settings that yield good performance across a range of experimental conditions. However, frequency remains a major influence of prediction quality, and we also identify more robust parameter settings suitable for applications with many infrequent items.
Katrin Erk, Sebastian Padó, Ulrike Padó
Comput. Linguistics2
2009 Robust Machine Translation Evaluation with Entailment Features
Sebastian Padó, Michel Galley, Daniel Jurafsky, Christopher D. Manning
ACL/IJCNLP1
2009 Cross-lingual Annotation Projection for Semantic Roles
abstract
This article considers the task of automatically inducing role-semantic annotations in the FrameNet paradigm for new languages. We propose a general framework that is based on annotation projection, phrased as a graph optimization problem. It is relatively inexpensive and has the potential to reduce the human effort involved in creating role-semantic resources. Within this framework, we present projection models that exploit lexical and syntactic information. We provide an experimental evaluation on an English-German parallel corpus which demonstrates the feasibility of inducing high-precision German semantic role annotation both for manually and automatically annotated English data.
Sebastian Padó, Mirella Lapata
J. Artif. Intell. Res.1
2009 Measuring machine translation quality as semantic equivalence: A metric based on entailment features
Sebastian Padó, Daniel M. Cer, Michel Galley, Daniel Jurafsky, Christopher D. Manning
Mach. Transl.1
2008 Semantic Role Assignment for Event Nominalisations by Leveraging Verbal Data
Sebastian Padó, Marco Pennacchiotti, Caroline Sporleder
COLING1
2008 A Structured Vector Space Model for Word Meaning in Context
Katrin Erk, Sebastian Padó
EMNLP2
2008 Formalising Multi-layer Corpora in OWL DL - Lexicon Modelling, Querying and Consistency Control
Aljoscha Burchardt, Sebastian Padó, Dennis Spohr, Anette Frank, Ulrich Heid
IJCNLP2
2007 Flexible, Corpus-Based Modelling of Human Plausibility Judgements
Sebastian Padó, Ulrike Padó, Katrin Erk
EMNLP-CoNLL1
2007 Dependency-Based Construction of Semantic Space Models
abstract
Traditionally, vector-based semantic space models use word co-occurrence counts from large corpora to represent lexical meaning. In this article we present a novel framework for constructing semantic spaces that takes syntactic relations into account. We introduce a formalization for this class of models, which allows linguistic knowledge to guide the construction process. We evaluate our framework on a range of tasks relevant for cognitive science and natural language processing: semantic priming, synonymy detection, and word sense disambiguation. In all cases, our framework obtains results that are comparable or superior to the state of the art.
Sebastian Padó, Mirella Lapata
Comput. Linguistics1
2006 Optimal Constituent Alignment with Edge Covers for Semantic Projection
abstract
Given a parallel corpus, semantic projection attempts to transfer semantic role annotations from one language to another, typically by exploiting word alignments. In this paper, we present an improved method for obtaining constituent alignments between parallel sentences to guide the role projection task. Our extensions are twofold: (a) we model constituent alignment as minimum weight edge covers in a bipartite graph, which allows us to find a globally optimal solution efficiently; (b) we propose tree pruning as a promising strategy for reducing alignment noise. Experimental results on an English-German parallel corpus demonstrate improvements over state-of-the-art models.
Sebastian Padó, Mirella Lapata
ACL1
2006 SALTO - A Versatile Multi-Level Annotation Tool
Aljoscha Burchardt, Katrin Erk, Anette Frank, Andrea Kowalski, Sebastian Padó
LREC5
2006 The SALSA Corpus: a German Corpus Resource for Lexical Semantics
Aljoscha Burchardt, Katrin Erk, Anette Frank, Andrea Kowalski, Sebastian Padó
LREC5
2006 Shalmaneser - A Toolchain For Shallow Semantic Parsing
Katrin Erk, Sebastian Padó
LREC2
2005 Cross-Lingual Bootstrapping of Semantic Lexicons: The Case of FrameNet
Sebastian Padó, Mirella Lapata
AAAI1
2004 Semantic Role Labelling With Chunk Sequences
Ulrike Padó, Katrin Erk, Sebastian Padó, Detlef Prescher
CoNLL3
2004 The Influence of Argument Structure on Semantic Role Assignment
Sebastian Padó, Gemma Boleda
EMNLP1
2004 A Powerful and Versatile XML Format for Representing Role-semantic Annotation
Katrin Erk, Sebastian Padó
LREC2
2004 Querying Both Time-aligned and Hierarchical Corpora with NXT Search
Ulrich Heid, Holger Voormann, Jan-Torsten Milde, Ulrike Gut, Katrin Erk, Sebastian Padó
LREC6
2003 Towards a Resource for Lexical Semantics: A Large German Corpus with Extensive Semantic Annotation
abstract
We describe the ongoing construction of a large, semantically annotated corpus resource as reliable basis for the large-scale acquisition of word-semantic information, e.g. the construction of domain-independent lexica. The backbone of the annotation are semantic roles in the frame semantics paradigm. We report experiences and evaluate the annotated data from the first project stage. On this basis, we discuss the problems of vagueness and ambiguity in semantic annotation.
Katrin Erk, Andrea Kowalski, Sebastian Padó, Manfred Pinkal
ACL3
2003 Constructing Semantic Space Models from Parsed Corpora
abstract
Traditional vector-based models use word co-occurrence counts from large corpora to represent lexical meaning. In this paper we present a novel approach for constructing semantic spaces that takes syntactic relations into account. We introduce a formalisation for this class of models and evaluate their adequacy on two modelling tasks: semantic priming and automatic discrimination of lexical relations.
Sebastian Padó, Mirella Lapata
ACL1