Simone Paolo Ponzetto

dblp:04/2532 · DBLP profile ↗
← Back
84ranked-venue papers
10as first author
27since 2021 · last 2026
0000-0001-7484-2049ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 68 · 10 first-author · 19 since 2021Databases, data management, data science and information retrieval · 13 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 6 · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Appeal, Align, Divide? Stance Detection for Group-Directed Messages in German Parliamentary Debates
Ines Rehbein, Maris Leander Buttmann, Julian Schlenker, Simone Paolo Ponzetto
LREC4
2026 GePaDeU - a Multi-layer Corpus of German Parliamentary Debates with Rich Semantic and Pragmatic Annotations
Ines Rehbein, Julian Schlenker, Lars Ostertag, Simone Paolo Ponzetto
LREC4
2026 GePaDeSE: A New Resource for Clause-Level Aspect in German Parliamentary Debates
Julian Schlenker, Ines Rehbein, Lilly Brauner, Florian Ertz, Ines Reinig, Simone Paolo Ponzetto
LREC6
2025 Closing the Loop between User Stories and GUI Prototypes: An LLM-Based Assistant for Cross-Functional Integration in Software Development
abstract
Figure 1: GUI prototype (1) and three views (2-4) of our assistant for GUI prototype designers integrated as a plug-in into a prototyping tool.Our assistant displays user stories (2) imported from collaboration tools (e.g., JIRA) for prototype designers to reference while working.It detects whether a user story is implemented (3, 4), identifies relevant GUI components (3), and generates GUI components for user stories (4). Figure uses Google Material 3 Design Kit [24] components under CC BY 4.0.
Felix Kretzer, Kristian Kolthoff, Christian Bartelt, Simone Paolo Ponzetto, Alexander Maedche
CHI4
2025 Steering Language Models in Multi-Token Generation: A Case Study on Tense and Aspect
abstract
Large language models (LLMs) are able to generate grammatically well-formed text, but how do they encode their syntactic knowledge internally?While prior work has focused largely on binary grammatical contrasts, in this work, we study the representation and control of two multidimensional hierarchical grammar phenomena-verb tense and aspect-and for each, identify distinct, orthogonal directions in residual space using linear discriminant analysis.Next, we demonstrate causal control over both grammatical features through concept steering across three generation tasks.Then, we use these identified features in a case study to investigate factors influencing effective steering in multi-token generation.We find that steering strength, location, and duration are crucial parameters for reducing undesirable side effects such as topic shift and degeneration.Our findings suggest that models encode tense and aspect in structurally organized, human-like ways, but effective control of such features during generation is sensitive to multiple factors and requires manual tuning or automated optimization.1
Alina Klerings, Jannik Brinkmann, Daniel Ruffinelli, Simone Paolo Ponzetto
EMNLP4
2025 Moral Framing in Politics (MFiP): A new resource and models for moral framing
abstract
The construct of morality permeates our entire lives and influences our behavior and how we perceive others.It therefore comes at no surprise that morality also plays an important role in politics, as morally framed arguments are perceived as more appealing and persuasive.Thus, being able to identify moral framing in political communication and to detect subtle differences in politicians' moral framing can provide the basis for many interesting analyses in the political sciences.In the paper, we release MoralFramingInPolitics (MFiP), a new corpus of German parliamentary debates where the speakers' moral framing has been coded, using the framework of Moral Foundations Theory (MFT).Our fine-grained annotations distinguish different types of moral frames and also include narrative roles, together with the moral foundations for each frame.We then present models for frame type and moral foundation classification and explore the benefits of data augmentation (DA) and contrastive learning (CL) for the two tasks.All data and code will be made available to the research community.
Ines Rehbein, Ines Reinig, Simone Paolo Ponzetto
EMNLP3
2025 Randomly Removing 50% of Dimensions in Text Embeddings has Minimal Impact on Retrieval and Classification Tasks
abstract
In this paper, we study the surprising impact that truncating text embeddings has on downstream performance.We consistently observe across 6 state-of-the-art text encoders and 26 downstream tasks, that randomly removing up to 50% of embedding dimensions results in only a minor drop in performance, less than 10%, in retrieval and classification tasks.Given the benefits of using smaller-sized embeddings, as well as the potential insights about text encoding, we study this phenomenon and find that, contrary to what is suggested in prior work, this is not the result of an ineffective use of representation space.Instead, we find that a large number of uniformly distributed dimensions actually cause an increase in performance when removed.This would explain why, on average, removing a large number of embedding dimensions results in a marginal drop in performance.We make similar observations when truncating the embeddings used by large language models to make next-token predictions on generative tasks, suggesting that this phenomenon is not isolated to classification or retrieval tasks.Our code is attached to the submission. 1
Sotaro Takeshita, Yurina Takeshita, Daniel Ruffinelli, Simone Paolo Ponzetto
EMNLP4
2025 Tikzero: Zero-Shot Text-Guided Graphics Program Synthesis
Jonas Belouadi, Eddy Ilg, Margret Keuper, Hideki Tanaka, Masao Utiyama, Raj Dabre, Steffen Eger, Simone Paolo Ponzetto
ICCV8
2025 GUI-ReRank: Enhancing GUI Retrieval with Multi-Modal LLM-based Reranking
abstract
Graphical User Interface (GUI) prototyping is a fundamental component in the development of modern interactive systems, which are now ubiquitous across diverse application domains. GUI prototypes play a critical role in requirements elicitation by enabling stakeholders to visualize, assess, and refine system concepts collaboratively. Moreover, prototypes serve as effective tools for early testing, iterative evaluation, and validation of design ideas with both end users and development teams. Despite these advantages, the process of constructing GUI prototypes remains resource-intensive and time-consuming, frequently demanding substantial effort and expertise. Recent research has sought to alleviate this burden through natural language (NL)-based GUI retrieval approaches, which typically rely on embedding-based retrieval or tailored ranking models for specific GUI repositories. However, these methods often suffer from limited retrieval performance and struggle to generalize across arbitrary GUI datasets. In this work, we present GUI-ReRank, a novel framework that integrates rapid embedding-based constrained retrieval models with highly effective multi-modal (M)LLM-based reranking techniques. GUI-ReRank further introduces a fully customizable GUI repository annotation and embedding pipeline, enabling users to effortlessly make their own GUI repositories searchable, which allows for rapid discovery of relevant GUIs for inspiration or seamless integration into customized LLM-based retrieval-augmented generation (RAG) workflows. We evaluated our approach on an established NL-based GUI retrieval benchmark, demonstrating that GUI-ReRank significantly outperforms state-of-the-art (SOTA) tailored Learning-to-Rank (LTR) models in both retrieval accuracy and generalizability. Additionally, we conducted a comprehensive cost and efficiency analysis of employing MLLMs for reranking, providing valuable insights regarding the trade-offs between retrieval effectiveness and computational resources. Video presentation of GUI-ReRank available at: https://youtu.be/7x9UCh82ug
Kristian Kolthoff, Felix Kretzer, Alexander Maedche, Simone Paolo Ponzetto, Christian Bartelt
ASE4
2024 Out of the Mouths of MPs: Speaker Attribution in Parliamentary Debates
abstract
This paper presents GePaDe_SpkAtt , a new corpus for speaker attribution in German parliamentary debates, with more than 7,700 manually annotated events of speech, thought and writing. Our role inventory includes the sources, addressees, messages and topics of the speech event and also two additional roles, medium and evidence. We report baseline results for the automatic prediction of speech events and their roles, with high scores for both, event triggers and roles. Then we apply our model to predict speech events in 20 years of parliamentary debates and investigate the use of factives in the rhetoric of MPs.
Ines Rehbein, Josef Ruppenhofer, Annelen Brunner, Simone Paolo Ponzetto
LREC/COLING4
2024 How to Do Politics with Words: Investigating Speech Acts in Parliamentary Debates
abstract
This paper presents a new perspective on framing through the lens of speech acts and investigates how politicians make use of different pragmatic speech act functions in political debates. To that end, we created a new resource of German parliamentary debates, annotated with fine-grained speech act types. Our hierarchical annotation scheme distinguishes between cooperation and conflict communication, further structured into six subtypes, such as informative, declarative or argumentative-critical speech acts, with 14 fine-grained classes at the lowest level. We present classification baselines on our new data and show that the fine-grained classes in our schema can be predicted with an avg. F1 of around 82.0%. We then use our classifier to analyse the use of speech acts in a large corpus of parliamentary debates over a time span from 2003–2023.
Ines Reinig, Ines Rehbein, Simone Paolo Ponzetto
LREC/COLING3
2024 Self-Elicitation of Requirements with Automated GUI Prototyping
abstract
Requirements Elicitation (RE) is a crucial activity especially in the early stages of software development. GUI prototyping has widely been adopted as one of the most effective RE techniques for user-facing software systems. However, GUI prototyping requires (i) the availability of experienced requirements analysts, (ii) typically necessitates conducting multiple joint sessions with customers and (iii) creates considerable manual effort. In this work, we propose SERGUI, a novel approach enabling the Self-Elicitation of Requirements (SER) based on an automated GUI prototyping assistant. SERGUI exploits the vast prototyping knowledge embodied in a large-scale GUI repository through Natural Language Requirements (NLR) based GUI retrieval and facilitates fast feedback through GUI prototypes. The GUI retrieval approach is closely integrated with a Large Language Model (LLM) driving the prompting-based recommendation of GUI features for the current GUI prototyping context and thus stimulating the elicitation of additional requirements. We envision SERGUI to be employed in the initial RE phase, creating an initial GUI prototype specification to be used by the analyst as a means for communicating the requirements. To measure the effectiveness of our approach, we conducted a preliminary evaluation. Video presentation of SERGUI at: https://youtu.be/pzAAB9Uht80
Kristian Kolthoff, Christian Bartelt, Simone Paolo Ponzetto, Kurt Schneider
ASE3
2024 ACLSum: A New Dataset for Aspect-based Summarization of Scientific Publications
abstract
Sotaro Takeshita, Tommaso Green, Ines Reinig, Kai Eckert, Simone Ponzetto. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Sotaro Takeshita, Tommaso Green, Ines Reinig, Kai Eckert 0001, Simone Paolo Ponzetto
NAACL-HLT5
2024 DeTikZify: Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZ
abstract
Creating high-quality scientific figures can be time-consuming and challenging, even though sketching ideas on paper is relatively easy. Furthermore, recreating existing figures that are not stored in formats preserving semantic information is equally complex. To tackle this problem, we introduce DeTikZify, a novel multimodal language model that automatically synthesizes scientific figures as semantics-preserving TikZ graphics programs based on sketches and existing figures. To achieve this, we create three new datasets: DaTikZv2, the largest TikZ dataset to date, containing over 360k human-created TikZ graphics; SketchFig, a dataset that pairs hand-drawn sketches with their corresponding scientific figures; and MetaFig, a collection of diverse scientific figures and associated metadata. We train DeTikZify on MetaFig and DaTikZv2, along with synthetically generated sketches learned from SketchFig. We also introduce an MCTS-based inference algorithm that enables DeTikZify to iteratively refine its outputs without the need for additional training. Through both automatic and human evaluation, we demonstrate that DeTikZify outperforms commercial Claude 3 and GPT-4V in synthesizing TikZ programs, with the MCTS algorithm effectively boosting its performance. We make our code, models, and datasets publicly available.
Jonas Belouadi, Simone Paolo Ponzetto, Steffen Eger
NeurIPS2
2024 Interlinking User Stories and GUI Prototyping: A Semi-Automatic LLM-Based Approach
abstract
Interactive systems are omnipresent today and the need to create graphical user interfaces (GUIs) is just as ubiq-uitous. For the elicitation and validation of requirements, GUI prototyping is a well-known and effective technique, typically employed after gathering initial user requirements represented in natural language (NL) (e.g., in the form of user stories). Un-fortunately, G UI prototyping often requires extensive resources, resulting in a costly and time-consuming process. Despite various easy-to-use prototyping tools in practice, there is often a lack of adequate resources for developing G UI prototypes based on given user requirements. In this work, we present a novel Large Language Model (LLM)-based approach providing assistance for validating the implementation of functional NL- based require-ments in a GUI prototype embedded in a prototyping tool. In particular, our approach aims to detect functional user stories that are not implemented in a G UI prototype and provides recommendations for suitable GUI components directly imple-menting the requirements. We collected requirements for existing GUIs in the form of user stories and evaluated our proposed validation and recommendation approach with this dataset. The obtained results are promising for user story validation and we demonstrate feasibility for the GUI component recommendations.
Kristian Kolthoff, Felix Kretzer, Christian Bartelt, Alexander Maedche, Simone Paolo Ponzetto
RE5
2023 Massively Multilingual Lexical Specialization of Multilingual Transformers
abstract
While pretrained language models (PLMs) primarily serve as general-purpose text encoders that can be fine-tuned for a wide variety of downstream tasks, recent work has shown that they can also be rewired to produce highquality word representations (i.e., static word embeddings) and yield good performance in type-level lexical tasks.While existing work primarily focused on the lexical specialization of single monolingual PLMs, in this work we expose massively multilingual transformers (MMTs, e.g., mBERT or XLM-R) to multilingual lexical knowledge at scale, leveraging BabelNet as the readily available rich source of multilingual and cross-lingual type-level lexical knowledge.Concretely, we use Babel-Net's multilingual synsets to create synonym pairs (or synonym-gloss pairs) across 50 languages and then subject the MMTs (mBERT and XLM-R) to a lexical specialization procedure guided by a contrastive objective.We show that such multilingual lexical specialization brings substantial gains in two standard cross-lingual lexical tasks, bilingual lexicon induction and cross-lingual word similarity, as well as in cross-lingual sentence retrieval.Crucially, we observe gains for languages unseen in specialization, indicating that multilingual lexical specialization enables generalization to languages with no lexical constraints.In a series of controlled experiments, we show that the number of specialization constraints plays a much greater role than the set of languages from which they originate.
Tommaso Green, Simone Paolo Ponzetto, Goran Glavas
ACL (1)2
2023 Vicinal Risk Minimization for Few-Shot Cross-lingual Transfer in Abusive Language Detection
abstract
Cross-lingual transfer learning from highresource to medium and low-resource languages has shown encouraging results.However, the scarcity of resources in target languages remains a challenge.In this work, we resort to data augmentation and continual pre-training for domain adaptation to improve cross-lingual abusive language detection.For data augmentation, we analyze two existing techniques based on vicinal risk minimization and propose MIXAG, a novel data augmentation method which interpolates pairs of instances based on the angle of their representations.Our experiments involve seven languages typologically distinct from English and three different domains.The results reveal that the data augmentation strategies can enhance fewshot cross-lingual abusive language detection.Specifically, we observe that consistently in all target languages, MIXAG improves significantly in multidomain and multilingual environments.Finally, we show through an error analysis how the domain adaptation can favour the class of abusive texts (reducing false negatives), but at the same time, declines the precision of the abusive language detection model.
Gretel Liz De la Peña Sarracén, Paolo Rosso, Robert Litschko, Goran Glavas, Simone Paolo Ponzetto
EMNLP5
2023 Data-driven prototyping via natural-language-based GUI retrieval
abstract
Abstract Rapid GUI prototyping has evolved into a widely applied technique in early stages of software development to facilitate the clarification and refinement of requirements. Especially high-fidelity GUI prototyping has shown to enable productive discussions with customers and mitigate potential misunderstandings, however, the benefits of applying high-fidelity GUI prototypes are accompanied by the disadvantage of being expensive and time-consuming in development and requiring experience to create. In this work, we showRaWi, a data-driven GUI prototyping approach that effectively retrieves GUIs for reuse from a large-scale semi-automatically created GUI repository for mobile apps on the basis of Natural Language (NL) searches to facilitate GUI prototyping and improve its productivity by leveraging the vast GUI prototyping knowledge embodied in the repository. Retrieved GUIs can directly be reused and adapted in the graphical editor ofRaWi. Moreover, we present a comprehensive evaluation methodology to enable (i) the systematic evaluation of NL-based GUI ranking methods through a novel high-quality gold standard and conduct an in-depth evaluation of traditional IR and state-of-the-art BERT-based models for GUI ranking, and (ii) the assessment of GUI prototyping productivity accompanied by an extensive user study in a practical GUI prototyping environment.
Kristian Kolthoff, Christian Bartelt, Simone Paolo Ponzetto
Autom. Softw. Eng.3
2023 Correction to: Data-driven prototyping via natural-language-based GUI retrieval
Kristian Kolthoff, Christian Bartelt, Simone Paolo Ponzetto
Autom. Softw. Eng.3
2022 Fair and Argumentative Language Modeling for Computational Argumentation
abstract
Although much work in NLP has focused on measuring and mitigating stereotypical bias in semantic spaces, research addressing bias in computational argumentation is still in its infancy.In this paper, we address this research gap and conduct a thorough investigation of bias in argumentative language models.To this end, we introduce AB BA , a novel resource for bias measurement specifically tailored to argumentation.We employ our resource to assess the effect of argumentative fine-tuning and debiasing on the intrinsic bias found in transformer-based language models using a lightweight adapter-based approach that is more sustainable and parameterefficient than full fine-tuning.Finally, we analyze the potential impact of language model debiasing on the performance in argument quality prediction, a downstream task of computational argumentation.Our results show that we are able to successfully and sustainably remove bias in general and argumentative language models while preserving (and sometimes improving) model performance in downstream tasks.We make all experimental code and data available at https://github.com/ umanlp/FairArgumentativeLM.
Carolin Holtermann, Anne Lauscher, Simone Paolo Ponzetto
ACL (1)3
2022 The Robotic Surgery Procedural Framebank
abstract
Robot-Assisted minimally invasive robotic surgery is the gold standard for the surgical treatment of many pathological conditions, and several manuals and academic papers describe how to perform these interventions. These high-quality, often peer-reviewed texts are the main study resource for medical personnel and consequently contain essential procedural domain-specific knowledge. The procedural knowledge therein described could be extracted, e.g., on the basis of semantic parsing models, and used to develop clinical decision support systems or even automation methods for some procedure’s steps. However, natural language understanding algorithms such as, for instance, semantic role labelers have lower efficacy and coverage issues when applied to domain others than those they are typically trained on (i.e., newswire text). To overcome this problem, starting from PropBank frames, we propose a new linguistic resource specific to the robotic-surgery domain, named Robotic Surgery Procedural Framebank (RSPF). We extract from robotic-surgical texts verbs and nouns that describe surgical actions and extend PropBank frames by adding any of new lemmas, frames or role sets required to cover missing lemmas, specific frames describing the surgical significance, or new semantic roles used in procedural surgical language. Our resource is publicly available and can be used to annotate corpora in the surgical domain to train and evaluate Semantic Role Labeling (SRL) systems in a challenging fine-grained domain setting.
Marco Bombieri, Marco Rospocher, Simone Paolo Ponzetto, Paolo Fiorini
LREC3
2022 Multi2WOZ: A Robust Multilingual Dataset and Conversational Pretraining for Task-Oriented Dialog
abstract
Chia-Chien Hung, Anne Lauscher, Ivan Vulić, Simone Ponzetto, Goran Glavaš. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Chia-Chien Hung, Anne Lauscher, Ivan Vulic, Simone Paolo Ponzetto, Goran Glavas
NAACL-HLT4
2022 On cross-lingual retrieval with multilingual text encoders
abstract
Abstract Pretrained multilingual text encoders based on neural transformer architectures , such as multilingual BERT (mBERT) and XLM, have recently become a default paradigm for cross-lingual transfer of natural language processing models, rendering cross-lingual word embedding spaces (CLWEs) effectively obsolete. In this work we present a systematic empirical study focused on the suitability of the state-of-the-art multilingual encoders for cross-lingual document and sentence retrieval tasks across a number of diverse language pairs. We first treat these models as multilingual text encoders and benchmark their performance in unsupervised ad-hoc sentence- and document-level CLIR. In contrast to supervised language understanding, our results indicate that for unsupervised document-level CLIR—a setup with no relevance judgments for IR-specific fine-tuning—pretrained multilingual encoders on average fail to significantly outperform earlier models based on CLWEs. For sentence-level retrieval, we do obtain state-of-the-art performance: the peak scores, however, are met by multilingual encoders that have been further specialized, in a supervised fashion, for sentence understanding tasks, rather than using their vanilla ‘off-the-shelf’ variants. Following these results, we introduce localized relevance matching for document-level CLIR, where we independently score a query against document sections. In the second part, we evaluate multilingual encoders fine-tuned in a supervised fashion (i.e., we learn to rank ) on English relevance data in a series of zero-shot language and domain transfer CLIR experiments. Our results show that, despite the supervision, and due to the domain and language shift, supervised re-ranking rarely improves the performance of multilingual transformers as unsupervised base rankers. Finally, only with in-domain contrastive fine-tuning (i.e., same domain, only language transfer), we manage to improve the ranking quality. We uncover substantial empirical differences between cross-lingual retrieval results and results of (zero-shot) cross-lingual transfer for monolingual retrieval in target languages, which point to “monolingual overfitting” of retrieval models trained on monolingual (English) data, even if they are based on multilingual transformers.
Robert Litschko, Ivan Vulic, Simone Paolo Ponzetto, Goran Glavas
Inf. Retr. J.3
2021 FakeFlow: Fake News Detection by Modeling the Flow of Affective Information
abstract
Fake news articles often stir the readers' attention by means of emotional appeals that arouse their feelings.Unlike in short news texts, authors of longer articles can exploit such affective factors to manipulate readers by adding exaggerations or fabricating events, in order to affect the readers' emotions.To capture this, we propose in this paper to model the flow of affective information in fake news articles using a neural architecture.The proposed model, FakeFlow, learns this flow by combining topic and affective information extracted from text.We evaluate the model's performance with several experiments on four real-world datasets.The results show that FakeFlow achieves superior results when compared against state-ofthe-art methods, thus confirming the importance of capturing the flow of the affective information in news articles.
Bilal Ghanem, Simone Paolo Ponzetto, Paolo Rosso, Francisco M. Rangel Pardo
EACL2
2021 Evaluating Multilingual Text Encoders for Unsupervised Cross-Lingual Retrieval
Robert Litschko, Ivan Vulic, Simone Paolo Ponzetto, Goran Glavas
ECIR (1)3
2021 Come hither or go away? Recognising pre-electoral coalition signals in the news
abstract
In this paper, we introduce the task of political coalition signal prediction from text, that is, the task of recognizing from the news coverage leading up to an election the (un)willingness of political parties to form a government coalition.We decompose our problem into two related, but distinct tasks: (i) predicting whether a reported statement from a politician or a journalist refers to a potential coalition and (ii) predicting the polarity of the signal -namely, whether the speaker is in favour of or against the coalition.For this, we explore the benefits of multi-task learning and investigate which setup and task formulation is best suited for each sub-task.We evaluate our approach, based on hand-coded newspaper articles, covering elections in three countries (Ireland, Germany, Austria) and two languages (English, German).Our results show that the multi-task learning approach can further improve results over a strong monolingual transfer learning baseline.* election * AND ( * coalition * OR * pact * OR * collaboration * OR * alliance * ) 1The keyword-based sampling was necessary so that the coders did not spend too much time on articles that did not contain any signals.The keywords were optimised for recall by domain experts who,
Ines Rehbein, Simone Paolo Ponzetto, Anna Adendorf, Oke Bahnsen, Lukas Stoetzer, Heiner Stuckenschmidt
EMNLP (1)2
2021 Automated Retrieval of Graphical User Interface Prototypes from Natural Language Requirements
Kristian Kolthoff, Christian Bartelt, Simone Paolo Ponzetto
NLDB3
2020 A General Framework for Implicit and Explicit Debiasing of Distributional Word Vector Spaces
abstract
Distributional word vectors have recently been shown to encode many of the human biases, most notably gender and racial biases, and models for attenuating such biases have consequently been proposed. However, existing models and studies (1) operate on under-specified and mutually differing bias definitions, (2) are tailored for a particular bias (e.g., gender bias) and (3) have been evaluated inconsistently and non-rigorously. In this work, we introduce a general framework for debiasing word embeddings. We operationalize the definition of a bias by discerning two types of bias specification: explicit and implicit. We then propose three debiasing models that operate on explicit or implicit bias specifications and that can be composed towards more robust debiasing. Finally, we devise a full-fledged evaluation framework in which we couple existing bias metrics with newly proposed ones. Experimental findings across three embedding methods suggest that the proposed debiasing models are robust and widely applicable: they often completely remove the bias both implicitly and explicitly without degradation of semantic information encoded in any of the input distributional spaces. Moreover, we successfully transfer debiasing models, by means of cross-lingual embedding spaces, and remove or attenuate biases in distributional word vector spaces of languages that lack readily available bias specifications.
Anne Lauscher, Goran Glavas, Simone Paolo Ponzetto, Ivan Vulic
AAAI3
2020 GUI2WiRe: Rapid Wireframing with a Mined and Large-Scale GUI Repository using Natural Language Requirements
abstract
High-fidelity Graphical User Interface (GUI) prototyping is a well-established and suitable method for enabling fruitful discussions, clarification and refinement of requirements formulated by customers. GUI prototypes can help to reduce misunderstandings between customers and developers, which may occur due to the ambiguity comprised in informal Natural Language (NL). However, a disadvantage of employing high-fidelity GUI prototypes is their time-consuming and expensive development. Common GUI prototyping tools are based on combining individual GUI components or manually crafted templates. In this work, we present GUI2WiRe, a tool that enables users to retrieve GUI prototypes from a semiautomatically created large-scale GUI repository for mobile applications matching user requirements specified in Natural Language (NLR). We extract multiple text segments from the GUI hierarchy data and employ various Information Retrieval (IR) models and Automatic Query Expansion (AQE) techniques to achieve ad-hoc GUI retrieval from NLR. Retrieved GUI prototypes mined from applications can be inserted in the graphical editor of GUI2WiRe to rapidly create wireframes. GUI components are extracted automatically from the GUI screenshots and basic editing functionality is provided to the user. Finally, a preview of the application is created from the wireframe to allow interactive exploration of the current design. We evaluated the applied IR and AQE approaches for their effectiveness in terms of GUI retrieval relevance on a manually annotated collection of NLR and discuss our planned user studies. Video presentation of GUI2WiRe: https://youtu.be/2nN-Xr2Hk7I
Kristian Kolthoff, Christian Bartelt, Simone Paolo Ponzetto
ASE3
2020 Word Sense Disambiguation for 158 Languages using Word Embeddings Only
abstract
Disambiguation of word senses in context is easy for humans, but is a major challenge for automatic approaches. Sophisticated supervised and knowledge-based models were developed to solve this task. However, (i) the inherent Zipfian distribution of supervised training instances for a given word and/or (ii) the quality of linguistic knowledge representations motivate the development of completely unsupervised and knowledge-free approaches to word sense disambiguation (WSD). They are particularly useful for under-resourced languages which do not have any resources for building either supervised and/or knowledge-based models. In this paper, we present a method that takes as input a standard pre-trained word embedding model and induces a fully-fledged word sense inventory, which can be used for disambiguation in context. We use this method to induce a collection of sense inventories for 158 languages on the basis of the original pre-trained fastText word embeddings by Grave et al., (2018), enabling WSD in these languages. Models and system are available online.
Varvara Logacheva, Denis Teslenko, Artem Shelmanov, Steffen Remus, Dmitry Ustalov, Andrey Kutuzov, Ekaterina Artemova, Chris Biemann, Simone Paolo Ponzetto, Alexander Panchenko
LREC9
2019 Multilingual and Cross-Lingual Graded Lexical Entailment
abstract
Grounded in cognitive linguistics, graded lexical entailment (GR-LE) is concerned with finegrained assertions regarding the directional hierarchical relationships between concepts on a continuous scale.In this paper, we present the first work on cross-lingual generalisation of GR-LE relation.Starting from Hyper-Lex, the only available GR-LE dataset in English, we construct new monolingual GR-LE datasets for three other languages, and combine those to create a set of six cross-lingual GR-LE datasets termed CL-HYPERLEX.We next present a novel method dubbed CLEAR (Cross-Lingual Lexical Entailment Attract-Repel) for effectively capturing graded (and binary) LE, both monolingually in different languages as well as across languages (i.e., on CL-HYPERLEX).Coupled with a bilingual dictionary, CLEAR leverages taxonomic LE knowledge in a resource-rich language (e.g., English) and propagates it to other languages.Supported by cross-lingual LE transfer, CLEAR sets competitive baseline performance on three new monolingual GR-LE datasets and six cross-lingual GR-LE datasets.In addition, we show that CLEAR outperforms current state-ofthe-art on binary cross-lingual LE detection by a wide margin for diverse language pairs.x y z en beagle es perro es mamífero en animal es organismo en roadster es coche en vehicle es transporte
Ivan Vulic, Simone Paolo Ponzetto, Goran Glavas
ACL (1)2
2019 Unmasking Bias in News
Javier Sánchez-Junquera, Paolo Rosso, Manuel Montes-y-Gómez, Simone Paolo Ponzetto
CICLing (1)4
2019 Policy Preference Detection in Parliamentary Debate Motions
abstract
Debate motions (proposals) tabled in the UK Parliament contain information about the stated policy preferences of the Members of Parliament who propose them, and are key to the analysis of all subsequent speeches given in response to them.We attempt to automatically label debate motions with codes from a pre-existing coding scheme developed by political scientists for the annotation and analysis of political parties' manifestos.We develop annotation guidelines for the task of applying these codes to debate motions at two levels of granularity and produce a dataset of manually labelled examples.We evaluate the annotation process and the reliability and utility of the labelling scheme, finding that inter-annotator agreement is comparable with that of other studies conducted on manifesto data.Moreover, we test a variety of ways of automatically labelling motions with the codes, ranging from similarity matching to neural classification methods, and evaluate them against the gold standard labels.From these experiments, we note that established supervised baselines are not always able to improve over simple lexical heuristics.At the same time, we detect a clear and evident benefit when employing BERT, a state-of-the-art deep language representation model, even in classification scenarios with over 30 different labels and limited amounts of training data.
Gavin Abercrombie, Federico Nanni, Riza Theresa Batista-Navarro, Simone Paolo Ponzetto
CoNLL4
2019 Watset
abstract
We present a detailed theoretical and computational analysis of the Watset meta-algorithm for fuzzy graph clustering, which has been found to be widely applicable in a variety of domains. This algorithm creates an intermediate representation of the input graph, which reflects the “ambiguity” of its nodes. Then, it uses hard clustering to discover clusters in this “disambiguated” intermediate graph. After outlining the approach and analyzing its computational complexity, we demonstrate that Watset shows competitive results in three applications: unsupervised synset induction from a synonymy graph, unsupervised semantic frame induction from dependency triples, and unsupervised semantic class induction from a distributional thesaurus. Our algorithm is generic and can also be applied to other networks of linguistic data.
Dmitry Ustalov, Alexander Panchenko, Chris Biemann, Simone Paolo Ponzetto
Comput. Linguistics4
2019 Improving lexical coverage of text simplification systems for Spanish
Sanja Stajner, Horacio Saggion, Simone Paolo Ponzetto
Expert Syst. Appl.3
2018 Investigating the Role of Argumentation in the Rhetorical Analysis of Scientific Publications with Neural Multi-Task Learning Models
abstract
Exponential growth in the number of scientific publications yields the need for effective automatic analysis of rhetorical aspects of scientific writing. Acknowledging the argumentative nature of scientific text, in this work we investigate the link between the argumentative structure of scientific publications and rhetorical aspects such as discourse categories or citation contexts. To this end, we (1) augment a corpus of scientific publications annotated with four layers of rhetoric annotations with argumentation annotations and (2) investigate neural multi-task learning architectures combining argument extraction with a set of rhetorical classification tasks. By coupling rhetorical classifiers with the extraction of argumentative components in a joint multi-task learning setting, we obtain significant performance gains for different rhetorical analysis tasks.
Anne Lauscher, Goran Glavas, Simone Paolo Ponzetto, Kai Eckert 0001
EMNLP3
2018 Efficient Pruning of Large Knowledge Graphs
abstract
In this paper we present an efficient and highly accurate algorithm to prune noisy or over-ambiguous knowledge graphs given as input an extensional definition of a domain of interest, namely as a set of instances or concepts. Our method climbs the graph in a bottom-up fashion, iteratively layering the graph and pruning nodes and edges in each layer while not compromising the connectivity of the set of input nodes. Iterative layering and protection of pre-defined nodes allow to extract semantically coherent DAG structures from noisy or over-ambiguous cyclic graphs, without loss of information and without incurring in computational bottlenecks, which are the main problem of state-of-the-art methods for cleaning large, i.e., Web-scale, knowledge graphs. We apply our algorithm to the tasks of pruning automatically acquired taxonomies using benchmarking data from a SemEval evaluation exercise, as well as the extraction of a domain-adapted taxonomy from the Wikipedia category hierarchy. The results show the superiority of our approach over state-of-art algorithms in terms of both output quality and computational efficiency.
Stefano Faralli 0001, Irene Finocchi, Simone Paolo Ponzetto, Paola Velardi
IJCAI3
2018 MIsA: Multilingual "IsA" Extraction from Corpora
Stefano Faralli 0001, Els Lefever, Simone Paolo Ponzetto
LREC3
2018 Enriching Frame Representations with Distributionally Induced Senses
Stefano Faralli 0001, Alexander Panchenko, Chris Biemann, Simone Paolo Ponzetto
LREC4
2018 Building a Web-Scale Dependency-Parsed Corpus from CommonCrawl
Alexander Panchenko, Eugen Ruppert, Stefano Faralli 0001, Simone Paolo Ponzetto, Chris Biemann
LREC4
2018 Improving Hypernymy Extraction with Distributional Semantic Classes
Alexander Panchenko, Dmitry Ustalov, Stefano Faralli 0001, Simone Paolo Ponzetto, Chris Biemann
LREC4
2018 CATS: A Tool for Customized Alignment of Text Simplification Corpora
Sanja Stajner, Marc Franco-Salvador, Paolo Rosso, Simone Paolo Ponzetto
LREC4
2018 An Unsupervised Word Sense Disambiguation System for Under-Resourced Languages
Dmitry Ustalov, Denis Teslenko, Alexander Panchenko, Mikhail Chernoskutov, Chris Biemann, Simone Paolo Ponzetto
LREC6
2018 Unsupervised Cross-Lingual Information Retrieval Using Monolingual Data Only
abstract
We propose a fully unsupervised framework for ad-hoc cross-lingual information retrieval (CLIR) which requires no bilingual data at all. The framework leverages shared cross-lingual word embedding spaces in which terms, queries, and documents can be represented, irrespective of their actual language. The shared embedding spaces are induced solely on the basis of monolingual corpora in two languages through an iterative process based on adversarial neural networks. Our experiments on the standard CLEF CLIR collections for three language pairs of varying degrees of language similarity (English-Dutch/Italian/Finnish) demonstrate the usefulness of the proposed fully unsupervised approach. Our CLIR models with unsupervised cross-lingual embeddings outperform baselines that utilize cross-lingual embeddings induced relying on word-level and document-level alignments. We then demonstrate that further improvements can be achieved by unsupervised ensemble CLIR models. We believe that the proposed framework is the first step towards development of effective CLIR models for language pairs and domains where parallel data are scarce or non-existent.
Robert Litschko, Goran Glavas, Simone Paolo Ponzetto, Ivan Vulic
SIGIR3
2018 Knowledge-rich image gist understanding beyond literal meaning
Lydia Weiland, Ioana Hulpus, Simone Paolo Ponzetto, Wolfgang Effelsberg, Laura Dietz
Data Knowl. Eng.3
2018 CrumbTrail: An efficient methodology to reduce multiple inheritance in knowledge graphs
Stefano Faralli 0001, Irene Finocchi, Simone Paolo Ponzetto, Paola Velardi
Knowl. Based Syst.3
2018 A resource-light method for cross-lingual semantic textual similarity
Goran Glavas, Marc Franco-Salvador, Simone Paolo Ponzetto, Paolo Rosso
Knowl. Based Syst.3
2018 A framework for enriching lexical semantic resources with distributional semantics
abstract
Abstract We present an approach to combining distributional semantic representations induced from text corpora with manually constructed lexical semantic networks. While both kinds of semantic resources are available with high lexical coverage, our aligned resource combines the domain specificity and availability of contextual information from distributional models with the conciseness and high quality of manually crafted lexical networks. We start with a distributional representation of induced senses of vocabulary terms, which are accompanied with rich context information given by related lexical items. We then automatically disambiguate such representations to obtain a full-fledged proto-conceptualization, i.e. a typed graph of induced word senses. In a final step, this proto-conceptualization is aligned to a lexical ontology, resulting in a hybrid aligned resource. Moreover, unmapped induced senses are associated with a semantic type in order to connect them to the core resource. Manual evaluations against ground-truth judgments for different stages of our method as well as an extrinsic evaluation on a knowledge-based Word Sense Disambiguation benchmark all indicate the high quality of the new hybrid resource. Additionally, we show the benefits of enriching top-down lexical knowledge resources with bottom-up distributional information from text for addressing high-end knowledge acquisition tasks such as cleaning hypernym graphs and learning taxonomies from scratch.
Chris Biemann, Stefano Faralli 0001, Alexander Panchenko, Simone Paolo Ponzetto
Nat. Lang. Eng.4
2017 Automatic Detection of Uncertain Statements in the Financial Domain
Christoph Kilian Theil, Sanja Stajner, Heiner Stuckenschmidt, Simone Paolo Ponzetto
CICLing (2)4
2017 The ContrastMedium Algorithm: Taxonomy Induction From Noisy Knowledge Graphs With Just A Few Links
abstract
Stefano Faralli, Alexander Panchenko, Chris Biemann, Simone Paolo Ponzetto. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017.
Stefano Faralli 0001, Alexander Panchenko, Chris Biemann, Simone Paolo Ponzetto
EACL (1)4
2017 Unsupervised Does Not Mean Uninterpretable: The Case for Word Sense Induction and Disambiguation
abstract
Alexander Panchenko, Eugen Ruppert, Stefano Faralli, Simone Paolo Ponzetto, Chris Biemann. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017.
Alexander Panchenko, Eugen Ruppert, Stefano Faralli 0001, Simone Paolo Ponzetto, Chris Biemann
EACL (1)4
2017 Dual Tensor Model for Detecting Asymmetric Lexico-Semantic Relations
abstract
Detection of lexico-semantic relations is one of the central tasks of computational semantics.Although some fundamental relations (e.g., hypernymy) are asymmetric, most existing models account for asymmetry only implicitly and use the same concept representations to support detection of symmetric and asymmetric relations alike.In this work, we propose the Dual Tensor model, a neural architecture with which we explicitly model the asymmetry and capture the translation between unspecialized and specialized word embeddings via a pair of tensors.Although our Dual Tensor model needs only unspecialized embeddings as input, our experiments on hypernymy and meronymy detection suggest that it can outperform more complex and resource-intensive models.We further demonstrate that the model can account for polysemy and that it exhibits stable performance across languages.
Goran Glavas, Simone Paolo Ponzetto
EMNLP2
2017 Topic-Based Agreement and Disagreement in US Electoral Manifestos
abstract
We present a topic-based analysis of agreement and disagreement in political manifestos, which relies on a new method for topic detection based on key concept clustering.Our approach outperforms both standard techniques like LDA and a state-of-the-art graph-based method, and provides promising initial results for this new task in computational social science.
Stefano Menini, Federico Nanni, Simone Paolo Ponzetto, Sara Tonelli
EMNLP3
2017 Automatic Assessment of Absolute Sentence Complexity
abstract
Lexically and syntactically simpler sentences result in shorter reading time and better understanding in many people. However, no reliable systems for automatic assessment of absolute sentence complexity have been proposed so far. Instead, the assessment is usually done manually, requiring expert human annotators. To address this problem, we first define the sentence complexity assessment as a five-level classification task, and build a ‘gold standard’ dataset. Next, we propose robust systems for sentence complexity assessment, using a novel set of features based on leveraging lexical properties of freely available corpora, and investigate the impact of the feature type and corpus size on the classification performance.
Sanja Stajner, Simone Paolo Ponzetto, Heiner Stuckenschmidt
IJCAI2
2017 Using Object Detection, NLP, and Knowledge Bases to Understand the Message of Images
Lydia Weiland, Ioana Hulpus, Simone Paolo Ponzetto, Laura Dietz
MMM (2)3
2017 Global RDF Vector Space Embeddings
Michael Cochez, Petar Ristoski, Simone Paolo Ponzetto, Heiko Paulheim
ISWC (1)3
2017 Large-scale taxonomy induction using entity and word embeddings
abstract
Taxonomies are an important ingredient of knowledge organization, and serve as a backbone for more sophisticated knowledge representations in intelligent systems, such as formal ontologies. However, building taxonomies manually is a costly endeavor, and hence, automatic methods for taxonomy induction are a good alternative to build large-scale taxonomies. In this paper, we propose TIEmb, an approach for automatic unsupervised class subsumption axiom extraction from knowledge bases using entity and text embeddings. We apply the approach on the WebIsA database, a database of subsumption relations extracted from the large portion of the World Wide Web, to extract class hierarchies in the Person and Place domain.
Petar Ristoski, Stefano Faralli 0001, Simone Paolo Ponzetto, Heiko Paulheim
WI3
2016 Finding Relevant Relations in Relevant Documents
Michael Schuhmacher, Benjamin Roth 0001, Simone Paolo Ponzetto, Laura Dietz
ECIR3
2016 Detecting Meaningful Compounds in Complex Class Labels
Heiner Stuckenschmidt, Simone Paolo Ponzetto, Christian Meilicke
EKAW2
2016 A Large DataBase of Hypernymy Relations Extracted from the Web
Julian Seitner, Christian Bizer, Kai Eckert 0001, Stefano Faralli 0001, Robert Meusel, Heiko Paulheim, Simone Paolo Ponzetto
LREC7
2016 Linked Disambiguated Distributional Semantic Networks
Stefano Faralli 0001, Alexander Panchenko, Chris Biemann, Simone Paolo Ponzetto
ISWC (2)4
2015 Ranking Entities for Web Queries Through Text and Knowledge
abstract
When humans explain complex topics, they naturally talk about involved entities, such as people, locations, or events. In this paper, we aim at automating this process by retrieving and ranking entities that are relevant to understand free-text web-style queries like Argentine British relations, which typically demand a set of heterogeneous entities with no specific target type like, for instance, Falklands_-War} or Margaret-_Thatcher, as answer. Standard approaches to entity retrieval rely purely on features from the knowledge base. We approach the problem from the opposite direction, namely by analyzing web documents that are found to be query-relevant. Our approach hinges on entity linking technology that identifies entity mentions and links them to a knowledge base like Wikipedia. We use a learning-to-rank approach and study different features that use documents, entity mentions, and knowledge base entities -- thus bridging document and entity retrieval. Since established benchmarks for this problem do not exist, we use TREC test collections for document ranking and collect custom relevance judgments for entities. Experiments on TREC Robust04 and TREC Web13/14 data show that: i) single entity features, like the frequency of occurrence within the top-ranke documents, or the query retrieval score against a knowledge base, perform generally well; ii) the best overall performance is achieved when combining different features that relate an entity to the query, its document mentions, and its knowledge base representation.
Michael Schuhmacher, Laura Dietz, Simone Paolo Ponzetto
CIKM3
2014 A Probabilistic Approach for Integrating Heterogeneous Knowledge Sources
Arnab Dutta 0001, Christian Meilicke, Simone Paolo Ponzetto
ESWC3
2014 DBpedia Domains: augmenting DBpedia with domain information
Gregor Titze, Volha Bryl, Cäcilia Zirn, Simone Paolo Ponzetto
LREC4
2014 Knowledge-based graph document modeling
abstract
We propose a graph-based semantic model for representing document content. Our method relies on the use of a semantic network, namely the DBpedia knowledge base, for acquiring fine-grained information about entities and their semantic relations, thus resulting in a knowledge-rich document model. We demonstrate the benefits of these semantic representations in two tasks: entity ranking and computing document semantic similarity. To this end, we couple DBpedia's structure with an information-theoretic measure of concept association, based on its explicit semantic relations, and compute semantic similarity using a Graph Edit Distance based measure, which finds the optimal matching between the documents' entities using the Hungarian method. Experimental results show that our general model outperforms baselines built on top of traditional methods, and achieves a performance close to that of highly specialized methods that have been tuned to these specific tasks.
Michael Schuhmacher, Simone Paolo Ponzetto
WSDM2
2013 Editorial
Eduard H. Hovy, Roberto Navigli, Simone Paolo Ponzetto
Artif. Intell.3
2013 Collaboratively built semi-structured content and Artificial Intelligence: The story so far
Eduard H. Hovy, Roberto Navigli, Simone Paolo Ponzetto
Artif. Intell.3
2012 BabelRelate! A Joint Multilingual Approach to Computing Semantic Relatedness
abstract
We present a knowledge-rich approach to computing semantic relatedness which exploits the joint contribution of different languages. Our approach is based on the lexicon and semantic knowledge of a wide-coverage multilingual knowledge base, which is used to compute semantic graphs in a variety of languages. Complementary information from these graphs is then combined to produce a 'core' graph where disambiguated translations are connected by means of strong semantic relations. We evaluate our approach on standard monolingual and bilingual datasets, and show that: i) we outperform a graph-based approach which does not use multilinguality in a joint way; ii) we achieve uniformly competitive results for both resource-rich and resource-poor languages.
Roberto Navigli, Simone Paolo Ponzetto
AAAI2
2012 Joining Forces Pays Off: Multilingual Joint Word Sense Disambiguation
Roberto Navigli, Simone Paolo Ponzetto
EMNLP-CoNLL2
2012 BabelNet: The automatic construction, evaluation and application of a wide-coverage multilingual semantic network
Roberto Navigli, Simone Paolo Ponzetto
Artif. Intell.2
2011 Taxonomy induction based on a collaboratively built knowledge repository
Simone Paolo Ponzetto, Michael Strube 0001
Artif. Intell.1
2010 BabelNet: Building a Very Large Multilingual Semantic Network
Roberto Navigli, Simone Paolo Ponzetto
ACL2
2010 Knowledge-Rich Word Sense Disambiguation Rivaling Supervised Systems
Simone Paolo Ponzetto, Roberto Navigli
ACL1
2010 Extending BART to Provide a Coreference Resolution System for German
Samuel Broscheit, Simone Paolo Ponzetto, Yannick Versley, Massimo Poesio
LREC2
2009 Large-Scale Taxonomy Mapping for Restructuring and Integrating Wikipedia
Simone Paolo Ponzetto, Roberto Navigli
IJCAI1
2008 WikiTaxonomy: A Large Scale Knowledge Resource
abstract
We present a taxonomy automatically generated from the system of categories in Wikipedia. Categories in the resource are identified as either classes or instances and included in a large subsumption, i.e. isa, hierarchy. The taxonomy is made available in RDFS format to the research community, e.g. for direct use within AI applications or to bootstrap the process of manual ontology creation.
Simone Paolo Ponzetto, Michael Strube 0001
ECAI1
2008 BART: A modular toolkit for coreference resolution
Yannick Versley, Simone Paolo Ponzetto, Massimo Poesio, Vladimir Eidelman, Alan Jern, Jason Smith 0006, Alessandro Moschitti
LREC2
2007 Deriving a Large-Scale Taxonomy from Wikipedia
Simone Paolo Ponzetto, Michael Strube 0001
AAAI1
2007 An API for Measuring the Relatedness of Words in Wikipedia
Simone Paolo Ponzetto, Michael Strube 0001
ACL1
2007 Knowledge Derived From Wikipedia For Computing Semantic Relatedness
abstract
Wikipedia provides a semantic network for computing semantic relatedness in a more structured fashion than a search engine and with more coverage than WordNet. We present experiments on using Wikipedia for computing semantic relatedness and compare it to WordNet on various benchmarking datasets. Existing relatedness measures perform better using Wikipedia than a baseline given by Google counts, and we show that Wikipedia outperforms WordNet on some datasets. We also address the question whether and how Wikipedia can be integrated into NLP applications as a knowledge base. Including Wikipedia improves the performance of a machine learning based coreference resolution system, indicating that it represents a valuable resource for NLP applications. Finally, we show that our method can be easily used for languages other than English by computing semantic relatedness for a German dataset.
Simone Paolo Ponzetto, Michael Strube 0001
J. Artif. Intell. Res.1
2006 WikiRelate! Computing Semantic Relatedness Using Wikipedia
Michael Strube 0001, Simone Paolo Ponzetto
AAAI2
2006 Semantic Role Labeling for Coreference Resolution
Simone Paolo Ponzetto, Michael Strube 0001
EACL1
2006 Exploiting Semantic Role Labeling, WordNet and Wikipedia for Coreference Resolution
Simone Paolo Ponzetto, Michael Strube 0001
HLT-NAACL1
2005 Semantic Role Labeling Using Lexical Statistical Information
Simone Paolo Ponzetto, Michael Strube 0001
CoNLL1