Felice Dell'Orletta

dblp:34/8114 · DBLP profile ↗
← Back
43ranked-venue papers
3as first author
20since 2021 · last 2026
0000-0003-3454-9387ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 35 · 3 first-author · 16 since 2021Software engineering, systems software and programming languages · 5 · 2 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Linguistic Profiling of Transformer Embedding Geometry
abstract
Transformer language models embed tokens in high-dimensional spaces, but whether geometry reflects linguistic structure remains unclear.We analyse token representations in BERT and GPT-2, selected as canonical encoder-only and decoder-only Transformer architectures, through a linguistically-grounded geometric lens.We partition tokens from the Universal Dependencies English Web treebank by surface and syntactic features (position, length, POS, head distance and arity) and examine how their representational geometry evolves across layers.We employ complementary diagnostic metrics, including isotropy, linear and nonlinear intrinsic dimensionality, to capture distinct aspects of embedding structure.Our findings reveal that BERT maintains more isotropic and higher-dimensional subspaces, whereas GPT-2 exhibits stronger anisotropy driven by a compact cluster of sentence-initial tokens.Across models, open-class words, longer tokens, and predicates with several dependents occupy more isotropic, higher-dimensional manifolds than short function words and pre-head modifiers, indicating that semantic richness and syntactic centrality play a key role in structuring embedding space.Our analysis provides a reusable framework for profiling how linguistic abstractions organize the geometry of Transformer embeddings. 1
Lucia Domenichelli, Dominique Brunato, Felice Dell'Orletta
CoNLL3
2026 Controllable Sentence Simplification in Italian: Fine-Tuning Large Language Models on Automatically Generated Resources
Michele Papucci, Giulia Venturi, Felice Dell'Orletta
LREC3
2026 On the impact of pretraining data ordering in transformer encoder- and decoder-only language models
abstract
Pretraining large language models typically relies on randomly ordered corpora, implicitly assuming that data order has limited impact on learning. However, curriculum learning suggests that the sequence of training examples can influence optimization and representation dynamics. In this work, we systematically examine pretraining data ordering as an independent design variable for transformer-based language models, analyzing how curriculum-inspired strategies affect learning trajectories, representations, and transfer performance. We pretrain encoder-only and decoder-only models under controlled conditions, varying only the ordering of training data according to readability-based complexity proxies and their inverted variants, alongside multiple random baselines. Beyond final accuracy, we adopt a multi-dimensional evaluation framework combining intrinsic metrics, linguistic probing across training stages, downstream tasks, and geometric analyses of embedding spaces. Results indicate architecture-dependent tendencies in response to data ordering. Encoder models generally exhibit stronger sensitivity to curriculum strategies, with noticeable differences in optimization behavior, probing dynamics, and representation geometry. Decoder models appear comparatively more stable under forward curricula, with more pronounced effects emerging under inverted orderings. Probing analyses suggest that early improvements reflect differences in data exposure rather than accelerated linguistic acquisition, while later-stage effects selectively mirror properties emphasized by specific curricula. Geometric analyses show that data ordering reshapes global variance structure, often increasing anisotropy, without substantially altering nonlinear intrinsic dimensionality. Overall, data ordering functions as a selective inductive bias during pretraining, influencing learning dynamics and representational emphasis rather than consistently improving performance. These findings clarify how curriculum design interacts with transformer architectures and delineate its practical impact on pretraining outcomes.
Luca Dini, Lucia Domenichelli, Dominique Brunato, Felice Dell'Orletta
Knowl. Based Syst.4
2026 Teaming Up with Artificial Agents in Non-routine Analytical Tasks
abstract
Although AI systems are becoming increasingly common in the workplace, research on their integration into human teams remains limited. In particular, little is known about how the embodiment of artificial agents shapes collaboration and performance in non-routine analytical tasks. To address this gap, we examine how different degrees of embodiment affect team performance and conversational dynamics in a real-life escape room. Teams composed of either three humans or two humans and an artificial agent (a Box, an Avatar, or a hyper-realistic humanoid) worked together to escape the room within a time limit. Our findings show that artificial agents have an uneven impact on team outcomes, with some mixed human–AI teams performing exceptionally well and others markedly worse. Human-only teams, by contrast, display more consistent performance: they are more likely to complete all tasks successfully, although they take longer and commit more errors. We also document a suggestive non-linear relationship between embodiment and team performance. Teams interacting with more embodied agents display conversational patterns that more closely resemble human–human dialogue. Together, these findings show that embodied AI shapes collaboration in complex ways, reinforcing evidence that social cues critically guide teamwork dynamics.
Lorenzo Cominelli, Federico A. Galatolo, Caterina Giannetti, Felice Dell'Orletta, Cristiano Ciaccio, Philipp Chapkovski, Giulia Venturi
ACM Trans. Hum. Robot Interact.4
2025 Evaluating Lexical Proficiency in Neural Language Models
abstract
We present a novel evaluation framework designed to assess the lexical proficiency and linguistic creativity of Transformer-based Language Models (LMs).We validate the framework by analyzing the performance of a set of LMs of different sizes, in both mono-and multilingual configuration, across tasks involving the generation, definition, and contextual usage of lexicalized words, neologisms, and nonce words.To support these evaluations, we developed a novel dataset of lexical entries for the Italian language, including curated definitions and usage examples sourced from various online platforms.The results highlight the robustness and effectiveness of our framework in evaluating multiple dimensions of LMs' linguistic understanding and offer an insight, through the assessment of their linguistic creativity, on the lexical generalization abilities of LMs 1 .
Cristiano Ciaccio, Alessio Miaschi, Felice Dell'Orletta
ACL (1)3
2025 From Human Reading to NLM Understanding: Evaluating the Role of Eye-Tracking Data in Encoder-Based Models
abstract
Cognitive signals, particularly eye-tracking data, offer valuable insights into human language processing.Leveraging eye-gaze data from the Ghent Eye-Tracking Corpus, we conducted a series of experiments to examine how integrating knowledge of human reading behavior impacts Neural Language Models (NLMs) across multiple dimensions: task performance, attention mechanisms, and the geometry of their embedding space.We explored several fine-tuning methodologies to inject eyetracking features into the models.Our results reveal that incorporating these features does not degrade downstream task performance, enhances alignment between model attention and human attention patterns, and compresses the geometry of the embedding space 1 .
Luca Dini, Lucia Domenichelli, Dominique Brunato, Felice Dell'Orletta
ACL (1)4
2025 TEXT-CAKE: Challenging Language Models on Local Text Coherence
abstract
We present a deep investigation of encoder-based Language Models (LMs) on their abilities to detect text coherence across four languages and four text genres using a new evaluation benchmark, TEXT-CAKE. We analyze both multilingual and monolingual LMs with varying architectures and parameters in different finetuning settings. Our findings demonstrate that identifying subtle perturbations that disrupt local coherence is still a challenging task. Furthermore, our results underline the importance of using diverse text genres during pre-training and of an optimal pre-traning objective and large vocabulary size. When controlling for other parameters, deep LMs (i.e., higher number of layers) have an advantage over shallow ones, even when the total number of parameters is smaller.
Luca Dini, Dominique Brunato, Felice Dell'Orletta, Tommaso Caselli
COLING3
2025 Contextualized Counterspeech: Strategies for Adaptation, Personalization, and Evaluation
abstract
AI-generated counterspeech offers a promising and scalable strategy to curb online toxicity through direct replies that promote civil discourse. However, current counterspeech is one-size-fits-all, lacking adaptation to the moderation context and the users involved. We propose and evaluate multiple strategies for generating tailored counterspeech that is adapted to the moderation context and personalized for the moderated user. We instruct a LLaMA2-13B model to generate counterspeech, experimenting with various configurations based on different contextual information and fine-tuning strategies. We identify the configurations that generate persuasive counterspeech through a combination of quantitative indicators and human evaluations collected via a pre-registered mixed-design crowdsourcing experiment. Results show that contextualized counterspeech can significantly outperform state-of-the-art generic counterspeech in adequacy and persuasiveness, without compromising other characteristics. Our findings also reveal a poor correlation between quantitative indicators and human evaluations, suggesting that these methods assess different aspects and highlighting the need for nuanced evaluation methodologies. The effectiveness of contextualized AI-generated counterspeech and the divergence between human and algorithmic evaluations underscore the importance of increased human-AI collaboration in content moderation.
Lorenzo Cima, Alessio Miaschi, Amaury Trujillo, Marco Avvenuti, Felice Dell'Orletta, Stefano Cresci
WWW5
2025 Leveraging encoder-only large language models for mobile app review feature extraction
Quim Motger, Alessio Miaschi, Felice Dell'Orletta, Xavier Franch, Jordi Marco
Empir. Softw. Eng.3
2025 In the eyes of a language model: A comprehensive examination through eye-tracking data
abstract
Cognitive signals, particularly eye-tracking data, offer a unique lens for understanding human sentence processing. Leveraging eye-gaze data from the English and Italian section of the Multilingual Eye-Movement Corpus (MECO), we designed a series of experiments aiming at exploring whether pre-trained neural language models (NLMs) encode patterns representative of human reading behaviour and if directly incorporating this information through a fine-tuning process influences the cognitive plausibility of the model. Additionally, we sought to determine if such an impact persists through a downstream task. Our findings reveal that transformers encode eye-gaze-related information during pretraining and that explicitly integrating eye-tracking features increases model alignment with human attention. When investigating the effect of intermediate fine-tuning on eye-tracking data on the model’s performance on a downstream task, we observe that this intermediate step does not result in catastrophic forgetting, despite the very different nature of the considered downstream task. In addition, the attention mechanism of models undergoing intermediate fine-tuning remains closely aligned with human attention. In conclusion, our comprehensive evaluation of NLMs informed by human attention patterns offers great potential for advancing the growing field of eXplainable Artificial Intelligence (XAI). Grounding language models in real-world cognitive processes enables the creation of systems that not only replicate human language output but also align with the cognitive mechanisms behind reading and comprehension. This alignment with human behaviour enhances model adaptability, interpretability, and effectiveness, fostering more human-centric, transparent, and reliable AI applications across various domains. 1
Luca Dini, Luca Moroni, Dominique Brunato, Felice Dell'Orletta
Neurocomputing4
2025 Cross-lingual distillation for domain knowledge transfer with sentence transformers
abstract
Recent advancements in Natural Language Processing (NLP) have substantially enhanced language understanding. However, non-English languages, especially in specialized and low-resource domains like biomedicine, remain largely underrepresented. Bridging this gap is essential for promoting inclusivity and expanding the global applicability of NLP technologies. This study presents a cross-lingual knowledge distillation framework that utilizes sentence transformers to improve domain-specific NLP capabilities in non-English languages. Specifically, the framework focuses on biomedical text classification tasks. By aligning sentence embeddings between a teacher model trained on English biomedical corpora and a multilingual student model, the proposed method effectively transfers both domain-specific and task-specific knowledge. This alignment allows the student model to efficiently process and adapt to biomedical texts in Spanish, French, and German, particularly in low-resource settings with limited tuning data. Extensive experiments with domain-adapted models like BioBERT and multilingual BERT with machine-translated text pairs demonstrate substantial performance improvements in downstream biomedical NLP tasks. The proposed framework proves highly effective in scenarios characterized by limited training data availability. The results highlight the scalability and effectiveness of this approach, facilitating the development of robust multilingual models tailored to the biomedical domain, thus advancing global accessibility and impact in biomedical NLP applications.
Ruben Piperno, Luca Bacco, Felice Dell'Orletta, Mario Merone, Leandro Pecchia
Knowl. Based Syst.3
2024 AI 'News' Content Farms Are Easy to Make and Hard to Detect: A Case Study in Italian
abstract
Giovanni Puccetti, Anna Rogers, Chiara Alzetta, Felice Dell’Orletta, Andrea Esuli. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Giovanni Puccetti 0002, Anna Rogers, Chiara Alzetta, Felice Dell'Orletta, Andrea Esuli
ACL (1)4
2024 Linguistic Knowledge Can Enhance Encoder-Decoder Models (If You Let It)
abstract
In this paper, we explore the impact of augmenting pre-trained Encoder-Decoder models, specifically T5, with linguistic knowledge for the prediction of a target task. In particular, we investigate whether fine-tuning a T5 model on an intermediate task that predicts structural linguistic properties of sentences modifies its performance in the target task of predicting sentence-level complexity. Our study encompasses diverse experiments conducted on Italian and English datasets, employing both monolingual and multilingual T5 models at various sizes. Results obtained for both languages and in cross-lingual configurations show that linguistically motivated intermediate fine-tuning has generally a positive impact on target task performance, especially when applied to smaller models and in scenarios with limited data availability.
Alessio Miaschi, Felice Dell'Orletta, Giulia Venturi
LREC/COLING2
2024 Evaluating Large Language Models via Linguistic Profiling
abstract
Large Language Models (LLMs) undergo extensive evaluation against various benchmarks collected in established leaderboards to assess their performance across multiple tasks.However, to the best of our knowledge, there is a lack of comprehensive studies evaluating these models' linguistic abilities independent of specific tasks.In this paper, we introduce a novel evaluation methodology designed to test LLMs' sentence generation abilities under specific linguistic constraints.Drawing on the 'linguistic profiling' approach, we rigorously investigate the extent to which five LLMs of varying sizes, tested in both zero-and few-shot scenarios, effectively adhere to (morpho)syntactic constraints.Our findings shed light on the linguistic proficiency of LLMs, revealing both their capabilities and limitations in generating linguistically-constrained sentences 1 .
Alessio Miaschi, Felice Dell'Orletta, Giulia Venturi
EMNLP2
2024 T-FREX: A Transformer-based Feature Extraction Method from Mobile App Reviews
abstract
Mobile app reviews are a large-scale data source for software-related knowledge generation activities, including software maintenance, evolution and feedback analysis. Effective extraction of features (i.e., functionalities or characteristics) from these reviews is key to support analysis on the acceptance of these features, identification of relevant new feature requests and prioritization of feature development, among others. Traditional methods focus on syntactic pattern-based approaches, typically context-agnostic, evaluated on a closed set of apps, difficult to replicate and limited to a reduced set and domain of apps. Mean-while, the pervasiveness of Large Language Models (LLMs) based on the Transformer architecture in software engineering tasks lays the groundwork for empirical evaluation of the performance of these models to support feature extraction. In this study, we present T-FREX, a Transformer-based, fully automatic approach for mobile app review feature extraction. First, we collect a set of ground truth features from users in a real crowdsourced software recommendation platform and transfer them automatically into a dataset of app reviews. Then, we use this newly created dataset to fine-tune multiple LLMs on a named entity recognition task under different data configurations. We assess the performance of T- FREX with respect to this ground truth, and we complement our analysis by comparing T- FREX with a baseline method from the field. Finally, we assess the quality of new features predicted by T- FREX through an external human evaluation. Results show that T- FREX outperforms on average the traditional syntactic-based method, especially when discovering new features from a domain for which the model has been fine-tuned.
Quim Motger, Alessio Miaschi, Felice Dell'Orletta, Xavier Franch, Jordi Marco
SANER3
2024 From pre-training to fine-tuning: An in-depth analysis of Large Language Models in the biomedical domain
abstract
In this study, we delve into the adaptation and effectiveness of Transformer-based, pre-trained Large Language Models (LLMs) within the biomedical domain, a field that poses unique challenges due to its complexity and the specialized nature of its data. Building on the foundation laid by the transformative architecture of Transformers, we investigate the nuanced dynamics of LLMs through a multifaceted lens, focusing on two domain-specific tasks, i.e., Natural Language Inference (NLI) and Named Entity Recognition (NER). Our objective is to bridge the knowledge gap regarding how these models’ downstream performances correlate with their capacity to encapsulate task-relevant information. To achieve this goal, we probed and analyzed the inner encoding and attention mechanisms in LLMs, both encoder- and decoder-based, tailored for either general or biomedical-specific applications. This examination occurs before and after the models are fine-tuned across various data volumes. Our findings reveal that the models’ downstream effectiveness is intricately linked to specific patterns within their internal mechanisms, shedding light on the nuanced ways in which LLMs process and apply knowledge in the biomedical context. The source code for this paper is available at https://github.com/agnesebonfigli99/LLMs-in-the-Biomedical-Domain . • Comparison between encoder/decoder LLMs and their domain-adapted versions. • Assessment of the impact of different data volumes on Fine-Tuning. • Probing and analysis of LLMs’ internal representations and attention mechanisms. • Identification of key internal patterns linked to LLMs’ performances.
Agnese Bonfigli, Luca Bacco, Mario Merone, Felice Dell'Orletta
Artif. Intell. Medicine4
2023 A text style transfer system for reducing the physician-patient expertise gap: An analysis with automatic and human evaluations
Luca Bacco, Felice Dell'Orletta, Huiyuan Lai, Mario Merone, Malvina Nissim
Expert Syst. Appl.2
2023 On Robustness and Sensitivity of a Neural Language Model: A Case Study on Italian L1 Learner Errors
abstract
In this paper, we propose a comprehensive linguistic study aimed at assessing the implicit behavior of one of the most prominent Neural Language Models (NLM) based on Transformer architectures, BERT Devlin et al., when dealing with a particular source of noisy data, namely essays written by L1 Italian learners containing a variety of errors targeting grammar, orthography and lexicon. Differently from previous works, we focus on the pre-training stage and we devise two complementary evaluation tasks aimed at assessing the impact of errors on sentence-level inner representations in terms of semantic robustness and linguistic sensitivity. While the first evaluation perspective is meant to probe the model's ability to encode the semantic similarity between sentences also in the presence of errors, the second type of probing task evaluates the influence of errors on BERT's implicit knowledge of a set of raw and morpho-syntactic properties of a sentence. Our experiments show that BERT's ability to compute sentence similarity and to correctly encode multi-leveled linguistic information of a sentence are differently modulated by the category of errors and that the error hierarchies in terms of robustness and sensitivity change across layer-wise representations.
Alessio Miaschi, Dominique Brunato, Felice Dell'Orletta, Giulia Venturi
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 How about Time? Probing a Multilingual Language Model for Temporal Relations
abstract
This paper presents a comprehensive set of probing experiments using a multilingual language model, XLM-R, for temporal relation classification between events in four languages. Results show an advantage of contextualized embeddings over static ones and a detrimen- tal role of sentence level embeddings. While obtaining competitive results against state-of-the-art systems, our probes indicate a lack of suitable encoded information to properly address this task.
Tommaso Caselli, Irene Dini, Felice Dell'Orletta
COLING3
2022 On the Nature of BERT: Correlating Fine-Tuning and Linguistic Competence
abstract
Several studies in the literature on the interpretation of Neural Language Models (NLM) focus on the linguistic generalization abilities of pre-trained models. However, little attention is paid to how the linguistic knowledge of the models changes during the fine-tuning steps. In this paper, we contribute to this line of research by showing to what extent a wide range of linguistic phenomena are forgotten across 50 epochs of fine-tuning, and how the preserved linguistic knowledge is correlated with the resolution of the fine-tuning task. To this end, we considered a quite understudied task where linguistic information plays the main role, i.e. the prediction of the evolution of written language competence of native language learners. In addition, we investigate whether it is possible to predict the fine-tuned NLM accuracy across the 50 epochs solely relying on the assessed linguistic competence. Our results are encouraging and show a high relationship between the model’s linguistic competence and its ability to solve a linguistically-based downstream task.
Federica Merendi, Felice Dell'Orletta, Giulia Venturi
COLING2
2020 Linguistic Profiling of a Neural Language Model
abstract
In this paper we investigate the linguistic knowledge learned by a Neural Language Model (NLM) before and after a fine-tuning process and how this knowledge affects its predictions during several classification problems.We use a wide set of probing tasks, each of which corresponds to a distinct sentence-level feature extracted from different levels of linguistic annotation.We show that BERT is able to encode a wide range of linguistic characteristics, but it tends to lose this information when trained on specific downstream tasks.We also find that BERT's capacity to encode different kind of linguistic properties has a positive influence on its predictions: the more it stores readable linguistic information of a sentence, the higher will be its capacity of predicting the expected label assigned to that sentence.
Alessio Miaschi, Dominique Brunato, Felice Dell'Orletta, Giulia Venturi
COLING3
2020 "Voices of the Great War": A Richly Annotated Corpus of Italian Texts on the First World War
abstract
“Voices of the Great War” is the first large corpus of Italian historical texts dating back to the period of First World War. This corpus differs from other existing resources in several respects. First, from the linguistic point of view it gives account of the wide range of varieties in which Italian was articulated in that period, namely from a diastratic (educated vs. uneducated writers), diaphasic (low/informal vs. high/formal registers) and diatopic (regional varieties, dialects) points of view. From the historical perspective, through a collection of texts belonging to different genres it represents different views on the war and the various styles of narrating war events and experiences. The final corpus is balanced along various dimensions, corresponding to the textual genre, the language variety used, the author type and the typology of conveyed contents. The corpus is fully annotated with lemmas, part-of-speech, terminology, and named entities. Significant corpus samples representative of the different “voices” have also been enriched with meta-linguistic and syntactic information. The layer of syntactic annotation forms the first nucleus of an Italian historical treebank complying with the Universal Dependencies standard. The paper illustrates the final resource, the methodology and tools used to build it, and the Web Interface for navigating it.
Federico Boschetti, Irene De Felice, Stefano Dei Rossi, Felice Dell'Orletta, Michele Di Giorgio, Martina Miliani, Lucia C. Passaro, Angelica Puddu, Giulia Venturi, Nicola Labanca, Alessandro Lenci, Simonetta Montemagni
LREC4
2020 Profiling-UD: a Tool for Linguistic Profiling of Texts
abstract
In this paper, we introduce Profiling–UD, a new text analysis tool inspired to the principles of linguistic profiling that can support language variation research from different perspectives. It allows the extraction of more than 130 features, spanning across different levels of linguistic description. Beyond the large number of features that can be monitored, a main novelty of Profiling–UD is that it has been specifically devised to be multilingual since it is based on the Universal Dependencies framework. In the second part of the paper, we demonstrate the effectiveness of these features in a number of theoretical and applicative studies in which they were successfully used for text and author profiling.
Dominique Brunato, Andrea Cimino, Felice Dell'Orletta, Giulia Venturi, Simonetta Montemagni
LREC3
2020 Invisible to People but not to Machines: Evaluation of Style-aware HeadlineGeneration in Absence of Reliable Human Judgment
abstract
We automatically generate headlines that are expected to comply with the specific styles of two different Italian newspapers. Through a data alignment strategy and different training/testing settings, we aim at decoupling content from style and preserve the latter in generation. In order to evaluate the generated headlines’ quality in terms of their specific newspaper-compliance, we devise a fine-grained evaluation strategy based on automatic classification. We observe that our models do indeed learn newspaper-specific style. Importantly, we also observe that humans aren’t reliable judges for this task, since although familiar with the newspapers, they are not able to discern their specific styles even in the original human-written headlines. The utility of automatic evaluation goes therefore beyond saving the costs and hurdles of manual annotation, and deserves particular care in its design.
Lorenzo De Mattei, Michele Cafagna, Felice Dell'Orletta, Malvina Nissim
LREC3
2018 Is this Sentence Difficult? Do you Agree?
abstract
In this paper, we present a crowdsourcing-based approach to model the human perception of sentence complexity. We collect a large corpus of sentences rated with judgments of complexity for two typologically-different languages, Italian and English. We test our approach in two experimental scenarios aimed to investigate the contribution of a wide set of lexical, morpho-syntactic and syntactic phenomena in predicting i) the degree of agreement among annotators independently from the assigned judgment and ii) the perception of sentence complexity.
Dominique Brunato, Lorenzo De Mattei, Felice Dell'Orletta, Benedetta Iavarone, Giulia Venturi
EMNLP3
2018 Real-World Witness Detection in Social Media via Hybrid Crowdsensing
Stefano Cresci, Andrea Cimino, Marco Avvenuti, Maurizio Tesconi, Felice Dell'Orletta
ICWSM5
2018 Universal Dependencies and Quantitative Typological Trends. A Case Study on Word Order
Chiara Alzetta, Felice Dell'Orletta, Simonetta Montemagni, Giulia Venturi
LREC2
2016 PaCCSS-IT: A Parallel Corpus of Complex-Simple Sentences for Automatic Text Simplification
abstract
In this paper we present PaCCSS-IT, a Parallel Corpus of Complex-Simple Sentences for ITalian.To build the resource we develop a new method for automatically acquiring a corpus of complex-simple paired sentences able to intercept structural transformations and particularly suitable for text simplification.The method requires a wide amount of texts that can be easily extracted from the web making it suitable also for less-resourced languages.We test it on the Italian language making available the biggest Italian corpus for automatic text simplification.
Dominique Brunato, Andrea Cimino, Felice Dell'Orletta, Giulia Venturi
EMNLP3
2016 CItA: an L1 Italian Learners Corpus to Study the Development of Writing Competence
Alessia Barbagli, Pietro Lucisano, Felice Dell'Orletta, Simonetta Montemagni, Giulia Venturi
LREC3
2015 CMT and FDE: tools to bridge the gap between natural language documents and feature diagrams
abstract
A business subject who wishes to enter an established technological market is required to accurately analyse the features of the products of the different competitors. Such features are normally accessible through natural language (NL) brochures, or NL Web pages, which describe the products to potential customers. Building a feature model that hierarchically summarises the different features available in competing products can bring relevant benefits in market analysis. A company can easily visualise existing features, and reason about aspects that are not covered by the available solutions. However, designing a feature model starting from publicly available documents of existing products is a time consuming and error-prone task. In this paper, we present two tools, namely Commonality Mining Tool (CMT) and Feature Diagram Editor (FDE), which can jointly support the feature model definition process. CMT allows mining common and variant features from NL descriptions of existing products, by leveraging a natural language processing (NLP) approach based on contrastive analysis, which allows identifying domain-relevant terms from NL documents. FDE takes the commonalities and variabilities extracted by CMT, and renders them in a visual form. Moreover, FDE allows the graphical design and refinement of the final feature model, by means of an intuitive GUI.
Alessio Ferrari 0001, Giorgio Oronzo Spagnolo, Stefania Gnesi, Felice Dell'Orletta
SPLC4
2015 Crisis Mapping During Natural Disasters via Text Analysis of Social Media Messages
abstract
Recent disasters demonstrated the central role of social media during emergencies thus motivating the exploitation of such data for crisis mapping. We propose a crisis mapping system that addresses limitations of current state-of-the-art approaches by analyzing the textual content of disaster reports from a twofold perspective. A damage detection component employs a SVM classifier to detect mentions of damage among emergency reports. A novel geoparsing technique is proposed and used to perform message geolocation. We report on a case study to show how the information extracted through damage detection and message geolocation can be combined to produce accurate crisis maps. Our crisis maps clearly detect both highly and lightly damaged areas, thus opening up the possibility to prioritize rescue efforts where they are most needed.
Stefano Cresci, Andrea Cimino, Felice Dell'Orletta, Maurizio Tesconi
WISE (2)3
2014 T2K^2: a System for Automatically Extracting and Organizing Knowledge from Texts
Felice Dell'Orletta, Giulia Venturi, Andrea Cimino, Simonetta Montemagni
LREC1
2014 Measuring and Improving the Completeness of Natural Language Requirements
Alessio Ferrari 0001, Felice Dell'Orletta, Giorgio Oronzo Spagnolo, Stefania Gnesi
REFSQ2
2013 Mining commonalities and variabilities from natural language documents
abstract
A company who wishes to enter an established marked with a new, competitive product is required to analyse the product solutions of the competitors. Identifying and comparing the features provided by the other vendors might greatly help during the market analysis. However, mining common and variant features of from the publicly available documents of the competitors is a time consuming and error-prone task. In this paper, we suggest to employ a natural language processing approach based on contrastive analysis to identify commonalities and variabilities from the brochures of a group of vendors. We present a first step towards a practical application of the approach, in the the context of the market of Communications-Based Train Control (CBTC) systems.
Alessio Ferrari 0001, Giorgio Oronzo Spagnolo, Felice Dell'Orletta
SPLC3
2013 Automatic extraction of function-behaviour-state information from patents
Gualtiero Fantoni, Riccardo Apreda, Felice Dell'Orletta, M. Monge
Adv. Eng. Informatics3
2011 ULISSE: an Unsupervised Algorithm for Detecting Reliable Dependency Parses
Felice Dell'Orletta, Giulia Venturi, Simonetta Montemagni
CoNLL1
2010 A Contrastive Approach to Multi-word Extraction from Domain-specific Corpora
Francesca Bonin, Felice Dell'Orletta, Simonetta Montemagni, Giulia Venturi
LREC2
2010 Comparing the Influence of Different Treebank Annotations on Dependency Parsing
Cristina Bosco, Simonetta Montemagni, Alessandro Mazzei, Vincenzo Lombardo, Felice Dell'Orletta, Alessandro Lenci, Leonardo Lesmo, Giuseppe Attardi, Maria Simi, Alberto Lavelli, Johan Hall, Jens Nilsson 0001, Joakim Nivre
LREC5
2010 Improvements in Parsing the Index Thomisticus Treebank. Revision, Combination and a Feature Model for Medieval Latin
Marco Passarotti, Felice Dell'Orletta
LREC2
2009 Temporal Relations with Signals: The Case of Italian Temporal Prepositions
abstract
This paper presents a maximum entropy tagger for the identification of intra-sentential temporal relations between temporal expressions and eventualities mediated by temporal signals in constructions of the kind "eventuality + signal + temporal relation". The tagger reports an accuracy rate of 90.8%, outperforming the baseline (81.8%). One of the main results of this work is represented by the identification of a set of robust features which may be automatically obtained with a relative computational effort.
Tommaso Caselli, Felice Dell'Orletta, Irina Prodanof
TIME2
2008 DeSRL: A Linear-Time Semantic Role Labeling System
Massimiliano Ciaramita, Giuseppe Attardi, Felice Dell'Orletta, Mihai Surdeanu
CoNLL3
2007 Multilingual Dependency Parsing and Domain Adaptation using DeSR
Giuseppe Attardi, Felice Dell'Orletta, Maria Simi, Atanas Chanev, Massimiliano Ciaramita
EMNLP-CoNLL2
2006 Searching treebanks for functional constraints: cross-lingual experiments in grammatical relation assignment
Felice Dell'Orletta, Alessandro Lenci, Simonetta Montemagni, Vito Pirrelli
LREC1