Luciano Del Corro

dblp:127/0394 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Information extraction and text analysis · 34% Trustworthy machine learning · 18% Knowledge representation and reasoning · 14%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 18 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
open information extraction
0.832018
Facts That Matter · EMNLP 2018
MinIE: Minimizing Facts in Open Information Extraction · EMNLP 2017
ClausIE: clause-based open information extraction · WWW 2013
Machine learning › Trustworthy machine learning
fairness
0.812024
The Greatest Good Benchmark: Measuring LLMs' Alignment with Utilitarian Moral Dilemmas · EMNLP 2024
Machine learning › Trustworthy machine learning › fairness
moral alignment
0.812024
The Greatest Good Benchmark: Measuring LLMs' Alignment with Utilitarian Moral Dilemmas · EMNLP 2024
Computer vision › Image recognition and object detection › text recognition › optical character recognition
Post-OCR correction
0.512021
Unsupervised Multi-View Post-OCR Error Correction With Language Models · EMNLP (1) 2021
Natural language and speech › Language models and text generation
pre-trained language model
0.512021
Unsupervised Multi-View Post-OCR Error Correction With Language Models · EMNLP (1) 2021
Natural language and speech › Language models and text generation
large language model inference
0.312026
A BERTology View of LLM Orchestrations: Token- and Layer-Selective Probes for Efficient Single-Pass Classification · ACL (1) 2026
Natural language and speech › Information extraction and text analysis › data annotation
semantic annotation
0.312017
MinIE: Minimizing Facts in Open Information Extraction · EMNLP 2017
Natural language and speech › Information extraction and text analysis
relation extraction
0.322015
CORE: Context-Aware Open Relation Extraction with Factorization Machines · EMNLP 2015
ClausIE: clause-based open information extraction · WWW 2013
Natural language and speech › Language models and text generation
large language model
0.212024
The Greatest Good Benchmark: Measuring LLMs' Alignment with Utilitarian Moral Dilemmas · EMNLP 2024
Natural language and speech › Information extraction and text analysis › entity typing
fine-grained entity typing
0.212015
FINET: Context-Aware Fine-Grained Named Entity Typing · EMNLP 2015
Natural language and speech › Information extraction and text analysis
named entity recognition
0.212015
FINET: Context-Aware Fine-Grained Named Entity Typing · EMNLP 2015
Natural language and speech › Information extraction and text analysis › relation extraction
open relation extraction
0.212015
CORE: Context-Aware Open Relation Extraction with Factorization Machines · EMNLP 2015
Natural language and speech › Information extraction and text analysis › word sense disambiguation
verb sense disambiguation
0.212014
Werdy: Recognition and Disambiguation of Verbs and Verb Phrases with Syntactic and Semantic Pruning · EMNLP 2014
Natural language and speech › Information extraction and text analysis
word sense disambiguation
0.212014
Werdy: Recognition and Disambiguation of Verbs and Verb Phrases with Syntactic and Semantic Pruning · EMNLP 2014
Computer vision › Image recognition and object detection › text recognition
optical character recognition
0.112021
Unsupervised Multi-View Post-OCR Error Correction With Language Models · EMNLP (1) 2021
Natural language and speech › Language models and text generation
text summarization
0.112018
Facts That Matter · EMNLP 2018
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition
knowledge base construction
0.112015
CORE: Context-Aware Open Relation Extraction with Factorization Machines · EMNLP 2015
Knowledge, reasoning and agents › Knowledge representation and reasoning › ontology › lexical ontology
wordnet
0.112015
FINET: Context-Aware Fine-Grained Named Entity Typing · EMNLP 2015

Methods — techniques the papers use, named apart from their topics

adaptive breadth-depth retrieval · 2.0probing · 1.0attention aggregation · 1.0benchmark construction · 0.8domain adaptation · 0.5pagerank · 0.3clustering · 0.3word sense disambiguation · 0.2matrix factorization · 0.2factorization machines · 0.2
YearPublicationVenuePosition
2026 A BERTology View of LLM Orchestrations: Token- and Layer-Selective Probes for Efficient Single-Pass Classification
abstract
Production LLM systems often rely on separate models for safety and other classificationheavy steps, increasing latency, VRAM footprint, and operational complexity.We instead reuse computation already paid for by the serving LLM: we train lightweight probes on its hidden states and predict labels in the same forward pass used for generation.We frame classification as representation selection over the full token×layer hidden-state tensor, rather than committing to a fixed token or fixed layer (e.g., first-token logits or final-layer pooling).To implement this, we introduce a two-stage aggregator that (i) summarizes tokens within each layer and (ii) aggregates across layer summaries to form a single representation for classification.We instantiate this template with direct pooling, a 100K-parameter scoring-attention gate, and a downcast multi-head self-attention (MHA) probe with up to 35M trainable parameters.Across safety and sentiment benchmarks our probes improve over logit-only reuse (e.g., MULI) and are competitive with substantially larger task-specific baselines, while preserving near-serving latency and avoiding the VRAM and latency costs of a separate guardmodel pipeline.Multi-backbone experiments on dense and mixture-of-experts architectures (Llama-3.2-3B,GPT-OSS-20B, Qwen3-30B-A3B) confirm that these findings generalize beyond a single model family.
Gonzalo Ariel Meyoyan, Luciano Del Corro
ACL (1)2
2026 Autonomous Knowledge Graph Exploration with Adaptive Breadth-Depth Retrieval
abstract
Joaquin Polonuer, Lucas Vittor, Iñaki Arango, Ayush Noori, David A. Clifton, Luciano Del Corro, Marinka Zitnik. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Joaquín Polonuer, Lucas Vittor, Iñaki Arango, Ayush Noori, David A. Clifton, Luciano Del Corro, Marinka Zitnik
ACL (1)6
2024 The Greatest Good Benchmark: Measuring LLMs' Alignment with Utilitarian Moral Dilemmas
abstract
Giovanni Franco Gabriel Marraffini, Andrés Cotton, Noe Fabian Hsueh, Axel Fridman, Juan Wisznia, Luciano Del Corro. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Giovanni Marraffini, Andrés Cotton, Noe Hsueh, Axel Fridman, Juan Wisznia, Luciano Del Corro
EMNLP6
2021 Unsupervised Multi-View Post-OCR Error Correction With Language Models
abstract
We investigate post-OCR correction in a setting where we have access to different OCR views of the same document.The goal of this study is to understand if a pretrained language model (LM) can be used in an unsupervised way to reconcile the different OCR views such that their combination contains fewer errors than each individual view.This approach is motivated by scenarios in which unconstrained text generation for error correction is too risky.We evaluated different pretrained LMs on two datasets and found significant gains in realistic scenarios with up to 15% WER improvement over the best OCR view.We also show the importance of domain adaptation for post-OCR correction on out-of-domain documents.
Luciano Del Corro, Samuel Broscheit, Johannes Hoffart, Eliot Brenner
EMNLP (1)2
2018 Facts That Matter
abstract
This work introduces fact salience: The task of generating a machine-readable representation of the most prominent information in a text document as a set of facts.We also present SALIE, the first fact salience system.SALIE is unsupervised and knowledge agnostic, based on open information extraction to detect facts in natural language text, PageRank to determine their relevance, and clustering to promote diversity.We compare SALIE with several baselines (including positional, standard for saliency tasks), and in an extrinsic evaluation, with state-of-the-art automatic text summarizers.SALIE outperforms baselines and text summarizers showing that facts are an effective way to compress information.
Marco Ponza, Luciano Del Corro, Gerhard Weikum
EMNLP2
2017 MinIE: Minimizing Facts in Open Information Extraction
abstract
The goal of Open Information Extraction (OIE) is to extract surface relations and their arguments from naturallanguage text in an unsupervised, domainindependent manner.In this paper, we propose MinIE, an OIE system that aims to provide useful, compact extractions with high precision and recall.MinIE approaches these goals by (1) representing information about polarity, modality, attribution, and quantities with semantic annotations instead of in the actual extraction, and (2) identifying and removing parts that are considered overly specific.We conducted an experimental study with several real-world datasets and found that MinIE achieves competitive or higher precision and recall than most prior systems, while at the same time producing shorter, semantically enriched extractions.Pinocchio believes that the hero Superman was not actually born on beautiful Krypton.OLLIE 1 (Pinocchio, believes that, the hero [...] beautiful Krypton) 2 (Superman, was not actually born on, beautiful Krypton) 3 (Superman, was not actually born on beau.Krypton in, the hero) ClausIE 4 (Pinocchio, believes, that the hero [...] beautiful K.) 5 (the hero Superman, was not born, on beautiful Krypton) 6 (the hero Superman, was not born, on beautiful Krypton actually) Stanford OIE No extractions MinIE-C(om-7 (Superman, was born actually on, beautiful Krypton) plete) A.: fact.(-[not], CT), attrib.(Pinocchio, +, PS [believes]) 8 (Superman, was born on, beautiful Krypton) A.: fact.(-[not], CT), attrib.(Pinocchio, +, PS [believes]) 9 (Superman, "is", hero) A.: fact.(+, CT)
Kiril Gashteovski, Rainer Gemulla, Luciano Del Corro
EMNLP3
2015 FINET: Context-Aware Fine-Grained Named Entity Typing
abstract
We propose FINET, a system for detecting the types of named entities in short inputs-such as sentences or tweets-with respect to WordNet's super fine-grained type system.FINET generates candidate types using a sequence of multiple extractors, ranging from explicitly mentioned types to implicit types, and subsequently selects the most appropriate using ideas from word-sense disambiguation.FINET combats data scarcity and noise from existing systems: It does not rely on supervision in its extractors and generates training data for type selection from WordNet and other resources.FINET supports the most fine-grained type system so far, including types with no annotated training data.Our experiments indicate that FINET outperforms state-of-the-art methods in terms of recall, precision, and granularity of extracted types.
Luciano Del Corro, Abdalghani Abujabal, Rainer Gemulla, Gerhard Weikum
EMNLP1
2015 CORE: Context-Aware Open Relation Extraction with Factorization Machines
abstract
We propose CORE, a novel matrix factorization model that leverages contextual information for open relation extraction.Our model is based on factorization machines and integrates facts from various sources, such as knowledge bases or open information extractors, as well as the context in which these facts have been observed.We argue that integrating contextual information-such as metadata about extraction sources, lexical context, or type information-significantly improves prediction performance.Open information extractors, for example, may produce extractions that are unspecific or ambiguous when taken out of context.Our experimental study on a large real-world dataset indicates that CORE has significantly better prediction performance than state-ofthe-art approaches when contextual information is available.
Fabio Petroni, Luciano Del Corro, Rainer Gemulla
EMNLP2
2014 Werdy: Recognition and Disambiguation of Verbs and Verb Phrases with Syntactic and Semantic Pruning
abstract
Word-sense recognition and disambigua-tion (WERD) is the task of identifying word phrases and their senses in natural language text. Though it is well under-stood how to disambiguate noun phrases, this task is much less studied for verbs and verbal phrases. We present Werdy, a framework for WERD with particular focus on verbs and verbal phrases. Our framework first identifies multi-word ex-pressions based on the syntactic structure of the sentence; this allows us to recog-nize both contiguous and non-contiguous phrases. We then generate a list of can-didate senses for each word or phrase, us-ing novel syntactic and semantic pruning techniques. We also construct and lever-age a new resource of pairs of senses for verbs and their object arguments. Finally, we feed the so-obtained candidate senses into standard word-sense disambiguation (WSD) methods, and boost their precision and recall. Our experiments indicate that Werdy significantly increases the perfor-mance of existing WSD methods. 1
Luciano Del Corro, Rainer Gemulla, Gerhard Weikum
EMNLP1
2013 ClausIE: clause-based open information extraction
abstract
We propose ClausIE, a novel, clause-based approach to open information extraction, which extracts relations and their arguments from natural language text. ClausIE fundamentally differs from previous approaches in that it separates the detection of ``useful'' pieces of information expressed in a sentence from their representation in terms of extractions. In more detail, ClausIE exploits linguistic knowledge about the grammar of the English language to first detect clauses in an input sentence and to subsequently identify the type of each clause according to the grammatical function of its constituents. Based on this information, ClausIE is able to generate high-precision extractions; the representation of these extractions can be flexibly customized to the underlying application. ClausIE is based on dependency parsing and a small set of domain-independent lexica, operates sentence by sentence without any post-processing, and requires no training data (whether labeled or unlabeled). Our experimental study on various real-world datasets suggests that ClausIE obtains higher recall and higher precision than existing approaches, both on high-quality text as well as on noisy text as found in the web.
Luciano Del Corro, Rainer Gemulla
WWW1