VLDB 2026 Research / reviewers in the wild / expert
Luciano Del Corro
dblp:127/0394
· DBLP profile ↗
10ranked-venue papers
3as first author
4since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Information extraction and text analysis · 34% Trustworthy machine learning · 18% Knowledge representation and reasoning · 14% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 18 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis
open information extraction |
0.8 | 3 | 2018 | Facts That Matter · EMNLP 2018 MinIE: Minimizing Facts in Open Information Extraction · EMNLP 2017 ClausIE: clause-based open information extraction · WWW 2013 |
Machine learning › Trustworthy machine learning
fairness |
0.8 | 1 | 2024 | The Greatest Good Benchmark: Measuring LLMs' Alignment with Utilitarian Moral Dilemmas · EMNLP 2024 |
Machine learning › Trustworthy machine learning › fairness
moral alignment |
0.8 | 1 | 2024 | The Greatest Good Benchmark: Measuring LLMs' Alignment with Utilitarian Moral Dilemmas · EMNLP 2024 |
Computer vision › Image recognition and object detection › text recognition › optical character recognition
Post-OCR correction |
0.5 | 1 | 2021 | Unsupervised Multi-View Post-OCR Error Correction With Language Models · EMNLP (1) 2021 |
Natural language and speech › Language models and text generation
pre-trained language model |
0.5 | 1 | 2021 | Unsupervised Multi-View Post-OCR Error Correction With Language Models · EMNLP (1) 2021 |
Natural language and speech › Language models and text generation
large language model inference |
0.3 | 1 | 2026 | A BERTology View of LLM Orchestrations: Token- and Layer-Selective Probes for Efficient Single-Pass Classification · ACL (1) 2026 |
Natural language and speech › Information extraction and text analysis › data annotation
semantic annotation |
0.3 | 1 | 2017 | MinIE: Minimizing Facts in Open Information Extraction · EMNLP 2017 |
Natural language and speech › Information extraction and text analysis
relation extraction |
0.3 | 2 | 2015 | CORE: Context-Aware Open Relation Extraction with Factorization Machines · EMNLP 2015 ClausIE: clause-based open information extraction · WWW 2013 |
Natural language and speech › Language models and text generation
large language model |
0.2 | 1 | 2024 | The Greatest Good Benchmark: Measuring LLMs' Alignment with Utilitarian Moral Dilemmas · EMNLP 2024 |
Natural language and speech › Information extraction and text analysis › entity typing
fine-grained entity typing |
0.2 | 1 | 2015 | FINET: Context-Aware Fine-Grained Named Entity Typing · EMNLP 2015 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.2 | 1 | 2015 | FINET: Context-Aware Fine-Grained Named Entity Typing · EMNLP 2015 |
Natural language and speech › Information extraction and text analysis › relation extraction
open relation extraction |
0.2 | 1 | 2015 | CORE: Context-Aware Open Relation Extraction with Factorization Machines · EMNLP 2015 |
Natural language and speech › Information extraction and text analysis › word sense disambiguation
verb sense disambiguation |
0.2 | 1 | 2014 | Werdy: Recognition and Disambiguation of Verbs and Verb Phrases with Syntactic and Semantic Pruning · EMNLP 2014 |
Natural language and speech › Information extraction and text analysis
word sense disambiguation |
0.2 | 1 | 2014 | Werdy: Recognition and Disambiguation of Verbs and Verb Phrases with Syntactic and Semantic Pruning · EMNLP 2014 |
Computer vision › Image recognition and object detection › text recognition
optical character recognition |
0.1 | 1 | 2021 | Unsupervised Multi-View Post-OCR Error Correction With Language Models · EMNLP (1) 2021 |
Natural language and speech › Language models and text generation
text summarization |
0.1 | 1 | 2018 | Facts That Matter · EMNLP 2018 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition
knowledge base construction |
0.1 | 1 | 2015 | CORE: Context-Aware Open Relation Extraction with Factorization Machines · EMNLP 2015 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › ontology › lexical ontology
wordnet |
0.1 | 1 | 2015 | FINET: Context-Aware Fine-Grained Named Entity Typing · EMNLP 2015 |
Methods — techniques the papers use, named apart from their topics
adaptive breadth-depth retrieval · 2.0probing · 1.0attention aggregation · 1.0benchmark construction · 0.8domain adaptation · 0.5pagerank · 0.3clustering · 0.3word sense disambiguation · 0.2matrix factorization · 0.2factorization machines · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A BERTology View of LLM Orchestrations: Token- and Layer-Selective Probes for Efficient Single-Pass ClassificationabstractProduction LLM systems often rely on separate models for safety and other classificationheavy steps, increasing latency, VRAM footprint, and operational complexity.We instead reuse computation already paid for by the serving LLM: we train lightweight probes on its hidden states and predict labels in the same forward pass used for generation.We frame classification as representation selection over the full token×layer hidden-state tensor, rather than committing to a fixed token or fixed layer (e.g., first-token logits or final-layer pooling).To implement this, we introduce a two-stage aggregator that (i) summarizes tokens within each layer and (ii) aggregates across layer summaries to form a single representation for classification.We instantiate this template with direct pooling, a 100K-parameter scoring-attention gate, and a downcast multi-head self-attention (MHA) probe with up to 35M trainable parameters.Across safety and sentiment benchmarks our probes improve over logit-only reuse (e.g., MULI) and are competitive with substantially larger task-specific baselines, while preserving near-serving latency and avoiding the VRAM and latency costs of a separate guardmodel pipeline.Multi-backbone experiments on dense and mixture-of-experts architectures (Llama-3.2-3B,GPT-OSS-20B, Qwen3-30B-A3B) confirm that these findings generalize beyond a single model family. Gonzalo Ariel Meyoyan, Luciano Del Corro |
ACL (1) | 2 |
| 2026 | Autonomous Knowledge Graph Exploration with Adaptive Breadth-Depth RetrievalabstractJoaquin Polonuer, Lucas Vittor, Iñaki Arango, Ayush Noori, David A. Clifton, Luciano Del Corro, Marinka Zitnik. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Joaquín Polonuer, Lucas Vittor, Iñaki Arango, Ayush Noori, David A. Clifton, Luciano Del Corro, Marinka Zitnik |
ACL (1) | 6 |
| 2024 | The Greatest Good Benchmark: Measuring LLMs' Alignment with Utilitarian Moral DilemmasabstractGiovanni Franco Gabriel Marraffini, Andrés Cotton, Noe Fabian Hsueh, Axel Fridman, Juan Wisznia, Luciano Del Corro. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Giovanni Marraffini, Andrés Cotton, Noe Hsueh, Axel Fridman, Juan Wisznia, Luciano Del Corro |
EMNLP | 6 |
| 2021 | Unsupervised Multi-View Post-OCR Error Correction With Language ModelsabstractWe investigate post-OCR correction in a setting where we have access to different OCR views of the same document.The goal of this study is to understand if a pretrained language model (LM) can be used in an unsupervised way to reconcile the different OCR views such that their combination contains fewer errors than each individual view.This approach is motivated by scenarios in which unconstrained text generation for error correction is too risky.We evaluated different pretrained LMs on two datasets and found significant gains in realistic scenarios with up to 15% WER improvement over the best OCR view.We also show the importance of domain adaptation for post-OCR correction on out-of-domain documents. Luciano Del Corro, Samuel Broscheit, Johannes Hoffart, Eliot Brenner |
EMNLP (1) | 2 |
| 2018 | Facts That MatterabstractThis work introduces fact salience: The task of generating a machine-readable representation of the most prominent information in a text document as a set of facts.We also present SALIE, the first fact salience system.SALIE is unsupervised and knowledge agnostic, based on open information extraction to detect facts in natural language text, PageRank to determine their relevance, and clustering to promote diversity.We compare SALIE with several baselines (including positional, standard for saliency tasks), and in an extrinsic evaluation, with state-of-the-art automatic text summarizers.SALIE outperforms baselines and text summarizers showing that facts are an effective way to compress information. Marco Ponza, Luciano Del Corro, Gerhard Weikum |
EMNLP | 2 |
| 2017 | MinIE: Minimizing Facts in Open Information ExtractionabstractThe goal of Open Information Extraction (OIE) is to extract surface relations and their arguments from naturallanguage text in an unsupervised, domainindependent manner.In this paper, we propose MinIE, an OIE system that aims to provide useful, compact extractions with high precision and recall.MinIE approaches these goals by (1) representing information about polarity, modality, attribution, and quantities with semantic annotations instead of in the actual extraction, and (2) identifying and removing parts that are considered overly specific.We conducted an experimental study with several real-world datasets and found that MinIE achieves competitive or higher precision and recall than most prior systems, while at the same time producing shorter, semantically enriched extractions.Pinocchio believes that the hero Superman was not actually born on beautiful Krypton.OLLIE 1 (Pinocchio, believes that, the hero [...] beautiful Krypton) 2 (Superman, was not actually born on, beautiful Krypton) 3 (Superman, was not actually born on beau.Krypton in, the hero) ClausIE 4 (Pinocchio, believes, that the hero [...] beautiful K.) 5 (the hero Superman, was not born, on beautiful Krypton) 6 (the hero Superman, was not born, on beautiful Krypton actually) Stanford OIE No extractions MinIE-C(om-7 (Superman, was born actually on, beautiful Krypton) plete) A.: fact.(-[not], CT), attrib.(Pinocchio, +, PS [believes]) 8 (Superman, was born on, beautiful Krypton) A.: fact.(-[not], CT), attrib.(Pinocchio, +, PS [believes]) 9 (Superman, "is", hero) A.: fact.(+, CT) Kiril Gashteovski, Rainer Gemulla, Luciano Del Corro |
EMNLP | 3 |
| 2015 | FINET: Context-Aware Fine-Grained Named Entity TypingabstractWe propose FINET, a system for detecting the types of named entities in short inputs-such as sentences or tweets-with respect to WordNet's super fine-grained type system.FINET generates candidate types using a sequence of multiple extractors, ranging from explicitly mentioned types to implicit types, and subsequently selects the most appropriate using ideas from word-sense disambiguation.FINET combats data scarcity and noise from existing systems: It does not rely on supervision in its extractors and generates training data for type selection from WordNet and other resources.FINET supports the most fine-grained type system so far, including types with no annotated training data.Our experiments indicate that FINET outperforms state-of-the-art methods in terms of recall, precision, and granularity of extracted types. Luciano Del Corro, Abdalghani Abujabal, Rainer Gemulla, Gerhard Weikum |
EMNLP | 1 |
| 2015 | CORE: Context-Aware Open Relation Extraction with Factorization MachinesabstractWe propose CORE, a novel matrix factorization model that leverages contextual information for open relation extraction.Our model is based on factorization machines and integrates facts from various sources, such as knowledge bases or open information extractors, as well as the context in which these facts have been observed.We argue that integrating contextual information-such as metadata about extraction sources, lexical context, or type information-significantly improves prediction performance.Open information extractors, for example, may produce extractions that are unspecific or ambiguous when taken out of context.Our experimental study on a large real-world dataset indicates that CORE has significantly better prediction performance than state-ofthe-art approaches when contextual information is available. Fabio Petroni, Luciano Del Corro, Rainer Gemulla |
EMNLP | 2 |
| 2014 | Werdy: Recognition and Disambiguation of Verbs and Verb Phrases with Syntactic and Semantic PruningabstractWord-sense recognition and disambigua-tion (WERD) is the task of identifying word phrases and their senses in natural language text. Though it is well under-stood how to disambiguate noun phrases, this task is much less studied for verbs and verbal phrases. We present Werdy, a framework for WERD with particular focus on verbs and verbal phrases. Our framework first identifies multi-word ex-pressions based on the syntactic structure of the sentence; this allows us to recog-nize both contiguous and non-contiguous phrases. We then generate a list of can-didate senses for each word or phrase, us-ing novel syntactic and semantic pruning techniques. We also construct and lever-age a new resource of pairs of senses for verbs and their object arguments. Finally, we feed the so-obtained candidate senses into standard word-sense disambiguation (WSD) methods, and boost their precision and recall. Our experiments indicate that Werdy significantly increases the perfor-mance of existing WSD methods. 1 Luciano Del Corro, Rainer Gemulla, Gerhard Weikum |
EMNLP | 1 |
| 2013 | ClausIE: clause-based open information extractionabstractWe propose ClausIE, a novel, clause-based approach to open information extraction, which extracts relations and their arguments from natural language text. ClausIE fundamentally differs from previous approaches in that it separates the detection of ``useful'' pieces of information expressed in a sentence from their representation in terms of extractions. In more detail, ClausIE exploits linguistic knowledge about the grammar of the English language to first detect clauses in an input sentence and to subsequently identify the type of each clause according to the grammatical function of its constituents. Based on this information, ClausIE is able to generate high-precision extractions; the representation of these extractions can be flexibly customized to the underlying application. ClausIE is based on dependency parsing and a small set of domain-independent lexica, operates sentence by sentence without any post-processing, and requires no training data (whether labeled or unlabeled). Our experimental study on various real-world datasets suggests that ClausIE obtains higher recall and higher precision than existing approaches, both on high-quality text as well as on noisy text as found in the web. Luciano Del Corro, Rainer Gemulla |
WWW | 1 |