VLDB 2026 Research / reviewers in the wild / expert
Alberto Barrón-Cedeño
dblp:40/3383
· DBLP profile ↗
19ranked-venue papers in the field
8as first author
6since 2021 · last 2025
0000-0003-4719-3420ORCID · reported
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 18 (8 first)Data Mining & Knowledge Discovery · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Natural vs programming language in LLM knowledge graph constructionabstractResearch on knowledge graph construction (KGC) has recently shown great promise also thanks to the adoption of large language models (LLM) for the automatic extraction of structured information from raw text. However, most works rely on commercial, closed-source LLMs, hindering reproducibility and accessibility. We explore KGC with smaller, open-weight LLMs and investigate whether they can be used to improve upon the results obtained by systems leveraging bigger, closed-source models. Specifically, we focus on CodeKGC, a prompting framework based on GPT-3.5. We choose a variety of models either pre-trained primarily on natural language or on code and fine-tune them on three datasets used for information extraction. We fine-tune with prompts formatted either in natural language or as Python-like scripts. In addition, we optionally train the models with prompts including chain-of-thought sections. After fine-tuning, the choice of coding vs natural language prompts has a limited impact on performance, while chain-of-thought training mostly leads to a performance decrease. Moreover, we show that a LLM can be outperformed by much smaller versions on this task, after undergoing the same amount of training. We find that in general the selected lightweight LLMs outperform the much larger CodeKGC by as much as 15–20 absolute F 1 points after fine-tuning. The results show that state-of-the-art KGC systems can be developed using smaller and open-weight models, enhancing research transparency, lowering compute requirements, and decreasing third-party API reliance. Code: https://github.com/TinfFoil/natcode-llm-kgc Paolo Gajo, Alberto Barrón-Cedeño |
Inf. Process. Manag. | 2 |
| 2024 | The CLEF-2024 CheckThat! Lab: Check-Worthiness, Subjectivity, Persuasion, Roles, Authorities, and Adversarial Robustness
Alberto Barrón-Cedeño, Firoj Alam, Tanmoy Chakraborty 0002, Tamer Elsayed, Preslav Nakov, Piotr Przybyla, Julia Maria Struß, Fatima Haouari, Maram Hasanain, Federico Ruggeri, Xingyi Song, Reem Suwaileh |
ECIR (5) | 1 |
| 2023 | The CLEF-2023 CheckThat! Lab: Checkworthiness, Subjectivity, Political Bias, Factuality, and Authority
Alberto Barrón-Cedeño, Firoj Alam, Tommaso Caselli, Giovanni Da San Martino, Tamer Elsayed, Andrea Galassi, Fatima Haouari, Federico Ruggeri, Julia Maria Struß, Rabindra Nath Nandi, Gullal Singh Cheema, Dilshod Azizov, Preslav Nakov |
ECIR (3) | 1 |
| 2023 | Tailoring and evaluating the Wikipedia for in-domain comparable corpora extractionabstractAbstract We propose a language-independent graph-based method to build à-la-carte article collections on user-defined domains from the Wikipedia. The core model is based on the exploration of the encyclopedia’s category graph and can produce both mono- and multilingual comparable collections. We run thorough experiments to assess the quality of the obtained corpora in 10 languages and 743 domains. According to an extensive manual evaluation, our graph model reaches an average precision of $$84\%$$ 84 % on in-domain articles, outperforming an alternative model based on information retrieval techniques. As manual evaluations are costly, we introduce the concept of domainness and design several automatic metrics to account for the quality of the collections. Our best metric for domainness shows a strong correlation with human judgments, representing a reasonable automatic alternative to assess the quality of domain-specific corpora. We release the toolkit with the implementation of the extraction methods, the evaluation measures and several utilities. Cristina España-Bonet, Alberto Barrón-Cedeño, Lluís Màrquez |
Knowl. Inf. Syst. | 2 |
| 2022 | The CLEF-2022 CheckThat! Lab on Fighting the COVID-19 Infodemic and Fake News Detection
Preslav Nakov, Alberto Barrón-Cedeño, Giovanni Da San Martino, Firoj Alam, Julia Maria Struß, Thomas Mandl 0001, Rubén Míguez, Tommaso Caselli, Mucahid Kutlu, Wajdi Zaghouani, Chengkai Li 0001, Shaden Shaar, Gautam Kishore Shahi, Hamdy Mubarak, Alex Nikolov, Nikolay Babulkov, Yavuz Selim Kartal, Javier Beltrán |
ECIR (2) | 2 |
| 2021 | The CLEF-2021 CheckThat! Lab on Detecting Check-Worthy Claims, Previously Fact-Checked Claims, and Fake News
Preslav Nakov, Giovanni Da San Martino, Tamer Elsayed, Alberto Barrón-Cedeño, Rubén Míguez, Shaden Shaar, Firoj Alam, Fatima Haouari, Maram Hasanain, Nikolay Babulkov, Alex Nikolov, Gautam Kishore Shahi, Julia Maria Struß, Thomas Mandl 0001 |
ECIR (2) | 4 |
| 2020 | CheckThat! at CLEF 2020: Enabling the Automatic Identification and Verification of Claims in Social Media
Alberto Barrón-Cedeño, Tamer Elsayed, Preslav Nakov, Giovanni Da San Martino, Maram Hasanain, Reem Suwaileh, Fatima Haouari |
ECIR (2) | 1 |
| 2019 | CheckThat! at CLEF 2019: Automatic Identification and Verification of Claims
Tamer Elsayed, Preslav Nakov, Alberto Barrón-Cedeño, Maram Hasanain, Reem Suwaileh, Giovanni Da San Martino, Pepa Atanasova |
ECIR (2) | 3 |
| 2019 | Third International Workshop on Recent Trends in News Information Retrieval (NewsIR'19)abstractThe journalism industry has undergone a revolution in the past decade, leading to new opportunities as well as challenges. News consumption, production and delivery have all been affected and transformed by technology Readers require new mechanisms to cope with the vast volume of information in order to be informed about news events. Reporters have begun to use natural language processing (NLP) and (IR) techniques for investigative work. Publishers and aggregators are seeking new business models, and new ways to reach and retain their audience. A shift in business models has led to a gradual shift in styles of journalism in attempts to increase page views; and, far more concerning, to real mis- and dis-information, alongside allegations of "fake news" threatening the journalistic freedom and integrity of legitimate news outlets. Social media platforms drive viewership, creating filter bubbles and an increasingly polarized readership. News documents have always been a part of research on information access and retrieval methods. Over the last few years, the IR community has increasingly recognized these challenges in journalism and opened a conversation about how we might begin to address them. Evidence of this recognition is the participation in the two previous editions of our NewsIR workshop, held in ECIR 2016 and 2018. One of the most important outcomes of those workshops is an increasing awareness in the community about the changing nature of journalism and the IR challenges it entails. To move yet another step forward, the goal of the third edition of our workshop will be to create a multidisciplinary venue that brings together news experts from both technology and journalism. This would take NewsIR from a European forum targeting mainly IR researchers, into a more inclusive and influential international forum. We hope that this new format will foster further understanding for both news professionals and IR researchers, as well as producing better outcomes for news consumers. We will address the possibilities and challenges that technology offers to the journalists, the challenges that new developments in journalism create for IR researchers, and the complexity of information access tasks for news readers. M-Dyaa Albakour, Miguel Martinez, Sylvia Tippmann, Ahmet Aker, Jonathan Stray, Shiri Dori-Hacohen, Alberto Barrón-Cedeño |
SIGIR | 7 |
| 2019 | Proppy: Organizing the news based on their propagandistic content
Alberto Barrón-Cedeño, Israa Jaradat, Giovanni Da San Martino, Preslav Nakov |
Inf. Process. Manag. | 1 |
| 2019 | Language processing and learning models for community question answering in Arabic
Salvatore Romeo, Giovanni Da San Martino, Yonatan Belinkov, Alberto Barrón-Cedeño, Mohamed Eldesouki, Kareem Darwish, Hamdy Mubarak, James R. Glass, Alessandro Moschitti |
Inf. Process. Manag. | 4 |
| 2017 | A Multiple-Instance Learning Approach to Sentence Selection for Question Ranking
Salvatore Romeo, Giovanni Da San Martino, Alberto Barrón-Cedeño, Alessandro Moschitti |
ECIR | 3 |
| 2017 | On the Use of an Intermediate Class in Boolean Crowdsourced Relevance Annotations for Learning to Rank CommentsabstractIn many Information Retrieval tasks, the boundary between classes is not well defined, and assigning a document to a specific class may be complicated, even for humans. For instance, a document which is not directly related to the user's query may still contain relevant information. In this scenario, an option is to define an intermediate class collecting ambiguous instances. Yet some natural questions arise. Is this annotation strategy convenient? how should the intermediate class be treated? To answer these questions, we explored two community question answering datasets whose comments were originally annotated with three classes. We re-annotated a subset of instances considering a binary good vs bad setting. Our main contribution is to show empirically that the inclusion of an intermediate class to assess Boolean relevance is not useful. Moreover, in case the data is already annotated with a 3-class strategy, the instances from the intermediate class can be safely removed at training time. Alberto Barrón-Cedeño, Giovanni Da San Martino, Simone Filice, Alessandro Moschitti |
SIGIR | 1 |
| 2017 | Cross-Language Question Re-RankingabstractWe study how to find relevant questions in community forums when the language of the new questions is different from that of the existing questions in the forum. In particular, we explore the Arabic-English language pair. We compare a kernel-based system with a feed-forward neural network in a scenario where a large parallel corpus is available for training a machine translation system, bilingual dictionaries, and cross-language word embeddings. We observe that both approaches degrade the performance of the system when working on the translated text, especially the kernel-based system, which depends heavily on a syntactic kernel. We address this issue using a cross-language tree kernel, which compares the original Arabic tree to the English trees of the related questions. We show that this kernel almost closes the performance gap with respect to the monolingual system. On the neural network side, we use the parallel corpus to train cross-language embeddings, which we then use to represent the Arabic input and the English related questions in the same space. The results also improve to close to those of the monolingual neural network. Overall, the kernel system shows a better performance compared to the neural network in all cases. Giovanni Da San Martino, Salvatore Romeo, Alberto Barrón-Cedeño, Shafiq R. Joty, Lluís Màrquez, Alessandro Moschitti, Preslav Nakov |
SIGIR | 3 |
| 2016 | Learning to Re-Rank Questions in Community Question Answering Using Advanced FeaturesabstractWe study the impact of different types of features for question ranking in community Question Answering: bag-of-words models (BoW), syntactic tree kernels (TKs) and rank features. It should be noted that structural kernels have never been applied to the question reranking task, i.e., question to question similarity, where they have to model paraphrase relations. Additionally, the informal text, typically present in forums, poses new challenges to the use of TKs. We compare our learning to rank (L2R) algorithms against a strong baseline given by the Google rank (GR). The results show that (i) our shallow structures used in TKs are robust enough to noisy data and (ii) improving GR requires effective BoW features and TKs along with an accurate model of GR features in the used L2R algorithm. Giovanni Da San Martino, Alberto Barrón-Cedeño, Salvatore Romeo, Antonio Uva 0001, Alessandro Moschitti |
CIKM | 2 |
| 2014 | A Comparison of Approaches for Measuring Cross-Lingual Similarity of Wikipedia Articles
Alberto Barrón-Cedeño, Monica Lestari Paramita, Paul D. Clough, Paolo Rosso |
ECIR | 1 |
| 2011 | Towards the Detection of Cross-Language Source Code Reuse
Enrique Flores, Alberto Barrón-Cedeño, Paolo Rosso, Lidia Moreno |
NLDB | 2 |
| 2010 | On the mono- and cross-language detection of text reuse and plagiarismabstractPlagiarism, the unacknowledged reuse of text, has increased in recent years due to the large amount of texts readily available. For instance, recent studies claim that nowadays a high rate of student reports include plagiarism, making manual plagiarism detection practically infeasible. Alberto Barrón-Cedeño |
SIGIR | 1 |
| 2009 | On Automatic Plagiarism Detection Based on n-Grams Comparison
Alberto Barrón-Cedeño, Paolo Rosso |
ECIR | 1 |