VLDB 2026 Research / reviewers in the wild / expert
Cyril Labbé
dblp:l/CyrilLabbe
· DBLP profile ↗
22ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0003-4855-7038ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 10 · 2 since 2021Artificial intelligence and machine learning · 9 · 2 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 1Computer networks · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SciCiteVal: A Multi-Domain Dataset for Scientific Citation Verification
Qinyue Liu, Yongxin Zhou 0004, Cyril Labbé |
LREC | 3 |
| 2025 | From Acquiring to Suggesting DL Design Choices with Agility: A System Design
Gustavo Rodrigues dos Reis, Mario Cortes Cornax, Adrian Mos, Cyril Labbé |
RCIS (2) | 4 |
| 2024 | Data Selection Driven by Item Difficulty: On Investigating Data Efficient Practice for Hyperparameter SearchabstractFoundation Models shift the interest to adapting models instead of creating proprietary models from scratch. Despite this change, performing hyperparameter optimization (HPO) is still needed. Users adapting systems powered by those models on proprietary data should not considerably increase the overall resource footprint with extensive hyperparameter search. Given that this footprint is also proportional to the data used in HPO, we aim to investigate how a user can effectively reduce the amount of data used, leveraging the deep learning model's empirical facility to output the expected correct result for an item in the dataset. Gustavo Rodrigues dos Reis, Adrian Mos, Mario Cortes Cornax, Cyril Labbé |
CAIN | 4 |
| 2024 | Sneaked references: Fabricated reference metadata distort citation countsabstractAbstract We report evidence of an undocumented method to manipulate citation counts involving “sneaked” references. Sneaked references are registered as metadata for published scientific articles in which they do not appear. This manipulation exploits trusted relationships between various actors: publishers, the Crossref metadata registration agency, digital libraries, and bibliometric platforms. By collecting metadata from various sources, we show that extra undue references are actually sneaked in at Digital Object Identifier (DOI) registration time, resulting in artificially inflated citation counts. As a case study, focusing on three journals from a given publisher, we identified at least 9% sneaked references () mainly benefiting two authors. Despite not being present in the published articles, these sneaked references exist in metadata registries and inappropriately propagate to bibliometric dashboards. Furthermore, we discovered “lost” references: the studied bibliometric platform failed to index at least 56% () of the references present in the HTML version of the publications. This research led to an investigation by Crossref (confirming our findings) and to subsequent corrective actions. The extent of the distortion—due to sneaked and lost references—in the global literature remains unknown and requires further investigations. Bibliometric platforms producing citation counts should identify, quantify, and correct these flaws to provide accurate data to their patrons and prevent further citation gaming. Lonni Besançon, Guillaume Cabanac, Cyril Labbé, Alexander Magazinov |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2022 | Prototyping Deep Learning Applications with Non-Experts: An Assistant PropositionabstractMachine learning (ML) systems based on deep neural networks are more present than ever in software solutions for numerous industries. Their inner workings relying on models learning with data are as helpful as they are mysterious for non-expert people. There is an increasing need to make the design and development of those solutions accessible to a more general public while at the same time making them easier to explore. In this paper, to address this need, we discuss a proposition of a new assisted approach, centered on the downstream task to be performed, for helping practitioners to start using and applying Deep Learning (DL) techniques. This proposal, supported by an initial testbed UI prototype, uses an externalized form of knowledge, where JSON files compile different pipeline metadata information with their respective related artifacts (e.g., model code, the dataset to be loaded, good hyperparameter choices) that are presented as the user interacts with a conversational agent to suggest candidate solutions for a given task. Gustavo Rodrigues dos Reis, Adrian Mos, Mario Cortes Cornax, Cyril Labbé |
ASE | 4 |
| 2021 | A Method for Modeling Process Performance Indicators Variability Integrated to Customizable Processes Models
Diego Diaz, Mario Cortes Cornax, Agnès Front, Cyril Labbé, David Faure |
RCIS | 4 |
| 2021 | Prevalence of nonsensical algorithmically generated papers in the scientific literatureabstractAbstract In 2014 leading publishers withdrew more than 120 nonsensical publications automatically generated with the SCIgen program. Casual observations suggested that similar problematic papers are still published and sold, without follow‐up retractions. No systematic screening has been performed and the prevalence of such nonsensical publications in the scientific literature is unknown. Our contribution is 2‐fold. First, we designed a detector that combs the scientific literature for grammar‐based computer‐generated papers. Applied to SCIgen, it has a 83.6% precision. Second, we performed a scientometric study of the 243 detected SCIgen‐papers from 19 publishers. We estimate the prevalence of SCIgen‐papers to be 75 per million papers in Information and Computing Sciences. Only 19% of the 243 problematic papers were dealt with: formal retraction (12) or silent removal (34). Publishers still serve and sometimes sell the remaining 197 papers without any caveat. We found evidence of citation manipulation via edited SCIgen bibliographies. This work reveals metric gaming up to the point of absurdity: fraudsters publish nonsensical algorithmically generated papers featuring genuine references. It stresses the need to screen papers for nonsense before peer‐review and chase citation manipulation in published papers. Overall, this is yet another illustration of the harmful effects of the pressure to publish or perish. Guillaume Cabanac, Cyril Labbé |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2020 | Seq2SeqPy: A Lightweight and Customizable Toolkit for Neural Sequence-to-Sequence ModelingabstractWe present Seq2SeqPy a lightweight toolkit for sequence-to-sequence modeling that prioritizes simplicity and ability to customize the standard architectures easily. The toolkit supports several known architectures such as Recurrent Neural Networks, Pointer Generator Networks, and transformer model. We evaluate the toolkit on two datasets and we show that the toolkit performs similarly or even better than a very widely used sequence-to-sequence toolkit. Raheel Qader, François Portet, Cyril Labbé |
LREC | 3 |
| 2019 | Semi-Supervised Neural Text Generation by Joint Learning of Natural Language Generation and Natural Language Understanding ModelsabstractIn Natural Language Generation (NLG), Endto-End (E2E) systems trained through deep learning have recently gained a strong interest.Such deep models need a large amount of carefully annotated data to reach satisfactory performance.However, acquiring such datasets for every new NLG application is a tedious and time-consuming task.In this paper, we propose a semi-supervised deep learning scheme that can learn from non-annotated data and annotated data when available.It uses an NLG and a Natural Language Understanding (NLU) sequence-to-sequence models which are learned jointly to compensate for the lack of annotation.Experiments on two benchmark datasets show that, with limited amount of annotated data, the method can achieve very competitive results while not using any preprocessing or re-scoring tricks.These findings open the way to the exploitation of nonannotated datasets which is the current bottleneck for the E2E NLG system development to new applications. Raheel Qader, François Portet, Cyril Labbé |
INLG | 3 |
| 2019 | StreamPref: a query language for temporal conditional preferences on data streams
Marcos Roberto Ribeiro, Maria Camila Nardini Barioni, Sandra de Amo, Claudia Roncancio, Cyril Labbé |
J. Intell. Inf. Syst. | 5 |
| 2018 | Generation of Company descriptions using concept-to-text and text-to-text deep models: dataset collection and systems evaluationabstractIn this paper we study the performance of several state-of-the-art sequence-tosequence models applied to generation of short company descriptions.The models are evaluated on a newly created and publicly available company dataset that has been collected from Wikipedia.The dataset consists of around 51K company descriptions that can be used for both concept-to-text and text-to-text generation tasks.Automatic metrics and human evaluation scores computed on the generated company descriptions show promising results despite the difficulty of the task as the dataset (like most available datasets) has not been originally designed for machine learning.In addition, we perform correlation analysis between automatic metrics and human evaluations and show that certain automatic metrics are more correlated to human judgments. Raheel Qader, Khoder Jneid, François Portet, Cyril Labbé |
INLG | 4 |
| 2018 | Incremental evaluation of continuous preference queries
Marcos Roberto Ribeiro, Maria Camila Nardini Barioni, Sandra de Amo, Claudia Roncancio, Cyril Labbé |
Inf. Sci. | 5 |
| 2017 | Representing and Learning Human Behavior Patterns with Contextual Variability
Paula Andrea Lago, Claudia Roncancio, Claudia Jiménez-Guarín, Cyril Labbé |
DEXA (1) | 4 |
| 2017 | Temporal Conditional Preference Queries on Streams
Marcos Roberto Ribeiro, Maria Camila Nardini Barioni, Sandra de Amo, Claudia Roncancio, Cyril Labbé |
DEXA (1) | 5 |
| 2012 | Top-k Context-Aware Queries on Streams
Loïc Petit, Sandra de Amo, Claudia Roncancio, Cyril Labbé |
DEXA (1) | 4 |
| 2010 | iCALT: Intelligent Context-Aware Learning and Teaching EnvironmentabstractThis paper presents a novel framework for obtaining real-time feedback from students as well as enabling the teacher to leverage such feedback and understand the context and situation of the students. We present the iCALT system and discuss its educational rationale, its system operation and usage as well results from trial of the system in a classroom setting and a seminar setting. Shonali Krishnaswamy, Selby Markham, A. John Hurst, Steven Cunningham, Cyril Labbé, Behrang Saeedzadeh, Brett Gillick |
ICALT | 5 |
| 2010 | RnR: A System for Extracting Rationale from Online Reviews and Ratings
Dwi A. P. Rahayu, Shonali Krishnaswamy, Cyril Labbé, Oshadi Alahakoon |
ICSOC | 3 |
| 2010 | Adaptable cache service and application to grid cachingabstractAbstract Caching is an important element to tackle performance issues in largely distributed data management. However, caches are efficient only if they are well configured according to the context of use. As a consequence, they are usually built from scratch. Such an approach appears to be expensive and time consuming in grids where the various characteristics lead to many heterogeneous cache requirements. This paper proposes a framework facilitating the construction of sophisticated and dynamically adaptable caches for heterogeneous applications. Such a framework has enabled the evaluation of several configurations for distributed data querying systems and leads us to propose innovative approaches for semantic and cooperative caching. This paper also reports the results obtained in bioinformatics data management on grids showing the relevance of our proposals. Copyright © 2009 John Wiley & Sons, Ltd. Laurent d'Orazio, Claudia Roncancio, Cyril Labbé |
Concurr. Comput. Pract. Exp. | 3 |
| 2007 | Distributed Semantic Caching in Grid Middleware
Laurent d'Orazio, Fabrice Jouanot, Yves Denneulin, Cyril Labbé, Claudia Roncancio, Olivier Valentin |
DEXA | 4 |
| 2004 | PinS: Peer-to-Peer Interrogation and Indexing System
María-Del-Pilar Villamil, Claudia Roncancio, Cyril Labbé |
IDEAS | 3 |
| 2004 | Context Aware Mobile TransactionsabstractApplications in mobile environments are confronted to limitations imposed by wireless networks and mobile hosts - low/variable bandwidth, frequent disconnections, high communication prices, limited battery autonomy, etc. These limitations lead to several potential failures that affect data management (e.g., queries, replication, transactions) (Pitoura and Samaras, 1998). In the mobile context, in some cases these failures must be handled as "normal" or "weak performances" and not as misfunctions. In traditional environments, application designers do not care about dynamic variability of host/network characteristics. Nevertheless, in the mobile context, it is crucial to overcome infrastructure variability to better fit application/user requirements. To overcome context variability, systems need to be context aware. This work focuses on mobile transaction management. There exist several proposals for mobile transactions. They offer solutions in several aspects but few of them take into account the importance of mobile environment variability and context awareness. Patricia Serrano-Alvarado, Claudia Roncancio, Michel E. Adiba, Cyril Labbé |
Mobile Data Management | 4 |
| 2000 | Kalman and Neural Network Approaches for the Control of a VP Bandwidth in an ATM Network
Raphaël Féraud, Fabrice Clérot, Jean-Louis Simon, Daniel Pallou, Cyril Labbé, Serge Martin |
NETWORKING | 5 |