EDBT 2026 Demo / reviewers in the wild / expert
Guillaume Cabanac
dblp:c/GCabanac
· DBLP profile ↗
26ranked-venue papers in the field
11as first author
8since 2021 · last 2026
0000-0003-3060-6241ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 23 (9 first)Data Mining & Knowledge Discovery · 2 (1 first)Database Systems & Data Management · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Second International Workshop on Scholarly Information Access (SCOLIA 2026)
Ingo Frommholz, Christin Kreutz, Philipp Mayr 0001, Guillaume Cabanac |
ECIR (3) | 4 |
| 2025 | The First Workshop on Scholarly Information Access (SCOLIA)
Ingo Frommholz, Philipp Mayr 0001, Guillaume Cabanac, Suzan Verberne, Christin Kreutz |
ECIR (5) | 3 |
| 2024 | Bibliometric-Enhanced Information Retrieval: 14th International BIR Workshop (BIR 2024)
Ingo Frommholz, Philipp Mayr 0001, Guillaume Cabanac, Suzan Verberne |
ECIR (5) | 3 |
| 2024 | Sneaked references: Fabricated reference metadata distort citation countsabstractAbstract We report evidence of an undocumented method to manipulate citation counts involving “sneaked” references. Sneaked references are registered as metadata for published scientific articles in which they do not appear. This manipulation exploits trusted relationships between various actors: publishers, the Crossref metadata registration agency, digital libraries, and bibliometric platforms. By collecting metadata from various sources, we show that extra undue references are actually sneaked in at Digital Object Identifier (DOI) registration time, resulting in artificially inflated citation counts. As a case study, focusing on three journals from a given publisher, we identified at least 9% sneaked references () mainly benefiting two authors. Despite not being present in the published articles, these sneaked references exist in metadata registries and inappropriately propagate to bibliometric dashboards. Furthermore, we discovered “lost” references: the studied bibliometric platform failed to index at least 56% () of the references present in the HTML version of the publications. This research led to an investigation by Crossref (confirming our findings) and to subsequent corrective actions. The extent of the distortion—due to sneaked and lost references—in the global literature remains unknown and requires further investigations. Bibliometric platforms producing citation counts should identify, quantify, and correct these flaws to provide accurate data to their patrons and prevent further citation gaming. Lonni Besançon, Guillaume Cabanac, Cyril Labbé, Alexander Magazinov |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2023 | Bibliometric-Enhanced Information Retrieval: 13th International BIR Workshop (BIR 2023)
Ingo Frommholz, Philipp Mayr 0001, Guillaume Cabanac, Suzan Verberne |
ECIR (3) | 3 |
| 2022 | Bibliometric-enhanced Information Retrieval: 12th International BIR Workshop (BIR 2022)
Ingo Frommholz, Philipp Mayr 0001, Guillaume Cabanac, Suzan Verberne |
ECIR (2) | 3 |
| 2021 | Bibliometric-Enhanced Information Retrieval: 11th International BIR Workshop
Ingo Frommholz, Philipp Mayr 0001, Guillaume Cabanac, Suzan Verberne |
ECIR (2) | 3 |
| 2021 | Prevalence of nonsensical algorithmically generated papers in the scientific literatureabstractAbstract In 2014 leading publishers withdrew more than 120 nonsensical publications automatically generated with the SCIgen program. Casual observations suggested that similar problematic papers are still published and sold, without follow‐up retractions. No systematic screening has been performed and the prevalence of such nonsensical publications in the scientific literature is unknown. Our contribution is 2‐fold. First, we designed a detector that combs the scientific literature for grammar‐based computer‐generated papers. Applied to SCIgen, it has a 83.6% precision. Second, we performed a scientometric study of the 243 detected SCIgen‐papers from 19 publishers. We estimate the prevalence of SCIgen‐papers to be 75 per million papers in Information and Computing Sciences. Only 19% of the 243 problematic papers were dealt with: formal retraction (12) or silent removal (34). Publishers still serve and sometimes sell the remaining 197 papers without any caveat. We found evidence of citation manipulation via edited SCIgen bibliographies. This work reveals metric gaming up to the point of absurdity: fraudsters publish nonsensical algorithmically generated papers featuring genuine references. It stresses the need to screen papers for nonsense before peer‐review and chase citation manipulation in published papers. Overall, this is yet another illustration of the harmful effects of the pressure to publish or perish. Guillaume Cabanac, Cyril Labbé |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2020 | Bibliometric-Enhanced Information Retrieval 10th Anniversary Workshop Edition
Guillaume Cabanac, Ingo Frommholz, Philipp Mayr 0001 |
ECIR (2) | 1 |
| 2019 | Bibliometric-Enhanced Information Retrieval: 8th International BIR Workshop
Guillaume Cabanac, Ingo Frommholz, Philipp Mayr 0001 |
ECIR (2) | 1 |
| 2018 | #élysée2017fr: The 2017 French Presidential Campaign on Twitter
Ophélie Fraisier-Vannier, Guillaume Cabanac, Yoann Pitarch, Romaric Besançon, Mohand Boughanem |
ICWSM | 2 |
| 2017 | Users Are Known by the Company They Keep: Topic Models for Viewpoint Discovery in Social NetworksabstractSocial media platforms such as weblogs and social networking sites provide Internet users with an unprecedented means to express their opinions and debate on a wide range of issues. Concurrently with their growing importance in public communication, social media platforms may foster echo chambers and filter bubbles: homophily and content personalization lead users to be increasingly exposed to conforming opinions. There is therefore a need for unbiased systems able to identify and provide access to varied viewpoints. To address this task, we propose in this paper a novel unsupervised topic model, the Social Network Viewpoint Discovery Model (SNVDM). Given a specific issue (e.g., U.S. policy) as well as the text and social interactions from the users discussing this issue on a social networking site, SNVDM jointly identifies the issue's topics, the users' viewpoints, and the discourse pertaining to the different topics and viewpoints. In order to overcome the potential sparsity of the social network (i.e., some users interact with only a few other users), we propose an extension to SNVDM based on the Generalized Pólya Urn sampling scheme (SNVDM-GPU) to leverage "acquaintances of acquaintances" relationships. We benchmark the different proposed models against three baselines, namely TAM, SN-LDA, and VODUM, on a viewpoint clustering task using two real-world datasets. We thereby provide evidence that our model SNVDM and its extension SNVDM-GPU significantly outperform state-of-the-art baselines, and we show that utilizing social interactions greatly improves viewpoint clustering performance. Thibaut Thonet, Guillaume Cabanac, Mohand Boughanem, Karen Pinel-Sauvagnat |
CIKM | 2 |
| 2017 | Predicting Emotional Reaction in Social Networks
Jérémie Clos, Anil Bandhakavi, Nirmalie Wiratunga, Guillaume Cabanac |
ECIR | 4 |
| 2016 | Bibliometric-Enhanced Information Retrieval: 3rd International BIR Workshop
Philipp Mayr 0001, Ingo Frommholz, Guillaume Cabanac |
ECIR | 3 |
| 2016 | VODUM: A Topic Model Unifying Viewpoint, Topic and Opinion Discovery
Thibaut Thonet, Guillaume Cabanac, Mohand Boughanem, Karen Pinel-Sauvagnat |
ECIR | 2 |
| 2016 | Bibliogifts in LibGen? A study of a text-sharing platform driven by biblioleaks and crowdsourcingabstractResearch articles disseminate the knowledge produced by the scientific community. Access to this literature is crucial for researchers and the general public. Apparently, “bibliogifts” are available online for free from text‐sharing platforms. However, little is known about such platforms. What is the size of the underlying digital libraries? What are the topics covered? Where do these documents originally come from? This article reports on a study of the Library Genesis platform (LibGen). The 25 million documents (42 terabytes) it hosts and distributes for free are mostly research articles, textbooks, and books in English. The article collection stems from isolated, but massive, article uploads (71%) in line with a “biblioleaks” scenario, as well as from daily crowdsourcing (29%) by worldwide users of platforms such as Reddit Scholar and Sci‐Hub. By relating the DOIs registered at CrossRef and those cached at LibGen, this study reveals that 36% of all DOI articles are available for free at LibGen. This figure is even higher (68%) for three major publishers: Elsevier, Springer, and Wiley. More research is needed to understand to what extent researchers and the general public have recourse to such text‐sharing platforms and why. Guillaume Cabanac |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2015 | Research on tables and graphs in academic articles: Pitfalls and promisesabstractMany papers have appeared recently assessing the effects of using tables and graphs in scientific publications. In this brief communication, we assess some of the methodological difficulties that have arisen in this context. These difficulties encompass issues of data availability, suitability of indicators, nature and purpose of tables and graphs, and the role of supplementary information. James Hartley, Guillaume Cabanac, Marcin Kozak, Gilles Hubert 0001 |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2014 | Solo versus collaborative writing: Discrepancies in the use of tables and graphs in academic articlesabstractThe number of authors collaborating to write scientific articles has been increasing steadily, and with this collaboration, other factors have also changed, such as the length of articles and the number of citations. However, little is known about potential discrepancies in the use of tables and graphs between single and collaborating authors. In this article, we ask whether multiauthor articles contain more tables and graphs than single‐author articles, and we studied 5,180 recent articles published in six science and social sciences journals. We found that pairs and multiple authors used significantly more tables and graphs than single authors. Such findings indicate that there is a greater emphasis on the role of tables and graphs in collaborative writing, and we discuss some of the possible causes and implications of these findings. Guillaume Cabanac, Gilles Hubert 0001, James Hartley |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2013 | Issues of work-life balance among JASIST authors and editorsabstractMany dedicated scientists reject the concept of maintaining a “work–life balance.” They argue that work is actually a huge part of life. In the mind‐set of these scientists, weekdays and weekends are equally appropriate for working on their research. Although we all have encountered such people, we may wonder how widespread this condition is with other scientists in our field. This brief communication probes work–life balance issues among JASIST authors and editors. We collected and examined the publication histories for 1,533 of the 2,402 articles published in JASIST between 2001 and 2012. Although there is no rush to submit, revise, or accept papers, we found that 11% of these events happened during weekends and that this trend has been increasing since 2005. Our findings suggest that working during the weekend may be one of the ways that scientists cope with the highly demanding era of “publish or perish.” We hope that our findings will raise an awareness of the steady increases in work among scientists before it affects our work–life balance even more. Guillaume Cabanac, James Hartley |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2013 | Capitalizing on order effects in the bids of peer-reviewed conferences to secure reviews by expert refereesabstractPeer review supports scientific conferences in selecting high‐quality papers for publication. Referees are expected to evaluate submissions equitably according to objective criteria (e.g., originality of the contribution, soundness of the theory, validity of the experiments). We argue that the submission date of papers is a subjective factor playing a role in the way they are evaluated. Indeed, program committee (PC) chairs and referees process submission lists that are usually sorted by paperIDs. This order conveys chronological information, as papers are numbered sequentially upon reception. We show that order effects lead to unconscious favoring of early‐submitted papers to the detriment of later‐submitted papers. Our point is supported by a study of 42 peer‐reviewed conferences in Computer Science showing a decrease in the number of bids placed on submissions with higher paperIDs. It is advised to counterbalance order effects during the bidding phase of peer review by promoting the submissions with fewer bids to potential referees. This manipulation intends to better share bids out among submissions in order to attract qualified referees for all submissions. This would secure reviews from confident referees, who are keen on voicing sharp opinions and recommendations (acceptance or rejection) about submissions. This work contributes to the integrity of peer review, which is mandatory to maintain public trust in science. Guillaume Cabanac, Thomas Preuß |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2012 | Shaping the landscape of research in information systems from the perspective of editorial boards: A scientometric study of 77 leading journalsabstractCharacteristics of the Journal of the American Society for Information Science and Technology and 76 other journals listed in the Information Systems category of the Journal Citation Reports–Science edition 2009 were analyzed. Besides reporting usual bibliographic indicators, we investigated the human cornerstone of any peer‐reviewed journal: its editorial board. Demographic data about the 2,846 gatekeepers serving in information systems (IS) editorial boards were collected. We discuss various scientometric indicators supported by descriptive statistics. Our findings reflect the great variety of IS journals in terms of research output, author communities, editorial boards, and gatekeeper demographics (e.g., diversity in gender and location), seniority, authority, and degree of involvement in editorial boards. We believe that these results may help the general public and scholars (e.g., readers, authors, journal gatekeepers, policy makers) to revise and increase their knowledge of scholarly communication in the IS field. The EB_IS_2009 dataset supporting this scientometric study is released as online supplementary material to this article to foster further research on editorial boards. Guillaume Cabanac |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2011 | Query Operators Shown Beneficial for Improving Search Results
Gilles Hubert 0001, Guillaume Cabanac, Christian Sallaberry, Damien Palacio |
TPDL | 2 |
| 2010 | Opinion Detection in Blogs: What Is Still Missing?abstractIn recent years, a lot of work has been done in the field of Opinion Detection in blogs but most of the research is based on machine learning or lexical based approaches. The objective of this paper is to focus on Social Network based evidences that can be exploited for the task of Opinion Detection. We propose a framework that makes use of the major elements of the blogosphere for extracting opinions from blogs. Besides this, we highlight the tasks of opinion prediction and multidimensional ranking. In addition, we also discuss the challenges that researchers might face while realizing the proposed framework. At the end, we demonstrate the importance of social networking evidences by performing experimentation. Malik Muhammad Saad Missen, Mohand Boughanem, Guillaume Cabanac |
ASONAM | 3 |
| 2010 | Social validation of collective annotations: Definition and experimentabstractAbstract People taking part in argumentative debates through collective annotations face a highly cognitive task when trying to estimate the group's global opinion. In order to reduce this effort, we propose in this paper to model such debates prior to evaluating their “social validation.” Computing the degree of global confirmation (or refutation) enables the identification of consensual (or controversial) debates. Readers as well as prominent information systems may thus benefit from this information. The accuracy of the social validation measure was tested through an online study conducted with 121 participants. We compared their human perception of consensus in argumentative debates with the results of the three proposed social validation algorithms. Their efficiency in synthesizing opinions was demonstrated by the fact that they achieved an accuracy of up to 84%. Guillaume Cabanac, Max Chevalier, Claude Chrisment, Christine Julien 0002 |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2007 | An Annotation Management System for Multidimensional Databases
Guillaume Cabanac, Max Chevalier, Franck Ravat, Olivier Teste |
DaWaK | 1 |
| 2007 | An Original Usage-Based Metrics for Building a Unified View of Corporate Documents
Guillaume Cabanac, Max Chevalier, Claude Chrisment, Christine Julien 0002 |
DEXA | 1 |