Antoine Gauquier

dblp:352/3990 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
0009-0005-9573-6364ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Efficient Crawling for Scalable Web Data Acquisition
Antoine Gauquier, Ioana Manolescu, Pierre Senellart
EDBT1
2026 Efficient and Scalable Search for Statistics
abstract
International audience
Antoine Gauquier, Simon Ebel, Helena Galhardas, Théo Galizzi, Ioana Manolescu, Aurélien Peden, Pierre Senellart
ICDE1
2025 TheoremView: A Framework for Extracting Theorem-Like Environments from Raw PDFs
Shrey Mishra, Neil Sharma, Antoine Gauquier, Pierre Senellart
ECIR (5)3
2025 Expected Shapley Value is Shapley Value for Expected Utility Game
Pratik Karmakar, Antoine Gauquier, Pierre Senellart
ECSQARU2
2025 Finding meaningful paths in heterogeneous graphs with PathWays
Nelly Barret, Antoine Gauquier, Jia Jean Law, Ioana Manolescu
Inf. Syst.2
2023 Exploring Heterogeneous Data Graphs Through Their Entity Paths
Nelly Barret, Antoine Gauquier, Jia Jean Law, Ioana Manolescu
ADBIS2
2023 Automatically Inferring the Document Class of a Scientific Article
abstract
We consider the problem of automatically inferring the (LATEX) document class used to write a scientific article from its PDF representation. Applications include improving the performance of information extraction techniques that rely on the style used in each document class, or determining the publisher of a given scientific article. We introduce two approaches: a simple classifier based on hand-coded document style features, as well as a CNN-based classifier taking as input the bitmap representation of the first page of the PDF article. We experiment on a dataset of around 100k articles from arXiv, where labels come from the source LATEX document associated to each article. Results show the CNN approach significantly outperforms that based on simple document style features, reaching over 90% average F1-score on a task to distinguish among several dozens of the most common document classes.
Antoine Gauquier, Pierre Senellart
DocEng1