Thomas Gerald

dblp:208/4829 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0001-7510-9690ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Can Multimodal LLMs Generate Pedagogical Questions?
abstract
International audience
Thomas Gerald, Sahar Ghannay, Julie Lascar, Paul Lerner, Anne Vilnat
LREC1
2026 Format Matters: A Critical Evaluation of Output Formats for Prompting LLMs in SLU and NER
abstract
International audience
Pierre Lepagnol, Sahar Ghannay, Thomas Gerald, Christophe Servan, Sophie Rosset
LREC3
2025 Leveraging Information Retrieval to Enhance Spoken Language Understanding Prompts in Few-Shot Learning
abstract
International audience
Pierre Lepagnol, Sahar Ghannay, Thomas Gerald, Christophe Servan, Sophie Rosset
INTERSPEECH3
2024 Introducing CQuAE : A New French Contextualised Question-Answering Corpus for the Education Domain
abstract
We present a new question answering corpus in French designed to educational domain. To be useful in such domain, we have to propose more complex questions and to be able to justify the answers on validated material. We analyze some properties of this corpus. The last part of this paper will be devoted to present the first experiments we have carried out to demonstrate the value of this dataset for learning a Retrieval Augmented Genration framework. Different experiments are proposed, with an automatic evaluation. A human evaluation is proposed to confirm or infirm this automatic evaluation.
Thomas Gerald, Anne Vilnat, Sofiane Ettayeb, Louis Tamames, Patrick Paroubek
LREC/COLING1
2024 Small Language Models Are Good Too: An Empirical Study of Zero-Shot Classification
abstract
This study is part of the debate on the efficiency of large versus small language models for text classification by prompting. We assess the performance of small language models in zero-shot text classification, challenging the prevailing dominance of large models. Across 15 datasets, our investigation benchmarks language models from 77M to 40B parameters using different architectures and scoring functions. Our findings reveal that small models can effectively classify texts, getting on par with or surpassing their larger counterparts. We developed and shared a comprehensive open-source repository that encapsulates our methodologies. This research underscores the notion that bigger isn’t always better, suggesting that resource-efficient small models may offer viable solutions for specific data classification challenges.
Pierre Lepagnol, Thomas Gerald, Sahar Ghannay, Christophe Servan, Sophie Rosset
LREC/COLING2
2024 CQuAE: A new Contextualized QUestion Answering corpus on Education domain
Thomas Gerald, Louis Tamames, Sofiane Ettayeb, Ha Quang Le, Patrick Paroubek, Anne Vilnat
Data Knowl. Eng.1
2023 CoSPLADE: Contextualizing SPLADE for Conversational Information Retrieval
Thomas Gerald, Thibault Formal, Jian-Yun Nie, Benjamin Piwowarski, Laure Soulier
ECIR (1)2
2023 A hyperbolic approach for learning communities on graphs
Thomas Gerald, Hadi Zaatiti, Hatem Hajri, Nicolas Baskiotis, Olivier Schwander
Data Min. Knowl. Discov.1
2022 Does Structure Matter? Leveraging Data-to-Text Generation for Answering Complex Information Needs
Hanane Djeddal, Thomas Gerald, Laure Soulier, Karen Pinel-Sauvagnat, Lynda Tamine-Lechani
ECIR (2)2
2022 Continual Learning of Long Topic Sequences in Neural Information Retrieval
Thomas Gerald, Laure Soulier
ECIR (1)1
2017 Binary Stochastic Representations for Large Multi-class Classification
Thomas Gerald, Nicolas Baskiotis, Ludovic Denoyer
ICONIP (1)1