Emmanuel Malherbe

dblp:153/2864 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
5since 2021 · last 2026
0009-0006-0898-6873ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author
YearPublicationVenuePosition
2026 Learned Hallucination Detection in Black-Box LLMs Using Token-Level Entropy Production Rate
Charles Moslonka, Hicham Randrianarivo, Arthur Garnier, Emmanuel Malherbe
ECIR (1)4
2025 Fairness-Aware Grouping for Continuous Sensitive Variables: Application for Debiasing Face Analysis with Respect to Skin Tone
abstract
Within a legal framework, fairness in datasets and models is typically assessed by dividing observations into predefined groups and then computing fairness measures (e.g., Disparate Impact or Equality of Odds with respect to gender). However, when sensitive attributes such as skin color are continuous, dividing into default groups may overlook or obscure the discrimination experienced by certain minority subpopulations. To address this limitation, we propose a fairness-based grouping approach for continuous (possibly multidimensional) sensitive attributes. By grouping data according to observed levels of discrimination, our method identifies the partition that maximizes a novel criterion based on inter-group variance in discrimination, thereby isolating the most critical subgroups. We validate the proposed approach using multiple synthetic datasets and demonstrate its robustness under changing population distributions—revealing how discrimination is manifested within the space of sensitive attributes. Furthermore, we examine a specialized setting of monotonic fairness for the case of skin color. Our empirical results on both CelebA and FFHQ, leveraging the skin tone as predicted by an industrial proprietary algorithm, show that the proposed segmentation uncovers more nuanced patterns of discrimination than previously reported, and that these findings remain stable across datasets for a given model. Finally, we leverage our grouping model for debiasing purpose, aiming at predicting fair scores with group-by-group post-processing. The results demonstrate that our approach improves fairness while having minimal impact on accuracy, thus confirming our partition method and opening the door for industrial deployment.
Veronika Shilova, Emmanuel Malherbe, Giovanni Palma, Laurent Risser, Jean-Michel Loubes
ECAI2
2025 Better Capturing Interactions Between Products in Retail: Revisited Negative Sampling for Basket Choice Modeling
Antoine Désir, Vincent Auriau, Martin Mozina, Emmanuel Malherbe
ECML/PKDD (9)4
2025 Harnessing Mixed Features for Imbalance Data Oversampling: Application to Bank Customers Scoring
Abdoulaye Sakho, Emmanuel Malherbe, Carl-Erik Gauthier, Erwan Scornet
ECML/PKDD (9)2
2022 Skin Tone Diagnosis in the Wild: Towards More Robust and Inclusive User Experience Using Oriented Aleatoric Uncertainty
Emmanuel Malherbe, Michel Remise, Matthieu Perrot
ACCV (2)1
2016 Bridge the terminology gap between recruiters and candidates: A multilingual skills base built from social media and linked data
abstract
A major part of the job offers and candidates profiles are now available online. Leveraging this public data, Multiposting, a subsidiary of SAP, aims at providing in real-time an exhaustive job market analysis through the SmartSearch project. One big issue in this project, and more generally in the e-recruitment and the human resources management, is to extract the skills from the raw texts in order to associate a job or a candidate to its corresponding skills. This paper proposes to generate a multilingual base of skills in a novel bottom-up approach that finds its roots from the terminology used by candidates in professional social networks. The knowledge base is built by leveraging the Linked Open Data project DBpedia, as well as the tags of a Q&A website, StackOverflow. The large-scale experiments on real-world job offers show that the coverage and precision of the skills extraction are higher using this base than existing bases. The system has been implemented in industrial context and is used daily to extract the skills from thousands of documents, leading to advanced statistics as illustrated at the end this paper.
Emmanuel Malherbe, Marie-Aude Aufaure
ASONAM1
2015 A Case-Based Approach for Easing Schema Semantic Mapping
Emmanuel Malherbe, Thomas Iwaszko, Marie-Aude Aufaure
ICCBR1
2015 From a Ranking System to a Confidence Aware Semi-automatic Classifier
abstract
Many systems rank outcomes before suggesting them to a user, such as Recommender Systems or Information Retrieval Algo- rithms. These systems require manual validation, which is time consuming and costly in industrial context. As it is the case in our industrial applications, we assume that the user's needs can be fulfilled by only one relevant outcome. We thus consider an algorithm that systematically selects the top ranked outcome. This approach requires to compute a correctness, estimating the confidence of the automatic decision, or equivalently how likely the first outcome of the ranking system is to be correct. Based on this estimation, we can apply a threshold on the correctness, above which no manual action is required; the system avoids human validation in many cases. This paper proposes a novel method to estimate this correctness based on a supervised classification approach using the manual validations available in the base coupled with a representation of the system's scores. We conducted experiments on Multiposting real-world datasets generated by algorithms used in the industry; the first algorithm categorizes a job offer, the second recommends semantic equivalents for a given expression in a nomenclature. Our approach has thereby been evaluated and compared, and showed good results on our datasets, even with a limited training base. Moreover, in our experiments, for a given threshold, the better is the correctness estimation, the more performant is the semi-automatic system, showing that the correctness estimation leads thus to a crucial efficiency gain.
Emmanuel Malherbe, Yves Vanrompay, Marie-Aude Aufaure
KES1
2015 Bringing Order to the Job Market: Efficient Job Offer Categorization in E-Recruitment
abstract
E-recruitment uses a range of web-based technologies to find, evaluate, and hire new personnel for organizations. A crucial challenge in this arena lies in the categorization of job offers: candidates and operators often explore and analyze large numbers of offers and profiles through a set of job categories. To date, recruitment organizations define job categories top-down, relying on standardized vocabularies that often fail to capture new skills and requirements that emerge from dynamic labor markets. In order to support e-recruitment, this paper presents a dynamic, bottom-up method to automatically enrich and revise job categories. The method detects novel, highly characterizing terms in a corpus of job offers, leading to a more effective categorization, and is evaluated on real-world data by Multiposting (http://www.multiposting.fr/en), a large French e-recruitment firm.
Emmanuel Malherbe, Mario Cataldi, Andrea Ballatore
SIGIR1
2014 Field selection for job categorization and recommendation to social network users
abstract
Nowadays, in the Web 2.0 reality, one of the most challenging task for companies that aim to manage and recommend job offers is to convey this enormous amount of information in a succinct and intelligent manner such to increase the performances of matching operations against users profiles/curricula and optimize the time/space complexity of these processes. With this goal, this paper presents a novel method to formalize the textual content of job offers that aims at identifying the most relevant information and fields expressed by them and leverage this compact formalization for job recommendation and profile matching in social network environments. This method has been then developed and tested in the industrial environment represented by Multiposting and Work4, world leaders in digital solutions of e-recruitment problems. In this study three classes of documents are considered: job offers, job categories and social network user profiles (as potential job candidates); each class contains several fields with textual information. The proposed representation method permits to dynamically identify those text fields, for each class, that could help a cross-matching strategy in order to preserve, from one hand, the matching/recommendation performances and, on the other hand, reduce the cost of these operations (due to a straightforward dimensionality reduction mechanism). We then evaluated and compared the presented approach showing significant improvements on both categorization and recommendation tasks by also drastically reducing their computational costs.
Emmanuel Malherbe, Mamadou Diaby, Mario Cataldi, Emmanuel Viennet, Marie-Aude Aufaure
ASONAM1