VLDB 2026 Research / reviewers in the wild / expert
Nicolas Sidere
dblp:32/6317 · also Nicolas Sidère
· DBLP profile ↗
17ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0001-6719-5007ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 9 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Optimizing identity documents classification in online systems: A comparative analysis
Joris Voerman, Musab Al-Ghadi, Nicolas Sidere, Mickaël Coustaty, Olivier Lessard |
Int. J. Document Anal. Recognit. | 3 |
| 2025 | Multidisciplinary End-to-End Document-Level Relation Extraction from Scientific Literature
Julien Delaunay, Tran Thi Hong Hanh, Carlos E. González-Gallardo, Georgeta Bordea, Nicolas Sidere, Antoine Doucet, Olivier de Viron |
ICDAR (4) | 5 |
| 2024 | Experimental study of rehearsal-based incremental classification of document streams
Usman Malik, Muriel Visani, Nicolas Sidere, Mickaël Coustaty, Aurélie Joseph |
Int. J. Document Anal. Recognit. | 3 |
| 2024 | Identifying fraudulent identity documents by analyzing imprinted guilloche patterns
Musab Al-Ghadi, Tanmoy Mondal, Zuheng Ming, Petra Gomez-Krämer, Mickaël Coustaty, Nicolas Sidere, Jean-Christophe Burie |
Multim. Tools Appl. | 6 |
| 2023 | Incremental Learning and Ambiguity Rejection for Document Classification
Tri-Cong Pham, Mickaël Coustaty, Aurélie Joseph, Vincent Poulain D'Andecy, Muriel Visani, Nicolas Sidere |
ICDAR (5) | 6 |
| 2023 | Receipt Dataset for Document Forgery Detection
Beatriz Martínez Tornés, Théo Taburet, Emanuela Boros, Kais Rouis, Antoine Doucet, Petra Gomez-Krämer, Nicolas Sidere, Vincent Poulain D'Andecy |
ICDAR (3) | 7 |
| 2023 | Guilloche Detection for ID Authentication: A Dataset and BaselinesabstractIn cases of digital enrolment via mobile and online services, identity documents (IDs) verification is critical to efficiently detect forgery and therefore build user trust in the digital world. In this paper, we propose a copy-move public dataset, called FMIDV (forged mobile ID video dataset) containing forged IDs with respect to guilloche patterns. Also, we propose two fraud detection models on guilloche patterns of IDs, which are based on contrastive and adversarial learning. In the sequel, each proposed model manages to read the entire ID and to recognize the guilloche pattern to check its similarity to the pattern of an authentic ID. The objective of the similarity check is to validate its authenticity or its rejection. Experiments are conducted on MIDV and FMIDV datasets to analyze and identify the most proper parameters to achieve higher authentication performance. The code and the dataset are available at https://github.com/malghadi/CheckID. Musab Al-Ghadi, Zuheng Ming, Petra Gomez-Krämer, Jean-Christophe Burie, Mickaël Coustaty, Nicolas Sidere |
MMSP | 6 |
| 2023 | In-depth analysis of the impact of OCR errors on named entity recognition and linkingabstractAbstract Named entities (NEs) are among the most relevant type of information that can be used to properly index digital documents and thus easily retrieve them. It has long been observed that NEs are key to accessing the contents of digital library portals as they are contained in most user queries. However, most digitized documents are indexed through their optical character recognition (OCRed) version which include numerous errors. Although OCR engines have considerably improved over the last few years, OCR errors still considerably impact document access. Previous works were conducted to evaluate the impact of OCR errors on named entity recognition (NER) and named entity linking (NEL) techniques separately. In this article, we experimented with a variety of OCRed documents with different levels and types of OCR noise to assess in depth the impact of OCR on named entity processing. We provide a deep analysis of OCR errors that impact the performance of NER and NEL. We then present the resulting exhaustive study and subsequent recommendations on the adequate documents, the OCR quality levels, and the post-OCR correction strategies required to perform reliable NER and NEL. Ahmed Hamdi, Elvys Linhares Pontes, Nicolas Sidere, Mickaël Coustaty, Antoine Doucet |
Nat. Lang. Eng. | 3 |
| 2020 | Alleviating Digitization Errors in Named Entity Recognition for Historical DocumentsabstractEmanuela Boros, Ahmed Hamdi, Elvys Linhares Pontes, Luis Adrián Cabrera-Diego, Jose G. Moreno, Nicolas Sidere, Antoine Doucet. Proceedings of the 24th Conference on Computational Natural Language Learning. 2020. Emanuela Boros, Ahmed Hamdi, Elvys Linhares Pontes, Luis Adrián Cabrera-Diego, José G. Moreno 0001, Nicolas Sidere, Antoine Doucet |
CoNLL | 6 |
| 2020 | Assessing and Minimizing the Impact of OCR Quality on Named Entity Recognition
Ahmed Hamdi, Axel Jean-Caurant, Nicolas Sidere, Mickaël Coustaty, Antoine Doucet |
TPDL | 3 |
| 2019 | A Meaningful Information Extraction System for Interactive Analysis of DocumentsabstractThis paper is related to a project aiming at discovering weak signals from different streams of information, possibly sent by whistleblowers. The study presented in this paper tackles the particular problem of clustering topics at multi-levels from multiple documents, and then extracting meaningful descriptors, such as weighted lists of words for document representations in a multi-dimensions space. In this context, we present a novel idea which combines Latent Dirichlet Allocation and Word2vec (providing a consistency metric regarding the partitioned topics) as potential method for limiting the "a priori" number of cluster K usually needed in classical partitioning approaches. We proposed 2 implementations of this idea, respectively able to: (1) finding the best K for LDA in terms of topic consistency; (2) gathering the optimal clusters from different levels of clustering. We also proposed a non-traditional visualization approach based on a multi-agents system which combines both dimension reduction and interactivity. Julien Maitre, Michel Ménard, Guillaume Chiron, Alain Bouju, Nicolas Sidere |
ICDAR | 5 |
| 2019 | Security and PrIvacy foR the Internet of Things: an overview of the projectabstractAs the adoption of digital technologies expands, it becomes vital to build trust and confidence in the integrity of such technology. The SPIRIT project investigates the proof of concept of employing novel secure and privacy-ensuring techniques in services set-up in the Internet of Things (IoT) environment, aiming to increase the trust of users in IoTbased systems. The proposed system integrates three highly novel technology concepts developed by the consortium partners. Specifically, a technology, ermed ICMetrics, for deriving encryption keys directly from the operating characteristics of digital devices; secondly, a technology based on a contentbased signature of user data in order to ensure the integrity of sentdata upon arrival; a third technology, termed semantic firewall, which is able to allow or deny the transmission of data derived from an IoT device according to the information contained within the data and the information gathered about the requester. Sabrine Aroua, Julian Murphy, Mourad Rabah, Kais Rouis, Nicolas Sidere, Nouredine Tamani, Ronan Champagnat, Mickaël Coustaty, Gilles Falquet, Sami Ghadfi, Yacine Ghamri-Doudane, Petra Gomez-Krämer, Gareth Howells 0001, Klaus D. McDonald-Maier |
SMC | 5 |
| 2018 | Find it! Fraud Detection Contest ReportabstractThis paper describes the ICPR2018 fraud detection contest, its data set, its evaluation methodology, as well as the different methods submitted by the participants to tackle the predefined tasks. Forensics research is quite a sensitive topic. Data are either private or unlabeled and most of related works are evaluated on private datasets with a restricted access. This restriction has two major consequences: results cannot be reproduced and no benchmarking can be done between every approach. This contest was conceived in order to address these drawbacks. Two tasks were proposed: detecting documents containing at least one forgery in a flow of documents and spotting and localizing these forgeries within documents. An original dataset composed of images and texts of French receipts was provided to participants. The results they obtained are presented and discussed. Chloé Artaud, Nicolas Sidere, Antoine Doucet, Jean-Marc Ogier, Vincent Poulain D'Andecy |
ICPR | 2 |
| 2017 | Local Binary Patterns for Document Forgery DetectionabstractDocument forgery is an increasing problem for both the public administration and private companies. It represents substantial losses in time and economical resources. Classical solutions to this problem such as watermarks or other integrated security patterns can not be applied in general for any unknown incoming document due to the large variability on types of documents. In that scenario it is important to resort to forensic techniques to seek and analyze inconsistencies on the intrinsic features of the document image. In this paper we present a classification-based approach for forgery detection. We use uniform Local Binary Patterns (LBP) to capture discriminant texture features that are common on forged regions. Besides, we combine multiple descriptors from neighboring regions to model contextual information. Results using Support Vector Machines (SVM) for patch classification show that we are able to detect several types of forgeries in a wide range of types of documents. Francisco Cruz 0003, Nicolas Sidere, Mickaël Coustaty, Vincent Poulain D'Andecy, Jean-Marc Ogier |
ICDAR | 2 |
| 2016 | A Compliant Document Image Classification System Based on One-Class ClassifierabstractDocument image classification in a professional context requires to respect some constraints such as dealing with a large variability of documents and/or number of classes. Whereas most methods deal with all classes at the same time, we answer this problem by presenting a new compliant system based on the specialization of the features and the parametrization of the classifier separately, class per class. We first compute a generalized vector of features based on global image characterization and structural primitives. Then, for each class, the feature vector is specialized by ranking the features according a stability score. Finally, a one-class K-nn classifier is trained using these specific features. Conducted experiments reveal good classification rates, proving the ability of our system to deal with a large range of documents classes. Nicolas Sidere, Jean-Yves Ramel, Sabine Barrat, Vincent Poulain D'Andecy, Saddok Kebairi |
DAS | 1 |
| 2013 | Document Classification in a Non-stationary Environment: A One-Class SVM ApproachabstractIn this paper, we investigate a specific area of document classification in which the documents come as a flow over the time. Moreover, the exact number of classes of document to deal with is not known from the beginning and could evolve over the time. To be able to perform classification task in such area, we need specific classifiers that are able to perform incremental learning and change their modeling over the time. More specifically, we are focusing our study on SVM approaches, known to perform well, and for which incremental (i-SVM) procedures exist. Nevertheless, most of them are only able to deal with a fixed number of classes. So we designed a new incremental learning procedure based on one-class SVMs. This one is able to improve its classification accuracy over the time, with the arrival of new labeled data, without performing any complete retraining. Moreover, when instances are coming with a previously unknown label (appearance of a new class), the training procedure is able to modify the classifier model to recognize this corresponding new kind of documents. To investigate this area, waiting for collecting documents images as a flow, we did first experiments on the Optical Recognition of Handwritten Digits Data Set. These experiments show that our incremental approach is able: to perform, at each time, as well as a static one-class classifier fully retrained using all previously seen data, to model very quickly and efficiently new incoming classes. Anh Khoi Ngo Ho, Nicolas Ragot, Jean-Yves Ramel, Véronique Eglin, Nicolas Sidere |
ICDAR | 5 |
| 2009 | Vector Representation of Graphs: Application to the Classification of Symbols and LettersabstractIn this article we present a new approach for the classification of structured data using graphs. We suggest to solve the problem of complexity in measuring the distance between graphs by using a new graph signature. We present an extension of the vector representation based on pattern frequency, which integrates labeling information. In this paper, we compare the results achieved on public graph databases for the classification of symbols and letters using this graph signature with those obtained using the graph edit distance. Nicolas Sidere, Pierre Héroux, Jean-Yves Ramel |
ICDAR | 1 |