VLDB 2026 Research / reviewers in the wild / expert
Denis Coquenet
dblp:252/8048
· DBLP profile ↗
8ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0001-5203-9423ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 5 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | n-Gram Injection into Transformers for Dynamic Language Model Adaptation in Handwritten Text RecognitionabstractTransformer-based encoder-decoder networks have recently achieved impressive results in handwritten text recognition, partly thanks to their auto-regressive decoder which implicitly learns a language model. However, such networks suffer from a large performance drop when evaluated on a target corpus whose language distribution is shifted from the source text seen during training. To retain recognition accuracy despite this language shift, we propose an external n-gram injection (NGI) for dynamic adaptation of the network's language modeling at inference time. Our method allows switching to an n-gram language model estimated on a corpus close to the target distribution, therefore mitigating bias without any extra training on target image-text pairs. We opt for an early injection of the n-gram into the transformer decoder so that the network learns to fully leverage text-only data at the low additional cost of n-gram inference. Experiments on three handwritten datasets demonstrate that the proposed NGI significantly reduces the performance gap between source and target corpora. Florent Meyer, Laurent Guichard, Yann Soullard, Denis Coquenet, Guillaume Gravier, Bertrand Coüasnon |
ICDAR (2) | 4 |
| 2026 | Meta-DAN: Towards an efficient prediction strategy for page-level handwritten text recognitionabstract• We propose a novel autoregressive decoding strategy for end-to-end text recognition • We achieve state-of-the-art results on 8 full-page handwriting recognition datasets • The proposed strategy enables speeding up decoding while improving context modeling • We propose a confidence-based approach offering a speed/performance trade-off Recent advances in text recognition led to a paradigm shift for page-level recognition, from multi-step segmentation-based approaches to end-to-end attention-based ones. However, the naïve character-level autoregressive decoding process results in long prediction times: it requires several seconds to process a single page image on a modern GPU. We propose the Meta Document Attention Network (Meta-DAN) as a novel decoding strategy to reduce the prediction time while enabling better context modeling. It relies on two main components: windowed queries, to process several transformer queries altogether, enlarging the context modeling with the near future; and multi-token predictions, whose goal is to predict several tokens per query instead of only the next one. We evaluate the proposed approach on 10 full-page handwritten datasets and demonstrate state-of-the-art results on average in terms of character error rate. Source code and weights of trained models are available at https://github.com/FactoDeepLearning/meta_dan . Denis Coquenet |
Pattern Recognit. | 1 |
| 2025 | Relaxed Syntax Modeling in Transformers for Future-Proof License Plate Recognition
Florent Meyer, Laurent Guichard, Denis Coquenet, Guillaume Gravier, Yann Soullard, Bertrand Coüasnon |
ICDAR (4) | 3 |
| 2023 | Faster DAN: Multi-target Queries with Document Positional Encoding for End-to-End Handwritten Document Recognition
Denis Coquenet, Clément Chatelain 0001, Thierry Paquet |
ICDAR (4) | 1 |
| 2023 | End-to-End Handwritten Paragraph Text Recognition Using a Vertical Attention NetworkabstractUnconstrained handwritten text recognition remains challenging for computer vision systems. Paragraph text recognition is traditionally achieved by two models: the first one for line segmentation and the second one for text line recognition. We propose a unified end-to-end model using hybrid attention to tackle this task. This model is designed to iteratively process a paragraph image line by line. It can be split into three modules. An encoder generates feature maps from the whole paragraph image. Then, an attention module recurrently generates a vertical weighted mask enabling to focus on the current text line features. This way, it performs a kind of implicit line segmentation. For each text line features, a decoder module recognizes the character sequence associated, leading to the recognition of a whole paragraph. We achieve state-of-the-art character error rate at paragraph level on three popular datasets: 1.91% for RIMES, 4.45% for IAM and 3.59% for READ 2016. Our code and trained model weights are available at https://github.com/FactoDeepLearning/VerticalAttentionOCR. Denis Coquenet, Clément Chatelain 0001, Thierry Paquet |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | DAN: A Segmentation-Free Document Attention Network for Handwritten Document RecognitionabstractUnconstrained handwritten text recognition is a challenging computer vision task. It is traditionally handled by a two-step approach, combining line segmentation followed by text line recognition. For the first time, we propose an end-to-end segmentation-free architecture for the task of handwritten document recognition: the Document Attention Network. In addition to text recognition, the model is trained to label text parts using begin and end tags in an XML-like fashion. This model is made up of an FCN encoder for feature extraction and a stack of transformer decoder layers for a recurrent token-by-token prediction process. It takes whole text documents as input and sequentially outputs characters, as well as logical layout tokens. Contrary to the existing segmentation-based approaches, the model is trained without using any segmentation label. We achieve competitive results on the READ 2016 dataset at page level, as well as double-page level with a CER of 3.43% and 3.70%, respectively. We also provide results for the RIMES 2009 dataset at page level, reaching 4.54% of CER. We provide all source code and pre-trained model weights at https://github.com/FactoDeepLearning/DAN. Denis Coquenet, Clément Chatelain 0001, Thierry Paquet |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | SPAN: A Simple Predict & Align Network for Handwritten Paragraph RecognitionabstractInternational audience Denis Coquenet, Clément Chatelain 0001, Thierry Paquet |
ICDAR (3) | 1 |
| 2020 | Recurrence-free unconstrained handwritten text recognition using gated fully convolutional networkabstractUnconstrained handwritten text recognition is a major step in most document analysis tasks. This is generally processed by deep recurrent neural networks and more specifically with the use of Long Short-Term Memory cells. The main drawbacks of these components are the large number of parameters involved and their sequential execution during training and prediction. One alternative solution to using LSTM cells is to compensate the long time memory loss with an heavy use of convolutional layers whose operations can be executed in parallel and which imply fewer parameters. In this paper we present a Gated Fully Convolutional Network architecture that is a recurrence-free alternative to the well-known CNN+LSTM architectures. Our model is trained with the CTC loss and shows competitive results on both the RIMES and IAM datasets. We release all code to enable reproduction of our experiments: https://github.com/FactoDeepLearning/LinePytorchOCR. Denis Coquenet, Clément Chatelain 0001, Thierry Paquet |
ICFHR | 1 |