VLDB 2026 Research / reviewers in the wild / expert
Lei Kang 0002
dblp:85/7881-2
· DBLP profile ↗
17ranked-venue papers
12as first author
14since 2021 · last 2026
0000-0002-1962-3916ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 9 first-author · 11 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reading in the Dark: Low-Light Scene Text Recognition
Xuanshuo Fu, Lei Kang 0002, Ernest Valveny, Dimosthenis Karatzas, Javier Vazquez-Corral |
ICPR (11) | 2 |
| 2026 | Preserving privacy without compromising accuracy: Machine unlearning for handwritten text recognitionabstract• Encoder-only Transformer baseline for handwriting text recognition (HTR). • Added writer-style head tracks & controls user-identifiable memorization. • Neural-activation pruning removes writer clues while preserving HTR accuracy. • Writer-ID Confusion (WIC) enforces uniform IDs, boosting efficient machine unlearning. • Membership-inference tests confirm privacy gains without costly retraining. Handwritten Text Recognition (HTR) is crucial for document digitization, but handwritten data can contain user-identifiable features, like unique writing styles, posing privacy risks. Regulations such as the “right to be forgotten” require models to remove these sensitive traces without full retraining. We introduce a practical encoder-only transformer baseline as a robust reference for future HTR research. Building on this, we propose a two-stage unlearning framework for multihead transformer HTR models. Our method combines neural pruning with machine unlearning applied to a writer classification head, ensuring sensitive information is removed while preserving the recognition head. We also present Writer-ID Confusion (WIC), a method that forces the forget set to follow a uniform distribution over writer identities, unlearning user-specific cues while maintaining text recognition performance. We compare WIC to Random Labeling, Fisher Forgetting, Amnesiac Unlearning, and DELETE within our prune-unlearn pipeline and consistently achieve better privacy and accuracy trade-offs. This is the first systematic study of machine unlearning for HTR. Using metrics such as Accuracy, Character Error Rate (CER), Word Error Rate (WER), and Membership Inference Attacks (MIA) on the IAM and CVL datasets, we demonstrate that our method achieves state-of-the-art or superior performance for effective unlearning. These experiments show that our approach effectively safeguards privacy without compromising accuracy, opening new directions for document analysis research. Our code is publicly available at https://github.com/leitro/WIC-WriterIDConfusion-MachineUnlearning . Lei Kang 0002, Xuanshuo Fu, Lluís Gómez i Bigorda, Alicia Fornés, Ernest Valveny, Dimosthenis Karatzas |
Pattern Recognit. | 1 |
| 2025 | Position-Aware Stamp-Like Adversarial Attack for Document Classification
Lei Kang 0002, Maura Pintor, Dimosthenis Karatzas |
ICDAR (4) | 2 |
| 2025 | LLM-Driven Medical Document Analysis: Enhancing Trustworthy Pathology and Differential Diagnosis
Lei Kang 0002, Xuanshuo Fu, Oriol Ramos Terrades, Javier Vazquez-Corral, Ernest Valveny, Dimosthenis Karatzas |
ICDAR (3) | 1 |
| 2025 | AVIR: Adaptive Visual In-Document Retrieval for Efficient Multi-Page Document Question AnsweringabstractMulti‑page Document Visual Question Answering (MP‑DocVQA) remains challenging because long documents not only strain computational resources but also reduce the effectiveness of the attention mechanism in large vision–language models (LVLMs). We tackle these issues with an Adaptive Visual In‑document Retrieval (AVIR) framework. A lightweight retrieval model first scores each page for question relevance. Pages are then clustered according to the score distribution to adaptively select relevant content. The clustered pages are screened again by Top-K to keep the context compact. However, for short documents, clustering reliability decreases, so we use a relevance probability threshold to select pages. The selected pages alone are fed to a frozen LVLM for answer generation, eliminating the need for model fine‑tuning. The proposed AVIR framework reduces the average page count required for question answering by 70%, while achieving an ANLS of 84.58% on the MP-DocVQA dataset—surpassing previous methods with significantly lower computational cost. The effectiveness of the proposed AVIR is also verified on the SlideVQA and DUDE benchmarks. Our code will be made publicly available upon acceptance. Zongmin Li, Yachuan Li, Lei Kang 0002, Dimosthenis Karatzas, Wenkang Ma |
MMAsia | 3 |
| 2024 | Multi-page Document VQA with Recurrent Memory Transformer
Lei Kang 0002, Dimosthenis Karatzas |
DAS | 2 |
| 2024 | GRIF-DM: Generation of Rich Impression Fonts Using Diffusion ModelsabstractFonts are integral to creative endeavors, design processes, and artistic productions. The appropriate selection of a font can significantly enhance artwork and endow advertisements with a higher level of expressivity. Despite the availability of numerous diverse font designs online, traditional retrieval-based methods for font selection are increasingly being supplanted by generation-based approaches. These newer methods offer enhanced flexibility, catering to specific user preferences and capturing unique stylistic impressions. However, current impression font techniques based on Generative Adversarial Networks (GANs) necessitate the utilization of multiple auxiliary losses to provide guidance during generation. Furthermore, these methods commonly employ weighted summation for the fusion of impression-related keywords. This leads to generic vectors with the addition of more impression keywords, ultimately lacking in detail generation capacity. In this paper, we introduce a diffusion-based method, termed GRIF-DM, to generate fonts that vividly embody specific impressions, utilizing an input consisting of a single letter and a set of descriptive impression keywords. The core innovation of GRIF-DM lies in the development of dual cross-attention modules, which process the characteristics of the letters and impression keywords independently but synergistically, ensuring effective integration of both types of information. Our experimental results, conducted on the MyFonts dataset, affirm that this method is capable of producing realistic, vibrant, and high-fidelity fonts that are closely aligned with user specifications. This confirms the potential of our approach to revolutionize font generation by accommodating a broad spectrum of user-driven design requirements. Our code is publicly available at https://github.com/leitro/GRIF-DM. Lei Kang 0002, Fei Yang 0004, Kai Wang 0060, Mohamed Ali Souibgui, Lluís Gómez i Bigorda, Alicia Fornés, Ernest Valveny, Dimosthenis Karatzas |
ECAI | 1 |
| 2024 | Machine Unlearning for Document Classification
Lei Kang 0002, Mohamed Ali Souibgui, Fei Yang 0004, Lluís Gómez i Bigorda, Ernest Valveny, Dimosthenis Karatzas |
ICDAR (4) | 1 |
| 2024 | Multi-page Document Visual Question Answering Using Self-attention Scoring Mechanism
Lei Kang 0002, Rubèn Tito, Ernest Valveny, Dimosthenis Karatzas |
ICDAR (6) | 1 |
| 2024 | Privacy-Aware Document Visual Question Answering
Rubèn Tito, Marlon Tobaben, Raouf Kerkouche, Mohamed Ali Souibgui, Kangsoo Jung, Joonas Jälkö, Vincent Poulain D'Andecy, Aurélie Joseph, Lei Kang 0002, Ernest Valveny, Antti Honkela, Mario Fritz, Dimosthenis Karatzas |
ICDAR (6) | 10 |
| 2023 | Learning Robust Self-Attention Features for Speech Emotion Recognition with Label-Adaptive MixupabstractSpeech Emotion Recognition (SER) is to recognize human emotions in a natural verbal interaction scenario with machines, which is considered as a challenging problem due to the ambiguous human emotions. Despite the recent progress in SER, state-of-the-art models struggle to achieve a satisfactory performance. We propose a self-attention based method with combined use of label-adaptive mixup and center loss. By adapting label probabilities in mixup and fitting center loss to the mixup training scheme, our proposed method achieves a superior performance to the state-of-the-art methods. Lei Kang 0002, Lichao Zhang 0001, Dazhi Jiang |
ICASSP | 1 |
| 2022 | Content and Style Aware Generation of Text-Line Images for Handwriting RecognitionabstractHandwritten Text Recognition has achieved an impressive performance in public benchmarks. However, due to the high inter- and intra-class variability between handwriting styles, such recognizers need to be trained using huge volumes of manually labeled training data. To alleviate this labor-consuming problem, synthetic data produced with TrueType fonts has been often used in the training loop to gain volume and augment the handwriting style variability. However, there is a significant style bias between synthetic and real data which hinders the improvement of recognition performance. To deal with such limitations, we propose a generative method for handwritten text-line images, which is conditioned on both visual appearance and textual content. Our method is able to produce long text-line samples with diverse handwriting styles. Once properly trained, our method can also be adapted to new target data by only accessing unlabeled text-line images to mimic handwritten styles and produce images with any textual content. Extensive experiments have been done on making use of the generated samples to boost Handwritten Text Recognition performance. Both qualitative and quantitative results demonstrate that the proposed approach outperforms the current state of the art. Lei Kang 0002, Pau Riba, Marçal Rusiñol, Alicia Fornés, Mauricio Villegas |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Pay attention to what you read: Non-recurrent handwritten text-Line recognition
Lei Kang 0002, Pau Riba, Marçal Rusiñol, Alicia Fornés, Mauricio Villegas |
Pattern Recognit. | 1 |
| 2021 | Candidate fusion: Integrating language modelling into a sequence-to-sequence handwritten word recognition architecture
Lei Kang 0002, Pau Riba, Mauricio Villegas, Alicia Fornés, Marçal Rusiñol |
Pattern Recognit. | 1 |
| 2020 | GANwriting: Content-Conditioned Generation of Styled Handwritten Word Images
Lei Kang 0002, Pau Riba, Yaxing Wang, Marçal Rusiñol, Alicia Fornés, Mauricio Villegas |
ECCV (23) | 1 |
| 2020 | Distilling Content from Style for Handwritten Word RecognitionabstractDespite the latest transcription accuracies reached using deep neural network architectures, handwritten text recognition still remains a challenging problem, mainly because of the large inter-writer style variability. Both augmenting the training set with artificial samples using synthetic fonts, and writer adaptation techniques have been proposed to yield more generic approaches aimed at dodging style unevenness. In this work, we take a step closer to learn style independent features from handwritten word images. We propose a novel method that is able to disentangle the content and style aspects of input images by jointly optimizing a generative process and a handwritten word recognizer. The generator is aimed at transferring writing style features from one sample to another in an image-to-image translation approach, thus leading to a learned content-centric features that shall be independent to writing style attributes. Our proposed recognition model is able then to leverage such writer-agnostic features to reach better recognition performances. We advance over prior training strategies and demonstrate with qualitative and quantitative evaluations the performance of both the generative process and the recognition efficiency in the IAM dataset. Lei Kang 0002, Pau Riba, Marçal Rusiñol, Alicia Fornés, Mauricio Villegas |
ICFHR | 1 |
| 2020 | Unsupervised Adaptation for Synthetic-to-Real Handwritten Word RecognitionabstractHandwritten Text Recognition (HTR) is still a challenging problem because it must deal with two important difficulties: the variability among writing styles, and the scarcity of labelled data. To alleviate such problems, synthetic data generation and data augmentation are typically used to train HTR systems. However, training with such data produces encouraging but still inaccurate transcriptions in real words. In this paper, we propose an unsupervised writer adaptation approach that is able to automatically adjust a generic handwritten word recognizer, fully trained with synthetic fonts, towards a new incoming writer. We have experimentally validated our proposal using five different datasets, covering several challenges (i) the document source: modern and historic samples, which may involve paper degradation problems; (ii) different handwriting styles: single and multiple writer collections; and (iii) language, which involves different character combinations. Across these challenging collections, we show that our system is able to maintain its performance, thus, it provides a practical and generic approach to deal with new document collections without requiring any expensive and tedious manual annotation step. Lei Kang 0002, Marçal Rusiñol, Alicia Fornés, Pau Riba, Mauricio Villegas |
WACV | 1 |