Lucas Morin

dblp:309/7785 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0002-5829-5118ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Information extraction and text analysis · 46% Vision and language · 24% Graph learning · 16%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › document understanding › document image analysis
chemical structure recognition
1.522025
MarkushGrapher: Joint Visual and Textual Recognition of Markush Structures · CVPR 2025
MolGrapher: Graph-based Visual Recognition of Chemical Structures · ICCV 2023
Natural language and speech › Information extraction and text analysis
document understanding
0.912025
SmolDocling: An Ultra-Compact Vision-Language Model for End-To-End Multi-Modal Document Conversion · ICCV 2025
Computer vision › Vision and language › multimodal understanding
multimodal document understanding
0.912025
MarkushGrapher: Joint Visual and Textual Recognition of Markush Structures · CVPR 2025
Computer vision › Vision and language
vision-language model
0.912025
SmolDocling: An Ultra-Compact Vision-Language Model for End-To-End Multi-Modal Document Conversion · ICCV 2025
Natural language and speech › Information extraction and text analysis › document analysis
document information extraction
0.812024
ESG Accountability Made Easy: DocQA at Your Service · AAAI 2024
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
document question answering
0.812024
ESG Accountability Made Easy: DocQA at Your Service · AAAI 2024
Machine learning › Graph learning
graph neural network
0.712023
MolGrapher: Graph-based Visual Recognition of Chemical Structures · ICCV 2023
Machine learning › Graph learning › graph neural network › node classification
graph node classification
0.712023
MolGrapher: Graph-based Visual Recognition of Chemical Structures · ICCV 2023
Natural language and speech › Information extraction and text analysis › document understanding › document image analysis
optical chemical structure recognition
0.712023
MolGrapher: Graph-based Visual Recognition of Chemical Structures · ICCV 2023
Computer vision › Vision and language › vision-language model
efficient vision-language model
0.312025
SmolDocling: An Ultra-Compact Vision-Language Model for End-To-End Multi-Modal Document Conversion · ICCV 2025
Machine learning › Efficient and distributed learning
model compression
0.312025
SmolDocling: An Ultra-Compact Vision-Language Model for End-To-End Multi-Modal Document Conversion · ICCV 2025
Natural language and speech › Language models and text generation › large language model
large language model applications
0.212024
ESG Accountability Made Easy: DocQA at Your Service · AAAI 2024
Bioinformatics and computational biology › molecular informatics
cheminformatics
0.212023
MolGrapher: Graph-based Visual Recognition of Chemical Structures · ICCV 2023

Methods — techniques the papers use, named apart from their topics

graph neural network · 1.3deep keypoint detection · 1.3vision-text-layout encoder · 0.9vision-language model · 0.9optical chemical structure recognition · 0.9end-to-end document conversion · 0.9autoregressive generation · 0.9natural language processing · 0.8large language model · 0.8computer vision · 0.8synthetic data generation · 0.7
YearPublicationVenuePosition
2026 Identify, Locate, Link: End-to-End Key-Value Extraction from Document Images
A. Said Gurbuz, Ahmed S. Nassar, Christoph Auer, Maksym Lysak, Lucas Morin, Matteo Omenetti, Tim Strohmeyer, Panagiotis Vagenas, Nikolaos Livathinos, Michele Dolfi, Peter W. J. Staar
ICDAR (2)5
2025 MarkushGrapher: Joint Visual and Textual Recognition of Markush Structures
abstract
The automated analysis of chemical literature holds promise to accelerate discovery in fields such as material science and drug development. In particular, search capabilities for chemical structures and Markush structures (chemical structure templates) within patent documents are valuable, e.g., for prior-art search. Advancements have been made in the automatic extraction of chemical structures from text and images, yet the Markush structures remain largely unexplored due to their complex multi-modal nature. In this work, we present MarkushGrapher, a multimodal approach for recognizing Markush structures in documents. Our method jointly encodes text, image, and layout information through a Vision-Text-Layout encoder and an Optical Chemical Structure Recognition vision encoder. These representations are merged and used to autoregressively generate a sequential graph representation of the Markush structure along with a table defining its variable groups. To overcome the lack of real-world training data, we propose a synthetic data generation pipeline that produces a wide range of realistic Markush structures. Additionally, we present M2S, the first annotated benchmark of real-world Markush structures, to advance research on this challenging task. Extensive experiments demonstrate that our approach outperforms state-of-the-art chemistry-specific and general-purpose vision-language models in most evaluation settings. Code, models, and datasets are available1.
Lucas Morin, Valéry Weber, Gerhard Ingmar Meijer, Luc Van Gool, Yawei Li 0001, Peter W. J. Staar
CVPR1
2025 SmolDocling: An Ultra-Compact Vision-Language Model for End-To-End Multi-Modal Document Conversion
abstract
We introduce SmolDocling, an ultra-compact vision-language model targeting end-to-end document conversion. Our model comprehensively processes entire pages by generating DocTags, a new universal markup format that captures all page elements in their full context with location. Unlike existing approaches that rely on large foundational models, or ensemble solutions that rely on handcrafted pipelines of multiple specialized models, SmolDocling offers an end-to-end conversion for accurately capturing content, structure and spatial location of document elements in a 256M parameters vision-language model. SmolDocling exhibits robust performance in correctly reproducing document features such as code listings, tables, equations, charts, lists, and more across a diverse range of document types including business documents, academic papers, technical reports, patents, and forms -- significantly extending beyond the commonly observed focus on scientific papers. Additionally, we contribute novel publicly sourced datasets for charts, tables, equations, and code recognition. Experimental results demonstrate that SmolDocling competes with other Vision Language Models that are up to 27 times larger in size, while reducing computational requirements substantially. The model is currently available, datasets will be publicly available soon.
Ahmed S. Nassar, Matteo Omenetti, Maksym Lysak, Nikolaos Livathinos, Christoph Auer, Lucas Morin, Rafael Teixeira de Lima, Yusik Kim, A. Said Gurbuz, Michele Dolfi, Peter W. J. Staar
ICCV6
2024 ESG Accountability Made Easy: DocQA at Your Service
abstract
We present Deep Search DocQA. This application enables information extraction from documents via a question-answering conversational assistant. The system integrates several technologies from different AI disciplines consisting of document conversion to machine-readable format (via computer vision), finding relevant data (via natural language processing), and formulating an eloquent response (via large language models). Users can explore over 10,000 Environmental, Social, and Governance (ESG) disclosure reports from over 2000 corporations. The Deep Search platform can be accessed at: https://ds4sd.github.io.
Lokesh Mishra, Cesar Berrospi, Kasper Dinkla, Diego Antognini, Francesco Fusco, Benedikt Bothur, Maksym Lysak, Nikolaos Livathinos, Ahmed S. Nassar, Panagiotis Vagenas, Lucas Morin, Christoph Auer, Michele Dolfi, Peter W. J. Staar
AAAI11
2023 MolGrapher: Graph-based Visual Recognition of Chemical Structures
abstract
The automatic analysis of chemical literature has immense potential to accelerate the discovery of new materials and drugs. Much of the critical information in patent documents and scientific articles is contained in figures, depicting the molecule structures. However, automatically parsing the exact chemical structure is a formidable challenge, due to the amount of detailed information, the diversity of drawing styles, and the need for training data. In this work, we introduce MolGrapher to recognize chemical structures visually. First, a deep keypoint detector detects the atoms. Second, we treat all candidate atoms and bonds as nodes and put them in a graph. This construct allows a natural graph representation of the molecule. Last, we classify atom and bond nodes in the graph with a Graph Neural Network. To address the lack of real training data, we propose a synthetic data generation pipeline producing diverse and realistic results. In addition, we introduce a large-scale benchmark of annotated real molecule images, USPTO-30K, to spur research on this critical topic. Extensive experiments on five datasets show that our approach significantly outperforms classical and learning-based methods in most settings. Code, models, and datasets are available1.
Lucas Morin, Martin Danelljan, Maria Isabel Agea, Ahmed S. Nassar, Valéry Weber, Gerhard Ingmar Meijer, Peter W. J. Staar, Fisher Yu 0001
ICCV1