Aleksandr Utkov

dblp:425/2911 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
0009-0002-2817-7337ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Image recognition and object detection · 41% Vision and language · 23% Information extraction and text analysis · 18%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 3 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection
object recognition
0.912025
MEH: A Multi-Style Dataset and Toolkit for Advancing Egyptian Hieroglyph Recognition · ICCV 2025
Computer vision › Image recognition and object detection › text recognition
optical character recognition
0.312025
MuMMy: Multimodal Dataset supporting VLM-based Egyptology Research Assistant · ACM Multimedia 2025
Multimedia analysis and retrieval › image analysis
cultural heritage image analysis
0.312025
MEH: A Multi-Style Dataset and Toolkit for Advancing Egyptian Hieroglyph Recognition · ICCV 2025

Methods — techniques the papers use, named apart from their topics

transliteration · 0.9translation · 0.9style-aware learning · 0.9recognition model · 0.9deep learning · 0.9OCR · 0.9
YearPublicationVenuePosition
2025 MEH: A Multi-Style Dataset and Toolkit for Advancing Egyptian Hieroglyph Recognition
Maksim Golyadkin, Valeria Rubanova, Aleksandr Utkov, Dmitry Nikolotov, Ilya Makarov
ICCV3
2025 MuMMy: Multimodal Dataset supporting VLM-based Egyptology Research Assistant
abstract
We present the first multimodal dataset MuMMy, for developing research assistants that can interpret Egyptian hieroglyphic texts. It pairs images with Gardiner codes, transliteration, and English translation at two levels of granularity. We also evaluate several deep learning pipelines across OCR, transliteration, and translation tasks, revealing the complexity of the domain and the challenges posed by error accumulation.
Maksim Golyadkin, Innokentiy Humonen, Valeria Rubanova, Danil Kalin, Yanis Plevokas, Dmitry Nikolotov, Aleksandr Utkov, Nikita Sidelnikov, Petr Ivanov, Ekaterina Bureeva, Ekaterina Alexandrova, Ilya Makarov
ACM Multimedia7
2025 Evaluation of Egyptian Hieroglyph Classification Across Diverse Writing Styles
abstract
The classification of Egyptian hieroglyphs remains a challenging problem due to the vast variability in writing styles across time periods, regions, and individual scribes. In this work, we present a comprehensive evaluation of hieroglyph classification performance across diverse stylistic domains, highlighting the limitations of current models in generalizing beyond a single style. We introduce a dataset that spans multiple writing styles, ranging from monumental inscriptions to handwritten manuscripts, and assess several near state-of-the-art recognition models. Our analysis reveals significant discrepancies in model performance when exposed to unseen styles, underscoring the need for style-aware learning strategies. This study provides a framework for future research on hieroglyph recognition with a focus on stylistic diversity and serves as a first step toward building vision-language systems capable of analyzing Egyptian hieroglyphic writings.
Maksim Golyadkin, Valeria Rubanova, Aleksandr Utkov, Dmitry Nikolotov, Ilya Makarov
ACM Multimedia3