VLDB 2026 Research / reviewers in the wild / expert
Maruf A. Dhali
dblp:180/2767
· DBLP profile ↗
9ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0002-7548-3858ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Self-HTR: A Novel Self-supervised Handwritten Text Recognition Framework Using Generative Adversarial Networks
Lisa Koopmans, Maruf A. Dhali, Lambert Schomaker |
ICDAR (3) | 2 |
| 2025 | Oil Spill Segmentation Using Deep Encoder-Decoder ModelsabstractCrude oil is an integral component of the world economy and transportation sectors. With the growing demand for crude oil due to its widespread applications, accidental oil spills are unfortunate yet unavoidable. Even though oil spills are difficult to clean up, the first and foremost challenge is to detect them. In this research, the authors test the feasibility of deep encoder-decoder models that can be trained effectively to detect oil spills remotely. The work examines and compares the results from several segmentation models on high dimensional satellite Synthetic Aperture Radar (SAR) image data to pave the way for further in-depth research. Multiple combinations of models are used to run the experiments. The best-performing model is the one with the ResNet-50 encoder and DeepLabV3+ decoder. It achieves a mean Intersection over Union (IoU) of 64.868% and an improved class IoU of 61.549% for the “oil spill” class when compared with the previous benchmark model, which achieved a mean IoU of 65.05% and a class IoU of 53.38% for the “oil spill” class. Abhishek Ramanathapura Satyanarayana, Maruf A. Dhali |
ICPRAM | 2 |
| 2025 | Deep Learning for Effective Classification and Information Extraction of Financial DocumentsabstractThe financial and accounting sectors are encountering increased demands to effectively manage large volumes of documents in today’s digital environment. Meeting this demand is crucial for accurate archiving, maintaining efficiency and competitiveness, and ensuring operational excellence in the industry. This study proposes and analyzes machine learning-based pipelines to effectively classify and extract information from scanned and photographed financial documents, such as invoices, receipts, bank statements, etc. It also addresses the challenges associated with financial document processing using deep learning techniques. This research explores several models, including LeNet5, VGG19, and MobileNetV2 for document classification and RoBERTa, LayoutLMv3, and GraphDoc for information extraction. The models are trained and tested on financial documents from previously available benchmark datasets and a new dataset with financial documents in Romanian. Results show MobileNetV2 excels in classification tasks (with accuracies of 99.24% with data augmentation and 93.33% without augmentation), while RoBERTa and LayoutLMv3 lead in extraction tasks (with F1-scores of 0.7761 and 0.7426, respectively). Despite the challenges posed by the imbalanced dataset and cross-language documents, the proposed pipeline shows potential for automating the processing of financial documents in the relevant sectors. Valentin-Adrian Serbanescu, Maruf A. Dhali |
ICPRAM | 2 |
| 2025 | LostPaw: Finding Lost Pets Using a Contrastive Learning-Based Transformer with Visual InputabstractLosing pets can be highly distressing for pet owners, and finding a lost pet is often challenging and time-consuming. An artificial intelligence-based application can significantly improve the speed and accuracy of finding lost pets. To facilitate such an application, this study introduces a contrastive neural network model capable of accurately distinguishing between images of pets. The model was trained on a large dataset of dog images and evaluated through 3-fold cross-validation. Following 350 epochs of training, the model achieved a test accuracy of 90%. Furthermore, overfitting was avoided, as the test accuracy closely matched the training accuracy. Our findings suggest that contrastive neural network models hold promise as a tool for locating lost pets. This paper presents the foundational framework for a potential web application designed to assist users in locating their missing pets. The application will allow users to upload images of their lost pets and provide notificati ons when matching images are identified within its image database. This functionality aims to enhance the efficiency and accuracy with which pet owners can search for and reunite with their beloved animals. Andrei Voinea, Robin Kock, Maruf A. Dhali |
ICPRAM | 3 |
| 2023 | The Effects of Character-Level Data Augmentation on Style-Based Dating of Historical ManuscriptsabstractIdentifying the production dates of historical manuscripts is one of the main goals for paleographers when studying ancient documents. Automatized methods can provide paleographers with objective tools to estimate dates more accurately. Previously, statistical features have been used to date digitized historical manuscripts based on the hypothesis that handwriting styles change over periods. However, the sparse availability of such documents poses a challenge in obtaining robust systems. Hence, the research of this article explores the influence of data augmentation on the dating of historical manuscripts. Linear Support Vector Machines were trained with k-fold cross-validation on textural and grapheme-based features extracted from historical manuscripts of different collections, including the Medieval Paleographical Scale, early Aramaic manuscripts, and the Dead Sea Scrolls. Results show that training models with augmented data improve the performance of historical manuscripts datin g by 1% - 3% in cumulative scores. Additionally, this indicates further enhancement possibilities by considering models specific to the features and the documents’ scripts Lisa Koopmans, Maruf A. Dhali, Lambert Schomaker |
ICPRAM | 2 |
| 2023 | Image-Based Material Analysis of Ancient Historical DocumentsabstractResearchers continually perform corroborative tests to classify ancient historical documents based on the physical materials of their writing surfaces. However, these tests, often performed on-site, requires actual access to the manuscript objects. The procedures involve a considerable amount of time and cost, and can damage the manuscripts. Developing a technique to classify such documents using only digital images can be very useful and efficient. In order to tackle this problem, this study uses images from a famous historical collection, the Dead Sea Scrolls, to propose a novel method to classify the materials of the manuscripts. The proposed classifier uses the two-dimensional Fourier Transform to identify patterns within the manuscript surfaces. Combining a binary classification system employing the transform with a majority voting process is shown to be effective for this classification task. This pilot study shows a successful classification percentage of up to 97% for a confi ned amount of manuscripts produced from either parchment or papyrus material. Feature vectors based on Fourier-space grid representation outperformed a concentric Fourier-space format. Thomas Reynolds, Maruf A. Dhali, Lambert Schomaker |
ICPRAM | 2 |
| 2020 | Feature-extraction methods for historical manuscript dating based on writing style developmentabstractPaleographers and philologists perform significant research in finding the dates of ancient manuscripts to understand the historical contexts. To estimate these dates, the traditional process of using classical paleography is subjective, tedious, and often time-consuming. An automatic system based on pattern recognition techniques that infers these dates would be a valuable tool for scholars. In this study, the development of handwriting styles over time in the Dead Sea Scrolls, a collection of ancient manuscripts, is used to create a model that predicts the date of a query manuscript. In order to extract the handwriting styles, several dedicated feature-extraction techniques have been explored. Additionally, a self-organizing time map is used as a codebook. Support vector regression is used to estimate a date based on the feature vector of a manuscript. The date estimation from grapheme-based technique outperforms other feature-extraction techniques in identifying the chronological style development of handwriting in this study of the Dead Sea Scrolls. Maruf A. Dhali, Camilo Nathan Jansen, Jan Willem de Wit, Lambert Schomaker |
Pattern Recognit. Lett. | 1 |
| 2017 | A Digital Palaeographic Approach towards Writer Identification in the Dead Sea ScrollsabstractTo understand the historical context of an ancient manuscript, scholars rely on the prior knowledge of writer and date of that document. In this paper, we study the Dead Sea Scrolls, a collection of ancient manuscripts with immense historical, religious, and linguistic significance, which was discovered in the mid-20th century near the Dead Sea. Most of the manuscripts of this collection have become digitally available only recently and techniques from the pattern recognition field can be applied to revise existing hypotheses on the writers and dates of these scrolls. This paper presents our ongoing work which aims to introduce digital palaeography to the field and generate fresh empirical data by means of pattern recognition and artificial intelligence. Challenges in analyzing the Dead Sea Scrolls are highlighted by a pilot experiment identifying the writers using several dedicated features. Finally, we discuss whether to use specifically-designed shape features for writer identification or to use the Deep Learning methods on a relatively limited ancient manuscript collection which is degraded over the course of time and is not labeled, as in the case of the Dead Sea Scrolls. Maruf A. Dhali, Mladen Popovic, Eibert Tigchelaar, Lambert Schomaker |
ICPRAM | 1 |
| 2016 | Super-resolution spectral analysis for ultrasound scatter characterizationabstractParametric Bayesian spectral estimation methods have been previously utilized to improve frequency resolution. Ultrasound signals have been tested in such methods resulting in higher precision frequency detection compared to common non-parametric spectral estimation methods based on the Fourier transform. Such a technique using a reversible jump Markov Chain Monte Carlo algorithm has been developed to fully characterize signals and in addition to frequency, to provide amplitude and noise estimation. The analysis of this method is demonstrated with a real copper sphere ultrasound scatter signal. Based on typical diagnostic ultrasound data between 1.2–4.5 MHz the new spectral estimation achieves 110 kHz minimum frequency resolution. This is at least twice the resolution of Fourier based methods, resulting in revealing new frequencies. The method may be used in the entire range of ultrasound imaging modalities and may help provide improved sensitivity, reproducibility and spatial resolution. Konstantinos Diamantis, Maruf A. Dhali, Gavin Gibson, James R. Hopgood, Vassilis Sboros |
ICASSP | 2 |