VLDB 2026 Research / reviewers in the wild / expert
Utpal Garain
dblp:81/4230
· DBLP profile ↗
68ranked-venue papers
22as first author
11since 2021 · last 2026
0000-0001-7207-5018ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 57 · 19 first-author · 9 since 2021Databases, data management, data science and information retrieval · 24 · 14 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | KisMATH: Do LLMs Have Knowledge of Implicit Structures in Mathematical Reasoning?abstractAbstract Chain-of-thought (CoT) traces have been shown to improve performance of large language models on a plethora of reasoning tasks, yet there is no consensus on the mechanism by which this boost is achieved. To shed more light on this, we introduce Causal CoT Graphs (CCGraphs), which are directed acyclic graphs automatically extracted from reasoning traces that model finegrained causal dependencies in language-model outputs. A collection of 1671 mathematical reasoning problems from MATH500, GSM8K, and AIME, together with their associated CCGraphs, has been compiled into our dataset—KisMATH. Our detailed empirical analysis with 15 open-weight LLMs shows that (i) reasoning nodes in the CCGraphs are causal contributors to the final answer, which we argue is constitutive of reasoning; and (ii) LLMs emphasize the reasoning paths captured by the CCGraphs, indicating that the models internally realize structures similar to our graphs. KisMATH enables controlled, graph-aligned interventions and opens avenues for further investigation into the role of CoT in LLM reasoning. Soumadeep Saha, Akshay Chaturvedi, Saptarshi Saha, Utpal Garain, Nicholas Asher |
Trans. Assoc. Comput. Linguistics | 4 |
| 2026 | A Comprehensive Performance Evaluation of LLMs for Data-to-Text Generation and Divergence-Weighted TrainingabstractData-to-Text Generation (D2T) aims to transform semi-structured data (such as tables and graphs) into natural language text. With the exceptional capability of Large Language Models (LLMs), they have become ubiquitous as foundational models for D2T. This article presents a comprehensive evaluation of LLMs for D2T, focusing on three key qualities: readability (fluency and coherence), informativeness (content preservation), and faithfulness (factual accuracy). We evaluate 12 LLMs from five prominent open source families (BART, T5, BLOOM, OPT, and Llama 2) across five widely used D2T datasets using six established automatic metrics, complemented by human evaluation for deeper insight. Our findings reveal that larger model sizes generally improve readability and informativeness, with Llama 2 showing superior overall performance. However, increased model size does not consistently enhance faithfulness and may sometimes degrade it. Human evaluations indicate that larger models are generally preferred for their readability, informativeness, and faithfulness from the human readers’ perspective, as their minor faithfulness errors are assessed more selectively by automatic evaluation metrics. Through robustness analyses, we confirm that these trends remain stable across different fine-tuning (QLoRA vs. Prefix-Tuning) and decoding (Beam Search vs. Nucleus Sampling) strategies. Furthermore, our experiments show that performance consistently declines as source-reference divergence increases, regardless of model size. To mitigate this, we propose a source-reference divergence-weighted training that adaptively reweights training instances based on their source-reference divergence, achieving consistent improvements across all three key evaluation qualities. This comprehensive study provides practical insights into LLM behavior in D2T and introduces an effective training paradigm for improving performance in D2T. Joy Mahapatra, Utpal Garain |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2025 | Copula Based Trainable Calibration Error Estimator of Multi-Label Classification with Label InterdependenciesabstractA key challenge in calibrating Multi-Label Classification(MLC) problems is to consider the interdependencies among labels. To address this, in this research we propose an unbiased, differentiable, trainable calibration error estimator for MLC problems by using Copula. Unlike other methods for calibrating MLC tasks that focus on marginal calibration, this novel estimator takes label interdependencies into account and enables us to tackle the strictest notion of calibration that is canonical calibration. To design the estimator, we begin by leveraging the kernel trick to construct a continuous distribution from the discrete label space. Then we take a semiparametric approach to construct the estimator where the marginals are modeled non-parametrically and the Copula is modeled parametrically. Theoretically we show that our estimator is unbiased and converges to true $L^p$ calibration error. We also use our estimator as a regularizer at the time of training and observe that it reduces calibration error on test datasets significantly. Experiments on a well established dataset endorses our claims. Arkapal Panda, Utpal Garain |
AISTATS | 2 |
| 2025 | Label Dependency Aware Loss for Reliable Multi-Label Medical Image ClassificationabstractA key challenge in multi-label classification is to model the dependencies between the labels while ensuring proper calibration, as the assumption of label independence often results in inferior classification performance and poor calibration. However, most of the earlier works that modeled label dependencies have neglected the problem of ensuring calibrated results, which is crucial in safety-critical applications like medical image analysis. In this paper, we propose a novel training loss function, Label Dependency Aware Cross Entropy (LDACE), specifically designed for capturing pairwise label dependencies during the learning process. Additionally, we introduce an auxiliary loss, Canonical Calibration Loss (CCL), which when combined with LDACE ensures better calibration. We evaluate the effectiveness of the proposed loss function through experiments on three publicly available datasets – ChestMNIST, PTB-XL and RFMiD – by comparing it against the traditional multi-label losses using various deep learning models. The classification and calibration results demonstrate the superiority of the proposed loss function over the traditional loss functions in terms of Hamming loss, Area Under the Receiver Operating Characteristic Curve (AUC), Average Calibration Error (ACE) and Maximum Calibration Error (MCE). Furthermore, the experiments also show that the auxiliary loss not only improves the calibration performance but also retains the classification performance. The code is available at https://github.com/theimageprocessingguy/LDACE-CCL. Aditya Shankar Pal, Arkapal Panda, Utpal Garain |
ICASSP | 3 |
| 2025 | Language Models are Crossword SolversabstractSoumadeep Saha, Sutanoya Chakraborty, Saptarshi Saha, Utpal Garain. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Soumadeep Saha, Sutanoya Chakraborty, Saptarshi Saha, Utpal Garain |
NAACL (Long Papers) | 4 |
| 2025 | Cyclic Counterfactuals under Shift-Scale InterventionsabstractMost counterfactual inference frameworks traditionally assume acyclic structural causal models (SCMs), i.e. directed acyclic graphs (DAGs). However, many real-world systems (e.g. biological systems) contain feedback loops or cyclic dependencies that violate acyclicity. In this work, we study counterfactual inference in cyclic SCMs under shift–scale interventions, i.e., soft, policy-style changes that rescale and/or shift a variable’s mechanism. Saptarshi Saha, Dhruv Vansraj Rathore, Utpal Garain |
NeurIPS | 3 |
| 2025 | SEMI-CAVA: A Causal Variational Approach to Semi-Supervised LearningabstractDeep learning has advanced rapidly, but relies heavily on large-labeled datasets for effective training. This is particularly challenging in fields like medicine, where expert labeling is costly, labor-intensive, and prone to bias and error. Semi-supervised learning (SSL) addresses this challenge by reducing reliance on labeled data. SSL is closely tied to the concept of causation. However, recent works relating causality to SSL are limited by modeling only low-dimensional observations or designing a plug-in module to alleviate the class imbalance. In this paper, we take steps towards training causal generative models for semi-supervised learning, combining principles from causality and variational inference. We interpret the Mixup strategy as a stochastic intervention and introduce a consistency loss to promote coherent latent representations. Under reasonable assumptions, we provide theoretical guarantees that the learned latent representations align with true causal factors up to permissible ambiguities. The experimental results show the proposed approach achieves state-of-the-art performance on several medical datasets of different modalities. Additionally, we test our model on standard benchmarking datasets: CIFAR10, CIFAR100, and SVHN, where it achieves competitive performance. Saptarshi Saha, Pratyush Kumar Sahoo, Utpal Garain |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Analyzing Semantic Faithfulness of Language Models via Input Intervention on Question AnsweringabstractAbstract Transformer-based language models have been shown to be highly effective for several NLP tasks. In this article, we consider three transformer models, BERT, RoBERTa, and XLNet, in both small and large versions, and investigate how faithful their representations are with respect to the semantic content of texts. We formalize a notion of semantic faithfulness, in which the semantic content of a text should causally figure in a model’s inferences in question answering. We then test this notion by observing a model’s behavior on answering questions about a story after performing two novel semantic interventions—deletion intervention and negation intervention. While transformer models achieve high performance on standard question answering tasks, we show that they fail to be semantically faithful once we perform these interventions for a significant number of cases (∼ 50% for deletion intervention, and ∼ 20% drop in accuracy for negation intervention). We then propose an intervention-based training regime that can mitigate the undesirable effects for deletion intervention by a significant margin (from ∼ 50% to ∼ 6%). We analyze the inner-workings of the models to better understand the effectiveness of intervention-based training for deletion intervention. But we show that this training does not attenuate other aspects of semantic unfaithfulness such as the models’ inability to deal with negation intervention or to capture the predicate–argument structure of texts. We also test InstructGPT, via prompting, for its ability to handle the two interventions and to capture predicate–argument structure. While InstructGPT models do achieve very high performance on predicate–argument structure task, they fail to respond adequately to our deletion and negation interventions. Akshay Chaturvedi, Swarnadeep Bhar, Soumadeep Saha, Utpal Garain, Nicholas Asher |
Comput. Linguistics | 4 |
| 2021 | Exploring Structural Encoding for Data-to-Text GenerationabstractDue to efficient end-to-end training and fluency in generated texts, several encoderdecoder framework-based models are recently proposed for data-to-text generations.Appropriate encoding of input data is a crucial part of such encoder-decoder models.However, only a few research works have concentrated on proper encoding methods.This paper presents a novel encoder-decoder based datato-text generation model where the proposed encoder carefully encodes input data according to underlying structure of the data.The effectiveness of the proposed encoder is evaluated both extrinsically and intrinsically by shuffling input data without changing meaning of that data.For selecting appropriate content information in encoded data from encoder, the proposed model incorporates attention gates in the decoder.With extensive experiments on WikiBio and E2E dataset, we show that our model outperforms the state-of-the models and several standard baseline systems.Analysis of the model through component ablation tests and human evaluation endorse the proposed model as a well-grounded system. Joy Mahapatra, Utpal Garain |
INLG | 2 |
| 2021 | Pick-Object-Attack: Type-specific adversarial attack for object detection
Omid Mohamad Nezami, Akshay Chaturvedi, Mark Dras, Utpal Garain |
Comput. Vis. Image Underst. | 4 |
| 2021 | Mimic and Fool: A Task-Agnostic Adversarial AttackabstractAt present, adversarial attacks are designed in a task-specific fashion. However, for downstream computer vision tasks such as image captioning and image segmentation, the current deep-learning systems use an image classifier such as VGG16, ResNet50, and Inception-v3 as a feature extractor. Keeping this in mind, we propose Mimic and Fool (MaF), a task-agnostic adversarial attack. Given a feature extractor, the proposed attack finds an adversarial image, which can mimic the image feature of the original image. This ensures that the two images give the same (or similar) output regardless of the task. We randomly select 1000 MSCOCO validation images for experimentation. We perform experiments on two image captioning models, Show and Tell, Show Attend and Tell, and one visual question answering (VQA) model, namely, end-to-end neural module network (N2NMN). The proposed attack achieves a success rate of 74.0%, 81.0%, and 87.1% for Show and Tell, Show Attend and Tell, and N2NMN, respectively. We also propose a slight modification to our attack to generate natural-looking adversarial images. In addition, we also show the applicability of the proposed attack for invertible architecture. Since MaF only requires information about the feature extractor of the model, it can be considered as a gray-box attack. Akshay Chaturvedi, Utpal Garain |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Even big data is not enough: need for a novel reference modelling for forensic document authentication
Utpal Garain, Biswajit Halder |
Int. J. Document Anal. Recognit. | 1 |
| 2020 | NeuMorph: Neural Morphological Tagging for Low-Resource Languages - An Experimental Study for Indic LanguagesabstractThis article deals with morphological tagging for low-resource languages. For this purpose, five Indic languages are taken as reference. In addition, two severely resource-poor languages, Coptic and Kurmanji, are also considered. The task entails prediction of the morphological tag (case, degree, gender, etc.) of an in-context word. We hypothesize that to predict the tag of a word, considering its longer context such as the entire sentence is not always necessary. In this light, the usefulness of convolution operation is studied resulting in a convolutional neural network (CNN) based morphological tagger. Our proposed model (BLSTM-CNN) achieves insightful results in comparison to the present state-of-the-art. Following the recent trend, the task is carried out under three different settings: single language, across languages, and across keys. Whereas the previous models used only character-level features, we show that the addition of word vectors along with character-level embedding significantly improves the performance of all the models. Since obtaining high-quality word vectors for resource-poor languages remains a challenge, in that scenario, the proposed character-level BLSTM-CNN proves to be most effective. 1 Abhisek Chakrabarty, Akshay Chaturvedi, Utpal Garain |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2019 | ICDAR 2019 CROHME + TFD: Competition on Recognition of Handwritten Mathematical Expressions and Typeset Formula DetectionabstractWe summarize the tasks, protocol, and outcome for the 6th Competition on Recognition of Handwritten Mathematical Expressions (CROHME), which includes a new formula detection in document images task (+ TFD). For CROHME + TFD 2019, participants chose between two tasks for recognizing handwritten formulas from 1) online stroke data, or 2) images generated from the handwritten strokes. To compare LATEX strings and the labeled directed trees over strokes (label graphs) used in previous CROHMEs, we convert LATEX and stroke-based label graphs to label graphs defined over symbols (symbol-level label graphs, or symLG). More than thirty (33) participants registered for the competition, with nineteen (19) teams submitting results. The strongest formula recognition results were produced by the USTC-iFLYTEK research team, for both stroke-based (81%) and image-based (77%) input. For the new typeset formula detection task, the Samsung R&D Institute Ukraine (Team 2) obtained a very strong F-score (93%). System performance has improved since the last CROHME - still, the competition results suggest that recognition of handwritten formulae remains a difficult structural pattern recognition task. Mahshad Mahdavi, Richard Zanibbi, Harold Mouchère, Christian Viard-Gaudin, Utpal Garain |
ICDAR | 5 |
| 2018 | CapsDeMM: Capsule Network for Detection of Munro's Microabscess in Skin Biopsy Images
Anabik Pal, Akshay Chaturvedi, Utpal Garain, Aditi Chandra, Raghunath Chatterjee, Swapan Senapati |
MICCAI (2) | 3 |
| 2018 | Erratum to "JCLMM: A finite mixture model for clustering of circular-linear data and its application to psoriatic plaque segmentation" [Pattern Recognition 66 (2017) 160-173]
Anandarup Roy 0001, Anabik Pal, Utpal Garain |
Pattern Recognit. | 3 |
| 2017 | Context Sensitive Lemmatization Using Two Successive Bidirectional Gated Recurrent NetworksabstractWe introduce a composite deep neural network architecture for supervised and language independent context sensitive lemmatization.The proposed method considers the task as to identify the correct edit tree representing the transformation between a word-lemma pair.To find the lemma of a surface word, we exploit two successive bidirectional gated recurrent structures -the first one is used to extract the character level dependencies and the next one captures the contextual information of the given word.The key advantages of our model compared to the state-of-the-art lemmatizers such as Lemming and Morfette are -(i) it is independent of human decided features (ii) except the gold lemma, no other expensive morphological attribute is required for joint learning.We evaluate the lemmatizer on nine languages -Bengali, Catalan, Dutch, Hindi, Hungarian, Italian, Latin, Romanian and Spanish.It is found that except Bengali, the proposed method outperforms Lemming and Morfette on the other languages.To train the model on Bengali, we develop a gold lemma annotated dataset 1 (having 1, 702 sentences with a total of 20, 257 word tokens), which is an additional contribution of this work. Abhisek Chakrabarty, Onkar Pandit, Utpal Garain |
ACL (1) | 3 |
| 2017 | Identification of Reader Specific Difficult Words by Analyzing Eye Gaze and Document ContentabstractThis paper presents an approach for identifying reader specific difficult words while someone is reading a textual document. The work is motivated by the need of developing human-document interaction systems, in general and creating person-specific online educational content, in particular. Eye gaze information gives person specific behavior whereas textual content is analyzed to get general linguistic aspect of the document content. These two pieces of information are fused together through machine learning algorithms to identify the set of difficult words for a particular reader reading a particular document. An annotated dataset has been created where each word in a document is marked with its bounding box information and each reader identifies a set of difficult words while reading the document. The dataset consists of sixteen documents and each document is read by five subjects. The method is evaluated through recall-precision analysis. The impressive precision at high recall attests the feasibility of building a practical application based on this research. The experiment further brings out several interesting facts about human reading behaviour. Utpal Garain, Onkar Pandit, Olivier Augereau, Ayano Okoso, Koichi Kise |
ICDAR | 1 |
| 2017 | JCLMM: A finite mixture model for clustering of circular-linear data and its application to psoriatic plaque segmentation
Anandarup Roy 0001, Anabik Pal, Utpal Garain |
Pattern Recognit. | 3 |
| 2017 | Named Entity Recognition with Word Embeddings and Wikipedia Categories for a Low-Resource LanguageabstractIn this article, we propose a word embedding--based named entity recognition (NER) approach. NER is commonly approached as a sequence labeling task with the application of methods such as conditional random field (CRF). However, for low-resource languages without the presence of sufficiently large training data, methods such as CRF do not perform well. In our work, we make use of the proximity of the vector embeddings of words to approach the NER problem. The hypothesis is that word vectors belonging to the same name category, such as a person’s name, occur in close vicinity in the abstract vector space of the embedded words. Assuming that this clustering hypothesis is true, we apply a standard classification approach on the vectors of words to learn a decision boundary between the NER classes. Our NER experiments are conducted on a morphologically rich and low-resource language, namely Bengali. Our approach significantly outperforms standard baseline CRF approaches that use cluster labels of word embeddings and gazetteers constructed from Wikipedia. Further, we propose an unsupervised approach (that uses an automatically created named entity (NE) gazetteer from Wikipedia in the absence of training data). For a low-resource language, the word vectors obtained from Wikipedia are not sufficient to train a classifier. As a result, we propose to make use of the distance measure between the vector embeddings of words to expand the set of Wikipedia training examples with additional NEs extracted from a monolingual corpus that yield significant improvement in the unsupervised NER performance. In fact, our expansion method performs better than the traditional CRF-based (supervised) approach (i.e., F-score of 65.4% vs. 64.2%). Finally, we compare our proposed approach to the official submission for the IJCNLP-2008 Bengali NER shared task and achieve an overall improvement of F-score 11.26% with respect to the best official system. Arjun Das, Debasis Ganguly, Utpal Garain |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2016 | ICFHR2016 CROHME: Competition on Recognition of Online Handwritten Mathematical ExpressionsabstractThis paper presents an overview of the 5th Competition on Recognition of Online Handwritten Mathematical Expressions (CROHME). As in previous years, the main task is formula recognition from handwritten strokes (Task 1). Additional tasks include classification of isolated symbols (Task 2a), classification of isolated valid and invalid symbols (Task 2b), a new task on parsing formula structure from valid handwritten symbols (Task 3), and parsing expressions with matrices (Task 4, experimental). In total, eleven (11) research labs registered for the competition, with six (6) teams submitting results. Innovations for this CROHME included providing a corpus of formulae from Wikipedia to train language models, and an online system for result submission. The highest recognition rates were obtained by MyScript corporation (Task 1. 67.65%, 2a. 92.81%, 2b. 86.77%, 3. 84.38%, and 4. 68.40%). Using only provided training data, the highest recognition rates were obtained by WIRIS corporation (Task 1. 49.61%, Task 3. 78.80%, Task 4. 56.40%), the Tokyo University of Agriculture and Technology (Task 2a. 92.28%), and RIT (Task 2b. 83.34%). The competition results suggest that recognition of handwritten formulae remains a difficult structural pattern recognition task. Harold Mouchère, Christian Viard-Gaudin, Richard Zanibbi, Utpal Garain |
ICFHR | 4 |
| 2016 | Severity grading of psoriatic plaques using deep CNN based multi-task learningabstractThis paper addresses the problem of automatic machine analysis based severity scoring of psoriasis skin disease. Three different disease parameters namely, erythema, scaling and induration are considered for such severity grading. Given an image containing a psoriatic plaque the task is to predict severity scores for all the three parameters. This paper presents a novel deep CNN based architecture for achieving the task. Apart from viewing this task as three different single task learning (STL) problems (i.e. three different classification problems), a new multi-task learning (MTL) is also presented where the three classification tasks are treated as interdependent and thereby the neural net is trained accordingly. A new annotated dataset consisting of seven hundred and seven (707) images has been constructed on which the performance of the severity scoring algorithms have been reported. Several competing baselines are considered to compare the performance of STL and MTL approaches. Experimental result shows that the deep CNN based architectures (both the STL and MTL) achieve promising performances, MTL producing slightly superior results to that of STL. Anabik Pal, Akshay Chaturvedi, Utpal Garain, Aditi Chandra, Raghunath Chatterjee |
ICPR | 3 |
| 2016 | A Neural Lemmatizer for Bengali
Abhisek Chakrabarty, Akshay Chaturvedi, Utpal Garain |
LREC | 3 |
| 2016 | Advancing the state of the art for handwritten math recognition: the CROHME competitions, 2011-2014
Harold Mouchère, Richard Zanibbi, Utpal Garain, Christian Viard-Gaudin |
Int. J. Document Anal. Recognit. | 3 |
| 2016 | BenLem (A Bengali Lemmatizer) and Its Role in WSDabstractA lemmatization algorithm for Bengali has been developed and evaluated. Its effectiveness for word sense disambiguation (WSD) is also investigated. One of the key challenges for computer processing of highly inflected languages is to deal with the frequent morphological variations of the root words appearing in the text. Therefore, a lemmatizer is essential for developing natural language processing (NLP) tools for such languages. In this experiment, Bengali, which is the national language of Bangladesh and the second most popular language in the Indian subcontinent, has been taken as a reference. In order to design the Bengali lemmatizer (named as BenLem), possible transformations through which surface words are formed from lemmas are studied so that appropriate reverse transformations can be applied on a surface word to get the corresponding lemma back. BenLem is found to be capable of handling both inflectional and derivational morphology in Bengali. It is evaluated on a set of 18 news articles taken from the FIRE Bengali News Corpus consisting of 3,342 surface words (excluding proper nouns) and found to be 81.95% accurate. The role of the lemmatizer is then investigated for Bengali WSD. Ten highly polysemous Bengali words are considered for sense disambiguation. The FIRE corpus and a collection of Tagore’s short stories are considered for creating the WSD dataset. Different WSD systems are considered for this experiment, and it is noticed that BenLem improves the performance of all the WSD systems and the improvements are statistically significant. Abhisek Chakrabarty, Utpal Garain |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2015 | A Computational Approach for Corpus Based Analysis of Reduplicated Words in Bengali
Apurbalal Senapati, Utpal Garain |
CICLing (1) | 2 |
| 2015 | Unconstrained Bengali handwriting recognition with recurrent modelsabstractThis paper presents a pioneering attempt for developing a recurrent neural net based connectionist system for unconstrained Bengali offline handwriting recognition. The major challenge in configuring such a classification system for a complex script like Bengali is to effectively define the character classes. A novel way of defining character classes is introduced making the recognition problem suitable for using a recurrent model. Indeed, it has to deal with more than nine hundred character classes for which the occurrence probability is very skewed in the language. An off-the-shelf BLSTM-CTC recognizer is used. An open-source dataset is developed for unconstrained Bengali offline handwriting recognition. The dataset contains 2,338 handwritten text lines consisting of about 21,000 word. Experiment shows that with the new definition of character classes the BLSTM-CTC provides an impressive performance for unconstrained Bengali offline handwriting recognition. The character level recognition accuracy is 75.40% without doing any post-processing on the BLSTM-CTC output. Among the 24.60% character level errors, the substitution, deletion and insertion errors are 18.91%, 4.69% and 0.98%, respectively. Utpal Garain, Luc Mioulet, Bidyut B. Chaudhuri, Clément Chatelain 0001, Thierry Paquet |
ICDAR | 1 |
| 2015 | Language identification from handwritten documentsabstractThis paper presents a novel approach for language identification in handwritten documents. The approach is based on script identification followed by character recognition. BLSTM-CTC based handwriting recognizers are used and the OCR output is fed to a statistical language identifier for detecting the language of the input handwritten document. Documents in two scripts (Latin and Bengali) and four languages (English, French, Bengali and Assamese) are considered for evaluation. Several alternative frameworks have been explored, effects of handwriting recognition and text length on language detection have been studied. It is observed that with some empirical restrictions it is very much possible to achieve more that 80% language detection accuracy and based on the current research practical systems can be designed. Luc Mioulet, Utpal Garain, Clément Chatelain 0001, Philippine Barlas, Thierry Paquet |
ICDAR | 2 |
| 2015 | Machine-assisted authentication of paper currency: an experiment on Indian banknotes
Ankush Roy, Biswajit Halder, Utpal Garain, David S. Doermann |
Int. J. Document Anal. Recognit. | 3 |
| 2014 | A Maximum Entropy Based Honorificity Identification for Bengali Pronominal Anaphora Resolution
Apurbalal Senapati, Utpal Garain |
CICLing (1) | 2 |
| 2014 | ICFHR 2014 Competition on Recognition of On-Line Handwritten Mathematical Expressions (CROHME 2014)abstractWe present the outcome of the latest edition of the CROHME competition, dedicated to on-line handwritten mathematical expression recognition. In addition to the standard full expression recognition task from previous competitions, CROHME 2014 features two new tasks. The first is dedicated to isolated symbol recognition including a reject option for invalid symbol hypotheses, and the second concerns recognizing expressions that contain matrices. System performance is improving relative to previous competitions. Data and evaluation tools used for the competition are publicly available. Harold Mouchère, Christian Viard-Gaudin, Richard Zanibbi, Utpal Garain |
ICFHR | 4 |
| 2013 | Automatic Selection of Binarization Method for Robust OCRabstractMany algorithms are now available for doing the same task (e.g. binarization, page segmentation, character recognition, etc.) in document image analysis (DIA) and choosing a particular algorithm(s) for a particular task is often a non-trivial problem. This paper proposes a model for automatically selecting the correct algorithm(s) for a given problem. Binarization has been taken a reference to illustrate the proposed approach. Several previously unexplored issues are addressed in this work. For example, only one method may not be good for the binarization of an entire document whereas a particular method may produce desired result for a particular region. Therefore, for a given document image, our model selects a set of one or more binarization techniques suitable for different regions of the document. This selection is completely automatic and guided by the machine learning approaches. Formulation of a completely automatic way for generating the annotated data for training the learning algorithms is also a novel contribution of this work. Evaluation of the approach is done using ICDAR 2003 Robust Reading data set and results highlight the potential of the proposed approach for automatic selection of correct DIA algorithm(s) from a set of several alternatives. Tanushyam Chattopadhyay, V. Ramu Reddy, Utpal Garain |
ICDAR | 3 |
| 2013 | ICDAR 2013 CROHME: Third International Competition on Recognition of Online Handwritten Mathematical ExpressionsabstractWe report on the third international Competition on Handwritten Mathematical Expression Recognition (CROHME), in which eight teams from academia and industry took part. For the third CROHME, the training dataset was expanded to over 8000 expressions, and new tools were developed for evaluating performance at the level of strokes as well as expressions and symbols. As an informal measure of progress, the performance of the participating systems on the CROHME 2012 data set is also reported. Data and tools used for the competition will be made publicly available. Harold Mouchère, Christian Viard-Gaudin, Richard Zanibbi, Utpal Garain |
ICDAR | 4 |
| 2013 | A Probabilistic Model for Reconstruction of Torn Forensic DocumentsabstractIn this paper we present a probabilistic approach to reconstruct a document from its torn pieces which is extremely helpful for the legal system. The method investigates several previously unexplored issues and proposes probabilistic dependencies of different (available) parts (or pieces) of a document. It iteratively calculates the probability of an arrangement subject to some constraints and attempts to produce the best possible configuration (or solution) using low level image statistics. The reconstruction method also ranks different (final) arrangements in order to help decision making process. Two types of data were used in the investigation. A set of documents where documents were torn naturally (as we do it in our daily life) and the second set consisting of documents that were torn to destroy particular evidences. Evaluation shows that the method is quite robust to tackle both the problems. Ankush Roy, Utpal Garain |
ICDAR | 2 |
| 2012 | ICFHR 2012 Competition on Recognition of On-Line Mathematical Expressions (CROHME 2012)abstractThis paper presents an overview of the second Competition on Recognition of Online Handwritten Mathematical Expressions, CROHME 2012. The objective of the contest is to identify current advances in mathematical expression recognition using common evaluation performance measures and datasets. This paper describes the contest details including the evaluation measures used as well as the performance of the 7 submitted systems along with a short description of each system. Progress as compared to the 1st version of CROHME is also documented. Harold Mouchère, Christian Viard-Gaudin, Jin Hyung Kim, Utpal Garain |
ICFHR | 5 |
| 2012 | A probabilistic framework for logo detection and localization in natural scene images
Ankush Roy, Utpal Garain |
ICPR | 2 |
| 2012 | Realization of an embeddable light weight Video OCRabstractGrowing popularity of the connected TV i.e. the TV with Internet, leads to research on providing different value added services like web-TV mash-up to the user. It is very easy to get the textual context from a digital TV broadcast. But in developing countries like India still nearly 90% TV viewers are using analog Radio Frequency (RF) cable as the input TV signal and thus that textual information are difficult to obtain. This paper presents a light weight Video Optical Character Recognition (VOCR) system that can recognize the English texts from Indian TV videos with average 76.53% accuracy which is significantly higher than the 68.18% accuracy of Tesseract on the same set of videos. Moreover the proposed method is significantly less resource intensive compared to Tesseract so that the proposed method can be realized on any embedded platform. Tanushyam Chattopadhyay, Utpal Garain |
ISDA | 2 |
| 2011 | A Weighted Finite-State Transducer (WFST)-Based Language Model for Online Indic Script Handwriting RecognitionabstractThough designing of classifies for Indic script handwriting recognition has been researched with enough attention, use of language model has so far received little exposure. This paper attempts to develop a weighted finite-state transducer (WFST) based language model for improving the current recognition accuracy. Both the recognition hypothesis (i.e. the segmentation lattice) and the lexicon are modeled as two WFSTs. Concatenation of these two FSTs accept a valid word(s) which is (are) present in the recognition lattice. A third FST called error FST is also introduced to retrieve certain words which were missing in the previous concatenation operation. The proposed model has been tested for online Bangla handwriting recognition though the underlying principle can equally be applied for recognition of offline or printed words. Experiment on a part of ISI-Bangla handwriting database shows that while the present classifiers (without using any language model) can recognize about 73% word, use of recognition and lexicon FSTs improve this result by about 9% giving an average word-level accuracy of 82%. Introduction of error FST further improves this accuracy to 93%. This remarkable improvement in word recognition accuracy by using FST-based language model would serve as a significant revelation for the research in handwriting recognition, in general and Indic script handwriting recognition, in particular. Suhan Chowdhury, Utpal Garain, Tanushyam Chattopadhyay |
ICDAR | 2 |
| 2011 | CROHME2011: Competition on Recognition of Online Handwritten Mathematical ExpressionsabstractA competition on recognition of online handwritten mathematical expressions is organized. Recognition of mathematical expressions has been an attractive problem for the pattern recognition community because of the presence of enormous uncertainties and ambiguities as encountered during parsing of the two-dimensional structure of expressions. The goal of this competition is to bring out a state of the art for the related research. Three labs come together to organize the event and six other research groups participated the competition. The competition defines a standard format for presenting information, provides a training set of 921 expressions and supplies the underlying grammar for understanding the content of the training data. Participants were invited to submit their recognizers which were tested with a new set of 348 expressions. Systems are evaluated based on four different aspects of the recognition problem. However, the final rating of the systems is done based on their correct expression recognition accuracies. The best expression level recognition accuracy (on the test data) shown by the competing systems is 19.83% whereas a baseline system developed by one of the organizing groups reports an accuracy 22.41% on the same data set. Harold Mouchère, Christian Viard-Gaudin, Jin Hyung Kim, Utpal Garain |
ICDAR | 5 |
| 2011 | EMERS: a tree matching-based performance evaluation of mathematical expression recognition systems
Kunal Sain, Abhishek Dasgupta, Utpal Garain |
Int. J. Document Anal. Recognit. | 3 |
| 2010 | Mash up of Breaking News and Contextual Web Information: A Novel Service for Connected TelevisionabstractThe Connected TV can be described as an Internet enabled TV. In the current paper we have proposed a system for connected TV that mash up the information from internet and RSS feeds related to the breaking news aired over the TV. The proposed system initially localize the text regions from the streamed video in hybrid mode, then recognize them and spot the key words and finally fetch the related information from Internet. Our experimental results show that the localization of the text regions from the video can work with no misses but have some false positives which are taken care by data semantic analysis. The errors of the Optical Character recognition (OCR) module are taken care by using string comparing techniques like longest common subsequence matching and Leveinsthein distance while matching with the RSS feed or internet. Tanushyam Chattopadhyay, Arpan Pal 0001, Utpal Garain |
ICCCN | 3 |
| 2010 | Role of Synthetically Generated Samples on Speech Recognition in a Resource-Scarce LanguageabstractSpeech recognition systems that make use of statistical classifiers require a large number of training samples. However, collection of real samples has always been a difficult problem due to the involvement of substantial amount of human intervention and cost. Considering this problem, this paper presents a novel method for generating synthetic samples from a handful of real samples and investigates the role of these samples in designing a speech recognition system. Speaker dependent limited vocabulary isolated word recognition in an Indian language (i.e. Bengali) has been taken a reference to demonstrate the potential of the proposed framework. The role of synthetic samples is demonstrated by showing a significant improvement in recognition accuracy. A maximum improvement of 10% is achieved using the proposed approach. Rupayan Chakraborty, Utpal Garain |
ICPR | 2 |
| 2010 | Color Feature Based Approach for Determining Ink Age in Printed DocumentsabstractAnswering to a query like when a particular document was printed is quite helpful in practice especially forensic purposes. This study attempts to develop a general framework that makes use of image processing and pattern recognition principles for ink age determination in printed documents. The approach, at first, computationally extracts a set of suitable color features and then analyzes them to properly associate them with ink age. Finally, a neural net is designed and trained to determine ages of unknown samples. The dataset used for the present experiment consists of the cover pages of LIFE magazines published in between 1930's and 70's (five decades). Test results show that a viable framework for involving machines in assisting human experts for determining age of printed documents. Biswajit Halder, Utpal Garain |
ICPR | 2 |
| 2009 | Identification of Mathematical Expressions in Document ImagesabstractIdentification of mathematical expressions in document images is attempted. Several features starting from low level image features, shape features, linguistics information, etc. are extracted and combined to pinpoint the expressions. A performance index has been formulated to evaluate the identification results. Experiment uses a data set of 200 publicly available images containing 1163 embedded and 1039 displayed expressions. Test results show accuracies of 88.3% and 97.2%, respectively for extracting embedded and displayed expressions perfectly. Utpal Garain |
ICDAR | 1 |
| 2009 | Machine Authentication of Security DocumentsabstractThis paper presents a pioneering effort towards machine authentication of security documents like bank cheques, legal deeds, certificates, etc. that fall under the same class as far as security is concerned. The proposed method first computationally extracts the security features from the document images and then the notion of dasiagenuinepsila vs. dasiaduplicatepsila is defined in the feature space. Bank cheques are taken as a reference for conducting the present experiment. Support Vector Machines (SVMs) and neural networks (NN) are involved to verify authenticity of these cheques. Results on a test dataset of 200 samples show that the proposed approach achieves about 98% accuracy for discriminating duplicate cheques from genuine ones. This strongly attests the viability of involving machine in authenticating security documents. Utpal Garain, Biswajit Halder |
ICDAR | 1 |
| 2009 | Off-Line Multi-Script Writer Identification Using AR CoefficientsabstractThe problem of writer identification in a multi-script environment is attempted using a two-dimensional (2D) autoregressive (AR) modeling technique. Each writer is represented by a set of 2D AR model coefficients. A method to estimate AR model coefficients is proposed. This method is applied to an image of text written by a specific writer so that AR coefficients are obtained to characterize the writer. For a given sample, AR coefficients are computed and its L2distance with each of the stored (writer) prototypes identifies the writer for the sample. The method has been tested on datasets of two different scripts, namely RIMES containing 382 French writers and ISI consisting of samples from 40 Bengali writers. Modeling of writing styles using different context patterns at different image resolution has been investigated. Experimental results show that the technique achieves results comparable with that of the previous approaches. Utpal Garain, Thierry Paquet |
ICDAR | 1 |
| 2008 | Automatic Diagram Drawing Based on Natural Language Text Understanding
Utpal Garain |
Diagrams | 2 |
| 2008 | Stop word detection in compressed textual images: An experiment on indic script documentsabstractStop word detection is attempted in this work in the context of retrieval of document images in the compressed domain. Algorithms are presented to identify text lines and words and to cluster similar words to count word occurrence frequencies. A list of words with their occurrence frequencies is generated from a corpus of textual images. As stop words in any language show high occurrence frequencies, such words occupy the upper positions in the sorted word list. Experiments have been carried out on two major indic scripts (Devanagari (Hindi) and Bangla). Test results using 150 document images consisting of about 12 K words in each script show the promising potential of the proposed approach. Utpal Garain, Amit Kumar Das 0001 |
ICPR | 1 |
| 2008 | Machine reading of camera-held low quality text images: An ICA-based image enhancement approach for improving OCR accuracyabstractAn Independent Component Analysis (ICA)-based image enhancement technique is presented to improve the accuracy for machine reading of camera-based images. Images of inscriptions that are normally engraved on stones or other durable materials and found at the sites of historical monuments are taken as a reference for conducting the present experiments. Significant improvement in recognition rate of a commercial OCR system shows the potential of the proposed ICA-based method. It improves word and character recognition accuracies of the OCR system by 68.6% (from 11.2% to 79.8%) and 57.3% (from 34.8% to 92.1%), respectively. Utpal Garain, Atishay Jain, Anjan Maity, Bhabatosh Chanda |
ICPR | 1 |
| 2008 | Summarization of compressed text images: an experience on Indic script documentsabstractAutomatic summarization of JBIG2 coded textual images is discussed. Compressed images are partially decompressed to compute relevant features. The feature extraction method is free from using any character recognition module. Summary sentences are ranked. Experiment considers documents in Indic scripts that lack in having any efficient OCR systems. Script independent aspect of the approach is highlighted through use of two most popular Indic scripts. Sentence selection efficiency of about 61% is achieved when judged against man-made summarization. A nonparametric (distribution-free) rank statistic shows a correlation coefficient of 0.33 as a measure of the (minimum) strength of the associations between sentence ranking by machine and human. Utpal Garain |
SIGIR | 1 |
| 2008 | Prototype reduction using an artificial immune model
Utpal Garain |
Pattern Anal. Appl. | 1 |
| 2007 | Machine Dating of Handwritten ManuscriptsabstractThis paper presents a pioneering study on automatic dating of handwritten manuscripts. Analysis of handwriting style forms the core of the dating method. Initially, it is hypothesized that a manuscript can be dated, to a certain level of accuracy, by looking at the way it is written. The hypothesis is then verified with real samples of known dates. A general framework is proposed for machine dating of handwritten manuscripts. Experiments on a database containing manuscripts of Gustave Flaubert (1821- 1880), the famous French novelist reports about 62% accuracy when manuscripts are dated within a range of five calendar years with respect to their exact year of writing. Utpal Garain, Swapan K. Parui, Thierry Paquet, Laurent Heutte |
ICDAR | 1 |
| 2006 | On foreground - background separation in low quality document images
Utpal Garain, Thierry Paquet, Laurent Heutte |
Int. J. Document Anal. Recognit. | 1 |
| 2005 | Segmentation of Touching Symbols for OCR of Printed Mathematical Expressions: An Approach based on Multifactorial AnalysisabstractThis paper deals with segmentation and recognition of touching characters appearing in scanned mathematical expressions. The technique is based on multifactorial analysis that integrates several factors determining cut-positions in a touching character image. A predictive algorithm is developed for efficient selection of possible cut-positions for segmenting touching characters. Experiment has been carried out using a test-set of reasonable size and results show that a considerable improvement in recognition accuracy can be achieved with a modest increase in computations. Utpal Garain, Bidyut B. Chaudhuri |
ICDAR | 1 |
| 2005 | An Approach for Stemming in Symbolically Compressed Indian Language Imaged DocumentsabstractStemming is used in many information retrieval (IR) systems to reduce variant word forms to common roots, and thereby improving the overall retrieval efficiency. This paper presents an algorithm for stemming in the context of document image retrieval system. The algorithm assumes that the documents are symbolically compressed and stemming has been attempted in the compressed domain itself. Experiments have been conducted on Indian language imaged documents for which efficient OCR still remains a challenging task. Results obtained from a set 150 document images (in Bangla script, the second most popular script in the Indian sub-continent) consisting of about 12K word show a promising performance of the proposed approach. Utpal Garain, Alok Kumar Datta |
ICDAR | 1 |
| 2005 | On Foreground-Background Separation in Low Quality Color Document ImagesabstractThis paper proposes an adaptive method for separation of foreground and background in low quality color document images. A connected component labelling is initially implemented to capture the spatially connected similar color pixels. Next, dominant background components are determined to divide the entire image into number of grids each representing local uniformity in illumination, background, etc. Finally foreground parts are located using local information around them. Several color images of old historical documents including manuscripts of high importance are used in the experiment. Apart from a qualitative evaluation, results are quantitatively compared with one popular foreground/background separation technique. Utpal Garain, Thierry Paquet, Laurent Heutte |
ICDAR | 1 |
| 2005 | A corpus for OCR research on mathematical expressions
Utpal Garain, Bidyut B. Chaudhuri |
Int. J. Document Anal. Recognit. | 1 |
| 2004 | Recognition of online handwritten mathematical expressionsabstractThis paper aims at automatic understanding of online handwritten mathematical expressions (MEs) written on an electronic tablet. The proposed technique involves two major stages: symbol recognition and structural analysis. Combination of two different classifiers have been used to achieve high accuracy for the recognition of symbols. Several online and offline features are used in the structural analysis phase to identify the spatial relationships among symbols. A context-free grammar has been designed to convert the input expressions into their corresponding T(E)X strings which are subsequently converted into MathML format. Contextual information has been used to correct several structure interpretation errors. A new method for evaluating performance of the proposed system has been formulated. Experiments on a dataset of considerable size strongly support the feasibility of the proposed system. Utpal Garain, Bidyut B. Chaudhuri |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2003 | Compression of scan-digitized Indian language printed text: a soft pattern matching techniqueabstractIn this paper, a new compression scheme is presented for Indian Language (IL) textual document images. Since OCR technology for IL scripts is not matured enough, transcription of these documents into digital domain needs new techniques that achieve high degree of compression as well as suitable methods to perform various operations like document indexing, retrieval, etc. The proposed method is essentially based on symbolic compression technique, which has been realized with an efficient segmentation-based clustering approach. A soft pattern-matching technique has been implemented using two different feature sets that co-operate each other to build an efficient prototype library. Experiments have been done for documents printed in Devnagari (Hindi) and Bangla scripts, two mostly used script in Indian sub-continent. Test results show that the proposed technique outperforms several standard methods like CCITT Group-4, JBIG, etc. which are frequently used for compression of document images. Utpal Garain, S. Debnath, A. Mandal, Bidyut B. Chaudhuri |
ACM Symposium on Document Engineering | 1 |
| 2003 | On Machine Understanding of Online Handwritten Mathematical ExpressionsabstractThis paper aims at automatic recognition of online handwritten mathematical expressions written on an electronic tablet. The proposed technique involves two major stages: symbol recognition and structural analysis. A multiple-classifier consists of both parametric and nonparametric classifier has been used for recognition of symbols. Parametric classifier is based on hidden Markov model (HMM), whereas, non-parametric classifier uses Nearest Neighbor classification scheme. Structural analysis uses several online and offline features to identify the spatial relationships among symbols. A context free grammar has been designed to convert input expressions into their corresponding Latex strings. Contextual information has been used to correct several errors occurring at both recognition and structural analysis stage. A new method for evaluating performance of the proposed systems has been formulated. Experiments on a dataset of considerable size show high efficiency of the proposed system. Utpal Garain, Bidyut B. Chaudhuri |
ICDAR | 1 |
| 2003 | Automatic Understanding of Structures in Printed Mathematical ExpressionsabstractRecognizing mathematical expressions from document image is a key problem in automatic conversion of scientific documents into electronic form. In this paper, we propose a simple grammar-based approach to recognize complex two-dimensional structures of printed mathematical expressions with high accuracy. The proposed technique is based on the structural information of symbols in an expression. An efficient implementation of the grammar is presented. The system generates a TEX string for the input expression. A new criterion for defining structural complexity of a mathematical expression has been formulated to measure the performance of the proposed technique. Experiment using a good representative sample of mathematical expressions shows a reasonably high efficiency of the system. Joydip Mitra, Utpal Garain, Bidyut B. Chaudhuri, Kumar Swamy H. V., Tamaltaru Pal |
ICDAR | 2 |
| 2001 | Segmentation of Touching Characters in Printed Devnagari and Bangla Scripts Using Fuzzy Multifactorial AnalysisabstractExistence of touching characters in scanned documents is a major problem in designing an effective character segmentation procedure for OCR systems. In this paper, new techniques are presented for identification and segmentation of touching characters. The techniques are based on fuzzy multifactorial analysis. A predictive algorithm is developed for effectively selecting cut-points to segment touching characters. Initially, our proposed method has been applied for segmenting touching characters that appear in Devnagari (Hindi) and Bangla, two major scripts in the Indian sub-continent. The results obtained from a test-set of considerable size show that a high recognition rate can be achieved with a reasonable amount of computations. Utpal Garain, Bidyut B. Chaudhuri |
ICDAR | 1 |
| 2001 | Extraction of type style-based meta-information from imaged documents
Bidyut B. Chaudhuri, Utpal Garain |
Int. J. Document Anal. Recognit. | 2 |
| 2000 | A Syntactic Approach for Processing Mathematical Expressions in Printed DocumentsabstractWe propose an approach for understanding mathematical expressions in printed documents. The overall approach is divided into three main steps: (i) detection of mathematical expressions in a document, (ii) recognition of the symbols present in the expression and (iii) arrangement of the recognized symbols. The detection of mathematical expressions is done through recognition of a few most common symbols and exploiting some structural features of the expressions. A hybrid of feature based and a template-based technique is used for the recognition of symbols. A two-pass approach is used for arrangement of the symbols. The first pass (scanning or lexical analysis) performs a micro-level examination of the symbols in order to identify the symbol groups occurring in them and to determine their categories or descriptors. The second pass (parsing or syntax analysis) processes the descriptors synthesized in the first pass, to determine the syntactic structure of the expression. A set of predefined rules guides the activities in both the passes. Experiments conducted using this approach on a large number of documents show high accuracy. Utpal Garain, Bidyut B. Chaudhuri |
ICPR | 1 |
| 2000 | An Approach for Recognition and Interpretation of Mathematical Expressions in Printed Document
Bidyut B. Chaudhuri, Utpal Garain |
Pattern Anal. Appl. | 2 |
| 1999 | Extraction of Type Style based Meta-Information from Imaged DocumentsabstractExtraction of some meta-information from printed documents without an OCR approach is considered. It can be statistically verified that important terms in articles are printed in italic, bold and all capital style. Detection of these type styles helps in automatic extraction of the lines containing titles, authors' names, subtitles, references as well as sentences having important terms occurring in the text. It also helps in improving the OCR performance for reading the italic text. Some experimental results on the performance of the approach on good quality as well as degraded document images are presented. Utpal Garain, Bidyut B. Chaudhuri |
ICDAR | 1 |
| 1998 | An Approach for Processing Mathematical Expressions in Printed Document
Bidyut B. Chaudhuri, Utpal Garain |
Document Analysis Systems | 2 |
| 1998 | Automatic detection of italic, bold and all-capital words in document imagesabstractWe propose simple and fast algorithms for detection of italic, bold and all-capital words without doing actual character recognition. We present a statistical study which reveals that the detection of such words may play a key role in automatic information retrieval from documents. Moreover, detection of italic words can be used to improve the recognition accuracy of a text recognition system. Considerable number of document images have been tested and our algorithms give accurate results on all the tested images, and the algorithms are very easy to implement. Bidyut B. Chaudhuri, Utpal Garain |
ICPR | 2 |