Anjan Dutta 0001

dblp:91/8278-1 · DBLP profile ↗
← Back
44ranked-venue papers
12as first author
15since 2021 · last 2026
0000-0002-1667-2245ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 40 · 12 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author
YearPublicationVenuePosition
2026 CraftSVG: Multi-Object Text-to-SVG Synthesis via Layout Guided Diffusion
abstract
Generating SVGs from text is a challenging vision task, requiring diverse yet realistic depictions of the seen as well as unseen entities. Existing research has been mostly limited to generating single-object rather than comprehensive scenes comprising multiple elements. In response, CraftSVG introduces an end-to-end framework for creating SVGs depicting entire scenes from a textual description. Utilizing a pre-trained LLM for layout generation from text via iterative in-context learning, CraftSVG introduces a per-box mask latent mechanism for accurate object placement. A fusion mechanism is developed to integrate the attention maps, employing a diffusion U-Net for coherent composition, which accelerates stroke initialization. Recognizing the importance of abstract SVGs in communication, we incorporated an MLP-based mechanism to simplify SVGs, with alignment and perceptual loss via differential rendering and opacity modulation to improve aesthetics. CraftSVG outperforms previous methods in abstraction, recognizability, and detail, as depicted by its CLIP-T: 0.5013, Aesthetic: 7.0779, score. The code is available at github.com/CraftSVG.
Ayan Banerjee 0002, Nityanand Mathur, Josep Lladós 0001, Umapada Pal 0001, Anjan Dutta 0001
WACV5
2026 Beyond consistency: Preserving temporal structure in zero-shot video editing
Deyin Liu, Yisheng Ding, Zhe Jin 0001, Xiatian Zhu, Anjan Dutta 0001, Lin Wu 0001
Pattern Recognit.5
2025 OmniCount: Multi-label Object Counting with Semantic-Geometric Priors
abstract
Object counting is pivotal for understanding the composition of scenes. Previously, this task was dominated by class-specific methods, which have gradually evolved into more adaptable class-agnostic strategies. However, these strategies come with their own set of limitations, such as the need for manual exemplar input and multiple passes for multiple categories, resulting in significant inefficiencies. This paper introduces a more practical approach enabling simultaneous counting of multiple object categories using an open-vocabulary framework. Our solution, OmniCount, stands out by using semantic and geometric insights (priors) from pre-trained models to count multiple categories of objects as specified by users, all without additional training. OmniCount distinguishes itself by generating precise object masks and leveraging varied interactive prompts via the Segment Anything Model for efficient counting. To evaluate OmniCount, we created the OmniCount-191 benchmark, a first-of-its-kind dataset with multi-label object counts, including points, bounding boxes, and VQA annotations. Our comprehensive evaluation in OmniCount-191, alongside other leading benchmarks, demonstrates OmniCount's exceptional performance, significantly outpacing existing solutions.
Anindya Mondal, Sauradip Nag, Xiatian Zhu, Anjan Dutta 0001
AAAI4
2025 A Closer Look at Multimodal Representation Collapse
abstract
We aim to develop a fundamental understanding of modality collapse, a recently observed empirical phenomenon wherein models trained for multimodal fusion tend to rely only on a subset of the modalities, ignoring the rest. We show that modality collapse happens when noisy features from one modality are entangled, via a shared set of neurons in the fusion head, with predictive features from another, effectively masking out positive contributions from the predictive features of the former modality and leading to its collapse. We further prove that cross-modal knowledge distillation implicitly disentangles such representations by freeing up rank bottlenecks in the student encoder, denoising the fusion-head outputs without negatively impacting the predictive features from either modality. Based on the above findings, we propose an algorithm that prevents modality collapse through explicit basis reallocation, with applications in dealing with missing modalities. Extensive experiments on multiple multimodal benchmarks validate our theoretical claims. Project page: https://abhrac.github.io/mmcollapse/.
Abhra Chaudhuri, Anjan Dutta 0001, Tu Bui, Serban Georgescu
ICML2
2025 RespoDiff: Dual-Module Bottleneck Transformation for Responsible & Faithful T2I Generation
abstract
The rapid advancement of diffusion models has enabled high-fidelity and semantically rich text-to-image generation; however, ensuring fairness and safety remains an open challenge. Existing methods typically improve fairness and safety at the expense of semantic fidelity and image quality. In this work, we propose RespoDiff, a novel framework for responsible text-to-image generation that incorporates a dual-module transformation on the intermediate bottleneck representations of diffusion models. Our approach introduces two distinct learnable modules: one focused on capturing and enforcing responsible concepts, such as fairness and safety, and the other dedicated to maintaining semantic alignment with neutral prompts. To facilitate the dual learning process, we introduce a novel score-matching objective that enables effective coordination between the modules. Our method outperforms state-of-the-art methods in responsible generation by ensuring semantic alignment while optimizing both objectives without compromising image fidelity. Our approach improves responsible and semantically coherent generation by \textasciitilde20\% across diverse, unseen prompts. Moreover, it integrates seamlessly into large-scale models like SDXL, enhancing fairness and safety. The project page is available at https://vssilpa.github.io/respodiff_project_page.
Silpa Vadakkeeveetil Sreelatha, Sauradip Nag, Serge J. Belongie, Anjan Dutta 0001
NeurIPS5
2024 Learning Conditional Invariances through Non-Commutativity
abstract
Invariance learning algorithms that conditionally filter out domain-specific random variables as distractors, do so based only on the data semantics, and not the target domain under evaluation. We show that a provably optimal and sample-efficient way of learning conditional invariances is by relaxing the invariance criterion to be non-commutatively directed towards the target domain. Under domain asymmetry, i.e., when the target domain contains semantically relevant information absent in the source, the risk of the encoder $\varphi^*$ that is optimal on average across domains is strictly lower-bounded by the risk of the target-specific optimal encoder $\Phi^*_\tau$. We prove that non-commutativity steers the optimization towards $\Phi^*_\tau$ instead of $\varphi^*$, bringing the $\mathcal{H}$-divergence between domains down to zero, leading to a stricter bound on the target risk. Both our theory and experiments demonstrate that non-commutative invariance (NCI) can leverage source domain samples to meet the sample complexity needs of learning $\Phi^*_\tau$, surpassing SOTA invariance learning algorithms for domain adaptation, at times by over 2\%, approaching the performance of an oracle. Implementation is available at https://github.com/abhrac/nci.
Abhra Chaudhuri, Serban Georgescu, Anjan Dutta 0001
ICLR3
2024 DeNetDM: Debiasing by Network Depth Modulation
abstract
Neural networks trained on biased datasets tend to inadvertently learn spurious correlations, hindering generalization. We formally prove that (1) samples that exhibit spurious correlations lie on a lower rank manifold relative to the ones that do not; and (2) the depth of a network acts as an implicit regularizer on the rank of the attribute subspace that is encoded in its representations. Leveraging these insights, we present DeNetDM, a novel debiasing method that uses network depth modulation as a way of developing robustness to spurious correlations. Using a training paradigm derived from Product of Experts, we create both biased and debiased branches with deep and shallow architectures and then distill knowledge to produce the target debiased model. Our method requires no bias annotations or explicit data augmentation while performing on par with approaches that require either or both. We demonstrate that DeNetDM outperforms existing debiasing techniques on both synthetic and real-world datasets by 5\%. The project page is available at https://vssilpa.github.io/denetdm/.
Silpa Vadakkeeveetil Sreelatha, Adarsh K, Abhra Chaudhuri, Anjan Dutta 0001
NeurIPS4
2024 Relational Proxies: Fine-Grained Relationships as Zero-Shot Discriminators
abstract
Visual categories that largely share the same set of local parts cannot be discriminated based on part information alone, as they mostly differ in the way the local parts relate to the overall global structure of the object. We propose Relational Proxies, a novel approach that leverages the relational information between the global and local views of an object for encoding its semantic label, even for categories it has not encountered during training. Starting with a rigorous formalization of the notion of distinguishability between categories that share attributes, we prove the necessary and sufficient conditions that a model must satisfy in order to learn the underlying decision boundaries to tell them apart. We design Relational Proxies based on our theoretical findings and evaluate it on seven challenging fine-grained benchmark datasets and achieve state-of-the-art results on all of them, surpassing the performance of all existing works with a margin exceeding 4% in some cases. We additionally show that Relational Proxies also generalizes to the zero-shot setting, where it can efficiently leverage emergent relationships among attributes and image views to generalize to unseen categories, surpassing current state-of-the-art in both the non-generative and generative settings. Implementation is available at https://github.com/abhrac/relational-proxies.
Abhra Chaudhuri, Massimiliano Mancini, Zeynep Akata, Anjan Dutta 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Data-Free Sketch-Based Image Retrieval
abstract
Rising concerns about privacy and anonymity preservation of deep learning models have facilitated research in data-free learning (DFL). For the first time, we identify that for data-scarce tasks like Sketch-Based Image Retrieval (SBIR), where the difficulty in acquiring paired photos and hand-drawn sketches limits data-dependent cross-modal learning algorithms, DFL can prove to be a much more practical paradigm. We thus propose Data-Free (DF)-SBIR, where, unlike existing DFL problems, pre-trained, single-modality classification models have to be leveraged to learn a cross-modal metric-space for retrieval without access to any training data. The widespread availability of pre-trained classification models, along with the difficulty in acquiring paired photo-sketch datasets for SBIR justify the practicality of this setting. We present a methodology for DF-SBIR, which can leverage knowledge from models independently trained to perform classification on photos and sketches. We evaluate our model on the Sketchy, TU-Berlin, and QuickDraw benchmarks, designing a variety of baselines based on state-of-the-art DFL literature, and observe that our method surpasses all of them by significant margins. Our method also achieves mAPs competitive with data-dependent approaches, all the while requiring no training data. Implementation is available at https://github.com/abhrac/data-free-sbir.
Abhra Chaudhuri, Ayan Kumar Bhunia, Yi-Zhe Song, Anjan Dutta 0001
CVPR4
2023 Transitivity Recovering Decompositions: Interpretable and Robust Fine-Grained Relationships
abstract
Recent advances in fine-grained representation learning leverage local-to-global (emergent) relationships for achieving state-of-the-art results. The relational representations relied upon by such methods, however, are abstract. We aim to deconstruct this abstraction by expressing them as interpretable graphs over image views. We begin by theoretically showing that abstract relational representations are nothing but a way of recovering transitive relationships among local views. Based on this, we design Transitivity Recovering Decompositions (TRD), a graph-space search algorithm that identifies interpretable equivalents of abstract emergent relationships at both instance and class levels, and with no post-hoc computations. We additionally show that TRD is provably robust to noisy views, with empirical evidence also supporting this finding. The latter allows TRD to perform at par or even better than the state-of-the-art, while being fully interpretable. Implementation is available at https://github.com/abhrac/trd.
Abhra Chaudhuri, Massimiliano Mancini, Zeynep Akata, Anjan Dutta 0001
NeurIPS4
2023 Implicit and explicit attention mechanisms for zero-shot learning
Faisal Alamri, Anjan Dutta 0001
Neurocomputing2
2022 Cross-Modal Fusion Distillation for Fine-Grained Sketch-Based Image Retrieval
Abhra Chaudhuri, Massimiliano Mancini, Yanbei Chen, Zeynep Akata, Anjan Dutta 0001
BMVC5
2022 Abstracting Sketches Through Simple Primitives
Stephan Alaniz, Massimiliano Mancini, Anjan Dutta 0001, Diego Marcos, Zeynep Akata
ECCV (29)3
2022 Relational Proxies: Emergent Relationships as Fine-Grained Discriminators
abstract
Fine-grained categories that largely share the same set of parts cannot be discriminated based on part information alone, as they mostly differ in the way the local parts relate to the overall global structure of the object. We propose Relational Proxies, a novel approach that leverages the relational information between the global and local views of an object for encoding its semantic label. Starting with a rigorous formalization of the notion of distinguishability between fine-grained categories, we prove the necessary and sufficient conditions that a model must satisfy in order to learn the underlying decision boundaries in the fine-grained setting. We design Relational Proxies based on our theoretical findings and evaluate it on seven challenging fine-grained benchmark datasets and achieve state-of-the-art results on all of them, surpassing the performance of all existing works with a margin exceeding 4% in some cases. We also experimentally validate our theory on fine-grained distinguishability and obtain consistent results across multiple benchmarks. Implementation is available at https://github.com/abhrac/relational-proxies.
Abhra Chaudhuri, Massimiliano Mancini, Zeynep Akata, Anjan Dutta 0001
NeurIPS4
2022 BDA-SketRet: Bi-level domain adaptation for zero-shot SBIR
Ushasi Chaudhuri, Ruchika Chavan, Biplab Banerjee, Anjan Dutta 0001, Zeynep Akata
Neurocomputing4
2020 Learning Robust Representations via Multi-View Information Bottleneck
Marco Federici, Anjan Dutta 0001, Patrick Forré, Nate Kushman, Zeynep Akata
ICLR2
2020 Semantically Tied Paired Cycle Consistency for Any-Shot Sketch-Based Image Retrieval
abstract
Abstract Low-shot sketch-based image retrieval is an emerging task in computer vision, allowing to retrieve natural images relevant to hand-drawn sketch queries that are rarely seen during the training phase. Related prior works either require aligned sketch-image pairs that are costly to obtain or inefficient memory fusion layer for mapping the visual information to a semantic space. In this paper, we address any-shot,i.e. zero-shot and few-shot, sketch-based image retrieval (SBIR) tasks, where we introduce the few-shot setting for SBIR. For solving these tasks, we propose a semantically aligned paired cycle-consistent generative adversarial network (SEM-PCYC) for any-shot SBIR, where each branch of the generative adversarial network maps the visual information from sketch and image to a common semantic space via adversarial training. Each of these branches maintains cycle consistency that only requires supervision at the category level, and avoids the need of aligned sketch-image pairs. A classification criteria on the generators’ outputs ensures the visual to semantic space mapping to be class-specific. Furthermore, we propose to combine textual and hierarchical side information via an auto-encoder that selects discriminating side information within a same end-to-end model. Our results demonstrate a significant boost in any-shot SBIR performance over the state-of-the-art on the extended version of the challenging Sketchy, TU-Berlin and QuickDraw datasets.
Anjan Dutta 0001, Zeynep Akata
Int. J. Comput. Vis.1
2020 Hierarchical stochastic graphlet embedding for graph-based pattern recognition
abstract
Abstract Despite being very successful within the pattern recognition and machine learning community, graph-based methods are often unusable because of the lack of mathematical operations defined in graph domain. Graph embedding, which maps graphs to a vectorial space, has been proposed as a way to tackle these difficulties enabling the use of standard machine learning techniques. However, it is well known that graph embedding functions usually suffer from the loss of structural information. In this paper, we consider the hierarchical structure of a graph as a way to mitigate this loss of information. The hierarchical structure is constructed by topologically clustering the graph nodes and considering each cluster as a node in the upper hierarchical level. Once this hierarchical structure is constructed, we consider several configurations to define the mapping into a vector space given a classical graph embedding, in particular, we propose to make use of the stochastic graphlet embedding (SGE). Broadly speaking, SGE produces a distribution of uniformly sampled low-to-high-order graphlets as a way to embed graphs into the vector space. In what follows, the coarse-to-fine structure of a graph hierarchy and the statistics fetched by the SGE complements each other and includes important structural information with varied contexts. Altogether, these two techniques substantially cope with the usual information loss involved in graph embedding techniques, obtaining a more robust graph representation. This fact has been corroborated through a detailed experimental evaluation on various benchmark graph datasets, where we outperform the state-of-the-art methods.
Anjan Dutta 0001, Pau Riba, Josep Lladós 0001, Alicia Fornés
Neural Comput. Appl.1
2019 Semantically Tied Paired Cycle Consistency for Zero-Shot Sketch-Based Image Retrieval
abstract
Zero-shot sketch-based image retrieval (SBIR) is an emerging task in computer vision, allowing to retrieve natural images relevant to sketch queries that might not been seen in the training phase. Existing works either require aligned sketch-image pairs or inefficient memory fusion layer for mapping the visual information to a semantic space. In this work, we propose a semantically aligned paired cycle-consistent generative (SEM-PCYC) model for zero-shot SBIR, where each branch maps the visual information to a common semantic space via an adversarial training. Each of these branches maintains a cycle consistency that only requires supervision at category levels, and avoids the need of highly-priced aligned sketch-image pairs. A classification criteria on the generators' outputs ensures the visual to semantic space mapping to be discriminating. Furthermore, we propose to combine textual and hierarchical side information via a feature selection auto-encoder that selects discriminating side information within a same end-to-end model. Our results demonstrate a significant boost in zero-shot SBIR performance over the state-of-the-art on the challenging Sketchy and TU-Berlin datasets.
Anjan Dutta 0001, Zeynep Akata
CVPR1
2019 Doodle to Search: Practical Zero-Shot Sketch-Based Image Retrieval
abstract
In this paper, we investigate the problem of zero-shot sketch-based image retrieval (ZS-SBIR), where human sketches are used as queries to conduct retrieval of photos from unseen categories. We importantly advance prior arts by proposing a novel ZS-SBIR scenario that represents a firm step forward in its practical application. The new setting uniquely recognizes two important yet often neglected challenges of practical ZS-SBIR, (i) the large domain gap between amateur sketch and photo, and (ii) the necessity for moving towards large-scale retrieval. We first contribute to the community a novel ZS-SBIR dataset, QuickDraw-Extended, that consists of 330,000 sketches and 204,000 photos spanning across 110 categories. Highly abstract amateur human sketches are purposefully sourced to maximize the domain gap, instead of ones included in existing datasets that can often be semi-photorealistic. We then formulate a ZS-SBIR framework to jointly model sketches and photos into a common embedding space. A novel strategy to mine the mutual information among domains is specifically engineered to alleviate the domain gap. External semantic knowledge is further embedded to aid semantic transfer. We show that, rather surprisingly, retrieval performance significantly outperforms that of state-of-the-art on existing datasets that can already be achieved using a reduced version of our model. We further demonstrate the superior performance of our full model by comparing with a number of alternatives on the newly proposed dataset. The new dataset, plus all training and testing code of our model, will be publicly released to facilitate future research.
Sounak Dey, Pau Riba, Anjan Dutta 0001, Josep Lladós 0001, Yi-Zhe Song
CVPR3
2019 Table Detection in Invoice Documents by Graph Neural Networks
abstract
Tabular structures in documents offer a complementary dimension to the raw textual data, representing logical or quantitative relationships among pieces of information. In digital mail room applications, where a large amount of administrative documents must be processed with reasonable accuracy, the detection and interpretation of tables is crucial. Table recognition has gained interest in document image analysis, in particular in unconstrained formats (absence of rule lines, unknown information of rows and columns). In this work, we propose a graph-based approach for detecting tables in document images. Instead of using the raw content (recognized text), we make use of the location, context and content type, thus it is purely a structure perception approach, not dependent on the language and the quality of the text reading. Our framework makes use of Graph Neural Networks (GNNs) in order to describe the local repetitive structural information of tables in invoice documents. Our proposed model has been experimentally validated in two invoice datasets and achieved encouraging results. Additionally, due to the scarcity of benchmark datasets for this task, we have contributed to the community a novel dataset derived from the RVL-CDIP invoice data. It will be publicly released to facilitate future research.
Pau Riba, Anjan Dutta 0001, Lutz Goldmann, Alicia Fornés, Oriol Ramos Terrades, Josep Lladós 0001
ICDAR2
2019 Stochastic Graphlet Embedding
abstract
Graph-based methods are known to be successful in many machine learning and pattern classification tasks. These methods consider semistructured data as graphs where nodes correspond to primitives (parts, interest points, and segments) and edges characterize the relationships between these primitives. However, these nonvectorial graph data cannot be straightforwardly plugged into off-the-shelf machine learning algorithms without a preliminary step of-explicit/implicit-graph vectorization and embedding. This embedding process should be resilient to intraclass graph variations while being highly discriminant. In this paper, we propose a novel high-order stochastic graphlet embedding that maps graphs into vector spaces. Our main contribution includes a new stochastic search procedure that efficiently parses a given graph and extracts/samples unlimitedly high-order graphlets. We consider these graphlets, with increasing orders, to model local primitives as well as their increasingly complex interactions. In order to build our graph representation, we measure the distribution of these graphlets into a given graph, using particular hash functions that efficiently assign sampled graphlets into isomorphic sets with a very low probability of collision. When combined with maximum margin classifiers, these graphlet-based representations have a positive impact on the performance of pattern comparison and recognition as corroborated through extensive experiments using standard benchmark databases.
Anjan Dutta 0001, Hichem Sahbi
IEEE Trans. Neural Networks Learn. Syst.1
2018 Aligning Salient Objects to Queries: A Multi-modal and Multi-object Image Retrieval Framework
Sounak Dey, Anjan Dutta 0001, Suman K. Ghosh, Ernest Valveny, Josep Lladós 0001, Umapada Pal 0001
ACCV (2)2
2018 Learning Cross-Modal Deep Embeddings for Multi-Object Image Retrieval using Text and Sketch
abstract
In this work we introduce a cross modal image retrieval system that allows both text and sketch as input modalities for the query. A cross-modal deep network architecture is formulated to jointly model the sketch and text input modalities as well as the the image output modality, learning a common embedding between text and images and between sketches and images. In addition, an attention model is used to selectively focus the attention on the different objects of the image, allowing for retrieval with multiple objects in the query. Experiments show that the proposed method performs the best in both single and multiple object image retrieval in standard datasets.
Sounak Dey, Anjan Dutta 0001, Suman K. Ghosh, Ernest Valveny, Josep Lladós 0001, Umapada Pal 0001
ICPR2
2018 Product graph-based higher order contextual similarities for inexact subgraph matching
Anjan Dutta 0001, Josep Lladós 0001, Horst Bunke, Umapada Pal 0001
Pattern Recognit.1
2018 Rough-fuzzy based scene categorization for text detection and recognition in video
Sangheeta Roy, Palaiahnakote Shivakumara, Namita Jain, Vijeta Khare, Anjan Dutta 0001, Umapada Pal 0001, Tong Lu 0002
Pattern Recognit.5
2018 Subgraph spotting in graph representations of comic book images
Nam Le Thanh 0001, Muhammad Muzzamil Luqman, Anjan Dutta 0001, Pierre Héroux, Christophe Rigaud, Clément Guérin, Pasquale Foggia, Jean-Christophe Burie, Jean-Marc Ogier, Josep Lladós 0001, Sébastien Adam
Pattern Recognit. Lett.3
2017 Pyramidal Stochastic Graphlet Embedding for Document Pattern Classification
abstract
Document pattern classification methods using graphs have received a lot of attention because of its robust representation paradigm and rich theoretical background. However, the way of preserving and the process for delineating documents with graphs introduce noise in the rendition of underlying data, which creates instability in the graph representation. To deal with such unreliability in representation, in this paper, we propose Pyramidal Stochastic Graphlet Embedding (PSGE). Given a graph representing a document pattern, our method first computes a graph pyramid by successively reducing the base graph. Once the graph pyramid is computed, we apply Stochastic Graphlet Embedding (SGE) for each level of the pyramid and combine their embedded representation to obtain a global delineation of the original graph. The consideration of pyramid of graphs rather than just a base graph extends the representational power of the graph embedding, which reduces the instability caused due to noise and distortion. When plugged with support vector machine, our proposed PSGE has outperformed the state-of-the-art results in recognition of handwritten words as well as graphical symbols.
Anjan Dutta 0001, Pau Riba, Josep Lladós 0001, Alicia Fornés
ICDAR1
2017 Improving Information Retrieval in Multiwriter Scenario by Exploiting the Similarity Graph of Document Terms
abstract
Information Retrieval (IR) is the activity of obtaining information resources relevant to a questioned information. It usually retrieves a set of objects ranked according to the relevancy to the needed fact. In document analysis, information retrieval receives a lot of attention in terms of symbol and word spotting. However, through decades the community mostly focused either on printed or on single writer scenario, where the state-of-the-art results have achieved reasonable performance on the available datasets. Nevertheless, the existing algorithms do not perform accordingly on multiwriter scenario. A graph representing relations between a set of objects is a structure where each node delineates an individual element and the similarity between them is represented as a weight on the connecting edge. In this paper, we explore different analytics of graphs constructed from words or graphical symbols, such as diffusion, shortest path, etc. to improve the performance of information retrieval methods in multiwriter scenario.
Pau Riba, Anjan Dutta 0001, Sounak Dey, Josep Lladós 0001, Alicia Fornés
ICDAR2
2017 Large-scale graph indexing using binary embeddings of node contexts for information spotting in document image databases
Pau Riba, Josep Lladós 0001, Alicia Fornés, Anjan Dutta 0001
Pattern Recognit. Lett.4
2016 Compact correlated features for writer independent signature verification
abstract
This paper considers the offline signature verification problem which is considered to be an important research line in the field of pattern recognition. In this work we propose hybrid features that consider the local features and their global statistics in the signature image. This has been done by creating a vocabulary of histogram of oriented gradients (HOGs). We impose weights on these local features based on the height information of water reservoirs obtained from the signature. Spatial information between local features are thought to play a vital role in considering the geometry of the signatures which distinguishes the originals from the forged ones. Nevertheless, learning a condensed set of higher order neighbouring features based on visual words, e.g., doublets and triplets, continues to be a challenging problem as possible combinations of visual words grow exponentially. To avoid this explosion of size, we create a code of local pairwise features which are represented as joint descriptors. Local features are paired based on the edges of a graph representation built upon the Delaunay triangulation. We reveal the advantage of combining both type of visual codebooks (order one and pairwise) for signature verification task. This is validated through an encouraging result on two benchmark datasets viz. CEDAR and GPDS300.
Anjan Dutta 0001, Umapada Pal 0001, Josep Lladós 0001
ICPR1
2014 Multi-oriented scene text detection in video based on wavelet and angle projection boundary growing
Palaiahnakote Shivakumara, Anjan Dutta 0001, Chew Lim Tan, Umapada Pal 0001
Multim. Tools Appl.2
2013 Near Convex Region Adjacency Graph and Approximate Neighborhood String Matching for Symbol Spotting in Graphical Documents
abstract
This paper deals with a sub graph matching problem in Region Adjacency Graph (RAG) applied to symbol spotting in graphical documents. RAG is a very important, efficient and natural way of representing graphical information with a graph but this is limited to cases where the information is well defined with perfectly delineated regions. What if the information we are interested in is not confined within well defined regions? This paper addresses this particular problem and solves it by defining near convex grouping of oriented line segments which results in near convex regions. Pure convexity imposes hard constraints and can not handle all the cases efficiently. Hence to solve this problem we have defined a new type of convexity of regions, which allows convex regions to have concavity to some extend. We call this kind of regions Near Convex Regions (NCRs). These NCRs are then used to create the Near Convex Region Adjacency Graph (NCRAG) and with this representation we have formulated the problem of symbol spotting in graphical documents as a sub graph matching problem. For sub graph matching we have used the Approximate Edit Distance Algorithm (AEDA) on the neighborhood string, which starts working after finding a key node in the input or target graph and iteratively identifies similar nodes of the query graph in the neighborhood of the key node. The experiments are performed on artificial, real and distorted datasets.
Anjan Dutta 0001, Josep Lladós 0001, Horst Bunke, Umapada Pal 0001
ICDAR1
2013 A symbol spotting approach in graphical documents by hashing serialized graphs
Anjan Dutta 0001, Josep Lladós 0001, Umapada Pal 0001
Pattern Recognit.1
2012 Combination of product graph and random walk kernel for symbol spotting in graphical documents
Anjan Dutta 0001, Jaume Gibert, Josep Lladós 0001, Horst Bunke, Umapada Pal 0001
ICPR1
2012 CVC-MUSCIMA: a ground truth of handwritten music score images for writer identification and staff removal
Alicia Fornés, Anjan Dutta 0001, Albert Gordo, Josep Lladós 0001
Int. J. Document Anal. Recognit.2
2012 On the Influence of Word Representations for Handwritten Word Spotting in Historical Documents
abstract
Word spotting is the process of retrieving all instances of a queried keyword from a digital library of document images. In this paper we evaluate the performance of different word descriptors to assess the advantages and disadvantages of statistical and structural models in a framework of query-by-example word spotting in historical documents. We compare four word representation models, namely sequence alignment using DTW as a baseline reference, a bag of visual words approach as statistical model, a pseudo-structural model based on a Loci features representation, and a structural approach where words are represented by graphs. The four approaches have been tested with two collections of historical data: the George Washington database and the marriage records from the Barcelona Cathedral. We experimentally demonstrate that statistical representations generally give a better performance, however it cannot be neglected that large descriptors are difficult to be implemented in a retrieval scenario where word spotting requires the indexation of data with million word images.
Josep Lladós 0001, Marçal Rusiñol, Alicia Fornés, David Fernández Mota, Anjan Dutta 0001
Int. J. Pattern Recognit. Artif. Intell.5
2011 Symbol Spotting in Line Drawings through Graph Paths Hashing
abstract
In this paper we propose a symbol spotting technique through hashing the shape descriptors of graph paths (Hamiltonian paths). Complex graphical structures in line drawings can be efficiently represented by graphs, which ease the accurate localization of the model symbol. Graph paths are the factorized substructures of graphs which enable robust recognition even in the presence of noise and distortion. In our framework, the entire database of the graphical documents is indexed in hash tables by the locality sensitive hashing (LSH) of shape descriptors of the paths. The hashing data structure aims to execute an approximate k-NN search in a sub-linear time. The spotting method is formulated by a spatial voting scheme to the list of locations of the paths that are decided during the hash table lookup process. We perform detailed experiments with various dataset of line drawings and the results demonstrate the effectiveness and efficiency of the technique.
Anjan Dutta 0001, Josep Lladós 0001, Umapada Pal 0001
ICDAR1
2011 The ICDAR 2011 Music Scores Competition: Staff Removal and Writer Identification
abstract
In the last years, there has been a growing interest in the analysis of handwritten music scores. In this sense, our goal has been to foster the interest in the analysis of handwritten music scores by the proposal of two different competitions: Staff removal and Writer Identification. Both competitions have been tested on the CVC-MUSCIMA database: a ground-truth of handwritten music score images. This paper describes the competition details, including the dataset and ground-truth, the evaluation metrics, and a short description of the participants, their methods, and the obtained results.
Alicia Fornés, Anjan Dutta 0001, Albert Gordo, Josep Lladós 0001
ICDAR2
2011 A novel mutual nearest neighbor based symmetry for text frame classification in video
Palaiahnakote Shivakumara, Anjan Dutta 0001, Trung Quy Phan, Chew Lim Tan, Umapada Pal 0001
Pattern Recognit.2
2010 A new wavelet-median-moment based method for multi-oriented video text detection
abstract
In this paper, we present a new method based on wavelet-median-moments and a novel idea of angle projection for detecting multi-oriented text in video. The proposed method uses wavelet decomposition first to obtain three high frequency sub-bands (LH, HL and HH) and then median moments are computed on the average sub-bands of the three high frequency sub-bands to brighten the text pixels. K-means clustering (K=2) is used for obtaining text pixels from the wavelet-median-moments features (WMMF). Text candidates are obtained by mapping the output of K-means on Sobel edge map of the original input frame. To deal with multi-oriented text, we introduce a new idea of Angle Projection (AP) based on boundary growing and nearest neighbor concepts from the text candidates instead of conventional projection profiles. The proposed method is experimented on horizontal text data, non-horizontal text data, temporal data, non-text data and camera based images (scene text data of ICDAR 2003 competition) to show that the proposed method is superior to existing methods.
Palaiahnakote Shivakumara, Anjan Dutta 0001, Chew Lim Tan, Umapada Pal 0001
Document Analysis Systems2
2010 A New Method for Handwritten Scene Text Detection in Video
abstract
There are many video images where hand written text may appear. Therefore handwritten scene text detection in video is essential and useful for many applications for efficient indexing, retrieval etc. Also there are many video frames where text line may be multi-oriented in nature. To the best of our knowledge there is no work on handwritten text detection in video, which is multi-oriented in nature. In this paper, we present a new method based on maximum color difference and boundary growing method for detection of multi-oriented handwritten scene text in video. The method computes maximum color difference for the average of R, G and B channels of the original frame to enhance the text information. The output of maximum color difference is fed to a K-means algorithm with K = 2 to separate text and non-text clusters. Text candidates are obtained by intersecting the text cluster with the Sobel output of the original frame. To tackle the fundamental problem of different orientations and skews of handwritten text, boundary growing method based on a nearest neighbor concept is employed. We evaluate the proposed method by testing on our own handwritten text database and publicly available video data (Hua's data). Experimental results obtained from the proposed method are promising.
Palaiahnakote Shivakumara, Anjan Dutta 0001, Umapada Pal 0001, Chew Lim Tan
ICFHR2
2010 An Efficient Staff Removal Approach from Printed Musical Documents
abstract
Staff removal is an important preprocessing step of the Optical Music Recognition (OMR). The process aims to remove the stafflines from a musical document and retain only the musical symbols, later these symbols are used effectively to identify the music information. This paper proposes a simple but robust method to remove stafflines from printed musical scores. In the proposed methodology we have considered a staffline segment as a horizontal linkage of vertical black runs with uniform height. We have used the neighbouring properties of a staffline segment to validate it as a true segment. We have considered the dataset along with the deformations described in for evaluation purpose. From experimentation we have got encouraging results.
Anjan Dutta 0001, Umapada Pal 0001, Alicia Fornés, Josep Lladós 0001
ICPR1
2010 A New Symmetry Based on Proximity of Wavelet-Moments for Text Frame Classification in Video
abstract
This paper proposes the use of a new symmetry property based on proximity of the median moments in the wavelet domain. The method divides a given frame into 16 equally sized blocks to classify the true text frame. The average of high frequency subbands of a block is used for computing median moments to brighten the text pixel in a block of video frame. Then K-means clustering with K=2 is applied on the median moments of the block to classify it as a probable text block. For classified blocks, average wavelet median moments are computed for a sliding window. We introduce Max-Min cluster to classify the probable text pixel in each probable text block. The four quadrants are formed from the centroid of the probable text pixels. The new concept called symmetry is introduced to identify the true text block based on proximity between probable text pixels in each quadrant. If the frame produces at least one true text block, it is considered as a text frame otherwise a non-text frame. The method is tested on three datasets to evaluate the robustness of the method in classification of text frames in terms of recall and precision.
Palaiahnakote Shivakumara, Anjan Dutta 0001, Chew Lim Tan, Umapada Pal 0001
ICPR2