Pulak Purkait

dblp:63/9103 · DBLP profile ↗
← Back
25ranked-venue papers
16as first author
4since 2021 · last 2025
0000-0003-0684-1209ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 13 first-author · 3 since 2021Artificial intelligence and machine learning · 17 · 9 first-author · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Image recognition and object detection · 26% 3D vision · 23% Segmentation and scene understanding · 13%
Computer graphics and multimedia
6 papers
Image and video processing · 68% Visual content generation and editing · 22% Computational photography and imaging · 7%
Databases, data mining, and information retrieval
3 papers
Data mining · 57% Information retrieval · 43%
Theoretical computer science
2 papers
Algorithms and data structures · 78% Mathematical optimization · 22%

Topics — the 30 heaviest of 41, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video processing
super-resolution
1.232025
SRSR: Enhancing Semantic Accuracy in Real-World Image Super-Resolution with Spatially Re-Focused Text-Conditioning · NeurIPS 2025
A Fuzzy-Rule-Based Approach for Single Frame Super Resolution · IEEE Trans. Image Process. 2014
Super Resolution Image Reconstruction Through Bregman Iteration Using Morphologic Regularization · IEEE Trans. Image Process. 2012
Image and video processing › super-resolution › image super-resolution › generative image super-resolution
diffusion-based super-resolution
0.912025
SRSR: Enhancing Semantic Accuracy in Real-World Image Super-Resolution with Spatially Re-Focused Text-Conditioning · NeurIPS 2025
Visual content generation and editing
visual text generation
0.912025
SRSR: Enhancing Semantic Accuracy in Real-World Image Super-Resolution with Spatially Re-Focused Text-Conditioning · NeurIPS 2025
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
label noise robustness
0.712023
Semi-Supervised Semantic Segmentation under Label Noise via Diverse Learning Groups · ICCV 2023
Computer vision › Image recognition and object detection
object detection
0.712023
Knowledge Combination to Learn Rotated Detection without Rotated Annotation · CVPR 2023
Computer vision › Image recognition and object detection › object detection
oriented object detection
0.712023
Knowledge Combination to Learn Rotated Detection without Rotated Annotation · CVPR 2023
Computer vision › Segmentation and scene understanding
semantic segmentation
0.712023
Semi-Supervised Semantic Segmentation under Label Noise via Diverse Learning Groups · ICCV 2023
Computer vision › Segmentation and scene understanding › annotation-efficient segmentation
semi-supervised semantic segmentation
0.712023
Semi-Supervised Semantic Segmentation under Label Noise via Diverse Learning Groups · ICCV 2023
Computer vision › Image recognition and object detection › object detection
weakly supervised object detection
0.712023
Knowledge Combination to Learn Rotated Detection without Rotated Annotation · CVPR 2023
Computer vision › Image recognition and object detection
image classification
0.612022
Retrieval Augmented Classification for Long-Tail Visual Recognition · CVPR 2022
Machine learning › Learning paradigms
long-tailed recognition
0.612022
Retrieval Augmented Classification for Long-Tail Visual Recognition · CVPR 2022
Information retrieval › retrieval augmentation
retrieval-augmented classification
0.612022
Retrieval Augmented Classification for Long-Tail Visual Recognition · CVPR 2022
Algorithms and data structures › data structure design › search structures
search trees
0.522017
Efficient Globally Optimal Consensus Maximisation with Tree Search · IEEE Trans. Pattern Anal. Mach. Intell. 2017
Efficient globally optimal consensus maximisation with tree search · CVPR 2015
Data mining
clustering
0.522017
Clustering with Hypergraphs: The Case for Large Hyperedges · IEEE Trans. Pattern Anal. Mach. Intell. 2017
Clustering with Hypergraphs: The Case for Large Hyperedges · ECCV (4) 2014
Computer vision › 3D vision › 3d scene understanding › scene synthesis
indoor scene synthesis
0.412020
SG-VAE: Scene Grammar Variational Autoencoder to Generate New Indoor Scenes · ECCV (24) 2020
Robotics › Robot navigation and mapping
localization
0.412020
Resolving Marker Pose Ambiguity by Robust Rotation Averaging with Clique Constraints* · ICRA 2020
Robotics › Robot navigation and mapping › localization › landmark-based localization
marker-based localization
0.412020
Resolving Marker Pose Ambiguity by Robust Rotation Averaging with Clique Constraints* · ICRA 2020
Computer vision › 3D vision › pose estimation › rigid body pose estimation
planar pose estimation
0.412020
Resolving Marker Pose Ambiguity by Robust Rotation Averaging with Clique Constraints* · ICRA 2020
Computer vision › 3D vision
pose estimation
0.412020
Resolving Marker Pose Ambiguity by Robust Rotation Averaging with Clique Constraints* · ICRA 2020
Computer vision › 3D vision › structure from motion
rotation averaging
0.412020
NeuRoRA: Neural Robust Rotation Averaging · ECCV (24) 2020
Machine learning › Generative modeling
variational autoencoder
0.412020
SG-VAE: Scene Grammar Variational Autoencoder to Generate New Indoor Scenes · ECCV (24) 2020
Data mining › clustering › graph clustering
hypergraph clustering
0.312017
Clustering with Hypergraphs: The Case for Large Hyperedges · IEEE Trans. Pattern Anal. Mach. Intell. 2017
Computational photography and imaging › image signal processing
rolling shutter correction
0.312017
Rolling Shutter Correction in Manhattan World · ICCV 2017
Algorithms and data structures › search algorithms › heuristic search
a* search
0.312017
Efficient Globally Optimal Consensus Maximisation with Tree Search · IEEE Trans. Pattern Anal. Mach. Intell. 2017
Computer vision › 3D vision › robust estimation
maximum consensus
0.212015
Efficient globally optimal consensus maximisation with tree search · CVPR 2015
Computer vision › 3D vision
robust estimation
0.212015
Efficient globally optimal consensus maximisation with tree search · CVPR 2015
Mathematical optimization
combinatorial optimization
0.212015
Efficient globally optimal consensus maximisation with tree search · CVPR 2015
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels
0.212023
Semi-Supervised Semantic Segmentation under Label Noise via Diverse Learning Groups · ICCV 2023
Machine learning › Trustworthy machine learning
robustness
0.212023
Semi-Supervised Semantic Segmentation under Label Noise via Diverse Learning Groups · ICCV 2023
Machine learning › Graph learning › hypergraph learning
hypergraph clustering
0.212014
Clustering with Hypergraphs: The Case for Large Hyperedges · ECCV (4) 2014

Methods — techniques the papers use, named apart from their topics

retrieval module · 1.1non-parametric memory · 1.1external memory · 1.1segmentation masks · 0.9cross-attention · 0.9classifier-free guidance · 0.9a* search · 0.8LP-type methods · 0.8teacher-student learning · 0.7projection loss · 0.7label filtering · 0.7knowledge combination · 0.7diverse learning groups · 0.7co-training · 0.7random cluster models · 0.6guided sampling · 0.6neural network · 0.4vanishing direction estimation · 0.3
YearPublicationVenuePosition
2025 SRSR: Enhancing Semantic Accuracy in Real-World Image Super-Resolution with Spatially Re-Focused Text-Conditioning
abstract
Existing diffusion-based super-resolution approaches often exhibit semantic ambiguities due to inaccuracies and incompleteness in their text conditioning, coupled with the inherent tendency for cross-attention to divert towards irrelevant pixels. These limitations can lead to semantic misalignment and hallucinated details in the generated high-resolution outputs. To address these, we propose a novel, plug-and-play *spatially re-focused super-resolution (SRSR)* framework that consists of two core components: first, we introduce Spatially Re-focused Cross-Attention (SRCA), which refines text conditioning at inference time by applying visually-grounded segmentation masks to guide cross-attention. Second, we introduce a Spatially Targeted Classifier-Free Guidance (STCFG) mechanism that selectively bypasses text influences on ungrounded pixels to prevent hallucinations. Extensive experiments on both synthetic and real-world datasets demonstrate that SRSR consistently outperforms seven state-of-the-art baselines in standard fidelity metrics (PSNR and SSIM) across all datasets, and in perceptual quality measures (LPIPS and DISTS) on two real-world benchmarks, underscoring its effectiveness in achieving both high semantic fidelity and perceptual quality in super-resolution.
Majid Abdolshah, Violetta Shevchenko, Hongdong Li, Pulak Purkait
NeurIPS6
2023 Knowledge Combination to Learn Rotated Detection without Rotated Annotation
abstract
Rotated bounding boxes drastically reduce output ambiguity of elongated objects, making it superior to axis-aligned bounding boxes. Despite the effectiveness, rotated detectors are not widely employed. Annotating rotated bounding boxes is such a laborious process that they are not provided in many detection datasets where axis-aligned annotations are used instead. In this paper, we propose a framework that allows the model to predict precise rotated boxes only requiring cheaper axis-aligned annotation of the target dataset1. To achieve this, we leverage the fact that neural networks are capable of learning richer representation of the target domain than what is utilized by the task. The under-utilized representation can be exploited to address a more detailed task. Our framework combines task knowledge of an out-of-domain source dataset with stronger annotation and domain knowledge of the target dataset with weaker annotation. A novel assignment process and projection loss are used to enable the cotraining on the source and target datasets. As a result, the model is able to solve the more detailed task in the target domain, without additional computation overhead during inference. We extensively evaluate the method on various target datasets including fresh-produce dataset, HRSC2016 and SSDD. Results show that the proposed method consistently performs on par with the fully supervised approach.
Tianyu Zhu 0001, Bryce Ferenczi, Pulak Purkait, Tom Drummond, Seyed Hamid Rezatofighi, Anton van den Hengel
CVPR3
2023 Semi-Supervised Semantic Segmentation under Label Noise via Diverse Learning Groups
abstract
Semi-supervised semantic segmentation methods use a small amount of clean pixel-level annotations to guide the interpretation of a larger quantity of unlabelled image data. The challenges of providing pixel-accurate annotations at scale mean that the labels are typically noisy, and this contaminates the final results. In this work, we propose an approach that is robust to label noise in the annotated data. The method uses two diverse learning groups with different network architectures to effectively handle both label noise and unlabelled images. Each learning group consists of a teacher network, a student network and a novel filter module. The filter module of each learning group utilizes pixel-level features from the teacher network to detect incorrectly labelled pixels. To reduce confirmation bias, we employ the labels cleaned by the filter module from one learning group to train the other learning group. Experimental results on two different benchmarks and settings demonstrate the superiority of our method over state-of-the-art approaches.
Peixia Li, Pulak Purkait, Thalaiyasingam Ajanthan, Majid Abdolshah, Ravi Garg, Hisham Husain, Stephen Gould, Wanli Ouyang, Anton van den Hengel
ICCV2
2022 Retrieval Augmented Classification for Long-Tail Visual Recognition
abstract
We introduce Retrieval Augmented Classification (RAC), a generic approach to augmenting standard image classification pipelines with an explicit retrieval module. RAC consists of a standard base image encoder fused with a parallel retrieval branch that queries a non-parametric external memory of pre-encoded images and associated text snippets. We apply RAC to the problem of long-tail classification and demonstrate a significant improvement over previous state-of-the-art on Places365-LT and iNaturalist-2018 (14.5% and 6.7% respectively), despite using only the training datasets themselves as the external information source. We demonstrate that RAC's retrieval module, without prompting, learns a high level of accuracy on tail classes. This, in turn, frees the base encoder to focus on common classes, and improve its performance thereon. RAC represents an alternative approach to utilizing large, pretrained models without requiring fine-tuning, as well as a first step towards more effectively making use of external memory within common computer vision architectures.
Alexander Long, Wei Yin 0006, Thalaiyasingam Ajanthan, Pulak Purkait, Ravi Garg, Alan Blair 0001, Chunhua Shen, Anton van den Hengel
CVPR5
2020 NeuRoRA: Neural Robust Rotation Averaging
Pulak Purkait, Tat-Jun Chin, Ian D. Reid 0001
ECCV (24)1
2020 SG-VAE: Scene Grammar Variational Autoencoder to Generate New Indoor Scenes
Pulak Purkait, Christopher Zach, Ian D. Reid 0001
ECCV (24)1
2020 Resolving Marker Pose Ambiguity by Robust Rotation Averaging with Clique Constraints*
abstract
Planar markers are useful in robotics and computer vision for mapping and localisation. Given a detected marker in an image, a frequent task is to estimate the 6DOF pose of the marker relative to the camera, which is an instance of planar pose estimation (PPE). Although there are mature techniques, PPE suffers from a fundamental ambiguity problem, in that there can be more than one plausible pose solutions for a PPE instance. Especially when localisation of the marker corners is noisy, it is often difficult to disambiguate the pose solutions based on reprojection error alone. Previous methods choose between the possible solutions using a heuristic criterion, or simply ignore ambiguous markers.We propose to resolve the ambiguities by examining the consistencies of a set of markers across multiple views. Our specific contributions include a novel rotation averaging formulation that incorporates long-range dependencies between possible marker orientation solutions that arise from PPE ambiguities. We analyse the combinatorial complexity of the problem, and develop a novel lifted algorithm to effectively resolve marker pose ambiguities, without discarding any marker observations. Results on real and synthetic data show that our method is able to handle highly ambiguous inputs, and provides more accurate and/or complete marker-based mapping and localisation.
Shin-Fang Ch'ng, Naoya Sogi, Pulak Purkait, Tat-Jun Chin, Kazuhiro Fukui
ICRA3
2019 Seeing Behind Things: Extending Semantic Segmentation to Occluded Regions
abstract
Semantic segmentation and instance level segmentation made substantial progress in recent years due to the emergence of deep neural networks (DNNs). A number of deep architectures with Convolution Neural Networks (CNNs) were proposed that surpass the traditional machine learning approaches for segmentation by a large margin. These architectures predict the directly observable semantic category of each pixel by usually optimizing a cross-entropy loss. In this work we push the limit of semantic segmentation towards predicting semantic labels of directly visible as well as occluded objects or objects parts, where the network's input is a single depth image. We group the semantic categories into one background and multiple foreground object groups, and we propose a modification of the standard cross-entropy loss to cope with the settings. In our experiments we demonstrate that a CNN trained by minimizing the proposed loss is able to predict semantic categories for visible and occluded object parts without requiring to increase the network size (compared to a standard segmentation task). The results are validated on a newly generated dataset (augmented from SUNCG) dataset.
Pulak Purkait, Christopher Zach, Ian D. Reid 0001
IROS1
2018 Weakly Supervised Learning of Indoor Geometry by Dual Warping
abstract
A major element of depth perception and 3D understanding is the ability to predict the 3D layout of a scene and its contained objects for a novel pose. Indoor environments are particularly suitable for novel view prediction, since the set of objects in such environments is relatively restricted. In this work we address the task of 3D prediction especially for indoor scenes by leveraging only weak supervision. In the literature 3D scene prediction is usually solved via a 3D voxel grid. However, such methods are limited to estimating rather coarse 3D voxel grids, since predicting entire voxel spaces has large computational costs. Hence, our method operates in image-space rather than in voxel space, and the task of 3D estimation essentially becomes a depth image completion problem. We propose a novel approach to easily generate training data containing depth maps with realistic occlusions, and subsequently train a network for completing those occluded regions. Using multiple publicly available datasets we benchmark our method against existing approaches and are able to obtain superior performance. We further demonstrate the flexibility of our method by presenting results for new view synthesis of RGB-D images.
Pulak Purkait, Ujwal Bonde, Christopher Zach
3DV1
2018 Synthetic View Generation for Absolute Pose Regression and Image Synthesis
Pulak Purkait, Cheng Zhao 0002, Christopher Zach
BMVC1
2018 Learning Monocular Visual Odometry with Dense 3D Mapping from Dense 3D Flow
abstract
This paper introduces a fully deep learning approach to monocular SLAM, which can perform simultaneous localization using a neural network for learning visual odometry (L-VO) and dense 3D mapping. Dense 2D flow and a depth image are generated from monocular images by sub-networks, which are then used by a 3D flow associated layer in the L-VO network to generate dense 3D flow. Given this 3D flow, the dual-stream L-VO network can then predict the 6DOF relative pose and furthermore reconstruct the vehicle trajectory. In order to learn the correlation between motion directions, the Bivariate Gaussian modeling is employed in the loss function. The L-VO network achieves an overall performance of 2.68 % for average translational error and 0.0143°/m for average rotational error on the KITTI odometry benchmark. Moreover, the learned depth is leveraged to generate a dense 3D map. As a result, an entire visual SLAM system, that is, learning monocular odometry combined with dense 3D mapping, is achieved.
Cheng Zhao 0002, Li Sun 0005, Pulak Purkait, Tom Duckett, Rustam Stolkin
IROS3
2018 Minimal Solvers for Monocular Rolling Shutter Compensation Under Ackermann Motion
abstract
Modern automotive vehicles are often equipped with a budget commercial rolling shutter camera. These devices often produce distorted images due to the inter-row delay of the camera while capturing the image. Recent methods for monocular rolling shutter motion compensation utilize blur kernel and the straightness property of line segments. However, these methods are limited to handling rotational motion and also are not fast enough to operate in real time. In this paper, we propose a minimal solver for the rolling shutter motion compensation which assumes known vertical direction of the camera. Thanks to the Ackermann motion model of vehicles which consists of only two motion parameters, and two parameters for the simplified depth assumption that lead to a 4-line algorithm. The proposed minimal solver estimates the rolling shutter camera motion efficiently and accurately. The extensive experiments on real and simulated datasets demonstrate the benefits of our approach in terms of qualitative and quantitative results.
Pulak Purkait, Christopher Zach
WACV1
2017 Rolling Shutter Correction in Manhattan World
abstract
A vast majority of consumer cameras operate the rolling shutter mechanism, which often produces distorted images due to inter-row delay while capturing an image. Recent methods for monocular rolling shutter compensation utilize blur kernel, straightness of line segments, as well as angle and length preservation. However, they do not incorporate scene geometry explicitly for rolling shutter correction, therefore, information about the 3D scene geometry is often distorted by the correction process. In this paper we propose a novel method which leverages geometric properties of the scene-in particular vanishing directions-to estimate the camera motion during rolling shutter exposure from a single distorted image. The proposed method jointly estimates the orthogonal vanishing directions and the rolling shutter camera motion. We performed extensive experiments on synthetic and real datasets which demonstrate the benefits of our approach both in terms of qualitative and quantitative results (in terms of a geometric structure fitting) as well as with respect to computation time.
Pulak Purkait, Christopher Zach, Ales Leonardis
ICCV1
2017 Efficient Globally Optimal Consensus Maximisation with Tree Search
abstract
Maximum consensus is one of the most popular criteria for robust estimation in computer vision. Despite its widespread use, optimising the criterion is still customarily done by randomised sample-and-test techniques, which do not guarantee optimality of the result. Several globally optimal algorithms exist, but they are too slow to challenge the dominance of randomised methods. Our work aims to change this state of affairs by proposing an efficient algorithm for global maximisation of consensus. Under the framework of LP-type methods, we show how consensus maximisation for a wide variety of vision tasks can be posed as a tree search problem. This insight leads to a novel algorithm based on A* search. We propose efficient heuristic and support set updating routines that enable A* search to efficiently find globally optimal results. On common estimation problems, our algorithm is much faster than previous exact methods. Our work identifies a promising direction for globally optimal consensus maximisation.
Tat-Jun Chin, Pulak Purkait, Anders P. Eriksson, David Suter
IEEE Trans. Pattern Anal. Mach. Intell.2
2017 Clustering with Hypergraphs: The Case for Large Hyperedges
abstract
The extension of conventional clustering to hypergraph clustering, which involves higher order similarities instead of pairwise similarities, is increasingly gaining attention in computer vision. This is due to the fact that many clustering problems require an affinity measure that must involve a subset of data of size more than two. In the context of hypergraph clustering, the calculation of such higher order similarities on data subsets gives rise to hyperedges. Almost all previous work on hypergraph clustering in computer vision, however, has considered the smallest possible hyperedge size, due to a lack of study into the potential benefits of large hyperedges and effective algorithms to generate them. In this paper, we show that large hyperedges are better from both a theoretical and an empirical standpoint. We then propose a novel guided sampling strategy for large hyperedges, based on the concept of random cluster models. Our method can generate large pure hyperedges that significantly improve grouping accuracy without exponential increases in sampling costs. We demonstrate the efficacy of our technique on various higher-order grouping problems. In particular, we show that our approach improves the accuracy and efficiency of motion segmentation from dense, long-term, trajectories.
Pulak Purkait, Tat-Jun Chin, Alireza Sadri, David Suter
IEEE Trans. Pattern Anal. Mach. Intell.1
2015 Efficient globally optimal consensus maximisation with tree search
abstract
Maximum consensus is one of the most popular criteria for robust estimation in computer vision. Despite its widespread use, optimising the criterion is still customarily done by randomised sample-and-test techniques, which do not guarantee optimality of the result. Several globally optimal algorithms exist, but they are too slow to challenge the dominance of randomised methods. We aim to change this state of affairs by proposing a very efficient algorithm for global maximisation of consensus. Under the framework of LP-type methods, we show how consensus maximisation for a wide variety of vision tasks can be posed as a tree search problem. This insight leads to a novel algorithm based on A* search. We propose efficient heuristic and support set updating routines that enable A* search to rapidly find globally optimal results. On common estimation problems, our algorithm is several orders of magnitude faster than previous exact methods. Our work identifies a promising solution for globally optimal consensus maximisation.
Tat-Jun Chin, Pulak Purkait, Anders P. Eriksson, David Suter
CVPR2
2014 Clustering with Hypergraphs: The Case for Large Hyperedges
Pulak Purkait, Tat-Jun Chin, Hanno Ackermann, David Suter
ECCV (4)1
2014 A Fuzzy-Rule-Based Approach for Single Frame Super Resolution
abstract
In this paper, a novel fuzzy rule-based prediction framework is developed for high-quality image zooming. In classical interpolation-based image zooming, resolution is increased by inserting pixels using certain interpolation techniques. Here, we propose a patch-based image zooming technique, where each low-resolution (LR) image patch is replaced by an estimated high-resolution (HR) patch. Since an LR patch can be generated from any of the many possible HR patches, it would be natural to develop rules to find different possible HR patches and then to combine them according to rule strength to get the estimated HR patch. Here, we generate a large number of LR–HR patch pairs from a collection of natural images, group them into different clusters, and then generate a fuzzy rule for each of these clusters. The rule parameters are also learned from these LR-HR patch pairs. As a result, an efficient mapping from LR patch space to HR patch space can be formulated. The performance of the proposed method is tested on different images, and is also compared with other representative as well as state-of-the-art image zooming techniques. Experimental results show that the proposed method is better than the competing methods and is capable of reconstructing thin lines, edges, fine details, and textures in the image efficiently.
Pulak Purkait, Nikhil R. Pal, Bhabatosh Chanda
IEEE Trans. Image Process.1
2013 Fuzzy-rule based approach for single frame super resolution
abstract
We propose a novel high-quality image zooming technique using fuzzy rule based prediction framework. Our approach is based on the fundamental idea that a low-resolution (LR) image patch could be generated from any of the many possible high-resolution (HR) image patches. Therefore it would be natural to assign certain fuzzyness to each of the possibilities of HR patches. We develop a prediction system that learns the LR-HR patch correspondence and also the natural image patch prior from an external database. We do so by collecting a large amount of the LR-HR natural image patch pairs from an existing database, grouping them into different clusters and then generating fuzzy rules to get an efficient mapping from LR patch space to HR patch space. Experimental results show the efficacy of our method over existing state-of-art methods.
Pulak Purkait, Bhabatosh Chanda
FUZZ-IEEE1
2013 A fast and robust deblurring technique on high noise environment
abstract
To have a unique solution of an ill-posed inverse problem, the usual way is to embed prior information in terms of regularizer or smoothness criterion. In this work, both the inverse mechanism (the relationship of blur and sharp patches) and the smoothness prior are learned simultaneously from the image itself, in multiple scales. We have shown experimentally that the proposed method outperform the existing state-of-the-art techniques on high noise environment and produce comparable result otherwise; moreover, it is almost three times faster than existing ones.
Pulak Purkait, Bhabatosh Chanda
ICIP1
2012 Image Upscaling Using Multiple Dictionaries of Natural Image Patches
Pulak Purkait, Bhabatosh Chanda
ACCV (3)1
2012 Indian Classical Dance classification by learning dance pose bases
abstract
In this paper, we address an interesting application of computer vision technique, namely classification of Indian Classical Dance (ICD). With the best of our knowledge, the problem has not been addressed so far in computer vision domain. To deal with this problem, we use a sparse representation based dictionary learning technique. First, we represent each frame of a dance video by a pose descriptor based on histogram of oriented optical flow (HOOF), in a hierarchical manner. The pose basis is learned using an on-line dictionary learning technique. Finally each video is represented sparsely as a dance descriptor by pooling pose descriptor of all the frames. In this work, dance videos are classified using support vector machine (SVM) with intersection kernel. Our contribution here are two folds. First, to address dance classification as a new problem in computer vision and second, to present a new action descriptor to represent a dance video which overcomes the problem of the “Bags-of-Words” model. We have tested our algorithm on our own ICD dataset created from the videos collected from YouTube. An accuracy of 86.67% is achieved on this dataset. Since we have proposed a new action descriptor too, we have tested our algorithm on well known KTH dataset. The performance of the system is comparable to the state-of-the-art.
Soumitra Samanta, Pulak Purkait, Bhabatosh Chanda
WACV2
2012 Super Resolution Image Reconstruction Through Bregman Iteration Using Morphologic Regularization
abstract
Multiscale morphological operators are studied extensively in the literature for image processing and feature extraction purposes. In this paper, we model a nonlinear regularization method based on multiscale morphology for edge-preserving super resolution (SR) image reconstruction. We formulate SR image reconstruction as a deblurring problem and then solve the inverse problem using Bregman iterations. The proposed algorithm can suppress inherent noise generated during low-resolution image formation as well as during SR image estimation efficiently. Experimental results show the effectiveness of the proposed regularization and reconstruction method for SR image.
Pulak Purkait, Bhabatosh Chanda
IEEE Trans. Image Process.1
2010 Off-line Recognition of Hand-Written Bengali Numerals Using Morphological Features
abstract
This paper proposes a technique for automatic recognition of Bengali handwritten numerals using multiple feature sets. We discuss about some novel Morphological features and k-curvature feature extraction technique to recognize handwritten scripts. We use different multi-layer perceptron (MLP) classifiers to train this feature spaces and then fuse those classifiers using modified `Naive'-Bayes combination to increase accuracy of recognition result. The individual feature sets give reasonably high accuracy up-to 96.25%, while fused classifier gives accuracy of 97.75%.
Pulak Purkait, Bhabatosh Chanda
ICFHR1
2010 Writer Identification for Handwritten Telugu Documents Using Directional Morphological Features
abstract
Linking a person based on handwritten documents is one of the oldest techniques that is used by crime investigators and forensic scientists. The importance of writer recognition in anthrax letter cases has made this examination popular in recent years. In this paper we propose four feature set namely directional opening, directional closing, direction erosion and k-curvature features for writer recognition on Telugu handwritten documents. Each of the features is extracted from the words after dividing them into a number of cells and then subjected to a nearest neighbor classifier for writer recognition. Although the results of each of the feature set is quite encouraging, the directional opening feature outperforms other feature sets.
Pulak Purkait, Rajesh Kumar 0008, Bhabatosh Chanda
ICFHR1