Sujoy Kumar Biswas

dblp:36/9241 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
0since 2021 · last 2017
0000-0003-0942-9199ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-authorArtificial intelligence and machine learning · 3 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Image recognition and object detection · 37% Video understanding and tracking · 26% Deep learning architectures and training · 20%
Computer graphics and multimedia
1 paper
Image and video processing · 100%
Theoretical computer science
1 paper
Algorithms and data structures · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection
pedestrian detection
0.312017
Linear Support Tensor Machine With LSK Channels: Pedestrian Detection in Thermal Infrared Images · IEEE Trans. Image Process. 2017
Machine learning › Deep learning architectures and training › tensor learning
support tensor machine
0.312017
Linear Support Tensor Machine With LSK Channels: Pedestrian Detection in Thermal Infrared Images · IEEE Trans. Image Process. 2017
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
locality preserving projection
0.212016
One Shot Detection with Laplacian Object and Fast Matrix Cosine Similarity · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Computer vision › Image recognition and object detection › object detection › few-shot object detection
one-shot object detection
0.212016
One Shot Detection with Laplacian Object and Fast Matrix Cosine Similarity · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Computer vision › Video understanding and tracking › action recognition
human action recognition
0.112011
Recognizing interaction between human performers using 'key pose doublet' · ACM Multimedia 2011
Computer vision › Video understanding and tracking › action recognition › human action recognition
human interaction recognition
0.112011
Recognizing interaction between human performers using 'key pose doublet' · ACM Multimedia 2011
Computer vision › Video understanding and tracking › action recognition
skeleton-based action recognition
0.112011
Recognizing interaction between human performers using 'key pose doublet' · ACM Multimedia 2011
Image and video processing › feature extraction › feature descriptor
local feature descriptor
0.112017
Linear Support Tensor Machine With LSK Channels: Pedestrian Detection in Thermal Infrared Images · IEEE Trans. Image Process. 2017
Algorithms and data structures
similarity search
0.112016
One Shot Detection with Laplacian Object and Fast Matrix Cosine Similarity · IEEE Trans. Pattern Anal. Mach. Intell. 2016

Methods — techniques the papers use, named apart from their topics

support vector machine · 0.6discrete fourier transform · 0.6laplacian eigenmaps · 0.5integral image · 0.5fourier transform · 0.5sliding-window detection · 0.3sliding window detection · 0.3visual words · 0.1key pose doublet · 0.1graph centrality · 0.1
YearPublicationVenuePosition
2017 Linear Support Tensor Machine With LSK Channels: Pedestrian Detection in Thermal Infrared Images
abstract
Pedestrian detection in thermal infrared images poses unique challenges because of the low resolution and noisy nature of the image. Here, we propose a mid-level attribute in the form of the multidimensional template, or tensor, using local steering kernel (LSK) as low-level descriptors for detecting pedestrians in far infrared images. LSK is specifically designed to deal with intrinsic image noise and pixel level uncertainty by capturing local image geometry succinctly instead of collecting local orientation statistics (e.g., histograms in histogram of oriented gradients). In order to learn the LSK tensor, we introduce a new image similarity kernel following the popular maximum margin framework of support vector machines facilitating a relatively short and simple training phase for building a rigid pedestrian detector. Tensor representation has several advantages, and indeed, LSK templates allow exact acceleration of the sluggish but de facto sliding window-based detection methodology with multichannel discrete Fourier transform, facilitating very fast and efficient pedestrian localization. The experimental studies on publicly available thermal infrared images justify our proposals and model assumptions. In addition, the proposed work also involves the release of our in-house annotations of pedestrians in more than 17 000 frames of OSU color thermal database for the purpose of sharing with the research community.
Sujoy Kumar Biswas, Peyman Milanfar
IEEE Trans. Image Process.1
2016 One Shot Detection with Laplacian Object and Fast Matrix Cosine Similarity
abstract
One shot, generic object detection involves searching for a single query object in a larger target image. Relevant approaches have benefited from features that typically model the local similarity patterns. In this paper, we combine local similarity (encoded by local descriptors) with a global context (i.e., a graph structure) of pairwise affinities among the local descriptors, embedding the query descriptors into a low dimensional but discriminatory subspace. Unlike principal components that preserve global structure of feature space, we actually seek a linear approximation to the Laplacian eigenmap that permits us a locality preserving embedding of high dimensional region descriptors. Our second contribution is an accelerated but exact computation of matrix cosine similarity as the decision rule for detection, obviating the computationally expensive sliding window search. We leverage the power of Fourier transform combined with integral image to achieve superior runtime efficiency that allows us to test multiple hypotheses (for pose estimation) within a reasonably short time. Our approach to one shot detection is training-free, and experiments on the standard data sets confirm the efficacy of our model. Besides, low computation cost of the proposed (codebook-free) object detector facilitates rather straightforward query detection in large data sets including movie videos.
Sujoy Kumar Biswas, Peyman Milanfar
IEEE Trans. Pattern Anal. Mach. Intell.1
2014 Laplacian object: One-shot object detection by locality preserving projection
abstract
One shot, generic object detection involves detecting a single query image in a target image. Relevant approaches have benefitted from features that typically model the local similarity patterns. Also important is the global matching of local features along the object detection process. In this paper, we consider such global information early in the feature extraction stage by combining local geodesic structure (encoded by LARK descriptors) with a global context (i.e., graph structure) of pairwise affinities among the local descriptors. The result is an embedding of the LARK descriptors (extracted from query image) into a discriminatory subspace (obtained using locality preserving projection [1]) that preserves the local intrinsic geometry of the query image patterns. Experiments on standard data sets demonstrate efficacy of our proposed approach.
Sujoy Kumar Biswas, Peyman Milanfar
ICIP1
2014 Recognizing interactions between human performers by 'Dominating Pose Doublet'
Snehasis Mukherjee, Sujoy Kumar Biswas, Dipti Prasad Mukherjee
Mach. Vis. Appl.2
2011 Recognizing interaction between human performers using 'key pose doublet'
abstract
In this paper, we propose a graph theoretic approach for recognizing interactions between two human performers present in a video clip. We watch primarily the human poses of each performer and derive descriptors that capture the motion patterns of the poses. From an initial dictionary of poses (visual words), we extract key poses (or key words) by ranking the poses on the centrality measure of graph connectivity. We argue that the key poses are graph nodes which share a close semantic relationship (in terms of some suitable edge weight function) with all other pose nodes and hence are said to be the central part of the graph. We apply the same centrality measure on all possible combinations of the key poses of the two performers to select the set of 'key pose doublets' that best represent the corresponding action. The results on standard interaction recognition dataset show the robustness of our approach when compared to the present state of the art method.
Snehasis Mukherjee, Sujoy Kumar Biswas, Dipti Prasad Mukherjee
ACM Multimedia2
2011 Recognizing Human Action at a Distance in Video by Key Poses
abstract
In this paper, we propose a graph theoretic technique for recognizing human actions at a distance in a video by modeling the visual senses associated with poses. The proposed methodology follows a bag-of-word approach that starts with a large vocabulary of poses (visual words) and derives a refined and compact codebook of key poses using centrality measure of graph connectivity. We introduce a “meaningful” threshold on centrality measure that selects key poses for each action type. Our contribution includes a novel pose descriptor based on histogram of oriented optical flow evaluated in a hierarchical fashion on a video frame. This pose descriptor combines both pose information and motion pattern of the human performer into a multidimensional feature vector. We evaluate our methodology on four standard activity-recognition datasets demonstrating the superiority of our method over the state-of-the-art.
Snehasis Mukherjee, Sujoy Kumar Biswas, Dipti Prasad Mukherjee
IEEE Trans. Circuits Syst. Video Technol.2
2010 Modeling Sense Disambiguation of Human Pose: Recognizing Action at a Distance by Key Poses
Snehasis Mukherjee, Sujoy Kumar Biswas, Dipti Prasad Mukherjee
ACCV (1)2