Soumava Paul

dblp:260/0294 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Transfer learning and domain adaptation · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Transfer learning and domain adaptation › cross-domain transfer
cross-domain retrieval
0.512021
Universal Cross-Domain Retrieval: Generalizing Across Classes and Domains · ICCV 2021
Machine learning › Transfer learning and domain adaptation
generalization to unseen classes
0.512021
Universal Cross-Domain Retrieval: Generalizing Across Classes and Domains · ICCV 2021
Information retrieval
cross-domain retrieval
0.512021
Universal Cross-Domain Retrieval: Generalizing Across Classes and Domains · ICCV 2021
Information retrieval › cross-domain retrieval
universal cross-domain retrieval
0.512021
Universal Cross-Domain Retrieval: Generalizing Across Classes and Domains · ICCV 2021

Methods — techniques the papers use, named apart from their topics

semantic neighborhood loss · 1.0mixup · 1.0mixture prediction loss · 1.0
YearPublicationVenuePosition
2021 Universal Cross-Domain Retrieval: Generalizing Across Classes and Domains
abstract
In this work, for the first time, we address the problem of universal cross-domain retrieval, where the test data can belong to classes or domains which are unseen during training. Due to dynamically increasing number of categories and practical constraint of training on every possible domain, which requires large amounts of data, generalizing to both unseen classes and domains is important. Towards that goal, we propose SnMpNet (Semantic Neighbourhood and Mixture Prediction Network), which incorporates two novel losses to account for the unseen classes and domains encountered during testing. Specifically, we introduce a novel Semantic Neighborhood loss to bridge the knowledge gap between seen and unseen classes and ensure that the latent space embedding of the unseen classes is semantically meaningful with respect to its neighboring classes. We also introduce a mix-up based supervision at image-level as well as semantic-level of the data for training with the Mixture Prediction loss, which helps in efficient retrieval when the query belongs to an unseen domain. These losses are incorporated on the SE-ResNet50 backbone to obtain SnMpNet. Extensive experiments on two large-scale datasets, Sketchy Extended and DomainNet, and thorough comparisons with state-of-the-art justify the effectiveness of the proposed model.
Soumava Paul, Titir Dutta, Soma Biswas
ICCV1
2021 Knowledge Distillation for Singing Voice Detection
abstract
Singing Voice Detection (SVD) has been an active area of research in music information retrieval (MIR). Currently, two deep neural network-based methods, one based on CNN and the other on RNN, exist in literature that learn optimized features for the voice detection (VD) task and achieve state-of-the-art performance on common datasets. Both these models have a huge number of parameters (1.4M for CNN and 65.7K for RNN) and hence not suitable for deployment on devices like smartphones or embedded sensors with limited capacity in terms of memory and computation power. The most popular method to address this issue is known as knowledge distillation in deep learning literature (in addition to model compression) where a large pre-trained network known as the teacher is used to train a smaller student network. Given the wide applications of SVD in music information retrieval, to the best of our knowledge, model compression for practical deployment has not yet been explored. In this paper, efforts have been made to investigate this issue using both conventional as well as ensemble knowledge distillation techniques.
Soumava Paul, Gurunath Reddy M, K. Sreenivasa Rao, Partha Pratim Das 0001
Interspeech1