Amit Awekar

dblp:136/7882 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
3since 2021 · last 2024
0000-0003-1559-8247ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 9 · 3 since 2021Artificial intelligence and machine learning · 6 · 3 since 2021
YearPublicationVenuePosition
2024 Compressed Models are NOT Miniature Versions of Large Models
abstract
Large neural models are often compressed before deployment. Model compression is necessary for many practical reasons, such as inference latency, memory footprint, and energy consumption. Compressed models are assumed to be miniature versions of corresponding large neural models. However, we question this belief in our work. We compare compressed models with corresponding large neural models using four model characteristics: prediction errors, data representation, data distribution, and vulnerability to adversarial attack. We perform experiments using the BERT-large model and its five compressed versions. For all four model characteristics, compressed models significantly differ from the BERT-large model. Even among compressed models, they differ from each other on all four model characteristics. Apart from the expected loss in model performance, there are major side effects of using compressed models to replace large neural models.
Rohit Raj Rai, Rishant Pal, Amit Awekar
CIKM3
2022 Scaling Up Mass-Based Clustering
abstract
This paper addresses the problem of scaling up the mass-based clustering paradigm to handle large datasets. The existing algorithm MBScan computes and stores all pairwise distances, resulting in quadratic time and space complexity. However, we observe that mass-based clustering requires information about only a tiny fraction of all possible data point pairs. We propose three optimizations to MBScan for quickly finding such pairs and computing their distances. We empirically evaluate our work on ten real-world and synthetic datasets. Our experiments show that our approach results in fast and memory-efficient clustering with no loss in the quality of clusters.
Nidhi Ahlawat, Amit Awekar
CIKM2
2022 Improving Relation Classification Using Relation Hierarchy
Akshay Parekh, Ashish Anand, Amit Awekar
NLDB3
2019 It's only Words and Words Are All I Have
Manash Pratim Barman, Kavish Dahekar, Abhinav Anshuman, Amit Awekar
ECIR (2)4
2019 Mining Strengths and Weaknesses of Cricket Players Using Short Text Commentary
abstract
Knowledge of strengths and weaknesses of players is the key for team selection and strategy planning in any team sport such as Cricket. Computationally, this problem is mostly unexplored. Existing methods focus only on aggregate and macroscopic statistics that ignore many details. The central idea of our paper is to mine strength and weakness rules using short text commentary data. This dataset is compact, semi-structured, accurate, and yet ignored by the machine learning community until now. We collect fine-grained information about each player from the short text commentary dataset and represent it using domain-specific features identified by us. We employ a dimensionality reduction method specific to discrete random variable case, namely correspondence analysis and construct semantic relation between bowler and batsman. This relation is plotted using biplots. Human readable strength and weakness rules are extracted from the biplots. We have performed experiments using a large dataset that describes over one million deliveries. We validate our extracted rules using both intrinsic and extrinsic validation.
Swarup Ranjan Behera, Parag Agrawal, Amit Awekar, Vijaya Saradhi Vedula
ICMLA3
2019 Decoding The Style And Bias of Song Lyrics
abstract
The central idea of this paper is to gain a deeper understanding of song lyrics computationally. We focus on two aspects: style and biases of song lyrics. All prior works to understand these two aspects are limited to manual analysis of a small corpus of song lyrics. In contrast, we analyzed more than half a million songs spread over five decades. We characterize the lyrics style in terms of vocabulary, length, repetitiveness, speed, and readability. We have observed that the style of popular songs significantly differs from other songs. We have used distributed representation methods and WEAT test to measure various gender and racial biases in the song lyrics. We have observed that biases in song lyrics correlate with prior results on human subjects. This correlation indicates that song lyrics reflect the biases that exist in society. Increasing consumption of music and the effect of lyrics on human emotions makes this analysis important.
Manash Pratim Barman, Amit Awekar, Sambhav Kothari
SIGIR2
2018 Deep Learning for Detecting Cyberbullying Across Multiple Social Media Platforms
Sweta Agrawal, Amit Awekar
ECIR2
2017 Fine-Grained Entity Type Classification by Jointly Learning Representations and Label Embeddings
abstract
Fine-grained entity type classification (FETC) is the task of classifying an entity mention to a broad set of types.Distant supervision paradigm is extensively used to generate training data for this task.However, generated training data assigns same set of labels to every mention of an entity without considering its local context.Existing FETC systems have two major drawbacks: assuming training data to be noise free and use of hand crafted features.Our work overcomes both drawbacks.We propose a neural network model that jointly learns entity mentions and their context representation to eliminate use of hand crafted features.Our model treats training data as noisy and uses non-parametric variant of hinge loss function.Experiments show that the proposed model outperforms previous stateof-the-art methods on two publicly available datasets, namely FIGER(GOLD) and BBN with an average relative improvement of 2.69% in micro-F1 score.Knowledge learnt by our model on one dataset can be transferred to other datasets while using same model or other FETC systems.These approaches of transferring knowledge further improve the performance of respective models.
Ashish Anand, Amit Awekar
EACL (1)3
2017 Batch Incremental Shared Nearest Neighbor Density Based Clustering Algorithm for Dynamic Datasets
Panthadeep Bhattacharjee, Amit Awekar
ECIR2
2017 Faster K-Means Cluster Estimation
Siddhesh Khandelwal, Amit Awekar
ECIR2
2013 Incremental shared nearest neighbor density-based clustering
abstract
Shared Nearest Neighbor Density-based clustering (SNN-DBSCAN) is a robust graph-based clustering algorithm and has wide applications from climate data analysis to network intrusion detection. We propose an incremental extension to this algorithm IncSNN-DBSCAN, capable of finding clusters on a dataset to which frequent inserts are made. For each data point, the algorithm maintains four properties: nearest neighbor list, strengths of shared links, total connection strength and topic property. Algorithm only targets points that undergo change to their properties. We prove that, to obtain the exact clustering it is sufficient to re-compute properties for only the targeted points, followed by possible cluster mergers on newly formed links and cluster splits on the deleted links.
Sumeet Singh, Amit Awekar
CIKM2