Akansha Bhardwaj

dblp:208/4944 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
4since 2021 · last 2023
0000-0002-3267-7265ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2023 Human-in-the-Loop Rule Discovery for Micropost Event Detection
abstract
Platforms such as Twitter are increasingly being used for real-world event detection. Recent work often leverages event-related keywords for training machine learning based event detection models. These approaches make strong assumptions on the distribution of the relevant microposts containing the keyword – referred to as the expectation – and use it as a posterior regularization parameter during model training. Such approaches are, however, limited by the informativeness of the keywords and by the accuracy of the expectation estimation for keywords. In this work, we introduce a human-in-the-loop approach to jointly discover informative rules for model training while estimating their expectation. Our approach iteratively leverages the crowd to estimate both rule-specific expectation and the disagreement between the crowd and the model in order to discover new rules that are most beneficial for model training. To identify such rules, we introduce a hybrid human-machine workflow that engages human workers in rule discovery through an interactive hypothesis creation and testing interface and leverages automatic methods for suggesting useful rules for human verification. We empirically demonstrate the merits of our approach, on multiple real-world datasets and show that our approach improves the state of the art by a margin of 25.63% in terms of AUC.
Akansha Bhardwaj, Jie Yang 0028, Philippe Cudré-Mauroux
IEEE Trans. Knowl. Data Eng.1
2021 MARTA: Leveraging Human Rationales for Explainable Text Classification
abstract
Explainability is a key requirement for text classification in many application domains ranging from sentiment analysis to medical diagnosis or legal reviews. Existing methods often rely on "attention" mechanisms for explaining classification results by estimating the relative importance of input units. However, recent studies have shown that such mechanisms tend to mis-identify irrelevant input units in their explanation. In this work, we propose a hybrid human-AI approach that incorporates human rationales into attention-based text classification models to improve the explainability of classification results. Specifically, we ask workers to provide rationales for their annotation by selecting relevant pieces of text. We introduce MARTA, a Bayesian framework that jointly learns an attention-based model and the reliability of workers while injecting human rationales into model training. We derive a principled optimization algorithm based on variational inference with efficient updating rules for learning MARTA parameters. Extensive validation on real-world datasets shows that our framework significantly improves the state of the art both in terms of classification explainability and accuracy.
Ines Arous, Ljiljana Dolamic, Jie Yang 0028, Akansha Bhardwaj, Giuseppe Cuccu, Philippe Cudré-Mauroux
AAAI4
2021 ConvTab: A Context-Preserving, Convolutional Model for Ad-Hoc Table Retrieval
abstract
Ad-hoc table retrieval, also known as table search, is the problem of finding tables relevant to a search query. This search query can be a keyword or a table itself, referred to as keyword-based and table-based search, respectively. With the vast amounts of tabular data available online, it has become essential for users to identify relevant tables that meet their search criteria. In this regard, there has been a wide variety of research on this problem using pure lexical features, semantic representation, embeddings, as well as intrinsic and extrinsic features of the tables. However, one of the significant limitations of most of the existing methods is that they do not keep the table’s structure and the globalized context intact when building semantic representations of tabular data. Deriving motivation from this fact, we propose an effective approach based on Convolutional Neural Networks (CNNs) – ConvTab – to train the embeddings of tabular data. Our approach is divided into two phases. First, we leverage the discriminating power of CNNs to train a table classifier. Next, the representations learned from this model are used to generate semantic features for query-table similarity. These query-table similarity features are then used as input to the learning algorithm. We evaluate our approach on the table retrieval task using standard NDCG, MAP, and MRR metrics. Experiments reveal that ConvTab significantly outperforms the state of the art in ad-hoc table retrieval by 16.9% and 8.37% using NDCG at cutoffs 5 and 20, respectively. For reproducibility purposes, we share our model as well as all details of our implementation1.
Vibhav Agarwal, Akansha Bhardwaj, Paolo Rosso, Philippe Cudré-Mauroux
IEEE BigData2
2021 Event Detection on Microposts: A Comparison of Four Approaches
abstract
Microblogging services such as Twitter are important, up-to-date, and live sources of information on a multitude of topics and events. An increasing number of systems use such services to detect and analyze events in real-time as they unfold. In this context, we recently proposed ArmaTweet-a system developed in collaboration among armasuisse and the Universities of Oxford and Fribourg to support semantic event detection on Twitter streams. Our experiments have shown that ArmaTweet is successful at detecting many complex events that cannot be detected by simple keyword-based search methods alone. Building up on this work, we explore in this paper several approaches for event detection on microposts. In particular, we describe and compare four different approaches based on keyword search (Plain-Seed-Query), information retrieval (Temporal Query Expansion), Word2Vec word embeddings (Embedding), and semantic retrieval (ArmaTweet). We provide an extensive empirical evaluation of these techniques using a benchmark dataset of about 200 million tweets on six event categories that we collected. While the performance of individual systems varies depending on the event category, our results show that ArmaTweet outperforms the other approaches on five out of six categories, and that a combined approach offers highest recall without adversely affecting precision of event detection.
Akansha Bhardwaj, Albert Blarer, Philippe Cudré-Mauroux, Vincent Lenders, Boris Motik, Axel Tanner, Alberto Tonon
IEEE Trans. Knowl. Data Eng.1
2020 A Human-AI Loop Approach for Joint Keyword Discovery and Expectation Estimation in Micropost Event Detection
abstract
Microblogging platforms such as Twitter are increasingly being used in event detection. Existing approaches mainly use machine learning models and rely on event-related keywords to collect the data for model training. These approaches make strong assumptions on the distribution of the relevant microposts containing the keyword – referred to as the expectation of the distribution – and use it as a posterior regularization parameter during model training. Such approaches are, however, limited as they fail to reliably estimate the informativeness of a keyword and its expectation for model training. This paper introduces a Human-AI loop approach to jointly discover informative keywords for model training while estimating their expectation. Our approach iteratively leverages the crowd to estimate both keyword-specific expectation and the disagreement between the crowd and the model in order to discover new keywords that are most beneficial for model training. These keywords and their expectation not only improve the resulting performance but also make the model training process more transparent. We empirically demonstrate the merits of our approach, both in terms of accuracy and interpretability, on multiple real-world datasets and show that our approach improves the state of the art by 24.3%.
Akansha Bhardwaj, Jie Yang 0028, Philippe Cudré-Mauroux
AAAI1
2020 Hydra: Cancer Detection Leveraging Multiple Heads and Heterogeneous Datasets
abstract
We propose an approach combining layer freezing and fine-tuning steps alternatively to train a neural network over multiple and diverse datasets in the context of cancer detection from medical images. Our method explicitly splits the network into two distinct but complementary components: the feature extractor and the decision maker. While the former remains constant throughout training, a different decision maker is used on each new dataset. This enables end-to-end training of the feature extractor on heterogeneous datasets (here MRIs and CT scans) and organs (here prostate, lung and brain). The feature extractor learns features across all images, with two major benefits: (i) extended training data pool, and (ii) enforced generalization across different data. We show the effectiveness of our method by detecting cancerous masses in the SPIE-AAPM-NCI Prostate MR Classification data. Our training process integrates the SPIE-AAPM-NCI Lung CT Classification dataset as well as the Kaggle Brain MRI dataset, each paired with a separate decision maker, improving the AUC of the base network architecture on the Prostate MR dataset by 0.12 (18% relative increase) versus training on the prostate dataset alone. We also compare against standard end-to-end Transfer Learning over the same datasets for reference, which only improves the results by 0.04 (6% relative increase).
Giuseppe Cuccu, Johan Jobin, Julien Clément 0001, Akansha Bhardwaj, Carolin Reischauer, Harriet Thöny, Philippe Cudré-Mauroux
IEEE BigData4
2018 SentiCite - An Approach for Publication Sentiment Analysis
abstract
With the rapid growth in the number of scientific publications, year after year, it is becoming increasingly difficult to identify quality authoritative work on a single topic. Though there is an availability of scientometric measures which promise to offer a solution to this problem, these measures are mostly quantitative and rely, for instance, only on the number of times an article is cited. With this approach, it becomes irrelevant if an article is cited 10 times in a positive, negative or neutral way. In this context, it is quite important to study the qualitative aspect of a citation to understand its significance. This paper presents a novel system for sentiment analysis of citations in scientific documents (SentiCite) and is also capable of detecting nature of citations by targeting the motivation behind a citation, e.g., reference to a dataset, reading reference. Furthermore, the paper also presents two datasets (SentiCiteDB and IntentCiteDB) containing about 2,600 citations with their ground truth for sentiment and nature of citation. SentiCite along with other state-of-the-art methods for sentiment analysis are evaluated on the presented datasets. Evaluation results reveal that SentiCite outperforms state-of-the-art methods for sentiment analysis in scientific publications by achieving a F1-measure of 0.71.
Dominique Mercier, Akansha Bhardwaj, Andreas Dengel 0001, Sheraz Ahmed
ICAART (2)2
2017 Academic Community Explorer (ACE) for Syntactic, Semantic and Pragmatic Document Analysis
abstract
This paper presents a novel Academic Community Explorer (ACE) which performs syntactic, semantic and pragmatic document analysis of scientific publications. Firstly, ACE uses syntactic structure to extract relevant information from a scientific document. Secondly, semantic analysis is performed to derive an article based co-authorship and citation network. Finally, ACE uses these document based networks to build a complete community network for pragmatic analysis. Furthermore, scientometric analysis is performed to extract the pragmatics by analyzing authors and publication community networks through micro and macro indicators. Two novel micro indicators Senti-Index, reflecting the sentiment present in citations and, Overlap index, reflecting community behavior have been introduced. This is a step in the direction of automatic qualitative assessment of scientific documents. In addition, ACE provides a rich visualization interface which helps in exploratory analysis of the community to identify hidden patterns, e.g, isolated small groups in the community which collaborate and cite each other frequently. A feasibility study is performed on the corpus of ICDAR publications from 1993-2015 to show the insights and benefits of the ACE framework. The results reveals that ICDAR is a highly collaborative community which has most likely arrived at its 'phase transition' stage with 70% of the community closely connected to each other.
Akansha Bhardwaj, Dominique Mercier, Hisham Hashmi, Sheraz Ahmed, Andreas Dengel 0001
ICDAR1
2017 DeepBIBX: Deep Learning for Image Based Bibliographic Data Extraction
Akansha Bhardwaj, Dominique Mercier, Andreas Dengel 0001, Sheraz Ahmed
ICONIP (2)1