Arvind Agarwal

dblp:18/7442 · DBLP profile ↗
← Back
17ranked-venue papers
8as first author
2since 2021 · last 2022
0000-0002-7052-653XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 7 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-authorDatabases, data management, data science and information retrieval · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Question answering and dialogue systems · 24% Information extraction and text analysis · 22% Knowledge representation and reasoning · 15%
Databases, data mining, and information retrieval
3 papers
Data mining · 50% Knowledge graphs · 35% Information retrieval · 15%

Topics — the 20 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Knowledge representation and reasoning › ontology
ontology matching
0.512021
VeeAlign: Multifaceted Context Representation Using Dual Attention for Ontology Alignment · EMNLP (1) 2021
Natural language and speech › Question answering and dialogue systems › response selection
multi-turn response selection
0.412019
A Practical Dialogue-Act-Driven Conversation Model for Multi-Turn Response Selection · EMNLP/IJCNLP (1) 2019
Natural language and speech › Question answering and dialogue systems › dialogue understanding
dialogue act classification
0.312018
Dialogue Act Sequence Labeling Using Hierarchical Encoder With CRF · AAAI 2018
Natural language and speech › Language models and text generation › text generation
paraphrase generation
0.312018
A Deep Generative Framework for Paraphrase Generation · AAAI 2018
Natural language and speech › Information extraction and text analysis
sequence labeling
0.312018
Dialogue Act Sequence Labeling Using Hierarchical Encoder With CRF · AAAI 2018
Natural language and speech › Information extraction and text analysis › topic model
supervised topic model
0.212015
Supervised Topic Models for Microblog Classification · ICDM 2015
Natural language and speech › Information extraction and text analysis
topic model
0.212015
Supervised Topic Models for Microblog Classification · ICDM 2015
Knowledge graphs › ontology
ontology matching
0.112021
VeeAlign: Multifaceted Context Representation Using Dual Attention for Ontology Alignment · EMNLP (1) 2021
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.112011
A Geometric View of Conjugate Priors · IJCAI 2011
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › prior distribution
conjugate priors
0.112011
A Geometric View of Conjugate Priors · IJCAI 2011
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › prior modeling
prior design
0.112011
A Geometric View of Conjugate Priors · IJCAI 2011
Natural language and speech › Question answering and dialogue systems › dialogue modeling
dialogue act modeling
0.112019
A Practical Dialogue-Act-Driven Conversation Model for Multi-Turn Response Selection · EMNLP/IJCNLP (1) 2019
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
manifold learning
0.112010
Learning Multiple Tasks using Manifold Regularization · NIPS 2010
Machine learning › Learning paradigms › semi-supervised learning › graph-based semi-supervised learning
manifold regularization
0.112010
Learning Multiple Tasks using Manifold Regularization · NIPS 2010
Machine learning › Learning paradigms
multi-task learning
0.112010
Learning Multiple Tasks using Manifold Regularization · NIPS 2010
Data mining
dimensionality reduction
0.112010
Universal multi-dimensional scaling · KDD 2010
Data mining › dimensionality reduction
multidimensional scaling
0.112010
Universal multi-dimensional scaling · KDD 2010
Machine learning › Deep learning architectures and training › recurrent neural network › deep recurrent network
hierarchical recurrent neural network
0.112018
Dialogue Act Sequence Labeling Using Hierarchical Encoder With CRF · AAAI 2018
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
exponential family
0.112009
Exponential Family Hybrid Semi-Supervised Learning · IJCAI 2009
Machine learning › Learning paradigms
semi-supervised learning
0.112009
Exponential Family Hybrid Semi-Supervised Learning · IJCAI 2009

Methods — techniques the papers use, named apart from their topics

dual attention · 1.0deep learning · 1.0nonparametric topic model · 0.4LDA · 0.4variational autoencoder · 0.3sequence-to-sequence model · 0.3hierarchical encoder · 0.3conditional random field · 0.3bidirectional LSTM · 0.3LSTM · 0.3iterative algorithm · 0.1convergence guarantees · 0.1
YearPublicationVenuePosition
2022 Toward Scientific Workflows in a Serverless World
abstract
Serverless computing and FaaS have gained popularity due to their ease of design, deployment, scaling and billing on clouds. However, when used to compose and orchestrate scientific workflows, they pose limitations due to cold starts, message indirection, vendor lock-in and lack of provenance support. Here, we propose a design for a Ser verless Scientific Workflow Orchestrator that overcomes these challenges using techniques like function fusion, pilot invocations and data fabrics.
Aakash Khochare, Yogesh L. Simmhan, Sameep Mehta, Arvind Agarwal
e-Science4
2021 VeeAlign: Multifaceted Context Representation Using Dual Attention for Ontology Alignment
abstract
Ontology Alignment is an important research problem applied to various fields such as data integration, data transfer, data preparation, etc. State-of-the-art (SOTA) Ontology Alignment systems typically use naive domain-dependent approaches with handcrafted rules or domainspecific architectures, making them unscalable and inefficient.In this work, we propose VeeAlign, a Deep Learning based model that uses a novel dual-attention mechanism to compute the contextualized representation of a concept which, in turn, is used to discover alignments.By doing this, not only is our approach able to exploit both syntactic and semantic information encoded in ontologies, it is also, by design, flexible and scalable to different domains with minimal effort.We evaluate our model on four different datasets from different domains and languages, and establish its superiority through these results as well as detailed ablation studies.The code and datasets used are available at https://github.com/Remorax/VeeAlign.
Vivek Iyer, Arvind Agarwal
EMNLP (1)2
2020 Budgeted Batch Mode Active Learning with Generalized Cost and Utility Functions
abstract
Active learning reduces the labeling cost by actively querying labels for the most valuable data points. Typical active learning methods select the most informative examples one-at-a-time, their batch variants exist which select a set of most informative points instead of one point at a time. These points are selected in such a way that when added to the training data along with their labels, they provide maximum benefit to the underlying model. In this paper, we present a learning framework that actively selects optimal set of examples (in a batch) within a given budget, based on given utility and cost functions. The framework is generic enough to incorporate any utility and any cost function defined on a set of examples. Furthermore, we propose a novel utility function based on the Facility Location problem that considers three important characteristics of utility i.e., diversity, density and point utility. We also propose a novel cost function, by formulating the cost computation problem as an optimization problem, the solution to which turns out to be the minimum spanning tree. Thus, our framework provides the optimal batch of points within the given budget based on the cost and utility functions. We evaluate our method on several data sets and show its superior performance over baseline methods.
Arvind Agarwal, Shashank Mujumdar, Nitin Gupta 0005, Sameep Mehta
ICPR1
2019 A Practical Dialogue-Act-Driven Conversation Model for Multi-Turn Response Selection
abstract
Harshit Kumar, Arvind Agarwal, Sachindra Joshi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Arvind Agarwal, Sachindra Joshi
EMNLP/IJCNLP (1)2
2018 A Deep Generative Framework for Paraphrase Generation
abstract
Paraphrase generation is an important problem in NLP, especially in question answering, information retrieval, information extraction, conversation systems, to name a few. In this paper, we address the problem of generating paraphrases automatically. Our proposed method is based on a combination of deep generative models (VAE) with sequence-to-sequence models (LSTM) to generate paraphrases, given an input sentence. Traditional VAEs when combined with recurrent neural networks can generate free text but they are not suitable for paraphrase generation for a given sentence. We address this problem by conditioning the both, encoder and decoder sides of VAE, on the original sentence, so that it can generate the given sentence's paraphrases. Unlike most existing models, our model is simple, modular and can generate multiple paraphrases, for a given sentence. Quantitative evaluation of the proposed method on a benchmark paraphrase dataset demonstrates its efficacy, and its performance improvement over the state-of-the-art methods by a significant margin, whereas qualitative human evaluation indicate that the generated paraphrases are well-formed, grammatically correct, and are relevant to the input sentence. Furthermore, we evaluate our method on a newly released question paraphrase dataset, and establish a new baseline for future research.
Arvind Agarwal, Prawaan Singh, Piyush Rai
AAAI2
2018 Dialogue Act Sequence Labeling Using Hierarchical Encoder With CRF
abstract
Dialogue Act recognition associate dialogue acts (i.e., semantic labels) to utterances in a conversation. The problem of associating semantic labels to utterances can be treated as a sequence labeling problem. In this work, we build a hierarchical recurrent neural network using bidirectional LSTM as a base unit and the conditional random field (CRF) as the top layer to classify each utterance into its corresponding dialogue act. The hierarchical network learns representations at multiple levels, i.e., word level, utterance level, and conversation level. The conversation level representations are input to the CRF layer, which takes into account not only all previous utterances but also their dialogue acts, thus modeling the dependency among both, labels and utterances, an important consideration of natural dialogue. We validate our approach on two different benchmark data sets, Switchboard and Meeting Recorder Dialogue Act, and show performance improvement over the state-of-the-art methods by 2.2% and 4.1% absolute points, respectively. It is worth noting that the inter-annotator agreement on Switchboard data set is 84%, and our method is able to achieve the accuracy of about 79% despite being trained on the noisy data.
Arvind Agarwal, Riddhiman Dasgupta, Sachindra Joshi
AAAI2
2018 Dialogue-act-driven Conversation Model : An Experimental Study
abstract
The utility of additional semantic information for the task of next utterance selection in an automated dialogue system is the focus of study in this paper. In particular, we show that additional information available in the form of dialogue acts –when used along with context given in the form of dialogue history– improves the performance irrespective of the underlying model being generative or discriminative. In order to show the model agnostic behavior of dialogue acts, we experiment with several well-known models such as sequence-to-sequence encoder-decoder model, hierarchical encoder-decoder model, and Siamese-based models with and without hierarchy; and show that in all models, incorporating dialogue acts improves the performance by a significant margin. We, furthermore, propose a novel way of encoding dialogue act information, and use it along with hierarchical encoder to build a model that can use the sequential dialogue act information in a natural way. Our proposed model achieves an MRR of about 84.8% for the task of next utterance selection on a newly introduced Daily Dialogue dataset, and outperform the baseline models. We also provide a detailed analysis of results including key insights that explain the improvement in MRR because of dialog act information.
Arvind Agarwal, Sachindra Joshi
COLING2
2016 Monsoon cloud observations using a low power vertical looking Ka Band Cloud radar
abstract
A low power zenith looking Ka Band dual polarization Cloud Radar is designed and fabricated by Society for Applied Microwave Electronics Engineering and research (SAMEER) to study the clouds during monsoon. This radar is installed on the terrace of the SAMEER building and conducted test operations in single polarisation mode during monsoon period(June to September, 2013-2015). The system is observed clouds upto 10 km during the observation period. Sample observations in different cloudy conditions are presented in this study.
Arvind Agarwal, J. S. Pillai, K. Aurobindo, J. D. Abhyankar, Giri Isola, Poornima Srivastava, Abhishek Kodilkar
IGARSS1
2016 A preliminary analysis of cloud classification results using Ka-band polarimetric radar signatures
abstract
Surface based millimeter wave radar systems play a substantial role in remote sensing of clouds. A preliminary analysis of the results obtained from the algorithm developed for cloud classification is presented. Our aim is to classify different cloud types (drizzling, precipitating, Mixed Phase, Ice clouds and non-meteorological targets ) solely based on Ka-band radar data. A fuzzy logic technique is adopted and is applied to few cases which infer useful information for cloud studies. Sample datasets from the Ka-band scanning Polarimetric radar (KASPR) of Indian Institute of Tropical Meteorology (IITM), Pune are used for the initial testing of the said algorithm. The developed algorithm would be implemented in Ka-band cloud profiling radar being indigenously developed by SAMEER.
Abhishek Kodilkar, Arvind Agarwal, Kalapureddy MCR, J. S. Pillai
IGARSS2
2015 Supervised Topic Models for Microblog Classification
abstract
In this paper we present a topic model based approach for classifying micro-blog posts into a given topics of interests. The short nature of micro-blog posts make them challenging for directly learning a classification model. To overcome this limitation, we use content of the links embedded in these posts to improve the topic learning. The hypothesis is that since the link content is far richer than the content of the post itself, using link content along with the content of the post will help learning. However, how this link content can be used to construct features for classification remains a challenging issue. Furthermore, in previous methods, user based information is utilized in an ad-hoc manner that only work for certain type of classification, such as characterizing content of microblogs. In this paper, we propose supervised topic model, User-Labeled-LDA and its nonparametric variant that can avoid the ad-hoc feature construction task and model the topics in a discriminative way. Our experiments on a Twitter dataset shows that modeling user interests and link information helps in learning quality topics for sparse tweets as well as helps significantly in classification task. Our experiments further show that modeling this information in a principled way through topic models helps more than simply adding this information through features.
Saurabh Kataria 0003, Arvind Agarwal
ICDM2
2014 WS2F: A weakly supervised framework for data stream filtering
abstract
In this paper we present a weakly supervised framework for relevant content filtering from social media platforms such as Twitter. Social media platforms are a rich source of information these days. However of all the available information, there is only a small fraction of which is of general interest. Most of the other information pertains to personal events, and is very specific to the users who are contributing that. It is therefore usually not of general interest. In this paper, we present a framework to filter out the topic-specific relevant information from the irrelevant information in the stream of text provided by social media platforms. Our framework does not depend on any labeled data, however it is capable of using domain knowledge in the form of rules and guidelines provided by domain experts. It is therefore easily extensible for new topics and events. The proposed framework is built keeping the streaming nature of social media platforms in mind, i.e., it is able to discover the content relevant to a specific event as it evolves in the text stream. Because of its adaptive nature, it is not only able to filter the relevant content, but also able to generate event story lines as the event evolves. We experiment on a dataset provided by TREC, and show that the framework not only filters relevant content for an event but also generates its story line effectively.
Cailing Dong 0002, Arvind Agarwal
IEEE BigData2
2012 Learning to rank for robust question answering
abstract
This paper aims to solve the problem of improving the ranking of answer candidates for factoid based questions in a state-of-the-art Question Answering system. We first provide an extensive comparison of 5 ranking algorithms on two datasets -- from the Jeopardy quiz show and a medical domain. We then show the effectiveness of a cascading approach, where the ranking produced by one ranker is used as input to the next stage. The cascading approach shows sizeable gains on both datasets. We finally evaluate several rank aggregation techniques to combine these algorithms, and find that Supervised Kemeny aggregation is a robust technique that always beats the baseline ranking approach used by Watson for the Jeopardy competition. We further corroborate our results on TREC Question Answering datasets.
Arvind Agarwal, Hema Raghavan, Karthik Subbian, Prem Melville, Richard D. Lawrence, David Gondek, James Fan
CIKM1
2011 A Geometric View of Conjugate Priors
Arvind Agarwal, Hal Daumé III
IJCAI1
2010 Universal multi-dimensional scaling
abstract
In this paper, we propose a unified algorithmic framework for solving many known variants of MDS. Our algorithm is a simple iterative scheme with guaranteed convergence, and is modular; by changing the internals of a single subroutine in the algorithm, we can switch cost functions and target spaces easily. In addition to the formal guarantees of convergence, our algorithms are accurate; in most cases, they converge to better quality solutions than existing methods in comparable time. Moreover, they have a small memory footprint and scale effectively for large data sets. We expect that this framework will be useful for a number of MDS variants that have not yet been studied.
Arvind Agarwal, Jeff M. Phillips, Suresh Venkatasubramanian
KDD1
2010 Learning Multiple Tasks using Manifold Regularization
abstract
We present a novel method for multitask learning (MTL) based on {\it manifold regularization}: assume that all task parameters lie on a manifold. This is the generalization of a common assumption made in the existing literature: task parameters share a common {\it linear} subspace. One proposed method uses the projection distance from the manifold to regularize the task parameters. The manifold structure and the task parameters are learned using an alternating optimization framework. When the manifold structure is fixed, our method decomposes across tasks which can be learnt independently. An approximation of the manifold regularization scheme is presented that preserves the convexity of the single task learning problem, and makes the proposed MTL framework efficient and easy to implement. We show the efficacy of our method on several datasets.
Arvind Agarwal, Hal Daumé III, Samuel Gerber
NIPS1
2010 A geometric view of conjugate priors
Arvind Agarwal, Hal Daumé III
Mach. Learn.1
2009 Exponential Family Hybrid Semi-Supervised Learning
Arvind Agarwal, Hal Daumé III
IJCAI1