Ayushi Dalmia

dblp:171/2113 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
1since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Information extraction and text analysis · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › query formulation
natural language querying
0.412020
ATHENA++: Natural Language Querying for Complex Nested SQL Queries · Proc. VLDB Endow. 2020
Natural language and speech › Information extraction and text analysis
semantic parsing
0.412019
Unified Semantic Parsing with Weak Supervision · ACL (1) 2019
Natural language and speech › Information extraction and text analysis › semantic parsing
text-to-SQL
0.112020
ATHENA++: Natural Language Querying for Complex Nested SQL Queries · Proc. VLDB Endow. 2020

Methods — techniques the papers use, named apart from their topics

natural language processing · 0.9weak supervision · 0.4multi-policy distillation · 0.4
YearPublicationVenuePosition
2024 TODM: Train Once Deploy Many Efficient Supernet-Based RNN-T Compression For On-Device ASR Models
abstract
Automatic Speech Recognition (ASR) models need to be optimized for specific hardware before they can be deployed on devices. This can be done by tuning the model’s hyperparameters or exploring variations in its architecture. Re-training and re-validating models after making these changes can be a resource-intensive task. This paper presents TODM (Train Once Deploy Many), a new approach to efficiently train many sizes of hardware-friendly on-device ASR models with comparable GPU-hours to that of a single training job. TODM leverages insights from prior work on Supernet, where Recurrent Neural Network Transducer (RNN-T) models share weights within a Supernet. It reduces layer sizes and widths of the Supernet to obtain subnetworks, making them smaller models suitable for all hardware types. We introduce a novel combination of three techniques to improve the outcomes of the TODM Supernet: adaptive dropout, an in-place Alpha-divergence knowledge distillation, and the use of ScaledAdam optimizer. We validate our approach by comparing Supernet-trained versus individually tuned Multi-Head State Space Model (MH-SSM) RNN-T using LibriSpeech. Results demonstrate that our TODM Supernet either matches or surpasses the performance of manually tuned models by up to a relative of 3% better in word error rate (WER), while efficiently keeping the cost of training many models at a small constant.
Yuan Shangguan, Haichuan Yang, Danni Li, Chunyang Wu, Yassir Fathullah, Dilin Wang, Ayushi Dalmia, Raghuraman Krishnamoorthi, Ozlem Kalinli, Junteng Jia, Jay Mahadeokar, Mike Seltzer, Vikas Chandra
ICASSP7
2020 ATHENA++: Natural Language Querying for Complex Nested SQL Queries
Jaydeep Sen, Chuan Lei, Abdul Quamar, Fatma Özcan 0001, Vasilis Efthymiou, Ayushi Dalmia, Greg Stager, Ashish R. Mittal, Diptikalyan Saha, Karthik Sankaranarayanan
Proc. VLDB Endow.6
2019 Unified Semantic Parsing with Weak Supervision
abstract
Semantic parsing over multiple knowledge bases enables a parser to exploit structural similarities of programs across the multiple domains.However, the fundamental challenge lies in obtaining high-quality annotations of (utterance, program) pairs across various domains needed for training such models.To overcome this, we propose a novel framework to build a unified multi-domain enabled semantic parser trained only with weak supervision (denotations).Weakly supervised training is particularly arduous as the program search space grows exponentially in a multi-domain setting.To solve this, we incorporate a multipolicy distillation mechanism in which we first train domain-specific semantic parsers (teachers) using weak supervision in the absence of the ground truth programs, followed by training a single unified parser (student) from the domain specific policies obtained from these teachers.The resultant semantic parser is not only compact but also generalizes better, and generates more accurate programs.It further does not require the user to provide a domain label while querying.On the standard OVERNIGHT dataset (containing multiple domains), we demonstrate that the proposed model improves performance by 20% in terms of denotation accuracy in comparison to baseline techniques.
Priyanka Agrawal, Ayushi Dalmia, Parag Jain, Abhishek Bansal 0004, Ashish R. Mittal, Karthik Sankaranarayanan
ACL (1)2
2015 Query-based Graph Cuboid Outlier Detection
abstract
Various projections or views of a heterogeneous information network can be modeled using the graph OLAP (On-line Analytical Processing) framework for effective decision making. Detecting anomalous projections of the network can help the analysts identify regions of interest from the graph specific to the projection attribute. While most previous studies on outlier detection in graphs deal with outlier nodes, edges or subgraphs, we are the first to propose detection of graph cuboid outliers. Further we perform this detection in a query sensitive way. Given a general subgraph query on a heterogeneous network, we study the problem of finding outlier cuboids from the graph OLAP lattice. A Graph Cuboid Outlier (GCOutlier) is a cuboid with exceptionally high density of matches for the query. The GCOutlier detection task is clearly challenging because: (1) finding matches for the query (subgraph isomorphism) is NP-hard; (2) number of matches for the query can be very high; and (3) number of cuboids can be large. We provide an approximate solution to the problem by computing only a fraction of the total matches originating from a select set of candidate nodes and including a select set of edges, chosen smartly. We perform extensive experiments on synthetic datasets to showcase the execution time versus accuracy trade-off. Experiments on real datasets like Four Area and Delicious containing thousands of nodes reveal interesting GCOutliers.
Ayushi Dalmia, Manish Gupta 0001, Vasudeva Varma
ASONAM1