Avik Ray

dblp:56/11157 · DBLP profile ↗
← Back
14ranked-venue papers
10as first author
3since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 5 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-author · 1 since 2021Systems, architecture and hardware · 3 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 first-authorComputer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Question answering and dialogue systems · 45% Probabilistic and Bayesian machine learning · 17% Speech recognition and synthesis · 12%
Network and information security
1 paper
Hardware security and side channels · 100%
Computer networks
1 paper
Cellular and mobile networks · 50% Wireless sensing and localization · 50%
Databases, data mining, and information retrieval
2 papers
Data mining · 62% Web and social media mining · 38%
Theoretical computer science
1 paper
Graph algorithms and graph theory · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Electronic design automation · 100%

Topics — the 20 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware security and side channels
hardware security verification
0.812024
NSPG: Natural language Processing-based Security Property Generator for Hardware Security Assurance · DAC 2024
Hardware security and side channels › hardware security verification
security property generation
0.812024
NSPG: Natural language Processing-based Security Property Generator for Hardware Security Assurance · DAC 2024
Natural language and speech › Question answering and dialogue systems
intent detection
0.512021
Enhancing the generalization for Intent Classification and Out-of-Domain Detection in SLU · ACL/IJCNLP (1) 2021
Natural language and speech › Question answering and dialogue systems › intent detection
out-of-domain detection
0.512021
Enhancing the generalization for Intent Classification and Out-of-Domain Detection in SLU · ACL/IJCNLP (1) 2021
Natural language and speech › Speech recognition and synthesis
spoken language understanding
0.512021
Enhancing the generalization for Intent Classification and Out-of-Domain Detection in SLU · ACL/IJCNLP (1) 2021
Natural language and speech › Question answering and dialogue systems › dialogue generation
dialogue response generation
0.412020
Generating Dialogue Responses from a Semantic Latent Space · EMNLP (1) 2020
Natural language and speech › Question answering and dialogue systems › dialogue generation
open-domain dialogue generation
0.412020
Generating Dialogue Responses from a Semantic Latent Space · EMNLP (1) 2020
Machine learning › Generative modeling › latent space model
semantic latent space
0.412020
Generating Dialogue Responses from a Semantic Latent Space · EMNLP (1) 2020
Natural language and speech › Language models and text generation › text generation
paraphrase generation
0.312018
Learning Out-of-Vocabulary Words in Intelligent Personal Agents · IJCAI 2018
Natural language and speech › Information extraction and text analysis
semantic parsing
0.312018
Learning Out-of-Vocabulary Words in Intelligent Personal Agents · IJCAI 2018
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
mixture model
0.312017
The Search Problem in Mixture Models · J. Mach. Learn. Res. 2017
Data mining
clustering
0.212016
Searching For A Single Community in a Graph · SIGMETRICS 2016
Cellular and mobile networks
cellular network analytics
0.212016
Localization of LTE measurement records with missing information · INFOCOM 2016
Wireless sensing and localization › localization algorithms
measurement record localization
0.212016
Localization of LTE measurement records with missing information · INFOCOM 2016
Graph algorithms and graph theory › graph clustering
community detection
0.212016
Searching For A Single Community in a Graph · SIGMETRICS 2016
Electronic design automation
hardware verification and test
0.212024
NSPG: Natural language Processing-based Security Property Generator for Hardware Security Assurance · DAC 2024
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models
0.212015
Improved Greedy Algorithms for Learning Graphical Models · IEEE Trans. Inf. Theory 2015
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
structure learning
0.212015
Improved Greedy Algorithms for Learning Graphical Models · IEEE Trans. Inf. Theory 2015
Web and social media mining
social network analysis
0.212014
Topic modeling from network spread · SIGMETRICS 2014
Data mining › text mining
topic modeling
0.112014
Topic modeling from network spread · SIGMETRICS 2014

Methods — techniques the papers use, named apart from their topics

natural language processing · 1.5BERT-based language model · 1.5method of moments · 0.7regression on latent space · 0.4human evaluation · 0.4sequence-to-sequence model · 0.3neural network · 0.3machine learning · 0.2greedy algorithm · 0.2convex optimization · 0.2spectral methods · 0.2
YearPublicationVenuePosition
2024 NSPG: Natural language Processing-based Security Property Generator for Hardware Security Assurance
abstract
The efficiency of validating complex System-on-Chips (SoCs) is contingent on the quality of the security properties provided. Generating security properties with traditional approaches often requires expert intervention and is limited to a few IPs, thereby resulting in a time-consuming and non-robust process. To address this issue, we, for the first time, propose a novel and automated Natural Language Processing (NLP)-based Security Property Generator (NSPG). Specifically, our approach utilizes hardware documentation in order to propose the first hardware security-specific language model, HS-BERT, for extracting security properties dedicated to hardware design. It is capable of phasing a significant amount of hardware specification, and the generated security properties can be easily converted into hardware assertions, thereby reducing the manual effort required for hardware verification. NSPG is trained using sentences from several SoC documentations and achieves up to 88% accuracy for property classification, outperforming ChatGPT. When assessed on five untrained OpenTitan hardware IP documents, NSPG aided in identifying eight security vulnerabilities in the buggy OpenTitan SoC presented in Hack@DAC 2022.
Amisha Srivastava, Ayush Arunachalam, Avik Ray, Pedro Henrique Silva, Rafail Psiakis, Yiorgos Makris, Kanad Basu
DAC4
2023 Compositional Generalization in Spoken Language Understanding
Avik Ray, Yilin Shen, Hongxia Jin
INTERSPEECH1
2021 Enhancing the generalization for Intent Classification and Out-of-Domain Detection in SLU
abstract
Yilin Shen, Yen-Chang Hsu, Avik Ray, Hongxia Jin. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yilin Shen, Yen-Chang Hsu, Avik Ray, Hongxia Jin
ACL/IJCNLP (1)3
2020 Generating Dialogue Responses from a Semantic Latent Space
abstract
Existing open-domain dialogue generation models are usually trained to mimic the gold response in the training set using cross-entropy loss on the vocabulary.However, a good response does not need to resemble the gold response, since there are multiple possible responses to a given prompt.In this work, we hypothesize that the current models are unable to integrate information from multiple semantically similar valid responses of a prompt, resulting in the generation of generic and uninformative responses.To address this issue, we propose an alternative to the end-to-end classification on vocabulary.We learn the pair relationship between the prompts and responses as a regression task on a latent space instead.In our novel dialog generation model, the representations of semantically related sentences are close to each other on the latent space.Human evaluation showed that learning the task on a continuous space can generate responses that are both relevant and informative.
Wei-Jen Ko, Avik Ray, Yilin Shen, Hongxia Jin
EMNLP (1)2
2019 Iterative Delexicalization for Improved Spoken Language Understanding
abstract
Recurrent neural network (RNN) based joint intent classification and slot tagging models have achieved tremendous success in recent years for building spoken language understanding and dialog systems.However, these models suffer from poor performance for slots which often encounter large semantic variability in slot values after deployment (e.g.message texts, partial movie/artist names).While greedy delexicalization of slots in the input utterance via substring matching can partly improve performance, it often produces incorrect input.Moreover, such techniques cannot delexicalize slots with out-of-vocabulary slot values not seen at training.In this paper, we propose a novel iterative delexicalization algorithm, which can accurately delexicalize the input, even with out-of-vocabulary slot values.Based on model confidence of the current delexicalized input, our algorithm improves delexicalization in every iteration to converge to the best input having the highest confidence.We show on benchmark and in-house datasets that our algorithm can greatly improve parsing performance for RNN based models, especially for out-of-distribution slot values.
Avik Ray, Yilin Shen, Hongxia Jin
INTERSPEECH1
2018 Learning Out-of-Vocabulary Words in Intelligent Personal Agents
abstract
Semantic parsers play a vital role in intelligent agents to convert natural language instructions to an actionable logical form representation. However, after deployment, these parsers suffer from poor accuracy on encountering out-of-vocabulary (OOV) words, or significant accuracy drop on previously supported instructions after retraining. Achieving both goals simultaneously is non-trivial. In this paper, we propose novel neural networks based parsers to learn OOV words; one incorporating a new hybrid paraphrase generation model, and an enhanced sequence-to-sequence model. Extensive experiments on both benchmark and custom datasets show our new parsers achieve significant accuracy gain on OOV words and phrases, and in the meanwhile learn OOV words while maintaining accuracy on previously supported instructions.
Avik Ray, Yilin Shen, Hongxia Jin
IJCAI1
2018 Robust Spoken Language Understanding via Paraphrasing
abstract
Learning intents and slot labels from user utterances is a fundamental step in all spoken language understanding (SLU) and dialog systems.State-of-the-art neural network based methods, after deployment, often suffer from performance degradation on encountering paraphrased utterances, and out-of-vocabulary words, rarely observed in their training set.We address this challenging problem by introducing a novel paraphrasing based SLU model which can be integrated with any existing SLU model in order to improve their overall performance.We propose two new paraphrase generators using RNN and sequence-to-sequence based neural networks, which are suitable for our application.Our experiments on existing benchmark and in house datasets demonstrate the robustness of our models to rare and complex paraphrased utterances, even under adversarial test distributions.
Avik Ray, Yilin Shen, Hongxia Jin
INTERSPEECH1
2018 Interactive recommendation via deep neural memory augmented contextual bandits
abstract
Personalized recommendation with user interactions has become increasingly popular nowadays in many applications with dynamic change of contents (news, media, etc.). Existing approaches model user interactive recommendation as a contextual bandit problem to balance the trade-off between exploration and exploitation. However, these solutions require a large number of interactions with each user to provide high quality personalized recommendations. To mitigate this limitation, we design a novel deep neural memory augmented mechanism to model and track the history state for each user based on his previous interactions. As such, the user's preferences on new items can be quickly learned within a small number of interactions. Moreover, we develop new algorithms to leverage large amount of all users' history data for offline model training and online model fine tuning for each user with the focus of policy evaluation. Extensive experiments on different synthetic and real-world datasets validate that our proposed approach consistently outperforms a variety of state-of-the-art approaches.
Yilin Shen, Yue Deng 0001, Avik Ray, Hongxia Jin
RecSys3
2017 The Search Problem in Mixture Models
Avik Ray, Joe Neeman, Sujay Sanghavi, Sanjay Shakkottai
J. Mach. Learn. Res.1
2016 Localization of LTE measurement records with missing information
abstract
As cellular networks like 4G LTE networks get more and more sophisticated, mobiles also measure and send enormous amount of mobile measurement data (in TBs/week/metropolitan) during every call and session. The mobile measurement records are saved in data center for further analysis and mining, however, these measurement records are not geo-tagged because the measurement procedures are implemented in mobile LTE stack. Geo-tagging (or localizing) the stored measurement record is a fundamental building block towards network analytics and troubleshooting since the measurement records contain rich information on call quality, latency, throughput, signal quality, error codes etc. In this work, our goal is to localize these mobile measurement records. Precisely, we answer the following question: what was the location of the mobile when it sent a given measurement record? We design and implement novel machine learning based algorithms to infer whether a mobile was outdoor and if so, it infers the latitude-longitude associated with the measurement record. The key technical challenge comes from the fact that measurement records do not contain sufficient information required for triangulation or RF fingerprinting based techniques to work by themselves. Experiments performed with real data sets from an operational 4G network in a major metropolitan show that, the median accuracy of our proposed solution is around 20 m for outdoor mobiles and outdoor classification accuracy is more than 98%.
Avik Ray, Supratim Deb, Pantelis Monogioudis
INFOCOM1
2016 Searching For A Single Community in a Graph
abstract
In standard graph clustering/community detection, one is interested in partitioning the graph into more densely connected subsets of nodes. In contrast, the search problem of this paper aims to only find the nodes in a single such community, the target, out of the many communities that may exist. To do so , we are given suitable side information about the target; for example, a very small number of nodes from the target are labeled as such. We consider a general yet simple notion of side information: all nodes are assumed to have random weights, with nodes in the target having higher weights on average. Given these weights and the graph, we develop a variant of the method of moments that identifies nodes in the target more reliably, and with lower computation, than generic community detection methods that do not use side information and partition the entire graph.
Avik Ray, Sujay Sanghavi, Sanjay Shakkottai
SIGMETRICS1
2015 Improved Greedy Algorithms for Learning Graphical Models
abstract
We propose new greedy algorithms for learning the structure of a graphical model of a probability distribution, given samples drawn from the distribution. While structure learning of graphical models is a widely studied problem with several existing methods, greedy approaches remain attractive due to their low computational cost. The most natural greedy algorithm would be one which, essentially, adds neighbors to a node in sequence until stopping; it would do this for each node. While it is fast, simple and parallel, this naive greedy algorithm has the tendency to add non-neighbors that show high correlations with the given node. Our new algorithms overcome this problem in three different ways. The recursive greedy algorithm iteratively recovers the neighbors by running the greedy algorithm in an inner loop, but each time only adding the last added node to the neighborhood set. The second forward-backward greedy algorithm includes a node deletion step in each iteration that allows non-neighbors to be removed from the neighborhood set which may have been added in the previous steps. Finally, the greedy algorithm with pruning runs the greedy algorithm until completion and then removes all the incorrect neighbors. We provide both analytical guarantees and empirical performance for our algorithms. We show that in graphical models with strong non-neighbor interactions, our greedy algorithms can correctly recover the graph, whereas the previous greedy and convex optimization-based algorithms do not succeed.
Avik Ray, Sujay Sanghavi, Sanjay Shakkottai
IEEE Trans. Inf. Theory1
2014 Topic modeling from network spread
abstract
Topic modeling refers to the task of inferring, only from data, the abstract ``topics" that occur in a collection of content. In this paper we look at latent topic modeling in a setting where unlike traditional topic modeling (a) there are no/few features (like words in documents) that are directly indicative of content topics (e.g. un-annotated videos and images, URLs etc.), but (b) users share and view content over a social network. We provide a new algorithm for inferring both the topics in which every user is interested, and thus also the topics in each content piece. We study its theoretical performance and demonstrate its empirical effectiveness over standard topic modeling algorithms.
Avik Ray, Sujay Sanghavi, Sanjay Shakkottai
SIGMETRICS1
2009 Ideal structure of the Silver code
abstract
The Silver code has captured a lot of attention in the recent past, because of its nice structure and fast decodability. In their recent paper, Hollanti et al. show that the Silver code forms a subset of the natural order of a particular cyclic division algebra (CDA). In this paper, the algebraic structure of this subset is characterized. It is shown that the Silver code is not an ideal in the natural order but a right ideal generated by two elements in a particular order of this CDA. The exact minimum determinant of the normalized Silver code is computed using the ideal structure of the code. The construction of Silver code is then extended to CDAs over other number fields.
Avik Ray, Ghaya Rekaya-Ben Othman, P. Vijay Kumar, K. Vinodh
ISIT1