EDBT 2026 Demo / reviewers in the wild / expert
Avik Ray
dblp:56/11157
· DBLP profile ↗
14ranked-venue papers
10as first author
3since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 5 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-author · 1 since 2021Systems, architecture and hardware · 3 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 first-authorComputer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Question answering and dialogue systems · 45% Probabilistic and Bayesian machine learning · 17% Speech recognition and synthesis · 12% | |
| Network and information security
1 paper |
Hardware security and side channels · 100% | |
| Computer networks
1 paper |
Cellular and mobile networks · 50% Wireless sensing and localization · 50% | |
| Databases, data mining, and information retrieval
2 papers |
Data mining · 62% Web and social media mining · 38% | |
| Theoretical computer science
1 paper |
Graph algorithms and graph theory · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Electronic design automation · 100% |
Topics — the 20 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware security and side channels
hardware security verification |
0.8 | 1 | 2024 | NSPG: Natural language Processing-based Security Property Generator for Hardware Security Assurance · DAC 2024 |
Hardware security and side channels › hardware security verification
security property generation |
0.8 | 1 | 2024 | NSPG: Natural language Processing-based Security Property Generator for Hardware Security Assurance · DAC 2024 |
Natural language and speech › Question answering and dialogue systems
intent detection |
0.5 | 1 | 2021 | Enhancing the generalization for Intent Classification and Out-of-Domain Detection in SLU · ACL/IJCNLP (1) 2021 |
Natural language and speech › Question answering and dialogue systems › intent detection
out-of-domain detection |
0.5 | 1 | 2021 | Enhancing the generalization for Intent Classification and Out-of-Domain Detection in SLU · ACL/IJCNLP (1) 2021 |
Natural language and speech › Speech recognition and synthesis
spoken language understanding |
0.5 | 1 | 2021 | Enhancing the generalization for Intent Classification and Out-of-Domain Detection in SLU · ACL/IJCNLP (1) 2021 |
Natural language and speech › Question answering and dialogue systems › dialogue generation
dialogue response generation |
0.4 | 1 | 2020 | Generating Dialogue Responses from a Semantic Latent Space · EMNLP (1) 2020 |
Natural language and speech › Question answering and dialogue systems › dialogue generation
open-domain dialogue generation |
0.4 | 1 | 2020 | Generating Dialogue Responses from a Semantic Latent Space · EMNLP (1) 2020 |
Machine learning › Generative modeling › latent space model
semantic latent space |
0.4 | 1 | 2020 | Generating Dialogue Responses from a Semantic Latent Space · EMNLP (1) 2020 |
Natural language and speech › Language models and text generation › text generation
paraphrase generation |
0.3 | 1 | 2018 | Learning Out-of-Vocabulary Words in Intelligent Personal Agents · IJCAI 2018 |
Natural language and speech › Information extraction and text analysis
semantic parsing |
0.3 | 1 | 2018 | Learning Out-of-Vocabulary Words in Intelligent Personal Agents · IJCAI 2018 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
mixture model |
0.3 | 1 | 2017 | The Search Problem in Mixture Models · J. Mach. Learn. Res. 2017 |
Data mining
clustering |
0.2 | 1 | 2016 | Searching For A Single Community in a Graph · SIGMETRICS 2016 |
Cellular and mobile networks
cellular network analytics |
0.2 | 1 | 2016 | Localization of LTE measurement records with missing information · INFOCOM 2016 |
Wireless sensing and localization › localization algorithms
measurement record localization |
0.2 | 1 | 2016 | Localization of LTE measurement records with missing information · INFOCOM 2016 |
Graph algorithms and graph theory › graph clustering
community detection |
0.2 | 1 | 2016 | Searching For A Single Community in a Graph · SIGMETRICS 2016 |
Electronic design automation
hardware verification and test |
0.2 | 1 | 2024 | NSPG: Natural language Processing-based Security Property Generator for Hardware Security Assurance · DAC 2024 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
0.2 | 1 | 2015 | Improved Greedy Algorithms for Learning Graphical Models · IEEE Trans. Inf. Theory 2015 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
structure learning |
0.2 | 1 | 2015 | Improved Greedy Algorithms for Learning Graphical Models · IEEE Trans. Inf. Theory 2015 |
Web and social media mining
social network analysis |
0.2 | 1 | 2014 | Topic modeling from network spread · SIGMETRICS 2014 |
Data mining › text mining
topic modeling |
0.1 | 1 | 2014 | Topic modeling from network spread · SIGMETRICS 2014 |
Methods — techniques the papers use, named apart from their topics
natural language processing · 1.5BERT-based language model · 1.5method of moments · 0.7regression on latent space · 0.4human evaluation · 0.4sequence-to-sequence model · 0.3neural network · 0.3machine learning · 0.2greedy algorithm · 0.2convex optimization · 0.2spectral methods · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | NSPG: Natural language Processing-based Security Property Generator for Hardware Security AssuranceabstractThe efficiency of validating complex System-on-Chips (SoCs) is contingent on the quality of the security properties provided. Generating security properties with traditional approaches often requires expert intervention and is limited to a few IPs, thereby resulting in a time-consuming and non-robust process. To address this issue, we, for the first time, propose a novel and automated Natural Language Processing (NLP)-based Security Property Generator (NSPG). Specifically, our approach utilizes hardware documentation in order to propose the first hardware security-specific language model, HS-BERT, for extracting security properties dedicated to hardware design. It is capable of phasing a significant amount of hardware specification, and the generated security properties can be easily converted into hardware assertions, thereby reducing the manual effort required for hardware verification. NSPG is trained using sentences from several SoC documentations and achieves up to 88% accuracy for property classification, outperforming ChatGPT. When assessed on five untrained OpenTitan hardware IP documents, NSPG aided in identifying eight security vulnerabilities in the buggy OpenTitan SoC presented in Hack@DAC 2022. Amisha Srivastava, Ayush Arunachalam, Avik Ray, Pedro Henrique Silva, Rafail Psiakis, Yiorgos Makris, Kanad Basu |
DAC | 4 |
| 2023 | Compositional Generalization in Spoken Language Understanding
Avik Ray, Yilin Shen, Hongxia Jin |
INTERSPEECH | 1 |
| 2021 | Enhancing the generalization for Intent Classification and Out-of-Domain Detection in SLUabstractYilin Shen, Yen-Chang Hsu, Avik Ray, Hongxia Jin. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yilin Shen, Yen-Chang Hsu, Avik Ray, Hongxia Jin |
ACL/IJCNLP (1) | 3 |
| 2020 | Generating Dialogue Responses from a Semantic Latent SpaceabstractExisting open-domain dialogue generation models are usually trained to mimic the gold response in the training set using cross-entropy loss on the vocabulary.However, a good response does not need to resemble the gold response, since there are multiple possible responses to a given prompt.In this work, we hypothesize that the current models are unable to integrate information from multiple semantically similar valid responses of a prompt, resulting in the generation of generic and uninformative responses.To address this issue, we propose an alternative to the end-to-end classification on vocabulary.We learn the pair relationship between the prompts and responses as a regression task on a latent space instead.In our novel dialog generation model, the representations of semantically related sentences are close to each other on the latent space.Human evaluation showed that learning the task on a continuous space can generate responses that are both relevant and informative. Wei-Jen Ko, Avik Ray, Yilin Shen, Hongxia Jin |
EMNLP (1) | 2 |
| 2019 | Iterative Delexicalization for Improved Spoken Language UnderstandingabstractRecurrent neural network (RNN) based joint intent classification and slot tagging models have achieved tremendous success in recent years for building spoken language understanding and dialog systems.However, these models suffer from poor performance for slots which often encounter large semantic variability in slot values after deployment (e.g.message texts, partial movie/artist names).While greedy delexicalization of slots in the input utterance via substring matching can partly improve performance, it often produces incorrect input.Moreover, such techniques cannot delexicalize slots with out-of-vocabulary slot values not seen at training.In this paper, we propose a novel iterative delexicalization algorithm, which can accurately delexicalize the input, even with out-of-vocabulary slot values.Based on model confidence of the current delexicalized input, our algorithm improves delexicalization in every iteration to converge to the best input having the highest confidence.We show on benchmark and in-house datasets that our algorithm can greatly improve parsing performance for RNN based models, especially for out-of-distribution slot values. Avik Ray, Yilin Shen, Hongxia Jin |
INTERSPEECH | 1 |
| 2018 | Learning Out-of-Vocabulary Words in Intelligent Personal AgentsabstractSemantic parsers play a vital role in intelligent agents to convert natural language instructions to an actionable logical form representation. However, after deployment, these parsers suffer from poor accuracy on encountering out-of-vocabulary (OOV) words, or significant accuracy drop on previously supported instructions after retraining. Achieving both goals simultaneously is non-trivial. In this paper, we propose novel neural networks based parsers to learn OOV words; one incorporating a new hybrid paraphrase generation model, and an enhanced sequence-to-sequence model. Extensive experiments on both benchmark and custom datasets show our new parsers achieve significant accuracy gain on OOV words and phrases, and in the meanwhile learn OOV words while maintaining accuracy on previously supported instructions. Avik Ray, Yilin Shen, Hongxia Jin |
IJCAI | 1 |
| 2018 | Robust Spoken Language Understanding via ParaphrasingabstractLearning intents and slot labels from user utterances is a fundamental step in all spoken language understanding (SLU) and dialog systems.State-of-the-art neural network based methods, after deployment, often suffer from performance degradation on encountering paraphrased utterances, and out-of-vocabulary words, rarely observed in their training set.We address this challenging problem by introducing a novel paraphrasing based SLU model which can be integrated with any existing SLU model in order to improve their overall performance.We propose two new paraphrase generators using RNN and sequence-to-sequence based neural networks, which are suitable for our application.Our experiments on existing benchmark and in house datasets demonstrate the robustness of our models to rare and complex paraphrased utterances, even under adversarial test distributions. Avik Ray, Yilin Shen, Hongxia Jin |
INTERSPEECH | 1 |
| 2018 | Interactive recommendation via deep neural memory augmented contextual banditsabstractPersonalized recommendation with user interactions has become increasingly popular nowadays in many applications with dynamic change of contents (news, media, etc.). Existing approaches model user interactive recommendation as a contextual bandit problem to balance the trade-off between exploration and exploitation. However, these solutions require a large number of interactions with each user to provide high quality personalized recommendations. To mitigate this limitation, we design a novel deep neural memory augmented mechanism to model and track the history state for each user based on his previous interactions. As such, the user's preferences on new items can be quickly learned within a small number of interactions. Moreover, we develop new algorithms to leverage large amount of all users' history data for offline model training and online model fine tuning for each user with the focus of policy evaluation. Extensive experiments on different synthetic and real-world datasets validate that our proposed approach consistently outperforms a variety of state-of-the-art approaches. Yilin Shen, Yue Deng 0001, Avik Ray, Hongxia Jin |
RecSys | 3 |
| 2017 | The Search Problem in Mixture Models
Avik Ray, Joe Neeman, Sujay Sanghavi, Sanjay Shakkottai |
J. Mach. Learn. Res. | 1 |
| 2016 | Localization of LTE measurement records with missing informationabstractAs cellular networks like 4G LTE networks get more and more sophisticated, mobiles also measure and send enormous amount of mobile measurement data (in TBs/week/metropolitan) during every call and session. The mobile measurement records are saved in data center for further analysis and mining, however, these measurement records are not geo-tagged because the measurement procedures are implemented in mobile LTE stack. Geo-tagging (or localizing) the stored measurement record is a fundamental building block towards network analytics and troubleshooting since the measurement records contain rich information on call quality, latency, throughput, signal quality, error codes etc. In this work, our goal is to localize these mobile measurement records. Precisely, we answer the following question: what was the location of the mobile when it sent a given measurement record? We design and implement novel machine learning based algorithms to infer whether a mobile was outdoor and if so, it infers the latitude-longitude associated with the measurement record. The key technical challenge comes from the fact that measurement records do not contain sufficient information required for triangulation or RF fingerprinting based techniques to work by themselves. Experiments performed with real data sets from an operational 4G network in a major metropolitan show that, the median accuracy of our proposed solution is around 20 m for outdoor mobiles and outdoor classification accuracy is more than 98%. Avik Ray, Supratim Deb, Pantelis Monogioudis |
INFOCOM | 1 |
| 2016 | Searching For A Single Community in a GraphabstractIn standard graph clustering/community detection, one is interested in partitioning the graph into more densely connected subsets of nodes. In contrast, the search problem of this paper aims to only find the nodes in a single such community, the target, out of the many communities that may exist. To do so , we are given suitable side information about the target; for example, a very small number of nodes from the target are labeled as such. We consider a general yet simple notion of side information: all nodes are assumed to have random weights, with nodes in the target having higher weights on average. Given these weights and the graph, we develop a variant of the method of moments that identifies nodes in the target more reliably, and with lower computation, than generic community detection methods that do not use side information and partition the entire graph. Avik Ray, Sujay Sanghavi, Sanjay Shakkottai |
SIGMETRICS | 1 |
| 2015 | Improved Greedy Algorithms for Learning Graphical ModelsabstractWe propose new greedy algorithms for learning the structure of a graphical model of a probability distribution, given samples drawn from the distribution. While structure learning of graphical models is a widely studied problem with several existing methods, greedy approaches remain attractive due to their low computational cost. The most natural greedy algorithm would be one which, essentially, adds neighbors to a node in sequence until stopping; it would do this for each node. While it is fast, simple and parallel, this naive greedy algorithm has the tendency to add non-neighbors that show high correlations with the given node. Our new algorithms overcome this problem in three different ways. The recursive greedy algorithm iteratively recovers the neighbors by running the greedy algorithm in an inner loop, but each time only adding the last added node to the neighborhood set. The second forward-backward greedy algorithm includes a node deletion step in each iteration that allows non-neighbors to be removed from the neighborhood set which may have been added in the previous steps. Finally, the greedy algorithm with pruning runs the greedy algorithm until completion and then removes all the incorrect neighbors. We provide both analytical guarantees and empirical performance for our algorithms. We show that in graphical models with strong non-neighbor interactions, our greedy algorithms can correctly recover the graph, whereas the previous greedy and convex optimization-based algorithms do not succeed. Avik Ray, Sujay Sanghavi, Sanjay Shakkottai |
IEEE Trans. Inf. Theory | 1 |
| 2014 | Topic modeling from network spreadabstractTopic modeling refers to the task of inferring, only from data, the abstract ``topics" that occur in a collection of content. In this paper we look at latent topic modeling in a setting where unlike traditional topic modeling (a) there are no/few features (like words in documents) that are directly indicative of content topics (e.g. un-annotated videos and images, URLs etc.), but (b) users share and view content over a social network. We provide a new algorithm for inferring both the topics in which every user is interested, and thus also the topics in each content piece. We study its theoretical performance and demonstrate its empirical effectiveness over standard topic modeling algorithms. Avik Ray, Sujay Sanghavi, Sanjay Shakkottai |
SIGMETRICS | 1 |
| 2009 | Ideal structure of the Silver codeabstractThe Silver code has captured a lot of attention in the recent past, because of its nice structure and fast decodability. In their recent paper, Hollanti et al. show that the Silver code forms a subset of the natural order of a particular cyclic division algebra (CDA). In this paper, the algebraic structure of this subset is characterized. It is shown that the Silver code is not an ideal in the natural order but a right ideal generated by two elements in a particular order of this CDA. The exact minimum determinant of the normalized Silver code is computed using the ideal structure of the code. The construction of Silver code is then extended to CDAs over other number fields. Avik Ray, Ghaya Rekaya-Ben Othman, P. Vijay Kumar, K. Vinodh |
ISIT | 1 |