EDBT 2026 Demo / reviewers in the wild / expert
Yash Verma
dblp:262/6430
· DBLP profile ↗
8ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computing education · 100% | |
| Artificial intelligence
1 paper |
Representation and self-supervised learning · 100% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computing education › AI education
AI literacy |
0.9 | 1 | 2025 | Word2Vec4Kids: Interactive Challenges to Introduce Middle School Students to Word Embeddings · AAAI 2025 |
Computing education › AI education
NLP education |
0.9 | 1 | 2025 | Word2Vec4Kids: Interactive Challenges to Introduce Middle School Students to Word Embeddings · AAAI 2025 |
Machine learning › Representation and self-supervised learning › word representation › word embedding
word2vec |
0.3 | 1 | 2025 | Word2Vec4Kids: Interactive Challenges to Introduce Middle School Students to Word Embeddings · AAAI 2025 |
Machine learning › Representation and self-supervised learning › word representation
word embedding |
0.3 | 1 | 2025 | Word2Vec4Kids: Interactive Challenges to Introduce Middle School Students to Word Embeddings · AAAI 2025 |
Methods — techniques the papers use, named apart from their topics
interactive application · 1.7game-based learning · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enabling Hardware-Software Co-Design for Large-Scale AI Systems via Fine-Grained GPU Memory TracingabstractMemory traces capturing GPU operations over modern scale-up fabrics (e.g., load/store accesses, ordering events, synchronization, and communication behavior) enable a broad set of research directions for next-generation AI infrastructure. Palak Mishra, Rajat Bhardwaj, Yash Verma, Ramanjeet Singh, Rinku Shah |
SIGCOMM | 3 |
| 2025 | Word2Vec4Kids: Interactive Challenges to Introduce Middle School Students to Word EmbeddingsabstractAs Artificial Intelligence (AI) continues to integrate into more aspects of society, equipping younger generations with foundational AI knowledge becomes increasingly critical. This paper presents Word2Vec4Kids (W2V4K), an interactive application designed to familiarize middle school students with word embeddings, a key aspect of Natural Language Processing (NLP). W2V4K leverages the Word2Vec model, allowing students to explore word associations, similarity, and vector arithmetic through engaging game modes. The application was tested with 38 middle school students aged 11-14 at a Science Technology Engineering Math (STEM)-focused charter school. Data were collected on students' interactions with the application, including screen recordings, audio, and survey responses. Results demonstrated that W2V4K effectively introduces NLP concepts to students. Qualitative observations revealed high levels of engagement with students expressing excitement and curiosity about word relationships. As they progressed through the game modes, students showed increasing confidence in predicting word associations, brainstorming relevant words, and connecting the concepts to real-world applications. Quantitative data from post-interaction surveys indicated positive learning outcomes with 44.5% of students achieving perfect scores on concept-related items. Additionally, students demonstrated an ability to critically think about language representation. This study suggests that W2V4K provides an effective and engaging method for introducing NLP concepts to middle school students, contributing to the broader goal of enhancing AI literacy among younger generations. Nathan Wiatrek, Yash Verma, Fred G. Martin |
AAAI | 2 |
| 2025 | Adopting Whisper for Confidence EstimationabstractRecent research on word-level confidence estimation for speech recognition systems has primarily focused on lightweight models known as Confidence Estimation Modules (CEMs), which rely on hand-engineered features derived from Automatic Speech Recognition (ASR) outputs. In contrast, we propose a novel end-to-end approach that leverages the ASR model itself (Whisper) to generate word-level confidence scores. Specifically, we introduce a method in which the Whisper model is fine-tuned to produce scalar confidence scores given an audio input and its corresponding hypothesis transcript. Our experiments demonstrate that the fine-tuned Whisper-tiny model, comparable in size to a strong CEM baseline, achieves similar performance on the in-domain dataset and surpasses the CEM baseline on eight out-of-domain datasets, whereas the fine-tuned Whisper-large model consistently outperforms the CEM baseline by a substantial margin across all the datasets. Vaibhav Aggarwal, Shabari S. Nair, Yash Verma, Yash Jogi |
ICASSP | 3 |
| 2024 | Improving Rare-Word Recognition of Whisper in Zero-Shot SettingsabstractWhisper, despite being trained on 680 K hours of web-scaled audio data, faces difficulty in recognizing rare words like domain-specific terms, with a solution being contextual biasing through prompting. To improve upon this method, in this paper, we propose a supervised learning strategy to fine-tune Whisper for contextual biasing instruction. We demonstrate that by using only 670 hours of Common Voice English set for fine-tuning, our model generalizes to 11 diverse opensource English datasets, achieving a 45.6% improvement in recognition of rare words and 60.8% improvement in recognition of words unseen during fine-tuning over the baseline method. Surprisingly, our model’s contextual biasing ability generalizes even to languages unseen during fine-tuning. Yash Jogi, Vaibhav Aggarwal, Shabari S. Nair, Yash Verma, Aayush Kubba |
SLT | 4 |
| 2023 | Large Scale Multi-Lingual Multi-Modal Summarization DatasetabstractSignificant developments in techniques such as encoder-decoder models have enabled us to represent information comprising multiple modalities.This information can further enhance many downstream tasks in the field of information retrieval and natural language processing; however, improvements in multi-modal techniques and their performance evaluation require large-scale multi-modal data which offers sufficient diversity.Multi-lingual modeling for a variety of tasks like multi-modal summarization, text generation, and translation leverages information derived from highquality multi-lingual annotated data.In this work, we present the current largest multilingual multi-modal summarization dataset (M3LS), and it consists of over a million instances of document-image pairs along with a professionally annotated multi-modal summary for each pair.It is derived from news articles published by British Broadcasting Corporation(BBC) over a decade and spans 20 languages, targeting diversity across five language roots, it is also the largest summarization dataset for 13 languages and consists of cross-lingual summarization data for 2 languages.We formally define the multi-lingual multi-modal summarization task utilizing our dataset and report baseline scores from various state-of-the-art summarization techniques in a multi-lingual setting.We also compare it with many similar datasets to analyze the uniqueness and difficulty of M3LS. Yash Verma, Anubhav Jangra, Raghvendra Verma, Sriparna Saha 0001 |
EACL | 1 |
| 2023 | Semi-Supervised Range-Based Anomaly Detection for Cloud SystemsabstractThe inherent characteristics of cloud systems often lead to anomalies, which pose challenges for high availability, reliability, and high performance. Detecting anomalies in cloud key performance indicators (KPI) is a critical step towards building a secure and trustworthy system with early mitigation features. This work is motivated by (i) the efficacy of recent reconstruction-based anomaly detection (AD), (ii) the misrepresentation of the accuracy of time series anomaly detection because point-basedPrecisionandRecallare used to evaluate the efficacy for range-based anomalies, and (iii) detects performance and security anomalies when distributions shift and overlaps. In this paper, we propose a novel semi-supervised dynamic density-based detection rule that uses the reconstruction error vectors in order to detect anomalies. We use long short-term memory networks based on encoder-decoder (LSTM-ED) architecture to reconstruct the normal KPI time series. We experiment with both testbed and a diverse set of real-world datasets. The experimental results show that the dynamic density approach exhibits better performance compared to other detection rules using both standard and range-based evaluation metrics. We also compare the performance of our approach with state-of-the-art methods, outperforms in detecting both performance and security anomalies. Pratyush Kr. Deka, Yash Verma, Adil Bin Bhutto, Erik Elmroth, Monowar Bhuyan |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2022 | MAKED: Multi-lingual Automatic Keyword Extraction DatasetabstractKeyword extraction is an integral task for many downstream problems like clustering, recommendation, search and classification. Development and evaluation of keyword extraction techniques require an exhaustive dataset; however, currently, the community lacks large-scale multi-lingual datasets. In this paper, we present MAKED, a large-scale multi-lingual keyword extraction dataset comprising of 540K+ news articles from British Broadcasting Corporation News (BBC News) spanning 20 languages. It is the first keyword extraction dataset for 11 of these 20 languages. The quality of the dataset is examined by experimentation with several baselines. We believe that the proposed dataset will help advance the field of automatic keyword extraction given its size, diversity in terms of languages used, topics covered and time periods as well as its focus on under-studied languages. Yash Verma, Anubhav Jangra, Sriparna Saha 0001, Adam Jatowt, Dwaipayan Roy 0001 |
LREC | 1 |
| 2020 | SentiInc: Incorporating Sentiment Information into Sentiment Transfer Without Parallel Data
Kartikey Pant, Yash Verma, Radhika Mamidi |
ECIR (2) | 2 |