EDBT 2026 Demo / reviewers in the wild / expert
Sushil Khyalia
dblp:284/0708
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2026
0009-0007-5688-3315ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Deep learning architectures and training · 37% Video understanding and tracking · 30% Language models and text generation · 16% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 50% Data integration and cleaning · 50% |
Topics — the 10 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking › multimodal video understanding
audio-visual video understanding |
1.0 | 1 | 2026 | MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX · AAAI 2026 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
1.0 | 1 | 2026 | MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX · AAAI 2026 |
Computer vision › Video understanding and tracking
video question answering |
1.0 | 1 | 2026 | MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX · AAAI 2026 |
Machine learning › Deep learning architectures and training
transformer |
0.8 | 1 | 2024 | Transformers Get Stable: An End-to-End Signal Propagation Theory for Language Models · ICML 2024 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.5 | 1 | 2021 | STING: Self-attention based Time-series Imputation Networks using GAN · ICDM 2021 |
Machine learning › Deep learning architectures and training › attention mechanism
self-attention |
0.5 | 1 | 2021 | STING: Self-attention based Time-series Imputation Networks using GAN · ICDM 2021 |
Data integration and cleaning › missing data
missing value imputation |
0.5 | 1 | 2021 | STING: Self-attention based Time-series Imputation Networks using GAN · ICDM 2021 |
Data mining
time series analysis |
0.5 | 1 | 2021 | STING: Self-attention based Time-series Imputation Networks using GAN · ICDM 2021 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.3 | 1 | 2026 | MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX · AAAI 2026 |
Machine learning › Generative modeling
generative adversarial network |
0.1 | 1 | 2021 | STING: Self-attention based Time-series Imputation Networks using GAN · ICDM 2021 |
Methods — techniques the papers use, named apart from their topics
self-attention · 1.0multimodal benchmark · 1.0human baseline · 1.0GAN · 1.0signal propagation theory · 0.8initialization scheme · 0.8bidirectional recurrent neural networks · 0.5bidirectional recurrent neural network · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeXabstractWe introduce MAVERIX (Multimodal Audio-Visual Evaluation and Recognition IndeX), a unified benchmark to probe video understanding in multimodal LLMs, encompassing video, audio, and text inputs with human performance baselines. Although recent advancements in audiovisual models have shown substantial progress, the field lacks a standardized evaluation framework to thoroughly assess their cross-modality comprehension performance. MAVERIX curates 2,556 questions from 700 videos, in the form of both multiple-choice and open-ended formats, explicitly designed to evaluate multimodal models through questions that necessitate tight integration of video and audio information, spanning a broad spectrum of agentic scenarios. MAVERIX uniquely provides models with questions that closely mimic the multimodal understanding experiences available to humans during decision-making processes. To our knowledge, MAVERIX is the first benchmark aimed explicitly at assessing comprehensive audiovisual integration in such granularity. Experiments with state-of-the-art models, including Qwen 2.5 Omni and Gemini 2.5 Flash-Lite, show performance around 64% accuracy, while human experts reach near-ceiling performance of 92.8%, exposing a substantial gap to human-level comprehension. With standardized evaluation protocols, a rigorously annotated pipeline, and a public toolkit, MAVERIX establishes a challenging testbed for advancing audiovisual multimodal intelligence, with the website publicly available below. Liuyue Xie, Avik Kuthiala, George Z. Wei, Ananya Bal, Mosam Dabhi, Liting Wen, Taru Rustagi, Ethan Lai, Sushil Khyalia, Rohan Choudhury, Morteza Ziyadi, László A. Jeni |
AAAI | 10 |
| 2024 | Data Driven Grapheme-to-Phoneme Representations for a Lexicon-Free Text-to-SpeechabstractGrapheme-to-Phoneme (G2P) is an essential first step in any modern, high-quality Text-to-Speech (TTS) system. Most of the current G2P systems rely on carefully hand-crafted lexicons developed by experts. This poses a two-fold problem. Firstly, the lexicons are generated using a fixed phoneme set, usually, ARPABET or IPA, which might not be the most optimal way to represent phonemes for all languages. Secondly, the man-hours required to produce such an expert lexicon are very high. In this paper, we eliminate both of these issues by using recent advances in self-supervised learning to obtain data-driven phoneme representations instead of fixed representations. We compare our lexicon-free approach against strong baselines that utilize a well-crafted lexicon. Furthermore, we show that our data-driven lexicon-free method performs as good or even marginally better than the conventional rule-based or lexicon-based neural G2Ps in terms of Mean Opinion Score (MOS) while using no prior language lexicon or phoneme set, i.e. no linguistic expertise. Abhinav Garg, Jiyeon Kim, Sushil Khyalia, Chanwoo Kim 0001, Dhananjaya Gowda |
ICASSP | 3 |
| 2024 | Transformers Get Stable: An End-to-End Signal Propagation Theory for Language ModelsabstractIn spite of their huge success, transformer models remain difficult to scale in depth. In this work, we develop a unified signal propagation theory and provide formulae that govern the moments of the forward and backward signal through the transformer model. Our framework can be used to understand and mitigate vanishing/exploding gradients, rank collapse, and instability associated with high attention scores. We also propose DeepScaleLM, an initialization and scaling scheme that conserves unit output/gradient moments throughout the model, enabling the training of very deep models with 1000 layers. We find that transformer models could be much deeper - our deep models with fewer parameters outperform shallow models in Language Modeling, Speech Translation, and Image Classification, across encoder-only, decoder-only and encoder-decoder variants, for both Pre-LN and Post-LN transformers, for multiple datasets and model sizes. These improvements also translate into improved performance on downstream Question Answering tasks and improved robustness for Image Classification. Akhil Kedia, Mohd Abbas Zaidi, Sushil Khyalia, Jungho Jung, Harshith Goka, Haejun Lee |
ICML | 3 |
| 2022 | PAC Mode Estimation using PPR Martingale Confidence SequencesabstractWe consider the problem of correctly identifying the mode of a discrete distribution $\mathcal{P}$ with sufficiently high probability by observing a sequence of i.i.d. samples drawn from $\mathcal{P}$. This problem reduces to the estimation of a single parameter when $\mathcal{P}$ has a support set of size $K = 2$. After noting that this special case is handled very well by prior-posterior-ratio (PPR) martingale confidence sequences (Waudby-Smith and Ramdas, 2020), we propose a generalisation to mode estimation, in which $\mathcal{P}$ may take $K \geq 2$ values. To begin, we show that the "one-versus-one" principle to generalise from $K = 2$ to $K \geq 2$ classes is more efficient than the "one-versus-rest" alternative. We then prove that our resulting stopping rule, denoted PPR-1v1, is asymptotically optimal (as the mistake probability is taken to 0). PPR-1v1 is simple and computationally light, and incurs significantly fewer samples than competitors even in the non-asymptotic regime. We demonstrate its gains in two practical applications of sampling: election forecasting and verification of smart contracts in blockchains. Shubham Anand Jain, Rohan Shah, Sanit Gupta, Denil Mehta, Inderjeet J. Nair, Jian Vora, Sushil Khyalia, Sourav Das 0001, Vinay J. Ribeiro, Shivaram Kalyanakrishnan |
AISTATS | 7 |
| 2021 | Meta-Learning for Effective Multi-task and Multilingual ModellingabstractIshan Tarunesh, Sushil Khyalia, Vishwajeet Kumar, Ganesh Ramakrishnan, Preethi Jyothi. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Ishan Tarunesh, Sushil Khyalia, Vishwajeet Kumar, Ganesh Ramakrishnan, Preethi Jyothi |
EACL | 2 |
| 2021 | STING: Self-attention based Time-series Imputation Networks using GANabstractTime series data are ubiquitous in real-world applications. However, one of the most common problems is that the time series could have missing values by the inherent nature of the data collection process. So imputing missing values from multivariate (correlated) time series is imperative to improve a prediction performance while making an accurate data-driven decision. Conventional works for imputation simply delete missing values or fill them based on mean/zero. Although recent works based on deep neural networks have shown remarkable results, they still have a limitation to capture the complex generation process of multivariate time series. In this paper, we propose a novel imputation method for multivariate time series, called STING (Self-attention based Time-series Imputation Networks using GAN). We take advantage of generative adversarial networks and bidirectional recurrent neural networks to learn the latent representations of time series. In addition, we introduce a novel attention mechanism to capture the weighted correlations of a whole sequence and avoid the potential bias brought by unrelated ones. The experimental results on three real-world datasets demonstrate that STING outperforms the existing state-of-the-art methods in terms of imputation accuracy as well as downstream tasks with the imputed values therein. Eunkyu Oh, Yunhu Ji, Sushil Khyalia |
ICDM | 4 |