VLDB 2026 Research / reviewers in the wild / expert
Mayank Aggarwal
dblp:223/4550
· DBLP profile ↗
3ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Speech recognition and synthesis · 100% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Speech recognition and synthesis › visual speech recognition
lip reading |
0.3 | 1 | 2018 | Harnessing AI for Speech Reconstruction using Multi-view Silent Video Feed · ACM Multimedia 2018 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
silent speech recognition |
0.3 | 1 | 2018 | Harnessing AI for Speech Reconstruction using Multi-view Silent Video Feed · ACM Multimedia 2018 |
Natural language and speech › Speech recognition and synthesis
speech reconstruction |
0.3 | 1 | 2018 | Harnessing AI for Speech Reconstruction using Multi-view Silent Video Feed · ACM Multimedia 2018 |
Multimedia analysis and retrieval › video content analysis
multi-view video analysis |
0.1 | 1 | 2018 | Harnessing AI for Speech Reconstruction using Multi-view Silent Video Feed · ACM Multimedia 2018 |
Methods — techniques the papers use, named apart from their topics
multi-view fusion · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Machine Learning Based Dynamic TTI Bundling for LTE VoIP Uplink Coverage ImprovementabstractTo decrease the end-to-end delays, Long-Term Evolution (LTE) radio access networks have considerably smaller Transmission Time Intervals (TTI) compared to WCDMA networks. In order to increase the uplink cell coverage over Voice Over LTE (VoLTE) calls for cell-edge User Equipments (UE), TTI Bundling (TTIB) was introduced in release 8 of the 3rd Generation Partnership Project (3GPP). It is based on the principle that the same packet with different redundancy versions (RVs) is transmitted in consecutive TTIs increasing the probability of successful packet decoding. The bundle size in TTI Bundling is always fixed to 4 without considering the UEs' dynamic channel conditions. For UEs that face better channel conditions, the Evolved Node B (eNB) might be able to decode the data packet with fewer RVs than all 4. UEs present in the cell edge need to transmit with a higher power compared to cell near UEs, and adding multiple redundancy version transmission, without considering the channel conditions of the UE, will not only add to the power requirement of the UE but also impact the number of Resource blocks (RBs) available to non-TTI Bundling UEs. In this work, we propose a Machine learning based approach for computing the optimal bundle size of each UE based on its current channel condition. We implement two sequential models, where the first model predicts the channel conditions of the UE for the next transmission. This prediction is then fed to the second model which classifies the prediction into a bundle size ranging from 1 to 4. Using this predicted bundle size for the TTI Bundling UE, our results show a$\mathbf{2 6. 6 7 \%}$reduction in UE's power consumption, and a$\mathbf{2 9. 4 4 \%}$improvement in the system throughput. Mayank Aggarwal, Shyamal Dhua, S. Ram Aditya, Vimal Bastin Edwin Joseph, Bharath Kumar Kalavala, Ballem Jeevan Prakash |
CCNC | 1 |
| 2024 | Foraging in a non-foraging task: Fitness maximization explains human risk preference dynamics under changing environmentabstractChanges in risk preference have been reported when making a series of independent risky choices or non-foraging economic decisions. Behavioral economics has put forward various explanations for specific changes in risk preference in non-foraging tasks, but a consensus regarding the general principle underlying these effects has not been reached. In contrast, recent studies have investigated human economic risky choices using tasks adapted from foraging theory, which require consideration of past choices and future opportunities to make optimal decisions. In these foraging tasks, human economic risky choices are explained by the ethological principle of fitness maximization, which naturally leads to dynamic risk preference. Here, we conducted two online experiments to investigate whether the principle of fitness maximization can explain risk preference dynamics in a non-foraging task. Participants were asked to make a series of independent risky economic decisions while the environmental richness changed. We found that participants' risk preferences were influenced by the current and past environments, making them more risk-averse during and after the rich environment compared to the poor environment. These changes in risk preference align with fitness maximization. Our findings suggest that the ethological principle of fitness maximization might serve as a generalizable principle for explaining dynamic preferences, including risk preference, in human economic decision-making. Yasuhiro Mochizuki, Norihiro Harasawa, Mayank Aggarwal, Haruaki Fukuda |
PLoS Comput. Biol. | 3 |
| 2018 | Harnessing AI for Speech Reconstruction using Multi-view Silent Video FeedabstractSpeechreading or lipreading is the technique of understanding and getting phonetic features from a speaker's visual features such as movement of lips, face, teeth and tongue. It has a wide range of multimedia applications such as in surveillance, Internet telephony, and as an aid to a person with hearing impairments. However, most of the work in speechreading has been limited to text generation from silent videos. Recently, research has started venturing into generating (audio) speech from silent video sequences but there have been no developments thus far in dealing with divergent views and poses of a speaker. Thus although, we have multiple camera feeds for the speech of a user, but we have failed in using these multiple video feeds for dealing with the different poses. To this end, this paper presents the world's first ever multi-view speech reading and reconstruction system. This work encompasses the boundaries of multimedia research by putting forth a model which leverages silent video feeds from multiple cameras recording the same subject to generate intelligent speech for a speaker. Initial results confirm the usefulness of exploiting multiple camera views in building an efficient speech reading and reconstruction system. It further shows the optimal placement of cameras which would lead to the maximum intelligibility of speech. Next, it lays out various innovative applications for the proposed system focusing on its potential prodigious impact in not just security arena but in many other multimedia analytics problems. Yaman Singla, Mayank Aggarwal, Pratham Nawal, Shin'ichi Satoh 0001, Rajiv Ratn Shah, Roger Zimmermann |
ACM Multimedia | 2 |