VLDB 2026 Research / reviewers in the wild / expert
Mayank Kumar Singh
dblp:145/6963
· DBLP profile ↗
9ranked-venue papers
4as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Comprehensive Real-World Assessment of Audio Watermarking Algorithms: Will They Survive Neural Codecs?
Yigitcan Özer, Woosung Choi, Joan Serrà, Mayank Kumar Singh, Wei-Hsiang Liao 0001, Yuki Mitsufuji |
INTERSPEECH | 4 |
| 2025 | A 2.9mW Inverter-based Quadrature Phase Clock Generator with ± 0.29° Phase ErrorabstractThe quadrature phase clocks are important elements in digital programmable transceivers in communication system applications. However, current solutions in quadrature clock generators for broad frequency ranges need more phase accuracy, and they suffer from substandard phase noise performance and excessive power consumption. To address these challenges, this paper proposes an inverter-based quadrature-phase clock (I-QPC) generator. The I-QPC generator utilizes inverters as delay elements to achieve the desired phase without using poly-phase type-1 filters because inverters are simpler to design and optimize for different phase delays. The system implements a phase-averaging mechanism using the delayed and interpolated signals, leading to quadrature-phase signals. The proposed technique has been validated in 28nm standard CMOS technology after post-layout parasitic extraction. The I-QPC generator operates over the broad frequency range (1GHz to 6GHz) and occupies an active area of 0.0005mm2. The post-layout simulation results show that the phase error is ±0.29°while operating at 6GHz. The phase noise is -131.7dBc/Hz at an offset of 1MHz with a power consumption of 2.9mW. The I-QPC generator’s figure of merit (FoM) is 217.3dBc/Hz at 1MHz offset frequency, which is better than state-of-the-art architectures. The performance of the I-QPC generator was further evaluated in hardware by implementing the circuit on a breadboard using the SN74HC04N inverter IC, and it demonstrated the successful generation of the quadrature signals at 1MHz frequency. Mayank Kumar Singh, M. Bhuvanesh, Rajasekhar Nagulapalli, Devarshi Mrinal Das, Mahendra Sakare |
ISCAS | 1 |
| 2024 | SilentCipher: Deep Audio Watermarking
Mayank Kumar Singh, Naoya Takahashi, Wei-Hsiang Liao 0001, Yuki Mitsufuji |
INTERSPEECH | 1 |
| 2023 | Nonparallel Emotional Voice Conversion for Unseen Speaker-Emotion Pairs Using Dual Domain Adversarial Network & Virtual Domain PairingabstractPrimary goal of an emotional voice conversion (EVC) system is to convert the emotion of a given speech signal from one style to another style without modifying the linguistic content of the signal. Most of the state-of-the-art approaches convert emotions for seen speaker-emotion combinations only. In this paper, we tackle the problem of converting the emotion of speakers whose only neutral data are present during the time of training and testing (i.e., unseen speaker-emotion combinations). To this end, we extend a recently proposed StartGANv2-VC architecture by utilizing dual encoders for learning the speaker and emotion style embeddings separately along with dual domain source classifiers. For achieving the conversion to unseen speaker-emotion combinations, we propose a Virtual Domain Pairing (VDP) training strategy, which virtually incorporates the speaker-emotion pairs that are not present in the real data without compromising the min-max game of a discriminator and generator in adversarial training. We evaluate the proposed method using a Hindi emotional database. Nirmesh J. Shah, Mayank Kumar Singh, Naoya Takahashi, Naoyuki Onoe |
ICASSP | 2 |
| 2023 | Hierarchical Diffusion Models for Singing Voice Neural VocoderabstractRecent progress in deep generative models has improved the quality of neural vocoders in speech domain. However, generating a high-quality singing voice remains challenging due to a wider variety of musical expressions in pitch, loudness, and pronunciations. In this work, we propose a hierarchical diffusion model for singing voice neural vocoders. The proposed method consists of multiple diffusion models operating in different sampling rates; the model at the lowest sampling rate focuses on generating accurate low-frequency components such as pitch, and other models progressively generate the waveform at higher sampling rates on the basis of the data at the lower sampling rate and acoustic features. Experimental results show that the proposed method produces high-quality singing voices for multiple singers, outperforming state-of-the-art neural vocoders with a similar range of computational costs. Naoya Takahashi, Mayank Kumar Singh, Yuki Mitsufuji |
ICASSP | 2 |
| 2023 | Iteratively Improving Speech Recognition and Voice Conversion
Mayank Kumar Singh, Naoya Takahashi, Naoyuki Onoe |
INTERSPEECH | 1 |
| 2023 | Log exponential shrinkage: a denoising technique for breast ultrasound images
Mayank Kumar Singh, Indu Saini, Neetu Sood |
Vis. Comput. | 1 |
| 2021 | Hierarchical disentangled representation learning for singing voice conversionabstractConventional singing voice conversion (SVC) methods often suffer from operating in high-resolution audio owing to a high dimensionality of data. In this paper, we propose a hierarchical representation learning that enables the learning of disentangled representations with multiple resolutions independently. With the learned disentangled representations, the proposed method progressively performs SVC from low to high resolutions. Experimental results show that the proposed method outperforms baselines that operate with a single resolution in terms of mean opinion score (MOS), similarity score, and pitch accuracy. Naoya Takahashi, Mayank Kumar Singh, Yuki Mitsufuji |
IJCNN | 2 |
| 2020 | Improving Voice Separation by Incorporating End-To-End Speech RecognitionabstractDespite recent advances in voice separation methods, many challenges remain in realistic scenarios such as noisy recording and the limits of available data. In this work, we propose to explicitly incorporate the phonetic and linguistic nature of speech by taking a transfer learning approach using an end-to-end automatic speech recognition (E2EASR) system. The voice separation is conditioned on deep features extracted from E2EASR to cover the long-term dependence of phonetic aspects. Experimental results on speech separation and enhancement task on the AVSpeech dataset show that the proposed method significantly improves the signal-to-distortion ratio over the baseline model and even outperforms an audio visual model, that utilizes visual information of lip movements. Naoya Takahashi, Mayank Kumar Singh, Sakya Basak, Sudarsanam Parthasaarathy, Sriram Ganapathy, Yuki Mitsufuji |
ICASSP | 2 |