VLDB 2026 Research / reviewers in the wild / expert
Sashi Novitasari
dblp:219/5250
· DBLP profile ↗
11ranked-venue papers
9as first author
8since 2021 · last 2025
0000-0001-7467-5682ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 8 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Granite-speech: open-source speech-aware LLMs with strong English ASR capabilitiesabstractGranite-speech LLMs are compact and efficient speech language models specifically designed for English ASR1and automatic speech translation (AST). The models were trained by modality aligning granite-3.3-instruct to speech on publicly available open-source corpora. Comprehensive benchmarking on English ASR shows that they outperform several competitors’ models that were trained on orders of magnitude more proprietary data, and they keep pace on English-to-X AST for major European languages, Japanese, and Mandarin. The speech-specific components are: a conformer acoustic encoder using block attention and self-conditioning trained with connectionist temporal classification, a windowed query-transformer speech modality adapter used to do temporal downsampling of the acoustic embeddings and map them to the LLM text embedding space, and LoRA adapters to further fine-tune the text LLM. The models are freely available on HuggingFace2under a permissive Apache 2.0 license.1The latest models (revision 3.3.2) support multilingual ASR in English, French, German, Spanish and Portuguese and bidirectional speech translation to and from English. This paper covers the initial English-only release.2https://huggingface.co/ibm-granite/granite-speech-3.3-2b (and…-8b). George Saon, Avihu Dekel, Alexi Brooks, Tohru Nagano, Abraham Daniels, Aharon Satt, Ashish R. Mittal, Brian Kingsbury, David Haws, Edmilson da Silva Morais, Gakuto Kurata, Hagai Aronowitz, Ibrahim Ibrahim, Hong-Kwang Jeff Kuo, Kate Soule, Luis A. Lastras, Masayuki Suzuki, Ron Hoory, Samuel Thomas 0001, Sashi Novitasari, Takashi Fukuda, Vishal Sunder, Zvi Kons |
ASRU | 20 |
| 2025 | Improving End-to-end Mixed-case ASR with Knowledge Distillation and Integration of Voice Activity Cues
Sashi Novitasari, Takashi Fukuda, Gakuto Kurata |
INTERSPEECH | 1 |
| 2025 | Voice Activity-based Text Segmentation for ASR Text Denormalization
Sashi Novitasari, Takashi Fukuda, Gakuto Kurata |
INTERSPEECH | 1 |
| 2023 | Self-Adaptive Incremental Machine Speech Chain for Lombard TTS with High-Granularity ASR Feedback in Dynamic Noise ConditionabstractA common approach for text-to-speech (TTS) in noisy conditions is offline fine-tuning, which is generally utilized on static noises and predefined conditions. We recently proposed a self-adaptive TTS in machine speech chain inference that enables TTS to control its voices in statically and dynamically noisy environments based on auditory feedback from automatic speech recognition (ASR) and speech-to-noise ratio (SNR) recognition. However, that study only investigated the system on synthetic Lombard speech data. Furthermore, the ASR feedback was at a lower granularity based only on the loss of the positive character class. In this paper, we improve the self-adaptive TTS using character-vocabulary level ASR feedback at higher granularity, considering the losses in the positive and negative classes. We focus on a self-adaptive incremental TTS (Adapt-ITTS) with a short-term feedback mechanism that aims for low latency adaptation for dynamically noisy situations. In experiments, our proposed Adapt- ITTS successfully improved intelligibility in noisy conditions based on synthetic and natural Lombard speech data on the Wall Street Journal and Hurricane datasets, respectively. Sashi Novitasari, Sakriani Sakti, Satoshi Nakamura 0001 |
ICASSP | 1 |
| 2022 | Improving ASR Robustness in Noisy Condition Through VAD Integration
Sashi Novitasari, Takashi Fukuda, Gakuto Kurata |
INTERSPEECH | 1 |
| 2022 | Improved Consistency Training for Semi-Supervised Sequence-to-Sequence ASR via Speech Chain Reconstruction and Self-TranscribingabstractConsistency regularization has recently been applied to semi-supervised sequence-to-sequence (S2S) automatic speech recognition (ASR).This principle encourages an ASR model to output similar predictions for the same input speech with different perturbations.The existing paradigm of semi-supervised S2S ASR utilizes SpecAugment as data augmentation and requires a static teacher model to produce pseudo transcripts for untranscribed speech.However, this paradigm fails to take full advantage of consistency regularization.First, the masking operations of SpecAugment may damage the linguistic contents of the speech, thus influencing the quality of pseudo labels.Second, S2S ASR requires both input speech and prefix tokens to make the next prediction.The static prefix tokens made by the offline teacher model cannot match dynamic pseudo labels during consistency training.In this work, we propose an improved consistency training paradigm of semi-supervised S2S ASR.We utilize speech chain reconstruction as the weak augmentation to generate high-quality pseudo labels.Moreover, we demonstrate that dynamic pseudo transcripts produced by the student ASR model benefit the consistency training.Experiments on LJSpeech and LibriSpeech corpora show that compared to supervised baselines, our improved paradigm achieves a 12.2% CER improvement in the single-speaker setting and 38.6% in the multi-speaker setting. Heli Qi, Sashi Novitasari, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2022 | A Machine Speech Chain Approach for Dynamically Adaptive Lombard TTS in Static and Dynamic Noise EnvironmentsabstractRecent end-to-end text-to-speech synthesis (TTS) systems have successfully synthesized high-quality speech. However, TTS speech intelligibility degrades in noisy environments because most of these systems were not designed to handle noisy environments. Several works attempted to address this problem by using offline fine-tuning to adapt their TTS to noisy conditions. Unlike machines, humans never perform offline fine-tuning. Instead, they speak with the Lombard effect in noisy places, where they dynamically adjust their vocal effort to improve the audibility of their speech. This ability is supported by the speech chain mechanism, which involves auditory feedback passing from speech perception to speech production. This paper proposes an alternative approach to TTS in noisy environments that is closer to the human Lombard effect. Specifically, we implement Lombard TTS in a machine speech chain framework to synthesize speech with dynamic adaptation. Our TTS performs adaptation by generating speech utterances based on the auditory feedback that consists of the automatic speech recognition (ASR) loss as the speech intelligibility measure and the speech-to-noise ratio (SNR) prediction as power measurement. Two versions of TTS are investigated: non-incremental TTS with utterance-level feedback and incremental TTS (ITTS) with short-term feedback to reduce the delay without significant performance loss. Furthermore, we evaluate the TTS systems in both static and dynamic noise conditions. Our experimental results show that auditory feedback enhanced the TTS speech intelligibility in noise. Sashi Novitasari, Sakriani Sakti, Satoshi Nakamura 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2021 | Dynamically Adaptive Machine Speech Chain Inference for TTS in Noisy Environment: Listen and Speak Louder
Sashi Novitasari, Sakriani Sakti, Satoshi Nakamura 0001 |
Interspeech | 1 |
| 2020 | Incremental Machine Speech Chain Towards Enabling Listening While Speaking in Real-TimeabstractInspired by a human speech chain mechanism, a machine speech chain framework based on deep learning was recently proposed for the semi-supervised development of automatic speech recognition (ASR) and text-to-speech synthesis (TTS) systems.However, the mechanism to listen while speaking can be done only after receiving entire input sequences.Thus, there is a significant delay when encountering long utterances.By contrast, humans can listen to what they speak in real-time, and if there is a delay in hearing, they won't be able to continue speaking.In this work, we propose an incremental machine speech chain towards enabling machine to listen while speaking in real-time.Specifically, we construct incremental ASR (ISR) and incremental TTS (ITTS) by letting both systems improve together through a short-term loop.Our experimental results reveal that our proposed framework is able to reduce delays due to long utterances while keeping a comparable performance to the non-incremental basic machine speech chain. Sashi Novitasari, Andros Tjandra, Tomoya Yanagita, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 1 |
| 2019 | Sequence-to-Sequence Learning via Attention Transfer for Incremental Speech RecognitionabstractAttention-based sequence-to-sequence automatic speech recognition (ASR) requires a significant delay to recognize long utterances because the output is generated after receiving entire input sequences. Although several studies recently proposed sequence mechanisms for incremental speech recognition (ISR), using different frameworks and learning algorithms is more complicated than the standard ASR model. One main reason is because the model needs to decide the incremental steps and learn the transcription that aligns with the current short speech segment. In this work, we investigate whether it is possible to employ the original architecture of attention-based ASR for ISR tasks by treating a full-utterance ASR as the teacher model and the ISR as the student model. We design an alternative student network that, instead of using a thinner or a shallower model, keeps the original architecture of the teacher model but with shorter sequences (few encoder and decoder states). Using attention transfer, the student network learns to mimic the same alignment between the current input short speech segments and the transcription. Our experiments show that by delaying the starting time of recognition process with about 1.7 sec, we can achieve comparable performance to one that needs to wait until the end. Sashi Novitasari, Andros Tjandra, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 1 |
| 2018 | Construction of English-French Multimodal Affective Conversational Corpus from TV Dramas
Sashi Novitasari, Quoc Truong Do, Sakriani Sakti, Dessi Puji Lestari, Satoshi Nakamura 0001 |
LREC | 1 |