Candy Olivia Mawalim

dblp:225/7510 · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
10since 2021 · last 2025
0000-0001-9853-8893ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Multimodal Classification of Co-speech Gesture Pragmatic Function in Storytelling
abstract
Gestures, as essential co-speech behaviors in human communication, carry rich pragmatic functions. Accurately recognizing these functions could enhance an agent’s ability to understand communicative behavior. In spontaneous storytelling scenarios—unlike lab-controlled settings—gesture functions exhibit high variability and are strongly influenced by individual speaker differences, making it difficult for unimodal systems to reliably capture their pragmatic intent. To address this challenge, we collected and annotated naturally occurring co-speech gestures in narrative dialogues, assigning each gesture one of six pragmatic function labels. We further propose a multimodal sequential classification model that encodes skeletal motion, acoustic prosody, and facial dynamics using separate bidirectional LSTM networks. These modality-specific encodings are fused via cross-modal attention and a gated mechanism to capture temporal dependencies and complementary information across modalities. Experimental results demonstrate that our tri-modal system achieves 62.5% accuracy and a weighted F1 score of 0.62 on the six-way classification task, outperforming uni-modal and bi-modal baselines by 3–11%. Ablation analysis reveals that skeletal features provide the most discriminative power for the majority of gesture functions, acoustic features are critical for specific categories, and facial features—though weak in isolation—substantially enhance overall performance when integrated via attention.
Jinqian Zhang, Sixia Li, Candy Olivia Mawalim, Shogo Okada
HAI3
2025 Fine-tuning TitaNet-Large Model for Speaker Anonymization Attacker Systems
abstract
Speaker anonymization techniques are crucial for safeguarding user privacy in voice-based applications. However, these methods are susceptible to adversarial attacks that can compromise their effectiveness. This paper proposes attacker systems that leverage the power of fine-tuned TitaNet-Large and ECAPA-TDNN models to identify the original speaker from anonymized speech generated by various anonymization methods. Both pre-trained models are renowned for their state-of-the-art ability to extract robust speaker embeddings. Finetuning these models with anonymized speech enables them to identify underlying patterns in anonymized speech. We evaluated the proposed attacker systems against multiple anonymization techniques that performed effectively in a series of voice privacy challenges. Our experimental results underscore the effectiveness of the fine-tuned TitaNet-Large model in breaking through these anonymization methods, as indicated by the reduced equal error rate (EER). This highlights the importance of robust and adaptive anonymization strategies to counter such emerging semiinformed threats.
Candy Olivia Mawalim, Aulia Adila, Masashi Unoki
ICASSP1
2025 Robust Multilingual Audio Deepfake Detection Through Hybrid Modeling
abstract
The increasing sophistication of AI-generated human voice poses a significant threat, demanding robust detection systems that can generalize effectively across diverse linguistic environments and synthesis techniques.In response to the SAFE Challenge, this paper introduces a novel approach to multilingual audio deepfake detection.Our primary contribution lies in the comprehensive study of deepfake detection using a multilingual speech corpus encompassing 17 languages and a broad spectrum of synthesis methods and acoustic conditions, designed to enable more realistic and challenging evaluations.To optimally utilize this diverse data, we propose a hybrid detection model that synergistically combines the strengths of end-to-end RawNet and AASIST architectures with language-agnostic representations learned from a multilingual selfsupervised learning model.Additionally, we explore the efficacy of RawBoost data augmentation in enhancing robustness against realworld noise.Our experimental evaluation demonstrates promising generalization in generated audio detection, achieving approximately 73% balanced accuracy across multilingual data and unseen synthesis algorithms.
Candy Olivia Mawalim, Aulia Adila, Shogo Okada, Masashi Unoki
IH&MMSec1
2025 WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines
abstract
Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan, David Anugraha, Rifki Afina Putri, Wang Yutong, Adam Nohejl, Ubaidillah Ariq Prathama, Nedjma Ousidhoum, Afifa Amriani, Anar Rzayev, Anirban Das, Ashmari Pramodya, Aulia Adila, Bryan Wilie, Candy Olivia Mawalim, Cheng Ching Lam, Daud Abolade, Emmanuele Chersoni, Enrico Santus, Fariz Ikhwantri, Garry Kuwanto, Hanyang Zhao, Haryo Akbarianto Wibowo, Holy Lovenia, Jan Christian Blaise Cruz, Jan Wira Gotama Putra, Junho Myung, Lucky Susanto, Maria Angelica Riera Machin, Marina Zhukova, Michael Anugraha, Muhammad Farid Adilazuarda, Natasha Christabelle Santosa, Peerat Limkonchotiwat, Raj Dabre, Rio Alexander Audino, Samuel Cahyawijaya, Shi-Xiong Zhang, Stephanie Yulia Salim, Yi Zhou, Yinxuan Gui, David Ifeoluwa Adelani, En-Shiun Annie Lee, Shogo Okada, Ayu Purwarianti, Alham Fikri Aji, Taro Watanabe, Derry Tanti Wijaya, Alice Oh, Chong-Wah Ngo. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan, David Anugraha, Rifki Afina Putri, Adam Nohejl, Ubaidillah Ariq Prathama, Nedjma Ousidhoum, Afifa Amriani, Anar Rzayev, Ashmari Pramodya, Aulia Adila, Bryan Wilie, Candy Olivia Mawalim, Cheng Ching Lam, Daud Abolade, Emmanuele Chersoni, Enrico Santus, Fariz Ikhwantri, Garry Kuwanto, Hanyang Zhao, Haryo Akbarianto Wibowo, Holy Lovenia, Jan Christian Blaise Cruz, Jan Wira Gotama Putra, Junho Myung, Lucky Susanto, Maria Angelica Riera Machin, Marina Zhukova, Michael Anugraha, Muhammad Farid Adilazuarda, Natasha Christabelle Santosa, Peerat Limkonchotiwat, Raj Dabre, Rio Alexander Audino, Samuel Cahyawijaya, Stephanie Yulia Salim, Yi Zhou 0019, Yinxuan Gui, David Ifeoluwa Adelani, Annie En-Shiun Lee, Shogo Okada, Ayu Purwarianti, Alham Fikri Aji, Taro Watanabe, Derry Wijaya, Alice Oh, Chong-Wah Ngo
NAACL (Long Papers)16
2024 Incremental Multimodal Sentiment Analysis for HAIs Based on Multitask Active Learning with Interannotator Agreement
abstract
Multimodal sentiment analysis (MSA) is critical in developing empathetic and adaptive multimodal dialogue systems or conversational agents that can naturally interact with users by recognizing sentiment and engagement. Addressing the challenges of collecting labeled data for MSA in human-agent interaction (HAI), this study introduces an innovative approach that combines active learning and multitask learning. Our efficient sentiment recognition model leverages active learning to select informative data for learning models, significantly reducing the labor-intensive data labeling process. Furthermore, we employ multitask learning to improve annotation (label) quality by evaluating alignment with true labels and interannotator agreement, thus enhancing the reliability of sentiment annotations. We evaluate the proposed multitask and active learning methods via a human-agent multimodal dialogue dataset that includes various types of sentiment annotations, which are publicly available. The experimental results demonstrate that by learning to predict the agreement score, multitask learning becomes better than singletask learning at capturing the uncertainties in the data. This study lays the groundwork for incremental learning strategies in MSA, aiming to adaptively understand user sentiments in human-agent interactions.
Thus Karnjanapatchara, Sixia Li, Candy Olivia Mawalim, Kazunori Komatani, Shogo Okada
ACII3
2024 Do We Need To Watch It All? Efficient Job Interview Video Processing with Differentiable Masking
abstract
With technological advancements in transmitting and storing large video files, more and more organizations are incorporating asynchronous video interviews as part of their personnel selection process. Automatic evaluation of these videos is a challenging machine learning setting because the samples are composed of time series input data but only one overall label is available. It is unclear which segments of the time series input (i.e., videos) are the most important ones for prediction. Not all nonverbal features, spoken words, and utterances contribute equally to the prediction; some segments of the videos might even introduce noise to the model. Processing all multimodal information is therefore inefficient. To address this challenge, we propose a framework to model the content of the answer via the full transcription and the speaking patterns of the interviewee via short clips. Our model learns to automatically select the most informative segment by previewing the acoustic modality using a technique called differentiable masking. The results show that our method outperforms existing approaches while being more efficient since only partial multimodal data are processed, and the interpretability of the model is enhanced.
Sixia Li, Candy Olivia Mawalim, Hung-Hsuan Huang, Chee Wee Leong, Shogo Okada
ICMI3
2024 Are Recent Deep Learning-Based Speech Enhancement Methods Ready to Confront Real-World Noisy Environments?
abstract
Recent advancements in speech enhancement techniques have ignited interest in improving speech quality and intelligibility.However, the effectiveness of recently proposed methods is unclear.In this paper, a comprehensive analysis of modern deep learning-based speech enhancement approaches is presented.Through evaluations using the Deep Suppression Noise and Clarity Enhancement Challenge datasets, we assess the performances of three methods: Denoiser, DeepFilterNet3, and FullSubNet+.Our findings reveal nuanced performance differences among these methods, with varying efficacy across datasets.While objective metrics offer valuable insights, they struggle to represent complex scenarios with multiple noise sources.Leveraging ASR-based methods for these scenarios shows promise but may induce critical hallucination effects.Our study emphasizes the need for ongoing research to refine techniques for diverse real-world environments.
Candy Olivia Mawalim, Shogo Okada, Masashi Unoki
INTERSPEECH1
2022 Multimodal Analysis for Communication Skill and Self-Efficacy Level Estimation in Job Interview Scenario
abstract
An interview for a job recruiting process requires applicants to demonstrate their communication skills. Interviewees sometimes become nervous about the interview because interviewees themselves do not know their assessed score. This study investigates the relationship between the communication skill (CS) and the self-efficacy level (SE) of interviewees through multimodal modeling. We also clarify the difference between effective features in the prediction of CS and SE labels. For this purpose, we collect a novel multimodal job interview data corpus by using a job interview agent system where users experience the interview using a virtual reality head-mounted display (VR-HMD). The data corpus includes annotations of CS by third-party experts and SE annotations by the interviewees. The data corpus also includes various kinds of multimodal data, including audio, biological (i.e., physiological), gaze, and language data. We present two types of regression models, linear regression and sequential-based regression models, to predict CS, SE, and the gap (GA) between skill and self-efficacy. Finally, we report that the model with acoustic, gaze, and linguistic features has the best regression accuracy in CS prediction (correlation coefficient r = 0.637). Furthermore, the regression model with biological features achieves the best accuracy in SE prediction (r = 0.330).
Tomoya Ohba, Candy Olivia Mawalim, Shun Katada, Haruki Kuroki, Shogo Okada
MUM2
2022 Speaker anonymization by modifying fundamental frequency and x-vector singular value
abstract
Speaker anonymization is a method of protecting voice privacy by concealing individual speaker characteristics while preserving linguistic information. The VoicePrivacy Challenge 2020 was initiated to generalize the task of speaker anonymization. In the challenge, two frameworks for speaker anonymization were introduced; in this study, we propose a method of improving the primary framework by modifying the state-of-the-art speaker individuality feature (namely, x-vector) in a neural waveform speech synthesis model. Our proposed method is constructed based on x-vector singular value modification with a clustering model. We also propose a technique of modifying the fundamental frequency and speech duration to enhance the anonymization performance. To evaluate our method, we carried out objective and subjective tests. The overall objective test results show that our proposed method improves the anonymization performance in terms of the speaker verifiability, whereas the subjective evaluation results show improvement in terms of the speaker dissimilarity. The intelligibility and naturalness of the anonymized speech with speech prosody modification were slightly reduced (less than 5% of word error rate) compared to the results obtained by the baseline system.
Candy Olivia Mawalim, Kasorn Galajit, Jessada Karnjana, Shunsuke Kidani, Masashi Unoki
Comput. Speech Lang.1
2021 Task-independent Recognition of Communication Skills in Group Interaction Using Time-series Modeling
abstract
Case studies of group discussions are considered an effective way to assess communication skills (CS). This method can help researchers evaluate participants’ engagement with each other in a specific realistic context. In this article, multimodal analysis was performed to estimate CS indices using a three-task-type group discussion dataset, the MATRICS corpus. The current research investigated the effectiveness of engaging both static and time-series modeling, especially in task-independent settings. This investigation aimed to understand three main points: first, the effectiveness of time-series modeling compared to nonsequential modeling; second, multimodal analysis in a task-independent setting; and third, important differences to consider when dealing with task-dependent and task-independent settings, specifically in terms of modalities and prediction models. Several modalities were extracted (e.g., acoustics, speaking turns, linguistic-related movement, dialog tags, head motions, and face feature sets) for inferring the CS indices as a regression task. Three predictive models, including support vector regression (SVR), long short-term memory (LSTM), and an enhanced time-series model (an LSTM model with a combination of static and time-series features), were taken into account in this study. Our evaluation was conducted by using the R 2 score in a cross-validation scheme. The experimental results suggested that time-series modeling can improve the performance of multimodal analysis significantly in the task-dependent setting (with the best R 2 = 0.797 for the total CS index), with word2vec being the most prominent feature. Unfortunately, highly context-related features did not fit well with the task-independent setting. Thus, we propose an enhanced LSTM model for dealing with task-independent settings, and we successfully obtained better performance with the enhanced model than with the conventional SVR and LSTM models (the best R 2 = 0.602 for the total CS index). In other words, our study shows that a particular time-series modeling can outperform traditional nonsequential modeling for automatically estimating the CS indices of a participant in a group discussion with regard to task dependency.
Candy Olivia Mawalim, Shogo Okada, Yukiko I. Nakano
ACM Trans. Multim. Comput. Commun. Appl.1
2020 X-Vector Singular Value Modification and Statistical-Based Decomposition with Ensemble Regression Modeling for Speaker Anonymization System
Candy Olivia Mawalim, Kasorn Galajit, Jessada Karnjana, Masashi Unoki
INTERSPEECH1
2017 Rule-based Reordering and Post-Processing for Indonesian-Korean Statistical Machine Translation
Candy Olivia Mawalim, Dessi Puji Lestari, Ayu Purwarianti
PACLIC1