Xin Jing 0001

dblp:03/11308-1 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-8803-9414ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Enhancing Emotional Text-to-Speech Controllability with Natural Language Guidance through Contrastive Learning and Diffusion Models
abstract
While current emotional text-to-speech (TTS) systems can generate highly intelligible emotional speech, achieving fine control over emotion rendering of the output speech still remains a significant challenge. In this paper, we introduce ParaEVITS, a novel emotional TTS framework that leverages the compositionality of natural language to enhance control over emotional rendering. By incorporating a text-audio encoder inspired by ParaCLAP, a contrastive language-audio pretraining (CLAP) model for computational paralinguistics, the diffusion model is trained to generate emotional embeddings based on textual emotional style descriptions. Our framework first trains on reference audio using the audio encoder, then fine-tunes a diffusion model to process textual inputs from ParaCLAP’s text encoder. During inference, speech attributes such as pitch, jitter, and loudness are manipulated using only textual conditioning. Our experiments demonstrate that ParaEVITS effectively control emotion rendering without compromising speech quality. Speech demos are publicly available1.
Xin Jing 0001, Kun Zhou 0003, Andreas Triantafyllopoulos, Björn W. Schuller
ICASSP1
2025 MADUV: The 1st INTERSPEECH Mice Autism Detection via Ultrasound Vocalization Challenge
Zijiang Yang 0007, Meishu Song, Xin Jing 0001, Kun Qian 0003, Bin Hu 0001, Kota Tamada, Toru Takumi, Björn W. Schuller, Yoshiharu Yamamoto
INTERSPEECH3
2025 Vishing: Detecting social engineering in spoken communication - A first survey & urgent roadmap to address an emerging societal challenge
abstract
Vishing – the use of voice calls for phishing – is a form of Social Engineering (SE) attacks. The latter have become a pervasive challenge in modern societies, with over 300,000 yearly victims in the US alone. An increasing number of those attacks is conducted via voice communication, be it through machine-generated ‘robocalls’ or human actors. The goals of ‘social engineers’ can be manifold, from outright fraud to more subtle forms of persuasion. Accordingly, social engineers adopt multi-faceted strategies for voice-based attacks, utilising a variety of ‘tricks’ to exert influence and achieve their goals. Importantly, while organisations have set in place a series of guardrails against other types of SE attacks, voice calls still remain ‘open ground’ for potential bad actors. In the present contribution, we provide an overview of the existing speech technology subfields that need to coalesce into a protective net against one of the major challenges to societies worldwide. Given the dearth of speech science and technology works targeting this issue, we have opted for a narrative review that bridges the gap between the existing psychological literature on the topic and research that has been pursued in parallel by the speech community on some of the constituent constructs. Our review reveals that very little literature exists on addressing this very important topic from a speech technology perspective, an omission further exacerbated by the lack of available data. Thus, our main goal is to highlight this gap and sketch out a roadmap to mitigate it, beginning with the psychological underpinnings of vishing, which primarily include deception and persuasion strategies, continuing with the speech-based approaches that can be used to detect those, as well as the generation and detection of AI-based vishing attempts, and close with a discussion of ethical and legal considerations. • Vishing is an emerging security problem. • Generative artificial intelligence will exacerbate the issue. • Speech-based detection is urgently needed. • Beyond detection performance, effective interventions are warranted.
Andreas Triantafyllopoulos, Anika A. Spiesberger, Iosif Tsangko, Xin Jing 0001, Verena Distler, Felix Dietz, Florian Alt, Björn W. Schuller
Comput. Speech Lang.4
2025 Audio-Based Kinship Verification Using Age Domain Conversion
abstract
Audio-based kinship verification (AKV) is important in many domains, such as home security monitoring, forensic identification, and social network analysis. A key challenge in the task arises from differences in age across samples from different individuals, which can be interpreted as a domain bias in a cross-domain verification task. To address this issue, we design the notion of an “age-standardised domain” wherein we utilise the optimised CycleGAN-VC3 network to perform age-audio conversion to generate the in-domain audio. The generated audio dataset is employed to extract a range of features, which are then fed into a metric learning architecture to verify kinship. Experiments are conducted on the KAN_AV audio dataset.The results demonstrate that the method markedly enhances the accuracy of kinship verification, while also offering novel insights for future kinship verification research.
Alican Akman, Xin Jing 0001, Manuel Milling, Björn W. Schuller
IEEE Signal Process. Lett.3
2025 STAA-Net: A Sparse and Transferable Adversarial Attack for Speech Emotion Recognition
abstract
Speech contains rich information on the emotions of humans, and Speech Emotion Recognition (SER) has been an important topic in the area of human-computer interaction. The robustness of SER models is crucial, particularly in privacy-sensitive and reliability-demanding domains like private healthcare. Recently, the vulnerability of deep neural networks in the audio domain to adversarial attacks has become a popular area of research. However, prior works on adversarial attacks in the audio domain primarily rely on iterative gradient-based techniques, which are time-consuming and prone to overfitting the specific threat model. Furthermore, the exploration of sparse perturbations, which have the potential for better stealthiness, remains limited in the audio domain. To address these challenges, we propose a generator-based attack method to generate sparse and transferable adversarial examples to deceive SER models in an end-to-end and efficient manner. We evaluate our method on two widely-used SER datasets, Database of Elicited Mood in Speech (DEMoS) and Interactive Emotional dyadic MOtion CAPture (IEMOCAP), and demonstrate its ability to generate successful sparse adversarial examples in an efficient manner. Moreover, our generated adversarial examples exhibit model-agnostic transferability, enabling effective adversarial attacks on advanced victim models.
Yi Chang 0004, Zhao Ren, Zixing Zhang 0001, Xin Jing 0001, Kun Qian 0003, Xi Shao, Bin Hu 0001, Tanja Schultz, Björn W. Schuller
IEEE Trans. Affect. Comput.4
2024 ParaCLAP - Towards a general language-audio model for computational paralinguistic tasks
abstract
Contrastive language-audio pretraining (CLAP) has recently emerged as a method for making audio analysis more generalisable.Specifically, CLAP-style models are able to 'answer' a diverse set of language queries, extending the capabilities of audio models beyond a closed set of labels.However, CLAP relies on a large set of (audio, query) pairs for pretraining.While such sets are available for general audio tasks, like captioning or sound event detection, there are no datasets with matched audio and text queries for computational paralinguistic (CP) tasks.As a result, the community relies on generic CLAP models trained for general audio with limited success.In the present study, we explore training considerations for ParaCLAP, a CLAP-style model suited to CP, including a novel process for creating audio-language queries.We demonstrate its effectiveness on a set of computational paralinguistic tasks, where it is shown to surpass the performance of open-source state-of-theart models.Our code and resources are publicly available at: https://github.com/KeiKinn/ParaCLAP
Xin Jing 0001, Andreas Triantafyllopoulos, Björn W. Schuller
INTERSPEECH1
2024 DB3V: A Dialect Dominated Dataset of Bird Vocalisation for Cross-corpus Bird Species Recognition
abstract
In ornithology, bird species are known to have variedit's widely acknowledged that bird species display diverse dialects in their calls across different regions.Consequently, computational methods to identify bird species onsolely through their calls face critsignificalnt challenges.There is growing interest in understanding the impact of species-specific dialects on the effectiveness of bird species recognition methods.Despite potential mitigation through the expansion of dialect datasets, the absence of publicly available testing data currently impedes robust benchmarking efforts.This paper presents the Dialect Dominated Dataset of Bird Vocalisation (D3BV), the first crosscorpus dataset that focuses on dialects in bird vocalisations.The D3BV comprises more than 25 hours of audio recordings from 10 bird species distributed across three distinct regions in the contiguous United States (CONUS).In addition to presenting the dataset, we conduct analyses and establish baseline models for cross-corpus bird recognition.The data and code are publicly available online: https://zenodo.org/
Xin Jing 0001, Jiangjian Xie, Alexander Gebhard 0001, Alice Baird, Björn W. Schuller
INTERSPEECH1
2023 Daily Mental Health Monitoring from Speech: A Real-World Japanese Dataset and Multitask Learning Analysis
abstract
Translating mental health recognition from clinical research into real-world application requires extensive data, yet existing emotion datasets are impoverished in terms of daily mental health monitoring, especially when aiming for self-reported anxiety and depression recognition. We introduce the Japanese Daily Speech Dataset (JDSD), a large in-the-wild daily speech emotion dataset consisting of 20,827 speech samples from 342 speakers and 54 hours of total duration. The data is annotated on the Depression and Anxiety Mood Scale (DAMS) – 9 self-reported emotions to evaluate mood state including "vigorous", "gloomy", "concerned", "happy", "unpleasant", "anxious", "cheerful", "depressed", and "worried". Our dataset possesses emotional states, activity, and time diversity, making it useful for training models to track daily emotional states for healthcare purposes. We partition our corpus and provide a multi-task benchmark across nine emotions, demonstrating that mental health states can be predicted reliably from self-reports with a Concordance Correlation Coefficient value of .547 on average. We hope that JDSD will become a valuable resource to further the development of daily emotional healthcare tracking.
Meishu Song, Andreas Triantafyllopoulos, Zijiang Yang 0007, Hiroki Takeuchi, Toru Nakamura, Akifumi Kishi, Tetsuro Ishizawa, Kazuhiro Yoshiuchi, Xin Jing 0001, Vincent Karas, Zhonghao Zhao, Kun Qian 0003, Bin Hu 0001, Björn W. Schuller, Yoshiharu Yamamoto
ICASSP9
2022 An Overview & Analysis of Sequence-to-Sequence Emotional Voice Conversion
abstract
Emotional voice conversion (EVC) focuses on converting a speech utterance from a source to a target emotion; it can thus be a key enabling technology for human-computer interaction applications and beyond. However, EVC remains an unsolved research problem with several challenges. In particular, as speech rate and rhythm are two key factors of emotional conversion, models have to generate output sequences of differing length. Sequence-to-sequence modelling is recently emerging as a competitive paradigm for models that can overcome those challenges. In an attempt to stimulate further research in this promising new direction, recent sequence-to-sequence EVC papers were systematically investigated and reviewed from six perspectives: their motivation, training strategies, model architectures, datasets, model inputs, and evaluation methods. This information is organised to provide the research community with an easily digestible overview of the current state-of-the-art. Finally, we discuss existing challenges of sequence-to-sequence EVC.
Zijiang Yang 0007, Xin Jing 0001, Andreas Triantafyllopoulos, Meishu Song, Ilhan Aslan, Björn W. Schuller
INTERSPEECH2