VLDB 2026 Research / reviewers in the wild / expert
Youxiang Zhu
dblp:250/2634
· DBLP profile ↗
16ranked-venue papers
7as first author
14since 2021 · last 2026
0009-0000-7294-1596ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Computer networks · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Assessing Privacy Preservation and Utility in Online Vision-Language Models
Karmesh Siddharam Chaudhari, Youxiang Zhu, Amy Feng |
ICC | 2 |
| 2025 | ReEvalMed: Rethinking Medical Report Evaluation by Aligning Metrics with Real-World Clinical JudgmentabstractAutomatically generated radiology reports often receive high scores from existing evaluation metrics but fail to earn clinicians' trust.This gap reveals fundamental flaws in how current metrics assess the quality of generated reports.We rethink the design and evaluation of these metrics and propose a clinically grounded Meta-Evaluation framework.We define clinically grounded criteria spanning clinical alignment and key metric capabilities, including discrimination, robustness, and monotonicity.Using a fine-grained dataset of ground truth and rewritten report pairs annotated with error types, clinical significance labels, and explanations, we systematically evaluate existing metrics and reveal their limitations in interpreting clinical semantics, such as failing to distinguish clinically significant errors, over-penalizing harmless variations, and lacking consistency across error severity levels.Our framework offers guidance for building more clinically reliable evaluation methods. Bailiang Jian, Youxiang Zhu |
EMNLP | 5 |
| 2025 | Unveil Multi-Picture Descriptions for Multilingual Mild Cognitive Impairment Detection via Contrastive LearningabstractDetecting Mild Cognitive Impairment from picture descriptions is critical yet challenging, especially in multilingual and multiple picture settings. Prior work has primarily focused on English speakers describing a single picture (e.g., the 'Cookie Theft'). The TAUKDIAL-2024 challenge expands this scope by introducing multilingual speakers and multiple pictures, which presents new challenges in analyzing picture-dependent content. To address these challenges, we propose a framework with three components: (1) enhancing discriminative representation learning via supervised contrastive learning, (2) involving image modality rather than relying solely on speech and text modalities, and (3) applying a Product of Experts (PoE) strategy to mitigate spurious correlations and overfitting. Our framework improves MCI detection performance, achieving a +7.1% increase in Unweighted Average Recall (UAR) (from 68.1% to 75.2%) and a +2.9% increase in F1 score (from 80.6% to 83.5%) compared to the text unimodal baseline. Notably, the contrastive learning component yields greater gains for the text modality compared to speech. These results highlight our framework's effectiveness in multilingual and multi-picture MCI detection. Kristin Qi, Jiali Cheng, Youxiang Zhu, Hadi Amiri, Xiaohui Liang 0002 |
GLOBECOM | 3 |
| 2025 | Cog-TiPRO: Iterative Prompt Refinement with LLMs to Detect Cognitive Decline via Longitudinal Voice Assistant CommandsabstractEarly detection of cognitive decline is crucial for enabling interventions that can slow neurodegenerative disease progression. Traditional diagnostic approaches rely on labor-intensive clinical assessments, which are impractical for frequent monitoring. Our pilot study investigates voice assistant systems (VAS) as non-invasive tools for detecting cognitive decline through longitudinal analysis of speech patterns in short and unstructured voice commands. Over an 18-month period, we collected voice commands from 35 older adults, with 15 participants providing daily at-home VAS interactions. To address the challenges of analyzing these short, unstructured and noisy commands, we propose Cog-TiPRO, a framework that combines (1) LLM-driven iterative prompt refinement for linguistic feature extraction, (2) HuBERT-based acoustic feature extraction, and (3) transformer-based temporal modeling. Using iTransformer, our approach achieves 73.80% accuracy and 72.67% F1-score in detecting MCI, outperforming its baseline by 27.13%. Through our LLM approach, we identify linguistic features that uniquely characterize everyday command usage patterns in individuals experiencing cognitive decline. Kristin Qi, Youxiang Zhu, Caroline Summerour, John A. Batsis, Xiaohui Liang 0002 |
GLOBECOM | 2 |
| 2024 | Adversarial Text Generation using Large Language Models for Dementia DetectionabstractAlthough large language models (LLMs) excel in various text classification tasks, regular prompting strategies (e.g., few-shot prompting) do not work well with dementia detection via picture description.The challenge lies in the language marks for dementia are unclear, and LLM may struggle with relating its internal knowledge to dementia detection.In this paper, we present an accurate and interpretable classification approach by Adversarial Text Generation (ATG), a novel decoding strategy that could relate dementia detection with other tasks.We further develop a comprehensive set of instructions corresponding to various tasks and use them to guide ATG, achieving the best accuracy of 85%, >10% improvement compared to the regular prompting strategies.In addition, we introduce feature context, a humanunderstandable text that reveals the underlying features of LLM used for classifying dementia.From feature contexts, we found that dementia detection can be related to tasks such as assessing attention to detail, language, and clarity with specific features of the environment, character, and other picture content or languagerelated features.Future work includes incorporating multi-modal LLMs to interpret speech and picture information.* The stool is falling over, which suggests that it's either old, unstable, or has been poorly maintained.This adds to the sense of chaos in the kitchen.* The cookies being taken out of the cookie jar by the boy and being asked by the little girl to be handed to her quietly implies that snacking is a common activity in this household.* The fact that the stool is falling while the cookies are being handled increases the risk of mess and potential accidents.**Sink Overflowing** * The sink is overflowing, which means the mother is either taking a long time to wash dishes or isn't paying attention to the water level.This could be a source of stress or annoyance for her, but for now, she's choosing to dry dishes (on the assumption that it's less pressing).* This detail also creates a sense of water-themed tension, with water running "into" and possibly outside, but no indication of what happens to it.**Wind Outside** * The curtains are drawn back to allow a gentle breeze into the house, and the scene appears to be on a relatively mild day, as you mentioned.The fact that you explicitly state there's nothing to note about the outside action (beyond the path, window, and building envelope) suggests that the wind is likely a gentle, harmless blow.**Kids' Actions** * The girl's finger to the mouth could be seen as a clever play on innocence, hiding mischief behind a cookie-stealing facade.Her reach for a cookie and the brother's slippage from the stool create a sense of partnership in their naughtiness.* Both kids seem comfortable snacking in the kitchen, almost like it's a normalized behavior.**Mother's Neglect (Daycare... er, laundry of attention?)*** You cleverly contrast the overflowing sink with the mother's drying of dishes, implying a degree of obliviousness or prioritization on her part.Her feet getting soaked or compromised hygiene aren't necessarily consequences of her actions, but it does add to the impression of a household where multiple tasks are happening simultaneously.**Summer Puddled Insight** * Your aside about not knowing whether you meant for it to be action-worthy mentions several extraneous details, but effectively highlights the fact that the outside scene doesn't quite gel with the chaotic interior.You raise an interesting question about the importance of showing consequences or actions related to the stool tipping, overflowing water, or blowing curtains.* The comment about shortsightedly standing in the mother's attention ("um the mother or don't see much.um mother wiping dishes ") serves as a wry commentary on the scene's pandemonium.Overall, this vivid scene showcases a busy, potentially messy household where multiple dynamics are at play.The juxtaposition of controlled chaos (curtains flowing, kids snacking) with underlying mayhem (stool tipping, overflowing sink) creates a engaging contrast.The outside weather serves as a brief interlude, acknowledging that not everything has to be resolved or escalating inside Youxiang Zhu, Nana Lin, Kiran Balivada, Daniel Haehn, Xiaohui Liang 0002 |
EMNLP | 1 |
| 2024 | Exploiting Privacy Preserving Prompt Techniques for Online Large Language Model UsageabstractOnline Large Language Models (LLMs) are widely employed across various tasks, including privacy-sensitive ones like financial advice or paragraph rewriting. Presently, users directly submit prompts to online LLM servers, inadvertently revealing sensitive keywords and facilitating server tracking to build user profiles. In this paper, we propose a local privacy-preserving prompt assistant (LPPA) that provides users with a usable method to balance privacy in the prompts and the utility of the LLM output. The LPPA will analyze the users' prompts, suggest modifying the prompts to protect the sensitive keywords, and provide an inference of the potential utility impact of the online LLM output. Specifically, we first propose a privacy module to identify the sensitive keywords in the prompt and adopt four privacy techniques, including remove, mask, replace, and rewrite to hide the keywords. While these techniques affect the utility of online LLM output, we measure such impact using the LLM output of the original prompt and modified prompts and discuss the cases with high, median, and low impact. In addition, we propose a utility inference model to infer the utility impact locally without disclosing the prompts to the online LLM. We evaluated LPPA on the real-world users' prompts and showed that the remove technique achieves the best performance, and it empowers users with meaningful ways to adjust their prompts to safeguard their privacy while still maintaining a satisfactory level of utility in online LLM usage. Youxiang Zhu, Xiaohui Liang 0002, Honggang Zhang 0003 |
GLOBECOM | 1 |
| 2024 | Analyzing Multimodal Features of Spontaneous Voice Assistant Commands for Mild Cognitive Impairment DetectionabstractMild cognitive impairment (MCI) is a major public health concern due to its high risk of progressing to dementia. This study investigates the potential of detecting MCI with spontaneous voice assistant (VA) commands from 35 older adults in a controlled setting. Specifically, a command-generation task is designed with pre-defined intents for participants to freely generate commands that are more associated with cognitive ability than read commands. We develop MCI classification and regression models with audio, textual, intent, and multimodal fusion features. We find the command-generation task outperforms the command-reading task with an average classification accuracy of 82%, achieved by leveraging multimodal fusion features. In addition, generated commands correlate more strongly with memory and attention subdomains than read commands. Our results confirm the effectiveness of the command-generation task and imply the promise of using longitudinal in-home commands for MCI detection. Nana Lin, Youxiang Zhu, Xiaohui Liang 0002, John A. Batsis, Caroline Summerour |
INTERSPEECH | 2 |
| 2023 | Early Detection of Cognitive Decline Using Voice Assistant CommandsabstractEarly detection of Alzheimer's Disease and Related Dementias (ADRD) is critical in treating the progression of the disease. Previous studies have shown that ADRD can be detected and classified using machine learning models trained on samples of spontaneous speech. We propose using Voice-Assistant Systems (VAS), e.g., Amazon Alexa, to monitor and collect data from at-risk adults, and we show that this data can be used to achieve functional accuracy in classifying their cognitive status. In this paper, we develop multiple unique feature sets from VAS data that can be used in the training of machine learning models. We then perform multi-class classification, binary classification, and regression using these features on our dataset of older adults with three varying stages of cognitive decline interacting with VAS. Our results show that the VAS data can be used to classify Dementia (DM), Mild Cognitive Impairment (MCI), and Healthy Control (HC) participants with an accuracy up to 74.7%, and classify between HC and MCI with accuracy up to 62.8%. Eli Kurtz, Youxiang Zhu, Tiffany M. Driesse, Bang Tran, John A. Batsis, Robert M. Roth, Xiaohui Liang 0002 |
ICASSP | 2 |
| 2023 | Exploiting Relevance of Speech to Sleepiness Detection via Attention MechanismabstractExcessive sleepiness in critical tasks and jobs can lead to adverse outcomes, such as work accidents and car crashes. Detecting and monitoring sleepiness levels can prevent these adverse events from happening. In this paper, we propose an attention-based sleepiness detection method using HuBERT embeddings and eGeMAPS features of human speech. Specifically, we propose an attention-based convolutional neural network (CNN) model that achieves accurate 82.57 % sleepiness detection using HuBERT embeddings plus age and gender as inputs. We also show that the embedded attention layers significantly improve the detection accuracy in different cases of inputs. We then explore the attention weights from the attention layers and observe that the long and semantically-different responses from “Picture description”, “Microphone test”, and “Free speech” tasks are more relevant to sleepiness detection when the model is trained with HuBERT only; the short and semantically-similar responses from “Sustained phonation” and “Diadochokinetic” tasks are more relevant when trained with HuBERT plus age and gender. The attention mechanism enables our model to take all responses as one input, simplifying the data pre-processing and identifying the relevant speech responses to sleepiness detection. Bang Tran, Youxiang Zhu, James W. Schwoebel, Xiaohui Liang 0002 |
ICC | 2 |
| 2022 | Speech Tasks Relevant to Sleepiness Determined With Deep Transfer LearningabstractExcessive sleepiness in attention-critical contexts can lead to adverse events, such as car crashes. Detecting and monitoring sleepiness can help prevent these adverse events from happening. In this paper, we use the Voiceome dataset to extract speech from 1,828 participants to develop a deep transfer learning model using Hidden-Unit BERT (HuBERT) speech representations to detect sleepiness from individuals. Speech is an under-utilized source of data in sleep detection, but as speech collection is easy, cost-effective, and non-invasive, it provides a promising resource for sleepiness detection. Two complementary techniques were conducted in order to seek converging evidence regarding the importance of individual speech tasks. Our first technique, masking, evaluated task importance by combining all speech tasks, masking selected responses in the speech, and observing systematic changes in model accuracy. Our second technique, separate training, compared the accuracy of multiple models, each of which used the same architecture, but was trained on a different subset of speech tasks. Our evaluation shows that the best-performing model utilizes the memory recall task and categorical naming task from the Boston Naming Test, which achieved an accuracy of 80.07% (F1-score of 0.85) and 81.13% (F1-score of 0.89), respectively. Bang Tran, Youxiang Zhu, Xiaohui Liang 0002, James W. Schwoebel, Lindsay A. Warrenburg |
ICASSP | 2 |
| 2022 | Towards Interpretability of Speech Pause in Dementia Detection Using Adversarial LearningabstractSpeech pause is an effective biomarker in dementia detection. Recent deep learning models have exploited speech pauses to achieve highly accurate dementia detection, but have not exploited the interpretability of speech pauses, i.e., what and how positions and lengths of speech pauses affect the result of dementia detection. In this paper, we will study the positions and lengths of dementia-sensitive pauses using adversarial learning approaches. Specifically, we first utilize an adversarial attack approach by adding the perturbation to the speech pauses of the testing samples, aiming to reduce the confidence levels of the detection model. Then, we apply an adversarial training approach to evaluate the impact of the perturbation in training samples on the detection model. We examine the interpretability from the perspectives of model accuracy, pause context, and pause length. We found that some pauses are more sensitive to dementia than other pauses from the model's perspective, e.g., speech pauses near to the verb "is". Increasing lengths of sensitive pauses or adding sensitive pauses leads the model inference to Alzheimer's Disease (AD), while decreasing the lengths of sensitive pauses or deleting sensitive pauses leads to non-AD. Youxiang Zhu, Bang Tran, Xiaohui Liang 0002, John A. Batsis, Robert M. Roth |
ICASSP | 1 |
| 2022 | Domain-aware Intermediate Pretraining for Dementia Detection with Limited DataabstractDetecting dementia using human speech is promising but faces a limited data challenge. While recent research has shown general pretrained models (e.g., BERT) can be applied to improve dementia detection, the pretrained model can hardly be fine-tuned with the available small dementia dataset as that would raise the overfitting problem. In this paper, we propose a domain-aware intermediate pretraining to enable a pretraining process using a domain-similar dataset that is selected by incorporating the knowledge from the dementia dataset. Specifically, we use pseudo-perplexity to find an effective pretraining dataset, and then propose dataset-level and sample-level domain-aware intermediate pretraining techniques. We further employ information units (IU) from previous dementia research and define an IU-pseudo-perplexity to reduce calculation complexity. We confirm the effectiveness of perplexity by showing a strong correlation between perplexity and accuracy using 9 datasets and models from the GLUE benchmark. We show that our domain-aware intermediate pretraining improves detection accuracy in almost all cases. Our results suggested that the difference in text-based perplexity values between patients with Alzheimer's Disease (AD) and Healthy Control (HC) is still small, and the perplexity incorporating acoustic features (e.g., pause) may make the pretraining more effective. Youxiang Zhu, Xiaohui Liang 0002, John A. Batsis, Robert M. Roth |
INTERSPEECH | 1 |
| 2022 | Evaluating voice-assistant commands for dementia detection
Xiaohui Liang 0002, John A. Batsis, Youxiang Zhu, Tiffany M. Driesse, Robert M. Roth, David Kotz, Brian MacWhinney |
Comput. Speech Lang. | 3 |
| 2021 | WavBERT: Exploiting Semantic and Non-Semantic Speech Using Wav2vec and BERT for Dementia DetectionabstractIn this paper, we exploit semantic and non-semantic information from patient's speech data using Wav2vec and Bidirectional Encoder Representations from Transformers (BERT) for dementia detection. We first propose a basic WavBERT model by extracting semantic information from speech data using Wav2vec, and analyzing the semantic information using BERT for dementia detection. While the basic model discards the non-semantic information, we propose extended WavBERT models that convert the output of Wav2vec to the input to BERT for preserving the non-semantic information in dementia detection. Specifically, we determine the locations and lengths of inter-word pauses using the number of blank tokens from Wav2vec where the threshold for setting the pauses is automatically generated via BERT. We further design a pre-trained embedding conversion network that converts the output embedding of Wav2vec to the input embedding of BERT, enabling the fine-tuning of WavBERT with non-semantic information. Our evaluation results using the ADReSSo dataset showed that the WavBERT models achieved the highest accuracy of 83.1% in the classification task, the lowest Root-Mean-Square Error (RMSE) score of 4.44 in the regression task, and a mean F1 of 70.91% in the progression task. We confirmed the effectiveness of WavBERT models exploiting both semantic and non-semantic speech. Youxiang Zhu, Abdelrahman Obyat, Xiaohui Liang 0002, John A. Batsis, Robert M. Roth |
Interspeech | 1 |
| 2020 | Learning Cascade Attention for fine-grained image classification
Youxiang Zhu, Yin Yang 0002, Ning Ye 0001 |
Neural Networks | 1 |
| 2019 | TA-CNN: Two-way attention models in deep convolutional neural network for plant recognition
Youxiang Zhu, Weiming Sun, Xiangying Cao, Chunyan Wang 0018, Dongyang Wu, Yin Yang 0002, Ning Ye 0001 |
Neurocomputing | 1 |