Xiaotian Zou

dblp:133/8094 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0006-8681-1358ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Collaborative Dialogue Analysis for Productive Problem Solving
abstract
Collaborative problem solving requires students to jointly reason, negotiate, and regulate their learning. Understanding collaborative problem solving through student dialogue can inform timely identification of productive and unproductive collaborative behaviors. In this study, we investigate the use of large language models to automatically classify collaborative problem-solving dialogue segments into two categories: Productive and Unproductive. To support deeper analysis, we additionally explore classification of eight detailed collaborative problem-solving sub-categories. We present an error-augmented few-shot prompting method that incorporates misclassified examples to refine model understanding of classification boundaries. Using dialogue data from a middle school collaborative game-based learning environment, our approach substantially improves classification accuracy over zero-shot baselines. Qualitative analysis of the resulting models further highlights which dialogue types are most frequently misclassified, suggesting design implications for adaptive scaffolding. These findings demonstrate that large language models, when guided with targeted prompting strategies, can effectively recognize productive and unproductive dialogue in collaborative learning.
Yeo Jin Kim, Daeun Hong, Xiaotian Zou, Cindy E. Hmelo-Silver, Wookhee Min, Snigdha Chaturvedi, James C. Lester
LAK3
2025 Refusal-Aware Red Teaming: Exposing Inconsistency in Safety Evaluations
abstract
The responsible deployment of Large Language Models (LLMs) necessitates rigorous safety evaluations.However, a critical challenge arises from inconsistencies between an LLM's internal refusal decisions and external safety assessments, hindering effective validation.This paper introduces the concept of the 'refusal gap' to formally define these discrepancies.We then present a novel, refusal-aware red teaming framework designed to automatically generate test cases that expose such gaps.Our framework employs 'refusal probes', which leverage the target model's hidden states, to detect internal model refusals.These are subsequently contrasted with judgments from an external safety evaluator.The identified discrepancy serves as a signal to guide a red-teaming model in crafting test cases that maximize this refusal gap.To further enhance test case diversity and address challenges related to sparse rewards, we introduce a hierarchical, curiositydriven mechanism that incentivizes both refusal gap maximization and broad topic exploration.Empirical results demonstrate that our method significantly outperforms existing reinforcement learning-based approaches in generating diverse test cases and achieves a substantially higher discovery rate of refusal gaps.
Xiaohu Du, Xiaotian Zou, Chongyang Zhao 0004, Xiaohui Kuang
EMNLP3
2025 FlowJD: Your Imagination Can Help You Jailbreak in Visual Language Models
abstract
Large Visual Language Models (VLMs), such as GPT-4V, have achieved impressive results in generating detailed and nuanced responses. Although researchers have proposed various benchmarks to evaluate VLM performance, they have often neglected the examination of inherent security capabilities, particularly by evaluating the logical comprehension of image information. To address this gap, this paper introduces a novel dataset, FlowJD, specifically designed to evaluate logical flowchart jailbreak capabilities in VLMs. We conduct a comprehensive evaluation on GPT-4o, GPT-4V, and seven other state-of-the-art VLMs, revealing jailbreak rates of up to 92.8%. Our findings reveal significant vulnerabilities in current VLMs concerning logical flowchart jailbreak, emphasizing the urgent need for robust and effective defenses in future VLM development.Warning: Some of the examples may be harmful!
Xiaotian Zou, Qianqian Han, Ke Li 0001
ICME1
2025 A Multimodal Classroom Video Question-Answering Framework for Automated Understanding of Collaborative Learning
Nithin Sivakumaran, Chia-Yu Yang, Abhaysinh Zala, Shoubin Yu, Daeun Hong, Xiaotian Zou, Elias Stengel-Eskin, Dan Carpenter, Wookhee Min, Cindy E. Hmelo-Silver, Jonathan P. Rowe, James C. Lester, Mohit Bansal
ICMI6
2025 Reward-Guided Many-Shot Jailbreaking
Xiaotian Zou, Tong Wang 0042, Jianwen Tian, Xiaohui Kuang
NLPCC (1)2
2025 Evolutionary multitasking with evolutionary trend alignment in subdomains
Wenhao Du, Jack Cole, Xiaotian Zou, Chaowen Wang
Expert Syst. Appl.4
2022 Secure and Private Coding for Edge Computing Against Cooperative Attack with Low Communication Cost and Computational Load
Xiaotian Zou, Jin Wang 0009, Lingzhi Li 0001, Fei Gu 0001, Guojing Li
CollaborateCom (1)1
2021 Causality extraction based on self-attentive BiLSTM-CRF with transferred embeddings
Zhaoning Li, Xiaotian Zou, Jiangtao Ren
Neurocomputing3
2019 Towards Helping Teachers Select Optimal Content for Students
Xiaotian Zou, Zhenjun Ma, Ryan Baker 0001
AIED (2)1