Zhuoxuan Jiang

dblp:183/2883 · DBLP profile ↗
← Back
22ranked-venue papers
10as first author
13since 2021 · last 2026
0009-0009-9517-1091ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 7 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GeoProblem Factory: A Visual Interaction System for Solvable and Controllable Geometric Problem Generation by Leveraging Symbolic Deduction Engine
abstract
We propose a novel system, GeoProblem Factory, designed to effectively generate high-quality geometry problems for intelligent education. The system enables to efficiently produce batches of geometry problems for teachers and students, either to save time and manual effort or to support personalized learning. Generating geometry problems is particularly challenging, as it requires ensuring both solvability and controllability from a pedagogical perspective. To address these issues, we adopt a state-of-the-art pipeline method based on a symbolic deduction engine and develop a visual interaction demo. This demo allows users to easily refine the generated problems through visual operations. It provides two modes for inputting controllable information: specifying knowledge points or supplying a reference problem. Moreover, the system can automatically generate a preliminary geometric diagram corresponding to each problem for further refinement. Through human–machine interaction, the system can more efficiently produce high-quality geometry problems than ever.
Zhuoxuan Jiang, Tianyang Zhang 0004, Mo Guang, Wen Si
AAAI1
2026 MERGE-PAG: Agent-based multimodal knowledge extraction and reasoning framework for pilot-action graph
Tiance Yang, Shanshan Feng 0001, Zhuoxuan Jiang, Zhensheng Zhang, Fan Li 0015
Adv. Eng. Informatics3
2026 Towards enterprise-specific question-answering for IT operations and maintenance based on retrieval-augmented generation mechanism
Zhuoxuan Jiang, Tianyang Zhang 0004, Shengguang Bai, Haotian Zhang 0029, Yinong Xun, Wen Si
Expert Syst. Appl.1
2025 MathMistake Checker: A Comprehensive Demonstration for Step-by-Step Math Problem Mistake Finding by Prompt-Guided LLMs
abstract
We propose a novel system, MathMistake Checker, designed to automate step-by-step mistake finding in mathematical problems with lengthy answers through a two-stage process. The system aims to simplify grading, increase efficiency, and enhance learning experiences from a pedagogical perspective. It integrates advanced technologies, including computer vision and the chain-of-thought capabilities of the latest large language models (LLMs). Our system supports open-ended grading without reference answers and promotes personalized learning by providing targeted feedback. We demonstrate its effectiveness across various types of math problems, such as calculation and word problems.
Tianyang Zhang 0004, Zhuoxuan Jiang, Haotian Zhang 0029
AAAI2
2025 LLMA4ITOps: A Lightweight LLM-Based Multi-Agent Framework for IT Operations and Maintenance
Zhuoxuan Jiang, Tianyang Zhang 0004, Haotian Zhang 0029, Yinong Xun, Yang Liu 0218, Dehua Feng, Wen Si
NLPCC (2)1
2025 A New Era in Human Factors Engineering: A Survey of the Applications and Prospects of Large Multimodal Models
abstract
In recent years, the potential applications of Large Multimodal Models (LMMs) in fields, such as healthcare, social psychology, and industrial design have attracted wide research attention, providing new directions for human factors research. For instance, LMM-based smart systems have become novel research subjects of human factors studies, and LMM introduces new research paradigms and methodologies to this field. Therefore, this article aims to explore the applications, challenges, and future prospects of LMM in the domain of human factors and ergonomics through an expert-LMM collaborated literature and patent review. Specifically, this article proposes a novel review method and introduces research studies and patents related to LMM-based accident analysis, human modeling, and intervention design. Subsequently, the article discusses future trends in research paradigm and challenges of human factors and ergonomics studies in the era of LMMs. It is expected that the review results offer valuable insights and serve as a reference for using LMMs in human factors research.
Fan Li 0015, Su Han, Ching-Hung Lee, Shanshan Feng 0001, Zhuoxuan Jiang, Zhu Sun 0001
Int. J. Hum. Comput. Interact.5
2024 LLMs Can Find Mathematical Reasoning Mistakes by Pedagogical Chain-of-Thought
Zhuoxuan Jiang, Haoyuan Peng, Shanshan Feng 0001, Fan Li 0015, Dongsheng Li 0002
IJCAI1
2023 Grafting Fine-Tuning and Reinforcement Learning for Empathetic Emotion Elicitation in Dialog Generation
abstract
For human-like dialogue systems, it is significant to inject the empathetic ability or elicit the opposite’s positive emotions, while existing studies mostly only focus on either of the above two research lines. In this work, we propose a novel and grafted task named Empathetic Emotion Elicitation Dialog to make a dialog system able to possess both aspects of ability simultaneously. We do not train an empathetic dialog system and an emotion elicitation dialog system separately and then simply concatenate the responses generated by these two systems, which will cause illogical and repetitive responses. Instead, we propose a unified solution: (1) To generate empathetic responses and emotion elicitation responses within the same semantic space, we design a unified framework. (2) The unified framework has three stages which first retrieve the empathetic and emotion elicitation exemplars as external knowledge, then fine-tune the emotion/action prediction on a pre-trained language model to enhance the empathetic ability, and finally model the user feedback by reinforcement learning to enhance the emotion elicitation ability. Experiments show that our method outperforms the baselines in the response generation quality and simultaneously empathizes with the user and elicits their positive emotions.
Bo Wang 0011, Zhuoxuan Jiang, Ruifang He, Yuexian Hou
ECAI5
2022 Gated Mechanism Enhanced Multi-Task Learning for Dialog Routing
abstract
Currently, human-bot symbiosis dialog systems, e.g. pre- and after-sales in E-commerce, are ubiquitous, and the dialog routing component is essential to improve the overall efficiency, reduce human resource cost and increase user experience. To satisfy this requirement, existing methods are mostly heuristic and cannot obtain high-quality performance. In this paper, we investigate the important problem by thoroughly mining both the data-to-task and task-to-task knowledge among various kinds of dialog data. To achieve the above target, we propose a comprehensive and general solution with multi-task learning framework, specifically including a novel dialog encoder and two tailored gated mechanism modules. The proposed Gated Mechanism enhanced Multi-task Model (G3M) can play the role of hierarchical information filtering and is non-invasive to the existing dialog systems. Experiments on two datasets collected from the real world demonstrate our method’s effectiveness and the results achieve the state-of-the-art performance by relatively increasing 8.7%/11.8% on RMSE metric and 2.2%/4.4% on F1 metric.
Ziming Huang, Zhuoxuan Jiang, Shanshan Feng 0001, Xianling Mao
COLING2
2022 OS-MSL: One Stage Multimodal Sequential Link Framework for Scene Segmentation and Classification
abstract
Scene segmentation and classification (SSC) serve as a critical step towards the field of video structuring analysis. Intuitively, jointly learning of these two tasks can promote each other by sharing common information. However, scene segmentation concerns more on the local difference between adjacent shots while classification needs the global representation of scene segments, which probably leads to the model dominated by one of the two tasks in the training phase. In this paper, from an alternate perspective to overcome the above challenges, we unite these two tasks into one task by a new form of predicting shots link: a link connects two adjacent shots, indicating that they belong to the same scene or category. To the end, we propose a general One Stage Multimodal Sequential Link Framework (OS-MSL) to both distinguish and leverage the two-fold semantics by reforming the two learning tasks into a unified one. Furthermore, we tailor a specific module called DiffCorrNet to explicitly extract the information of differences and correlations among shots. Extensive experiments on a brand-new large scale dataset collected from real-world applications, and MovieScenes are conducted. Both the results demonstrate the effectiveness of our proposed method against strong baselines. The code is made available.
Ye Liu 0013, Lingfeng Qiao, Zhuoxuan Jiang, Xinghua Jiang, Deqiang Jiang, Bo Ren 0002
ACM Multimedia4
2022 RAAT: Relation-Augmented Attention Transformer for Relation Modeling in Document-Level Event Extraction
abstract
In document-level event extraction (DEE) task, event arguments always scatter across sentences (across-sentence issue) and multiple events may lie in one document (multi-event issue).In this paper, we argue that the relation information of event arguments is of great significance for addressing the above two issues, and propose a new DEE framework which can model the relation dependencies, called Relation-augmented Document-level Event Extraction (ReDEE).More specifically, this framework features a novel and tailored transformer, named as Relation-augmented Attention Transformer (RAAT).RAAT is scalable to capture multi-scale and multi-amount argument relations.To further leverage relation information, we introduce a separate event relation prediction task and adopt multi-task learning method to explicitly enhance event extraction performance.Extensive experiments demonstrate the effectiveness of the proposed method, which can achieve state-ofthe-art performance on two public datasets.
Zhuoxuan Jiang, Bo Ren 0002
NAACL-HLT2
2021 Dialog Router: Automated Dialog Transition via Multi-Task Learning
abstract
Dialog Router is a general paradigm for human-bot symbiosis dialog systems to provide friendly customer care service. It is equipped with a multi-task learning model to automatically capture the underlying correlation between multiple related tasks, i.e. dialog classification and regression, and greatly reduce human labor work for system customization, which improves the accuracy of dialog transition. In addition, for learning the multi-task model, the training data and labels are easy to collect from human-to-human historical dialog logs, and the Dialog Router can be easily integrated into the majority of existing dialog systems by calling general APIs. We conduct experiments on real-world datasets for dialog classification and regression. The results show that our model achieves improvements on both tasks, which benefits the dialog transition application. The demo illustrates our method’s effectiveness in a real customer care service.
Ziming Huang, Zhuoxuan Jiang, Xue Han 0018, Yabin Dang
AAAI2
2021 Leveraging Tripartite Interaction Information from Live Stream E-Commerce for Improving Product Recommendation
abstract
Recently, a new form of online shopping becomes more and more popular, which combines live streaming with E-Commerce activity. The streamers introduce products and interact with their audiences, and hence greatly improve the performance of selling products. Despite of the successful applications in industries, the live stream E-commerce has not been well studied in the data science community. To fill this gap, we investigate this brand-new scenario and collect a real-world Live Stream E-Commerce (LSEC) dataset. Different from conventional E-commerce activities, the streamers play a pivotal role in the LSEC events. Hence, the key is to make full use of rich interaction information among streamers, users, and products. We first conduct data analysis on the tripartite interaction data and quantify the streamer's influence on users' purchase behavior. Based on the analysis results, we model the tripartite information as a heterogeneous graph, which can be decomposed to multiple bipartite graphs in order to better capture the influence. We propose a novel Live Stream E-Commerce Graph Neural Network framework (LSEC-GNN) to learn the node representations of each bipartite graph, and further design a multi-task learning approach to improve product recommendation. Extensive experiments on two real-world datasets with different scales show that our method can significantly outperform various baseline approaches.
Sanshi Yu, Zhuoxuan Jiang, Shanshan Feng 0001, Dongsheng Li 0002, Qi Liu 0003, Jinfeng Yi
KDD2
2020 When and Who? Conversation Transition Based on Bot-Agent Symbiosis Learning Network
abstract
In online customer service applications, multiple chatbots that are specialized in various topics are typically developed separately and are then merged with other human agents to a single platform, presenting to the users with a unified interface.Ideally the conversation can be transparently transferred between different sources of customer support so that domain-specific questions can be answered timely and this is what we coined as a Bot-Agent symbiosis.Conversation transition is a major challenge in such online customer service and our work formalises the challenge as two core problems, namely, when to transfer and which bot or agent to transfer to and introduces a deep neural networks based approach that addresses these problems.Inspired by the net promoter score (NPS), our research reveals how the problems can be effectively solved by providing user feedback and developing deep neural networks that predict the conversation category distribution and the NPS of the dialogues.Experiments on realistic data generated from an online service support platform demonstrate that the proposed approach outperforms state-of-the-art methods and shows promising perspective for transparent conversation transition.
Yipeng Yu, Ran Guan, Zhuoxuan Jiang, Jingchang Huang
COLING4
2019 A General Planning-Based Framework for Goal-Driven Conversation Assistant
abstract
We propose a general framework for goal-driven conversation assistant based on Planning methods. It aims to rapidly build a dialogue agent with less handcrafting and make the more interpretable and efficient dialogue management in various scenarios. By employing the Planning method, dialogue actions can be efficiently defined and reusable, and the transition of the dialogue are managed by a Planner. The proposed framework consists of a pipeline of Natural Language Understanding (intent labeler), Planning of Actions (with a World Model), and Natural Language Generation (learned by an attention-based neural network). We demonstrate our approach by creating conversational agents for several independent domains.
Zhuoxuan Jiang, GuangYuan Yu, Yipeng Yu, Shaochun Li
AAAI1
2019 Towards Automated Planning for Enterprise Services: Opportunities and Challenges
Maja Vukovic, Scott N. Gerard, Richard Hull 0001, Michael Katz 0001, Larisa Shwartz, Shirin Sohrabi, Christian J. Muise, John J. Rofrano, Anup K. Kalia, Jinho Hwang, Yabin Dang, Zhuoxuan Jiang
ICSOC13
2019 Towards End-to-End Learning for Efficient Dialogue Agent by Modeling Looking-ahead Ability
abstract
Learning an efficient manager of dialogue agent from data with little manual intervention is important, especially for goal-oriented dialogues.However, existing methods either take too many manual efforts (e.g.reinforcement learning methods) or cannot guarantee the dialogue efficiency (e.g.sequence-to-sequence methods).In this paper, we address this problem by proposing a novel end-to-end learning model to train a dialogue agent that can look ahead for several future turns and generate an optimal response to make the dialogue efficient.Our method is data-driven and does not require too much manual work for intervention during system design.We evaluate our method on two datasets of different scenarios and the experimental results demonstrate the efficiency of our model.
Zhuoxuan Jiang, Xianling Mao, Ziming Huang, Shaochun Li
SIGdial1
2017 A Novel Cascade Model for Learning Latent Similarity from Heterogeneous Sequential Data of MOOC
abstract
Recent years have witnessed the proliferation of Massive Open Online Courses (MOOCs).With massive learners being offered MOOCs, there is a demand that the forum contents within MOOCs need to be classified in order to facilitate both learners and instructors.Therefore we investigate a significant application, which is to associate forum threads to subtitles of video clips.This task can be regarded as a document ranking problem, and the key is how to learn a distinguishable text representation from word sequences and learners' behavior sequences.In this paper, we propose a novel cascade model, which can capture both the latent semantics and latent similarity by modeling MOOC data.Experimental results on two real-world datasets demonstrate that our textual representation outperforms state-of-the-art unsupervised counterparts for the application.
Zhuoxuan Jiang, Shanshan Feng 0001, Gao Cong, Chunyan Miao, Xiaoming Li 0001
EMNLP1
2017 Hierarchical Mixed Neural Network for Joint Representation Learning of Social-Attribute Network
Weizheng Chen, Jinpeng Wang 0001, Zhuoxuan Jiang, Yan Zhang 0004, Xiaoming Li 0001
PAKDD (1)3
2017 Unsupervised Embedding for Latent Similarity by Modeling Heterogeneous MOOC Data
Zhuoxuan Jiang, Shanshan Feng 0001, Weizheng Chen, Guangtao Wang, Xiaoming Li 0001
PAKDD (2)1
2016 Generating Semantic Concept Map for MOOCs
Zhuoxuan Jiang, Peng Li 0030, Yan Zhang 0004, Xiaoming Li 0001
EDM1
2015 Influence Analysis by Heterogeneous Network in MOOC Forums: What can We Discover?
Zhuoxuan Jiang, Yan Zhang 0004, Xiaoming Li 0001
EDM1