Xiaorui Wu

dblp:175/8416 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Computer networks · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Trustworthy machine learning · 33% Language models and text generation · 33% Efficient and distributed learning · 33%
Computer networks
1 paper
Routing and switching · 77% Network optimization and economics · 23%
Theoretical computer science
1 paper
Graph algorithms and graph theory · 100%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model
0.912025
TRIDENT: Enhancing Large Language Model Safety with Tri-Dimensional Diversified Red-Teaming Data Synthesis · ACL (1) 2025
Natural language and speech › Language models and text generation
large language model safety
0.912025
TRIDENT: Enhancing Large Language Model Safety with Tri-Dimensional Diversified Red-Teaming Data Synthesis · ACL (1) 2025
Machine learning › Trustworthy machine learning › safety evaluation
red teaming
0.912025
TRIDENT: Enhancing Large Language Model Safety with Tri-Dimensional Diversified Red-Teaming Data Synthesis · ACL (1) 2025
Machine learning › Trustworthy machine learning
robustness
0.912025
TRIDENT: Enhancing Large Language Model Safety with Tri-Dimensional Diversified Red-Teaming Data Synthesis · ACL (1) 2025
Machine learning › Efficient and distributed learning › distributed training › communication-efficient training
communication optimization
0.612022
Stanza: Layer Separation for Distributed Training in Deep Learning · IEEE Trans. Serv. Comput. 2022
Machine learning › Efficient and distributed learning
distributed training
0.612022
Stanza: Layer Separation for Distributed Training in Deep Learning · IEEE Trans. Serv. Comput. 2022
Machine learning › Efficient and distributed learning › distributed training › distributed training systems
parameter server
0.612022
Stanza: Layer Separation for Distributed Training in Deep Learning · IEEE Trans. Serv. Comput. 2022
Routing and switching
traffic engineering
0.412019
Node-Constrained Traffic Engineering: Theory and Applications · IEEE/ACM Trans. Netw. 2019
Graph algorithms and graph theory
graph algorithms
0.412019
Node-Constrained Traffic Engineering: Theory and Applications · IEEE/ACM Trans. Netw. 2019
Security and privacy of machine learning
adversarial attack
0.312025
TRIDENT: Enhancing Large Language Model Safety with Tri-Dimensional Diversified Red-Teaming Data Synthesis · ACL (1) 2025

Methods — techniques the papers use, named apart from their topics

data synthesis · 1.7polynomial-time algorithm · 0.8complexity analysis · 0.8layer separation · 0.6gradient exchange reduction · 0.6
YearPublicationVenuePosition
2026 GraphMamba: Graph-driven spatial order-aware Mamba for medical image segmentation
Chengjin Yu, Cailing Pu, Sangyin Lv, Xiaorui Wu, Dongsheng Ruan, Hanyu Xuan, Yuan-Ting Yan
Pattern Recognit.6
2026 Improving Emotion and Intent Understanding in Multimodal Conversations With Progressive Interaction
abstract
Emotion and intent joint understanding in multimodal conversations (MC-EIU) aims to model the semantic dependencies among multimodal conversations while inferring emotion and intent information. Despite making progress, existing methods overlook the differential contributions of modalities and rely on single-round interactions for emotion and intent recognition, resulting in suboptimal model understanding performance. To overcome these limitations, we propose MEIPro, a novel progressive interaction and adaptive weight fusion based multimodal joint understanding of emotion and intent framework. We first design a hierarchical denoising module to effectively remove noise and redundant information from multimodal data. Then, we propose an adaptive weight fusion mechanism that dynamically fuses multimodal features by taking the true classification probabilities of each modality as their respective contributions, thus enhancing the fusion process. Additionally, we present a progressive dual task interaction module to capture the deep seated interactions between emotion and intent through a step-by-step multi-round iteration. Experiments on the benchmark MC-EIU bilingual dataset demonstrate that our MEI-Pro framework significantly outperforms state-of-the-art baselines in both emotion and intent tasks. Specifically, on the English dataset, the F1-scores of the multimodal emotion and intent understanding tasks have increased by 6.12% and 7.25% respectively.
Tengyue Song, Yuzhe Ding 0001, Xiaorui Wu, Fei Li 0021, Dongdong Xie 0003, Chong Teng, Donghong Ji
IEEE Trans. Affect. Comput.4
2025 TRIDENT: Enhancing Large Language Model Safety with Tri-Dimensional Diversified Red-Teaming Data Synthesis
abstract
Xiaorui Wu, Xiaofeng Mao, Fei Li, Xin Zhang, Xuanhong Li, Chong Teng, Donghong Ji, Zhuang Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Xiaorui Wu, Xiaofeng Mao, Fei Li 0021, Xuanhong Li, Chong Teng, Donghong Ji, Zhuang Li 0001
ACL (1)1
2025 EVOREFUSE: Evolutionary Prompt Optimization for Evaluation and Mitigation of LLM Over-Refusal to Pseudo-Malicious Instructions
abstract
Large language models (LLMs) frequently refuse to respond to pseudo-malicious instructions: semantically harmless input queries triggering unnecessary LLM refusals due to conservative safety alignment, significantly impairing user experience. Collecting such instructions is crucial for evaluating and mitigating over-refusals, but existing instruction curation methods, like manual creation or instruction rewriting, either lack scalability or fail to produce sufficiently diverse and effective refusal-inducing prompts. To address these limitations, we introduce EVOREFUSE, a prompt optimization approach that generates diverse pseudo-malicious instructions consistently eliciting confident refusals across LLMs. EVOREFUSE employs an evolutionary algorithm exploring the instruction space in more diverse directions than existing methods via mutation strategies and recombination, and iteratively evolves seed instructions to maximize evidence lower bound on LLM refusal probability. Using EVOREFUSE, we create two novel datasets: EVOREFUSE-TEST, a benchmark of 582 pseudo-malicious instructions that outperforms the next-best benchmark with 85.34% higher average refusal triggering rate across 9 LLMs without a safety-prior system prompt, 34.86% greater lexical diversity, and 40.03% improved LLM response confidence scores; and EVOREFUSE-ALIGN, which provides 3,000 pseudo-malicious instructions with responses for supervised and preference-based alignment training. With supervised fine-tuning on EVOREFUSE-ALIGN, LLAMA3.1-8B-INSTRUCT achieves up to 29.85% fewer over-refusals than models trained on the second-best alignment dataset, without compromising safety. Our analysis with EVOREFUSE-TEST reveals models trigger over-refusals by overly focusing on sensitive keywords while ignoring broader context. Our code and datasets are available at https://github.com/FishT0ucher/EVOREFUSE .
Xiaorui Wu, Fei Li 0021, Xiaofeng Mao, Chong Teng, Donghong Ji, Zhuang Li 0001
NeurIPS1
2025 PantoPoseNet: A Two-Stage Framework for Real-Time Pantograph-Catenary Keypoint Detection
abstract
Accurate pantograph-catenary attitude detection is essential for high-speed railway safety, yet existing methods face significant challenges when dealing with sparse keypoints, small target regions, and complex operational environments. To address these limitations, we propose PantoPoseNet, a novel two-stage framework designed for real-time pantograph-catenary keypoint detection. Our approach introduces three key innovations in the first stage: (1) integration of VanillaNet with SIAF activation as the backbone network, achieving a 55.6% reduction in model parameters while preserving detection accuracy; (2) replacement of the conventional RepC3 module with CSP modules to enhance multiscale feature fusion capabilities; and (3) implementation of a hybrid loss function that combines GIoU and NWD metrics, specifically engineered to mitigate gradient vanishing issues inherent in small keypoint region detection. The second stage employs a specialized VanillaNet-KP network that processes 32×32 pixel regions to achieve precise keypoint localization. Comprehensive experiments conducted across various railway operational scenarios demonstrate that PantoPoseNet achieves superior performance with 98.2% mAP for keypoint region detection and 93.6% PCK for overall keypoint localization, while maintaining real-time processing at 25.94 FPS. These results significantly outperform current state-of-the-art methods, indicating strong potential for practical deployment in pantograph-catenary monitoring systems.
Yixuan Ma, Yutai Zhao, Xiaorui Wu
SMC4
2024 Pre-Training and Fine-Tuning for Efficient Routing in Opportunistic Networks
abstract
In opportunistic networks, it is a challenge to find the best relay node instead of blindly selecting from among available nodes to forward messages and effectively transmit them to their destinations. By comparing the similarity between nodes, the next hop node that is most similar to the destination node is found in time. Existing node similarity-based opportunistic networks routing algorithms only calculate similarity between nodes according to nodes' properties that can be directly obtained from the network, such as the historical meeting record between nodes, etc. However, this superficial similarity calculation is obviously inadequate to describe the inherent dynamic nature of opportunistic networks at both spatial and temporal levels, and ignores the salient movement characteristics of nodes, resulting in poor routing performance. Therefore, this paper proposes an opportunistic network routing strategy based on pre-training and fine-tuning (PTFT) model. Firstly, an autoencoder is added to the graph neural network model to encode node movement behavior. Then, a pre-training graph neural network model in large scale opportunistic network scenarios is applied to learn potential features of nodes through fine-tuning. Finally, we calculate the similarity between nodes based on the node potential feature, thereby assisting the nodes to achieve an efficient routing decision. The simulation results show that our PTFT-based algorithm not only has superior performance compared to traditional routing algorithms, but also faster learning speed than other machine learning-based routing algorithms.
Jia Hao 0006, Xiaorui Wu, Winston Khoon Guan Seah, Gang Xu 0007
SMC2
2022 Stanza: Layer Separation for Distributed Training in Deep Learning
abstract
The parameter server architecture is prevalently used for distributed deep learning. Each worker machine in a such system trains the complete model, which leads to a large amount of network data transfer between workers and servers. We empirically observe that the data transfer has a major impact on training time. We present a new distributed training system called Stanza to tackle this problem. Stanza exploits the fact that in many models such as convolution neural networks, most data exchange is attributed to the fully connected layers, while most computation is carried out in convolutional layers. Thus, we propose layer separation in distributed training: most nodes of the cluster train only the convolutional layers, while the rest train the fully connected layers. Gradients and parameters of the fully connected layers no longer need to be exchanged across the entire cluster, thereby substantially reducing the data transfer volume. We implement Stanza on PyTorch and evaluate its performance on Azure and EC2. Results show that Stanza accelerates training significantly over current parameter server systems: on EC2 instances with Tesla V100 GPU and 10Gb bandwidth for example, Stanza is 1.34x–13.9x faster for common deep learning models.
Xiaorui Wu, Hong Xu 0001, Bo Li 0001, Yongqiang Xiong
IEEE Trans. Serv. Comput.1
2021 Spatio-Temporal Topology Routing Algorithm for Opportunistic Network Based on Self-attention Mechanism
Xiaorui Wu, Baoqi Huang, Xiangyu Bai
ICA3PP (1)1
2021 Energy Balance and Cache Optimization Routing Algorithm Based on Communication Willingness
abstract
Existing opportunistic network routing algorithms usually have two main problems: excessive calculation of key nodes leads to the uneven energy consumption of nodes, and limited remaining cache of nodes leads to the loss of important messages. To solve the above problems, this paper proposed a new opportunistic network routing algorithm-EC-CW, which forwards messages according to the multi-copy mechanism and the communication willingness between nodes. The simulation results show that EC-CW reduces the average latency and the overhead rate in the nodes-sparse opportunistic network scenarios composed of high-cache nodes; EC-CW improves the delivery rate and reduces the overhead rate in the nodes-intensive opportunistic network scenarios composed of low-cache nodes.
JingJian Chen, Xiaorui Wu, Fengqi Wei, Liqiang He
WCNC3
2020 Irina: Accelerating DNN Inference with Efficient Online Scheduling
abstract
DNN inference is becoming prevalent for many real-world applications. Current machine learning frameworks usually schedule inference tasks with the goal of optimizing throughput under predictable workloads and task arrival patterns. Yet, inference workloads are becoming more dynamic with bursty queries generated by various video analytics pipelines which run expensive inference only on a fraction of video frames. Thus it is imperative to optimize the completion time of these unpredictable queries and improve customer experience.
Xiaorui Wu, Hong Xu 0001, Yi Wang 0004
APNet1
2019 Node-Constrained Traffic Engineering: Theory and Applications
abstract
Traffic engineering (TE) is a fundamental task in networking. Conventionally, traffic can take any path connecting the source and destination. Emerging technologies such as segment routing, however, use logical paths that are composed of shortest paths going through a predetermined set of middlepoints in order to reduce the flow table overhead of TE implementation. Inspired by this, in this paper, we introduce the problem of node-constrained TE, where the traffic must go through a set of middlepoints, and study its theoretical fundamentals. We show that the general node-constrained TE that allows the traffic to take any path going through one or more middlepoints is NP-hard for directed graphs but strongly polynomial for undirected graphs, unveiling a profound dichotomy between the two cases. We also investigate a variant of node-constrained TE that uses only shortest paths between middlepoints, and prove that the problem can now be solved in weakly polynomial time for a fixed number of middlepoints, which explains why existing work focuses on this variant. Yet, if we constrain the end-to-end paths to be acyclic, the problem can become NP-hard. An important application of our work concerns flow centrality, for which we are able to derive complexity results. Furthermore, we investigate the middlepoint selection problem in general node-constrained TE. We introduce and study group flow centrality as a solution concept, and show that it is monotone but not submodular. Our work provides a thorough theoretical treatment of node-constrained TE and sheds light on the development of the emerging node-constrained TE in practice.
George Trimponias, Yan Xiao 0002, Xiaorui Wu, Hong Xu 0001, Yanhui Geng
IEEE/ACM Trans. Netw.3
2018 Geometric Measurement of Interval Type-2 Fuzzy Sets and Application Based on $\lambda$-cuts
abstract
The aim of this paper is constructed a new combined fitting index for handling the decision-making problem of medical device under the interval type-2 fuzzy environment. Consider a matter of interval type-2 fuzzy sets geometric measurement with approximate positive and negative ideal, we use elasticity of decision making λ-cuts to obtain cutoff points and vectoring it. In contrast to the traditional nonlinear mathematical programming model which combined the common distance formula, this paper directly generates more intuitive geometric measurement formula which described the relationship between interval type-2 fuzzy numbers and approximate positive and negative ideal. Next, this task presents the concept of fitting-index using geometric measurement formula. Meanwhile, by using geometric measurement formula has been constructed to determine the attribute weights based on the minimizing overall uncertainty principle, a novel decision-making approach has been proposed with a comprehensive fitting index, the effectiveness and simplicity of method are reflect by a example analysis of medical device.
Junjun Mao, Xiaorui Wu
ICARCV3