Kai Xu 0014

dblp:30/495-14 · DBLP profile ↗
← Back
27ranked-venue papers
5as first author
18since 2021 · last 2026
0000-0002-7442-2383ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 7 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 4 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Natural Language-Guided Autonomous Agents for Counterterrorism Simulation via Deep Reinforcement Learning
Xinmeng Li, Kai Xu 0014, Yue Hu 0016, Miao Zhang 0037, Quanjun Yin
SIMULTECH2
2026 Behavior-Aware Consistent Distillation for Cold-Start Recommendation
Huan Gong, Hao Chen 0062, Lijia Chen, Feiran Huang, Kai Xu 0014, Yu Yang 0012, Fakhri Karray
IEEE Trans. Knowl. Data Eng.7
2025 CityEQA: A Hierarchical LLM Agent on Embodied Question Answering Benchmark in City Space
abstract
Yong Zhao, Kai Xu, Zhengqiu Zhu, Yue Hu, Zhiheng Zheng, Yingfeng Chen, Yatai Ji, Chen Gao, Yong Li, Jincai Huang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Kai Xu 0014, Zhengqiu Zhu, Yue Hu 0016, Zhiheng Zheng, Yatai Ji, Chen Gao 0001, Yong Li 0008, Jincai Huang 0001
EMNLP2
2025 Behavior Merging Graph Convolution Network for Multi-Behavior Recommendation
Hao Chen 0062, Yuanchen Bei, Kai Xu 0014, Feiran Huang, Yu Yang 0012, Huan Gong, Fakhri Karray
IEEE Trans. Knowl. Data Eng.4
2024 DAP: Domain-Aware Prompt Learning for Vision-and-Language Navigation
abstract
Following language instructions to navigate in unseen environments is a challenging task for autonomous embodied agents. With strong representation capabilities, pretrained vision-and-language models are widely used in VLN. However, most of them are trained on web-crawled generalpurpose datasets, which incurs a considerable domain gap when used for VLN tasks. To address the problem, we propose a novel and model-agnostic Domain-Aware Prompt learning (DAP) framework. For equipping the pretrained models with specific object-level and scene-level cross-modal alignment in VLN tasks, DAP applies a low-cost prompt tuning paradigm to learn soft visual prompts for extracting in-domain image semantics. Specifically, we first generate a set of in-domain image-text pairs with the help of the CLIP model. Then we introduce soft visual prompts in the input space of the visual encoder in a pretrained model. DAP injects in-domain visual knowledge into the visual encoder of the pretrained model in an efficient way. Experimental results on both R2R and REVERIE show the superiority of DAP compared to existing state-of-the-art methods.
Ting Liu 0018, Yue Hu 0016, Wansen Wu, Youkai Wang, Kai Xu 0014, Quanjun Yin
ICASSP5
2024 Enhancing Multimodal Sentiment Analysis via Learning from Large Language Model
abstract
Multimodal sentiment analysis (MSA) detects human sentiments by understanding data from multiple modalities, such as text and images. Existing research primarily strives for an effective multimodal fusion framework to derive informative representations. However, these methods neglect the necessity of exploiting external knowledge to aid in analyzing sentiments. As a result, the lack of external commonsense embarrasses these models when the opinion cues come in an implicit and obscure manner. To address the limitation, in this paper, we propose an Auxiliary Rationale Knowledge enhanced framework, namely ARK, which improves MSA models via learning from a multimodal large language model (MLLM). Specifically, based on text-image pairs, we employ Chain-of-Thought prompting to generate image descriptions and rationales from the MLLM as auxiliary knowledge, thus enriching the original samples with commonsense knowledge encoded within the MLLM. By combining the source text with image descriptions, we are able to effectively handle MSA through a Text+Text paradigm. In this paradigm, smaller pre-trained language models (LMs) can be tasked for sentiment classification via prompt-tuning. Besides, rationales are leveraged as additional supervision to facilitate the learning of reasoning abilities by LMs. Experimental results demonstrate that our proposed method outperforms current state-of-the-art approaches across four datasets. Our data and code are available at https://github.com/ningpang/ArkMSA.
Ning Pang, Wansen Wu, Yue Hu 0016, Kai Xu 0014, Quanjun Yin, Long Qin 0004
ICME4
2024 Learning High-Frequency Functions Made Easy with Sinusoidal Positional Encoding
abstract
Fourier features based positional encoding (PE) is commonly used in machine learning tasks that involve learning high-frequency features from low-dimensional inputs, such as 3D view synthesis and time series regression with neural tangent kernels. Despite their effectiveness, existing PEs require manual, empirical adjustment of crucial hyperparameters, specifically the Fourier features, tailored to each unique task. Further, PEs face challenges in efficiently learning high-frequency functions, particularly in tasks with limited data. In this paper, we introduce sinusoidal PE (SPE), designed to efficiently learn adaptive frequency features closely aligned with the true underlying function. Our experiments demonstrate that SPE, without hyperparameter tuning, consistently achieves enhanced fidelity and faster training across various tasks, including 3D view synthesis, Text-to-Speech generation, and 1D regression. SPE is implemented as a direct replacement for existing PEs. Its plug-and-play nature lets numerous tasks easily adopt and benefit from SPE.
Chuanhao Sun, Zhihang Yuan, Kai Xu 0014, Luo Mai, N. Siddharth 0001, Mahesh K. Marina
ICML3
2024 A Prototype Design of LLM-Based Autonomous Web Crowdsensing
Zhengqiu Zhu, Yatai Ji, Sihang Qiu, Kai Xu 0014, Rusheng Ju, Bin Chen 0003
ICWE5
2024 PANDA: Prompt-Based Context- and Indoor-Aware Pretraining for Vision and Language Navigation
Ting Liu 0018, Yue Hu 0016, Wansen Wu, Youkai Wang, Kai Xu 0014, Quanjun Yin
MMM (1)5
2024 Conversational Crowdsensing in the Age of Industry 5.0: A Parallel Intelligence and Large Models Powered Novel Sensing Approach
abstract
The transition from cyber-physical-system-based (CPS-based) Industry 4.0 to cyber-physical-social-system-based (CPSS-based) Industry 5.0 brings new requirements and opportunities to current sensing approaches, especially in light of recent progress in large language models (LLMs) and retrieval augmented generation (RAG). Therefore, the advancement of parallel intelligence powered crowdsensing intelligence (CSI) is witnessed, which is currently advancing toward linguistic intelligence. In this article, we propose a novel sensing paradigm, namely conversational crowdsensing, for Industry 5.0 (especially for social manufacturing). It can alleviate workload and professional requirements of individuals and promote the organization and operation of diverse workforce, thereby facilitating faster response and wider popularization of crowdsensing systems. Specifically, we design the architecture of conversational crowdsensing to effectively organize three types of participants (biological, robotic, and digital) from diverse communities. Through three levels of effective conversation (i.e., interhuman, human–AI, and inter-AI), complex interactions and service functionalities of different workers can be achieved to accomplish various tasks across three sensing phases (i.e., requesting, scheduling, and executing). Moreover, we explore the foundational technologies for realizing conversational crowdsensing, encompassing LLM-based multiagent systems, scenarios engineering and conversational human–AI cooperation. Finally, we present potential applications of conversational crowdsensing and discuss its implications. We envision that conversations in natural language will become the primary communication channel during crowdsensing process, enabling richer information exchange and cooperative problem-solving among humans, robots, and AI.
Zhengqiu Zhu, Sihang Qiu, Kai Xu 0014, Quanjun Yin, Jincai Huang 0001, Zhong Liu 0002, Fei-Yue Wang 0001
IEEE Trans. Comput. Soc. Syst.4
2023 Distributed Dynamic Data Driven Simulations: Basic Idea and an Illustration Example
abstract
The dynamic data driven simulation (DDDS) is a simulation paradigm where the simulation is continuously influenced by fresh data sampled from the real system for better analysis and prediction of the system under study. Traditional DDDS operates in the “computation away from data” mode, which would incur long response time, high network bandwidth requirement, and critical information loss. This paper proposes a novel distributed dynamic data driven simulation paradigm. In this paradigm, each participated simulation models a portion of a system, while all the participants collectively model the whole system. Additionally, each simulation assimilates data collected locally to produce local state estimations, which are aggregated somehow to generate global state estimation. The proposed simulation paradigm is supposed to have properties such as shorter response time, lower network bandwidth requirement, and less information loss. Finally, a case is studied to illustrate the effectiveness of the novel simulation paradigm.
Xu Xie 0005, Kai Xu 0014
DS-RT2
2023 Credit-based Differential Privacy Stochastic Model Aggregation Algorithm for Robust Federated Learning via Blockchain
abstract
By encapsulating model parameters in blocks when training machine learning models collaboratively, blockchain is recognized as a promising enabling technology to facilitate reliable federated learning under a distributed and untrusted environment. However, storing updated models of each worker in blockchain induces potential privacy risks, such as membership inference attacks. Besides, the volatile network conditions in the distributed environment may cause the deterioration of system robustness. This paper aims at addressing the privacy and robustness issues mentioned above. Specifically, a Credit-based Differential Privacy stochastic model aggregation algorithm combined with SIGN operation (Cre-DPSIGN) is adopted in our peer-to-peer network, which can realize the trade-off between privacy and accuracy. Furthermore, leveraging the transparency and tamper-proofing of blockchain, we design practical and reliable smart contracts for unbiased sampling based on the credit of workers to improve system robustness against Byzantine workers. In addition, we have demonstrated that the use of biased differential privacy mechanisms can lead to performance degradation. Therefore, we have introduced two unbiased differential privacy mechanisms and have proven their convergence and privacy guarantee. Extensive experiments conducted on MNIST datasets show that our algorithm can achieve byzantine fault tolerance rate with a private loss ϵ = 0.4. Compared with the state-of-the-art, aka DP-RSA (IJCAI-22), Cre-DPSIGN shows lower privacy loss consumption and better system robustness.
Mengyao Du, Miao Zhang 0037, Lin Liu 0018, Kai Xu 0014, Quanjun Yin
ICPP4
2023 Dynamic Multi-modal Prompting for Efficient Visual Grounding
Wansen Wu, Ting Liu 0018, Youkai Wang, Kai Xu 0014, Quanjun Yin, Yue Hu 0016
PRCV (7)4
2022 GenDT: mobile network drive testing made efficient with generative modeling
abstract
Drive testing continues to play a key role in mobile network optimization for operators but its high cost is a big concern. Alternative approaches like virtual drive testing (VDT) target device testing in the lab whereas MDT or crowdsourcing based approaches are limited by the incentives users have to participate and contribute measurements. With the aim of augmenting drive testing and significantly reducing its cost, we propose GenDT, a novel deep generative model that synthesizes high-fidelity time series of key radio network key performance indicators (KPIs). The training of GenDT relies on a relatively small amount of real-world measurement data along with corresponding and easily accessible network and environment context data. Through this, GenDT learns the relationship between context and radio network KPIs as they vary over time, and therefore trained GenDT model can subsequently be relied on to generate time series for different KPIs for new drive test routes (trajectories) without having to collect field measurements. GenDT represents an initial attempt at enabling efficient drive testing via generative modeling. Evaluations with real-world mobile network drive testing measurement datasets from two countries demonstrate that GenDT can synthesize significantly more dependable data than a range of baselines. We further show that GenDT has the potential to significantly reduce the drive testing related measurement effort, and that GenDT-generated data yields similar results to that with real data in the context of two downstream use cases - QoE prediction and handover analysis.
Chuanhao Sun, Kai Xu 0014, Mahesh K. Marina, Howard Benn
CoNEXT2
2022 CartaGenie: Context-Driven Synthesis of City-Scale Mobile Network Traffic Snapshots
abstract
Mobile network traffic data offers unprecedented opportunities for innovative studies within and beyond networking. However, progress is hindered by the very limited access that the research community at large has to the real-world mobile network data that is needed to develop and dependably test mobile traffic data-driven solutions. As a contribution to overcome this barrier, we propose CartaGenie, a generator of realistic mobile traffic snapshots at city scale. Taking a deep generative modeling approach and through a tailored conditional generator design, CartaGenie can synthesize high-fidelity and artifact-free spatial traffic snapshots using only contextual information about the target geographical region that is easily found in public repositories. Hence, CartaGenie allows researchers to create their own realistic datasets of spatial traffic from open data about their region of interest. Experiments with real-world mobile traffic measurements collected in multiple metropolitan areas show that CartaGenie can produce dependable network traffic loads for areas where no prior traffic information is available, significantly outperforming a comprehensive set of benchmarks. Moreover, tests with practical case studies demonstrate that the synthetic data generated by CartaGenie is as good as real data in supporting diverse research-oriented mobile traffic data-driven applications.
Kai Xu 0014, Rajkarn Singh, Hakan Bilen, Marco Fiore 0001, Mahesh K. Marina, Yue Wang 0008
PerCom1
2022 AppShot: A Conditional Deep Generative Model for Synthesizing Service-Level Mobile Traffic Snapshots at City Scale
abstract
Service-level mobile traffic data enables research studies and innovative applications with a potential to shape future service-oriented communication systems and beyond. However, real-world datasets reporting measurements at the individual service level are hard to access as such data is deemed commercially sensitive by operators. APPSHOT is a model for generating synthetic high-fidelity city-scale snapshots of service level mobile traffic. It can operate in any geographical region and relies solely on easily available spatial context information such as population density, thus allowing the generation of new and open traffic datasets for the research community. The design of APPSHOT is informed by an original characterization of service-level mobile traffic data. APPSHOT is a novel conditional GAN design instantiated by a convolutional neural network generator and two discriminators. The model features several other innovative mechanisms including multi-channel and overlapping patch based generation to address the unique challenges involved in generating mobile service traffic snapshots. Experiments with ground-truth data collected by a major European operator in multiple metropolitan areas show that APPSHOT can produce realistic network loads at the service level for areas where it has no prior traffic knowledge, and that such data can reliably support service-oriented networking studies.
Chuanhao Sun, Kai Xu 0014, Marco Fiore 0001, Mahesh K. Marina, Yue Wang 0008, Cezary Ziemlicki
IEEE Trans. Netw. Serv. Manag.2
2021 SpectraGAN: spectrum based generation of city scale spatiotemporal mobile network traffic data
abstract
City-scale spatiotemporal mobile network traffic data can support numerous applications in and beyond networking. However, operators are very reluctant to share their data, which is curbing innovation and research reproducibility. To remedy this status quo, we propose SpectraGAN, a novel deep generative model that, upon training with real-world network traffic measurements, can produce high-fidelity synthetic mobile traffic data for new, arbitrary sized geographical regions over long periods. To this end, the model only requires publicly available context information about the target region, such as population census data. SpectraGAN is an original conditional GAN design with the defining feature of generating spectra of mobile traffic at all locations of the target region based on their contextual features. Evaluations with mobile traffic measurement datasets collected by different operators in 13 cities across two European countries demonstrate that SpectraGAN can synthesize more dependable traffic than a range of representative baselines from the literature. We also show that synthetic data generated with SpectraGAN yield similar results to that with real data when used in applications like radio access network infrastructure power savings and resource allocation, or dynamic population mapping.
Kai Xu 0014, Rajkarn Singh, Marco Fiore 0001, Mahesh K. Marina, Hakan Bilen, Howard Benn, Cezary Ziemlicki
CoNEXT1
2021 Interpretable Goal Recognition for Path Planning with ART Networks
abstract
Goal recognition for path planning is an important task of intention identification and situation awareness, requiring an observer to predict the goal of an evader given observations of its movements. While existing models based on planning or Markov Decision Process (MDP) show superior performance over traditional library based methods, they require much effort in model design and can hardly provide legible decision rules for their users. To make the system more user-friendly while preserving accuracy of goal inference, this paper proposes a novel self-organizing neural network based inference model, which learns compact rule sets through generalizing the streaming observations of an evader. More critically, the system manifests a high level of interpretability with the linguistic if-then rule base, making it easily comprehensible for human decision makers. We conducted extensive experiments on a large-scale real-world road network. Results show that the proposed model produces accuracy comparable to those of two state-of-the-art methods while uniquely providing legible inference rules and strong robustness against multiple goals with missing data.
Yue Hu 0016, Kai Xu 0014, Budhitama Subagdja, Ah-Hwee Tan, Quanjun Yin
IJCNN2
2020 Toward A Thousand Lights: Decentralized Deep Reinforcement Learning for Large-Scale Traffic Signal Control
abstract
Traffic congestion plagues cities around the world. Recent years have witnessed an unprecedented trend in applying reinforcement learning for traffic signal control. However, the primary challenge is to control and coordinate traffic lights in large-scale urban networks. No one has ever tested RL models on a network of more than a thousand traffic lights. In this paper, we tackle the problem of multi-intersection traffic signal control, especially for large-scale networks, based on RL techniques and transportation theories. This problem is quite difficult because there are challenges such as scalability, signal coordination, data feasibility, etc. To address these challenges, we (1) design our RL agents utilizing ‘pressure’ concept to achieve signal coordination in region-level; (2) show that implicit coordination could be achieved by individual control agents with well-crafted reward design thus reducing the dimensionality; and (3) conduct extensive experiments on multiple scenarios, including a real-world scenario with 2510 traffic lights in Manhattan, New York City 1 2.
Chacha Chen, Hua Wei 0001, Guanjie Zheng, Yuanhao Xiong, Kai Xu 0014, Zhenhui Li
AAAI7
2019 CoLight: Learning Network-level Cooperation for Traffic Signal Control
abstract
Cooperation among the traffic signals enables vehicles to move through intersections more quickly. Conventional transportation approaches implement cooperation by pre-calculating the offsets between two intersections. Such pre-calculated offsets are not suitable for dynamic traffic environments. To enable cooperation of traffic signals, in this paper, we propose a model, CoLight, which uses graph attentional networks to facilitate communication. Specifically, for a target intersection in a network, CoLight can not only incorporate the temporal and spatial influences of neighboring intersections to the target intersection, but also build up index-free modeling of neighboring intersections. To the best of our knowledge, we are the first to use graph attentional networks in the setting of reinforcement learning for traffic signal control and to conduct experiments on the large-scale road network with hundreds of traffic signals. In experiments, we demonstrate that by learning the communication, the proposed model can achieve superior performance against the state-of-the-art methods.
Hua Wei 0001, Huichu Zhang, Guanjie Zheng, Xinshi Zang, Chacha Chen, Weinan Zhang 0001, Yanmin Zhu 0006, Kai Xu 0014, Zhenhui Li
CIKM9
2019 Learning Phase Competition for Traffic Signal Control
abstract
Increasingly available city data and advanced learning techniques have empowered people to improve the efficiency of our city functions. Among them, improving urban transportation efficiency is one of the most prominent topics. Recent studies have proposed to use reinforcement learning (RL) for traffic signal control. Different from traditional transportation approaches which rely heavily on prior knowledge, RL can learn directly from the feedback. However, without a careful model design, existing RL methods typically take a long time to converge and the learned models may fail to adapt to new scenarios. For example, a model trained well for morning traffic may not work for the afternoon traffic because the traffic flow could be reversed, resulting in very different state representation. In this paper, we propose a novel design called FRAP, which is based on the intuitive principle of phase competition in traffic signal control: when two traffic signals conflict, priority should be given to one with larger traffic movement (i.e., higher demand). Through the phase competition modeling, our model achieves invariance to symmetrical cases such as flipping and rotation in traffic flow. By conducting comprehensive experiments, we demonstrate that our model finds better solutions than existing RL methods in the complicated all-phase selection problem, converges much faster during training, and achieves superior generalizability for different road structures and traffic conditions.
Guanjie Zheng, Yuanhao Xiong, Xinshi Zang, Jie Feng 0002, Hua Wei 0001, Huichu Zhang, Yong Li 0008, Kai Xu 0014, Zhenhui Li
CIKM8
2019 PressLight: Learning Max Pressure Control to Coordinate Traffic Signals in Arterial Network
abstract
Traffic signal control is essential for transportation efficiency in road networks. It has been a challenging problem because of the complexity in traffic dynamics. Conventional transportation research suffers from the incompetency to adapt to dynamic traffic situations. Recent studies propose to use reinforcement learning (RL) to search for more efficient traffic signal plans. However, most existing RL-based studies design the key elements - reward and state - in a heuristic way. This results in highly sensitive performances and a long learning process. To avoid the heuristic design of RL elements, we propose to connect RL with recent studies in transportation research. Our method is inspired by the state-of-the-art method max pressure (MP) in the transportation field. The reward design of our method is well supported by the theory in MP, which can be proved to be maximizing the throughput of the traffic network, i.e., minimizing the overall network travel time. We also show that our concise state representation can fully support the optimization of the proposed reward function. Through comprehensive experiments, we demonstrate that our method outperforms both conventional transportation approaches and existing learning-based methods.
Hua Wei 0001, Chacha Chen, Guanjie Zheng, Vikash V. Gayah, Kai Xu 0014, Zhenhui Li
KDD6
2019 Targeted Knowledge Transfer for Learning Traffic Signal Plans
Guanjie Zheng, Kai Xu 0014, Yanmin Zhu 0006, Zhenhui Li
PAKDD (2)3
2019 Online probabilistic goal recognition and its application in dynamic shortest-path local network interdiction
Kai Xu 0014, Yunxiu Zeng, Qi Zhang 0017, Quanjun Yin, Lin Sun 0008, Kaiming Xiao
Eng. Appl. Artif. Intell.1
2017 Bridging the Gap between Observation and Decision Making: Goal Recognition and Flexible Resource Allocation in Dynamic Network Interdiction
abstract
Goal recognition, which is the task of inferring an agent’s goals given some or all of the agent’s observed actions, is one of the important approaches in bridging the gap between the observation and decision making within an observe-orient-decide-act cycle. Unfortunately, few researches focus on how to improve the utilization of knowledge produced by a goal recognition system. In this work, we propose a Markov Decision Process-based goal recognition approach tailored to a dynamic shortest-path local network interdiction (DSPLNI) problem. We first introduce a novel DSPLNI model and its solvable dual form so as to incorporate real-time knowledge acquired from goal recognition system. Then a Markov Decision Process-based goal recognition model along with its dynamic Bayesian network representation and the applied goal inference method is proposed to identify the evader’s real goal within the DSPLNI context. Based on that, we further propose an efficient scalable technique in maintaining action utility map used in fast goal inference, and develop a flexible resource assignment mechanism in DSPLNI using knowledge from goal recognition system. Experimental results show the effectiveness and accuracy of our methods both in goal recognition and dynamic network interdiction.
Kai Xu 0014, Kaiming Xiao, Quanjun Yin, Yabing Zha, Cheng Zhu 0002
IJCAI1
2015 Aggregating Opinions to Optimize Multi-objective Urban Tactical Position Selection
abstract
In this research-in-progress paper we present a new real-world domain for studying the aggregation of different opinions: optimal urban tactical position selection (TPS). This is an important and foreseeable real world application, not only because cities have been viewed as centers of gravity by military planners throughout history, but also because the military significance of cities has increased proportionally as the global urbanization does. We first present a mapping between the domain of engineering research and that of the agent models present in the literature and use genetic multi-objective optimization method to generate Pareto TPS plans. Further we study the importance of forming diverse teams when aggregating opinions of different problem solvers for tactical position selection, and also the relationships of the number of problem solvers with time ratio and the solution efficiency. We show that a diverse team of problem solvers is able to provide better force deployment plans for early-stage decision makers to choose from. We also find that opinion aggregation methods, like approval voting, help to allocate a difficult problem solving among several computing resources, and at the same time ensuring the efficiency of solutions. Finally, we present next steps for a deeper exploration of our questions.
Kai Xu 0014, Lin Sun 0008, Quanjun Yin
DS-RT1
2014 A Path Planning Algorithm Based on Parallel Particle Swarm Optimization
Weitao Dang, Kai Xu 0014, Quanjun Yin
ICIC (1)2