Jiayang Song

dblp:182/7194 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 7 · 1 first-author · 7 since 2021Computer networks · 3 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Mosaic: model-based safety analysis for AI-enabled cyber physical system
Xuan Xie 0001, Jiayang Song, Zhehua Zhou, Fuyuan Zhang, Lei Ma 0003
Empir. Softw. Eng.2
2026 ST-SSNet: Spatiotemporal Feature Fusion-Based DOA Estimation Network for Underwater Array Signals
abstract
For Underwater Internet of Things (UIoT) applications, the accurate and efficient estimation of the direction of arrival (DOA) is fundamental to technologies such as node localization and autonomous underwater vehicle (AUV) node cooperative communication. However, the low signal-to-noise ratio (SNR) and limited energy in underwater environments pose severe challenges to DOA estimation. Furthermore, existing methods typically require a large number of snapshots. To address these issues, this paper proposes the use of an adaptive wavelet denoising model to enhance the quality of underwater acoustic signals. Subsequently, a dual-branch space-time state space network (ST-SSNet) is proposed. This network consists of a time feature extraction branch (TFEB) and a space feature extraction branch (SFEB). The time branch incorporates gating units and time mixing functions into the state space model (SSM) within the Mamba framework to extract temporal features. The spatial branch uses one-dimensional convolutions in different directions and the convolutional block attention module (CBAM) to extract spatial features. Extensive simulation and sea trial experiments demonstrate that ST-SSNet outperforms other deep learning methods in various scenarios, while having lower computational complexity than other methods. Compared to ResNet18, the accuracy improves by 2.14%, and RMSE is reduced by 43.4%.
Jiayang Song, Qiuna Niu, Lingwei Xu, Yulei Yang, Shuzhuo Chen, Jingjing Wang 0003
IEEE Internet Things J.1
2026 Modeling and Optimal Control of Spatiotemporal Malware Propagation in Underwater Wireless Sensor Networks
abstract
Underwater Wireless Sensor Networks (UWSNs) have been shown to overcome the environmental extremes and energy dependency issues faced by traditional IoT in marine environments, leading to rapid development in fields such as environmental monitoring and disaster warning. Among these, Autonomous Underwater Vehicles (AUVs) play a pivotal role in UWSNs. However, the mobility of AUVs poses significant challenges in detecting and controlling infected nodes due to the randomization of malicious program cross-platform infection and propagation paths. Accordingly, a mathematical model centered on the framework of epidemic theory has been proposed to study the propagation patterns of malicious programs in two coupled networks (AUVs and UWSNs). This model utilizes the mutual infection coefficient between AUVs and UWSNs to represent the cross-infection of malicious programs. In order to investigate the impact of AUV mobility on malware propagation, an improved cellular automaton model is proposed. This model combines the state transitions of epidemic theory with the cellular automaton model to represent the spatio-temporal propagation of malware. Furthermore, to attain optimal decision-making under resource constraints, we propose an optimization problem combining a mathematical model with defense strategies and use the Sand Dune Cat Swarm Optimization Algorithm (SCSO) to obtain the optimal control strategy. Finally, simulation experiments demonstrate that AUV movement and expanded communication radii exacerbate malware propagation, while also validating the influence of the basic reproduction number (R0) on malware propagation and the inhibitory effect of optimal control strategies on malware propagation.
Yulei Yang, Zehua Du, Jingjing Wang 0003, Shuzhuo Chen, Jiayang Song
IEEE Internet Things J.5
2026 LeCov: Multi-level testing criteria for large language models
Xuan Xie 0001, Jiayang Song, Yuheng Huang 0004, Felix Juefei-Xu, Lei Ma 0003
J. Syst. Softw.2
2026 AcTracer: Active Testing of Large Language Model via Multi-Stage Sampling
abstract
Performance evaluation plays a crucial role in the development lifecycle of large language models (LLMs). It estimates the model’s capability, elucidates behavior characteristics, and facilitates the identification of potential issues and limitations, thereby guiding further improvement. Given that LLMs’ diverse task-handling abilities stem from large volumes of training data, a comprehensive evaluation also necessitates abundant, well-annotated, and representative test data to assess LLM performance across various downstream tasks. However, the demand for high-quality test data often entails substantial time, computational resources, and manual efforts, sometimes causing the evaluation to be inefficient or impractical. To address these challenges, researchers propose active testing, which estimates the overall performance by selecting a subset of test data. Nevertheless, the existing active testing methods tend to be inefficient, even inapplicable, given the unique new challenges of LLMs (e.g., diverse task types, increased model complexity, and unavailability of training data). To mitigate such limitations and expedite the development cycle of LLMs, in this work, we introduce AcTracer, an active testing framework tailored for LLMs that strategically selects a small subset of test data to achieve a more accurate performance estimation for LLMs. AcTracer utilizes both internal and external information from LLMs to guide the test sampling process, reducing variance through a multi-stage pool-based active selection. Our experiment results demonstrate that AcTracer achieves state-of-the-art performance compared to existing methods across various tasks.
Yuheng Huang 0004, Jiayang Song, Felix Juefei-Xu, Lei Ma 0003
ACM Trans. Softw. Eng. Methodol.2
2025 TrustVis: A Multi-Dimensional Trustworthiness Evaluation Framework for Large Language Models
abstract
As Large Language Models (LLMs) continue to revolutionize Natural Language Processing (NLP) applications, critical concerns about their trustworthiness persist, particularly in safety and robustness. To address these challenges, we introduce TrustVis, an automated evaluation framework that provides a comprehensive assessment of LLM trustworthiness. A key feature of our framework is its interactive user interface, designed to offer intuitive visualizations of trustworthiness metrics. By integrating well-known perturbation methods like AutoDAN and employing majority voting across various evaluation methods, TrustVis not only provides reliable results but also makes complex evaluation processes accessible to users. Preliminary case studies on models like Vicuna-7b, Llama2-7b, and GPT-3.5 demonstrate the effectiveness of our framework in identifying safety and robustness vulnerabilities, while the interactive interface allows users to explore results in detail, empowering targeted model improvements. Video Link: https://youtu.be/k1TrBqNVg8g
Ruoyu Sun 0011, Jiayang Song, Yuheng Huang 0004, Lei Ma 0003
ASE3
2025 GenSafe: A Generalizable Safety Enhancer for Safe Reinforcement Learning Algorithms Based on Reduced Order Markov Decision Process Model
abstract
Safe reinforcement learning (SRL) aims to realize a safe learning process for deep reinforcement learning (DRL) algorithms by incorporating safety constraints. However, the efficacy of SRL approaches often relies on accurate function approximations, which are notably challenging to achieve in the early learning stages due to data insufficiency. To address this issue, we introduce, in this work, a novel generalizable safety enhancer (GenSafe) that can overcome the challenge of data insufficiency and enhance the performance of SRL approaches. Leveraging model order reduction techniques, we first propose an innovative method to construct a reduced order Markov decision process (ROMDP) as a low-dimensional approximator of the original safety constraints. Then, by solving the reformulated ROMDP-based constraints, GenSafe refines the actions of the agent to increase the possibility of constraint satisfaction. Essentially, GenSafe acts as an additional safety layer for SRL algorithms. We evaluate GenSafe on multiple SRL approaches and benchmark problems. The results demonstrate its capability to improve safety performance, especially in the early learning phases, while maintaining satisfactory task performance. Our proposed GenSafe not only offers a novel measure to augment existing SRL methods but also shows broad compatibility with various SRL algorithms, making it applicable to a wide range of systems and SRL problems.
Zhehua Zhou, Xuan Xie 0001, Jiayang Song, Zhan Shu 0001, Lei Ma 0003
IEEE Trans. Neural Networks Learn. Syst.3
2025 Look Before You Leap: An Exploratory Study of Uncertainty Analysis for Large Language Models
abstract
The recent performance leap of Large Language Models (LLMs) opens up new opportunities across numerous industrial applications and domains. However, the potential erroneous behavior (e.g., the generation of misinformation and hallucination) has also raised severe concerns for the trustworthiness of LLMs, especially in safety-, security- and reliability-sensitive industrial scenarios, potentially hindering real-world adoptions. While uncertainty estimation has shown its potential for interpreting the prediction risks made by classic machine learning (ML) models, the unique characteristics of recent LLMs (e.g., adopting self-attention mechanism as its core, very largescale model size, often used in generative contexts) pose new challenges for the behavior analysis of LLMs. Up to the present, little progress has been made to better understand whether and to what extent uncertainty estimation can help characterize the capability boundary of an LLM, to counteract its undesired behavior, which is considered to be of great importance with the potential wide-range applications of LLMs across industry domains. To bridge the gap, in this paper, we initiate an early exploratory study of the risk assessment of LLMs from the lens of uncertainty. In particular, we conduct a large-scale study with as many as twelve uncertainty estimation methods and eight general LLMs on four NLP tasks and seven programming-capable LLMs on two code generation tasks to investigate to what extent uncertainty estimation techniques could help characterize the prediction risks of LLMs. Our findings confirm the potential of uncertainty estimation for revealing LLMs’ uncertain/nonfactual predictions. The insights derived from our study can pave the way for more advanced analysis and research on LLMs, ultimately aiming at enhancing their trustworthiness.
Yuheng Huang 0004, Jiayang Song, Zhijie Wang 0014, Shengming Zhao, Huaming Chen, Felix Juefei-Xu, Lei Ma 0003
IEEE Trans. Software Eng.2
2024 ISR-LLM: Iterative Self-Refined Large Language Model for Long-Horizon Sequential Task Planning
abstract
Motivated by the substantial achievements of Large Language Models (LLMs) in the field of natural language processing, recent research has commenced investigations into the application of LLMs for complex, long-horizon sequential task planning challenges in robotics. LLMs are advantageous in offering the potential to enhance the generalizability as task-agnostic planners and facilitate flexible interaction between human instructors and planning systems. However, task plans generated by LLMs often lack feasibility and correctness. To address this challenge, we introduce ISR-LLM, a novel framework that improves LLM-based planning through an iterative self-refinement process. The framework operates through three sequential steps: preprocessing, planning, and iterative self-refinement. During preprocessing, an LLM translator is employed to convert natural language input into a Planning Domain Definition Language (PDDL) formulation. In the planning phase, an LLM planner formulates an initial plan, which is then assessed and refined in the iterative self-refinement step by a validator. We examine the performance of ISR-LLM across three distinct planning domains. Our experimental results show that ISR-LLM is able to achieve markedly higher success rates in sequential task planning compared to state-of-the-art LLM-based planners. Moreover, it also preserves the broad applicability and generalizability of working with natural language instructions.
Zhehua Zhou, Jiayang Song, Kunpeng Yao, Zhan Shu 0001, Lei Ma 0003
ICRA2
2024 LUNA: A Model-Based Universal Analysis Framework for Large Language Models
abstract
Over the past decade, Artificial Intelligence (AI) has had great success recently and is being used in a wide range of academic and industrial fields. More recently, Large Language Models (LLMs) have made rapid advancements that have propelled AI to a new level, enabling and empowering even more diverse applications and industrial domains with intelligence, particularly in areas like software engineering and natural language processing. Nevertheless, a number of emerging trustworthiness concerns and issues exhibited in LLMs, e.g., robustness and hallucination, have already recently received much attention, without properly solving which the widespread adoption of LLMs could be greatly hindered in practice. The distinctive characteristics of LLMs, such as the self-attention mechanism, extremely large neural network scale, and autoregressive generation usage contexts, differ from classic AI software based on Convolutional Neural Networks and Recurrent Neural Networks and present new challenges for quality analysis. Up to the present, it still lacks universal and systematic analysis techniques for LLMs despite the urgent industrial demand across diverse domains. Towards bridging such a gap, in this paper, we initiate an early exploratory study and propose a universal analysis framework for LLMs, namedLUNA, which is designed to be general and extensible and enables versatile analysis of LLMs from multiple quality perspectives in a human-interpretable manner. In particular, we first leverage the data from desired trustworthiness perspectives to construct an abstract model as an auxiliary analysis asset and proxy, which is empowered by various abstract model construction methods built-inLUNA. To assess the quality of the abstract model, we collect and define a number of evaluation metrics, aiming at both the abstract model level and the semantics level. Then, the semantics, which is the degree of satisfaction of the LLM w.r.t. the trustworthiness perspective, is bound to and enriches the abstract model with semantics, which enables more detailed analysis applications for diverse purposes, e.g., abnormal behavior detection. To better understand the potential usefulness of our analysis frameworkLUNA, we conduct a large-scale evaluation, the results of which demonstrate that 1) the abstract model has the potential to distinguish normal and abnormal behavior in LLM, 2)LUNAis effective for the real-world analysis of LLMs in practice, and the hyperparameter settings influence the performance, 3) different evaluation metrics are in different correlations with the analysis performance. In order to encourage further studies in the quality assurance of LLMs, we made all of the code and more detailed experimental results data available on the supplementary website of this paperhttps://sites.google.com/view/llm-luna.
Xuan Xie 0001, Jiayang Song, Derui Zhu, Yuheng Huang 0004, Felix Juefei-Xu, Lei Ma 0003
IEEE Trans. Software Eng.3
2023 $\mathtt {SIEGE}$SIEGE: A Semantics-Guided Safety Enhancement Framework for AI-Enabled Cyber-Physical Systems
abstract
Cyber-Physical Systems (CPSs) have been widely adopted in various industry domains to support many important tasks that impact our daily lives, such as automotive vehicles, robotics manufacturing, and energy systems. As Artificial Intelligence (AI) has demonstrated its promising abilities in diverse tasks like decision-making, prediction, and optimization, a growing number of CPSs adopt AI components in the loop to further extend their efficiency and performance. However, these modern AI-enabled CPSs have to tackle pivotal problems that the AI-enabled control systems might need to compensate the balance acrossmultiple operation requirementsand avoid possible defections in advance to safeguard human lives and properties. Modular redundancy and ensemble method are two widely adopted solutions in the traditional CPSs and AI communities to enhance the functionality and flexibility of a system. Nevertheless, there is a lack of deep understanding of the effectiveness of such ensemble design on AI-CPSs across diverse industrial applications. Considering the complexity of AI-CPSs, existing ensemble methods fall short of handling such huge state space and sophisticated system dynamics. Furthermore, an ideal control solution should consider the multiple system specifications in real-time and avoid erroneous behaviors beforehand. Such that, a new specification-oriented ensemble control system is of urgent need for AI-CPSs. In this paper, we propose$\mathtt {SIEGE}$, a semantics-guided ensemble control framework to initiate an early exploratory study of ensemble methods on AI-CPSs and aim to construct an efficient, robust, and reliable control solution for multi-tasks AI-CPSs. We first utilize a semantic-based abstraction to decompose the large state space, capture the ongoing system status and predict future conditions in terms of the satisfaction of specifications. We propose a series of new semantics-aware ensemble strategies and an end-to-end Deep Reinforcement Learning (DRL) hierarchical ensemble method to improve the flexibility and reliability of the control systems. Our large-scale, comprehensive evaluations over five subject CPSs show that 1) the semantics abstraction can efficiently narrow the large state space and predict the semantics of incoming states, 2) our semantics-guided methods outperform state-of-the-art individual controllers and traditional ensemble methods, and 3) the DRL hierarchical ensemble approach shows promising capabilities to deliver a more robust, efficient, and safety-assured control system. To enable further research along this direction to build better AI-enabled CPS, we made all of the code and experimental results data publicly available athttps://sites.google.com/view/ai-cps-siege/home.
Jiayang Song, Xuan Xie 0001, Lei Ma 0003
IEEE Trans. Software Eng.1
2017 A Performance Analysis Model for TCP over Multiple Heterogeneous Paths in 5G Networks
abstract
The demand for multipath transmission is prominent in 5G networks with the deployment of multiple hierarchical access technologies. However, multipath schemes are still not widely adopted due to many reasons, such as deployment challenges and performance reduction under the circumstances of path heterogeneity. Thus, TCP is still in the dominant position of the transport layer protocol for now and for the foreseeable future. Link asymmetry, such as different latency and different bandwidth of different links, is considered to be the main reasons leading to packet reordering, and further result in TCP performance reduction. However, to the best of knowledge, no one has yet given a theoretical model to analyze the relationship between link asymmetry and TCP multipath performance. In this paper, we present a performance analysis model for TCP over multiple heterogeneous networks, which reveals the effect of link asymmetry on TCP throughput. Both bandwidth and delay asymmetry are taken into consideration in the proposed model. The evaluated throughput using the proposed model can accurately fit the simulation results.
Jiayang Song, Huachun Zhou, Tao Zheng 0003, Xiaojiang Du, Mohsen Guizani
GLOBECOM1
2016 Modeling Link Quality for High-Speed Railway Networks Based on Hidden Markov Chain
abstract
To design efficient high-speed railway (HSR) communication systems, it is essential to characterize the wireless link quality. In this paper, we made a large amount of field investigations on link quality of HSR network, and built a practical model to reflect the changing pattern of link quality along HSR lines in terms of round trip time (RTT) and packet loss rate (PLR). After analyzing a great number of collected dataset of RTT and PLR we excitedly found that their behaviors presented an obvious two-scale time- varying phenomenon. To this end, we analyzed the potential reasons and further characterized link quality of HSR network using a generalized reference model based on hidden Markov chain. An improved forward induction algorithm was proposed to simulate the two-time-scale phenomenon of RTT and PLR. Evaluation results show that the proposed model is able to well reflect the network link quality varying along the HSR line with accuracies of 71.2% and 63.5% in terms of PLR and RTT. The proposed model can be used to guide the HSR link quality prediction and evaluation.
Jiayang Song, Huachun Zhou, Wei Quan 0001
VTC Spring1