EDBT 2026 Demo / reviewers in the wild / expert
Xun Xiao
dblp:36/11343
· DBLP profile ↗
43ranked-venue papers
9as first author
36since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 13 since 2021Computer networks · 12 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TrustSeed: Lightweight Attestation Protocol for Ensuring LLM IntegrityabstractOver the last couple of years, large language models have increasingly been integrated into many computing applications. For privacy preservation, they are now deployed on edge devices. However, these deployments are vulnerable to bit flip attacks and backdoor attacks that compromise the integrity of the model. Traditional remote attestation techniques fail to detect such manipulations due to the large model size and the stealthiness of the attacks.In this paper, we present TrustSeed, a lightweight functional attestation protocol that uses a single inference to ensure large language models’ integrity. TrustSeed verifies integrity by applying deterministic, seed-based modifications to model weights within a Trusted Execution Environment and comparing the last intermediate activations and output distribution against a golden reference on the verifier. This approach prevents precomputed or forged responses, ensuring freshness and unpredictability in each attestation round. Our analysis shows that output distribution and last intermediate activations are effective indicators of integrity. We test TrustSeed against bit-flip, data poisoning, and weight poisoning attacks, reliably detecting even single-bit alterations. Extensive evaluations on edge platforms and an HPC system demonstrate minimal overhead and up to 127× faster attestation compared to state-of-the-art full-model hashing. Mohamed Alsharkawy, Mohamed Aboelenien Ahmed, Hassan Nassar, Jeferson González-Gómez, Heba Khdr, Osama Abboud, Xun Xiao, Jörg Henkel |
DATE | 7 |
| 2026 | A2SPM: A Memory Access Acceleration for Neuromorphic Processing with the SPMabstractOwing to their low power consumption and high energy efficiency, neuromorphic processors have found extensive applications across various intelligent computing domains, including artificial intelligence (AI) and low-power edge computing. Homogeneous neuromorphic processors demonstrate a unique integration of general-purpose computing versatility and neuromorphic computing acceleration capabilities. These processors effectively address the intensive memory access demands inherent in neuromorphic computing through the implementation of tightly coupled local memory, specifically Scratchpad Memory (SPM). This article presents the design of SPM-based memory access acceleration mechanisms, which have been specifically developed and optimized to accommodate the data characteristics and computational processes of typical Spiking Neural Networks (SNNs). The proposed scheme significantly enhances computational efficiency by reducing both memory access operations and computational instruction processing. Comparative performance evaluations reveal that the proposed design achieves a 28.99% speedup in Liquid State Machine operations and a 14.95% speedup in Spiking Convolutional Neural Network computations when benchmarked against a baseline homogeneous processor equipped with SPM. Xun Xiao, Lei Wang 0011 |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2026 | Modeling Frequent Event Recurrence in Manufacturing Logs With Markov-Modulated Renewal ProcessesabstractPrevious research on event log data analysis has primarily focused on identifying critical and frequent events, as well as qualitatively assessing correlations between event occurrences. However, the probabilistic behavior of frequently occurring events over time remains poorly understood. Through an in-depth exploratory analysis, we reveal that the (log) inter arrival times of events follow some mixture distributions with two modes, suggesting the presence of transitions between latent states. To better understand the data-generating mechanism underlying these frequent events, we employ Markov Modulated Renewal Processes (MMRPs), a type of hidden Markov model, to capture the patterns exhibited in the inter-arrival times between successive events. Due to limitations in record precision, some inter-arrival times are recorded as zero. To address this issue, we propose a simple data imputation algorithm to generate non zero inter-arrival times, facilitating inference on the inter-arrival time distributions and the underlying MMRPs. The effectiveness of the algorithm is validated using synthetic data. Finally, we evaluate the proposed model on real manufacturing system data, uncovering key insights into system states. Kangzhe He, Xun Xiao, Way Kuo, Min Xie 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2026 | Bayesian Analysis of Bivariate Degradation Data Using Hybrid Wiener-Inverse Gaussian Marginal Processes and a Shared FrailtyabstractIn engineering practice, it is common to observe simultaneous degradation of multiple performance characteristics in a system, in which these characteristics are correlated and exhibit differing degradation behaviors. This poses significant challenges to reliability modeling and analysis of multivariate degradation data. In this study, we propose a novel bivariate degradation model to meet the challenge. We employ the Wiener and inverse Gaussian processes to model the marginal processes, allowing for differing degradation patterns in the two dimensions. A shared frailty is then incorporated into the two marginal processes to capture their dependence structure. We derive the closed form of the reliability function for the proposed bivariate degradation model, and we develop an efficient Bayesian procedure for parameter estimation by combining the Gibbs sampler with the Metropolis-Hastings algorithm for posterior sampling. The performance of the Bayesian estimation method, along with the derived reliability formulas, is validated through comprehensive numerical simulations and a practical example involving a permanent magnet brake. Kai Song 0003, Xun Xiao, Zhisheng Ye 0001 |
IEEE Trans. Reliab. | 2 |
| 2025 | LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-SteeringabstractJinhe Bi, Yujun Wang, Haokun Chen, Xun Xiao, Artur Hecker, Volker Tresp, Yunpu Ma. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jinhe Bi, Xun Xiao, Artur Hecker, Volker Tresp, Yunpu Ma |
ACL (1) | 4 |
| 2025 | ASNPC: An Automated Generation Framework for SNN and Neuromorphic Processor Co-DesignabstractSpiking neural networks (SNNs) are promisingly considered as energy-efficient alternatives to traditional deep neural networks. At the same time, neuromorphic processors have garnered increasing development to support the efficient execution of large-scale SNNs. However, current works always separate their design to primarily prioritize a single criterion. Hardware-algorithm co-design allows for the simultaneous consideration of hardware and algorithm characteristics during the design process, effectively reducing resource usage while optimizing the algorithm's performance. In light of this, we developed a hardware-algorithm co-design framework named ASNPC for SNNs and neuromorphic processors. Considering the vast mixed-variable co-design space and the time-expensive function evaluations, we employed the surrogate-based multi-objective optimization algorithm MOTPE to identify Pareto solutions that balance algorithm performance and hardware costs. To rapidly obtain hardware results, we designed an end-to-end methodology that can automatically generate the Register-Transfer Level (RTL) code for neuromorphic processors corresponding to each candidate using templates from the hardware library. The evaluated hardware metrics, such as hardware resource and power consumption, are then fed back to MOTPE for the next candidate selection. Compared to existing works, the proposed approach can find better Pareto solutions within a limited search budget, making it effectively adapted to various application scenarios. Additionally, under the same hardware configuration, the neuromorphic processor we generated achieves lower hardware resource usage and higher throughput. Xun Xiao, Renzhi Chen |
DATE | 5 |
| 2025 | SICA: A Multicore Neuromorphic Processor Featuring Sparse Integration and Communication-Aware OptimizationabstractNeuromorphic computing has emerged as a promising paradigm due to its event-driven operation and energy efficiency, driving extensive research in neuromorphic processor development. When implementing spiking neural networks (SNNs) on such processors, two critical aspects must be addressed: neuron computation and spike communication. For neuron computation, previous work primarily relies on parallel accumulation via adder trees but fails to leverage the inherent sparsity in SNNs. For spike communication, conventional mesh topologies suffer from long-distance communication inefficiencies, while suboptimal mapping strategies further exacerbate latency issues. To address these challenges, we propose a low-overhead fast sparse detection mechanism that effectively exploits spike sparsity and optimizes the processor's workflow, thereby achieving efficient synaptic integration with minimal overhead. For spike communication, we employ an on-chip broadcast mechanism combined with a hybrid torus-mesh topology to significantly reduce communication latency, while systematically evaluating the impact of three distinct mapping strategies-random, sequential, and communication-aware mapping-on overall performance. Experimental results demonstrate significant improvements, with our solution delivering speedups of$1.26 \times$and$1.24 \times$compared to LSMCore on the N-MNIST and MNIST datasets, respectively. Furthermore, the communication-aware mapping strategy achieves a 24.87% reduction in communication latency, while the torus topology contributes an additional$\mathbf{1 4. 4 8} \boldsymbol{\%}$latency reduction. Junbo Tie, Xun Xiao, Yuanfeng Luo, Yang Guo 0003, Lei Wang 0011 |
HPCC | 8 |
| 2025 | MARQ: Engineering Mission-Critical AI-Based Software with Automated Result Quality AdaptationabstractAI-based mission-critical software exposes a blessing and a curse: its inherent statistical nature allows for flexibility in result quality, yet the mission-critical importance demands adherence to stringent constraints such as execution deadlines. This creates a space for trade-offs between the Quality of Result (QoR)-a metric that quantifies the quality of a computational outcome-and other application attributes like execution time and energy, particularly in real-time scenarios. Fluctuating resource constraints, such as data transfer to a remote server over unstable network connections, are prevalent in mobile and edge computing environments-encompassing use cases like Vehicle-to-Everything, drone swarms, or social-VR scenarios. We introduce a novel approach that enables software engineers to easily specify alternative AI service chains-sequences of AI services encapsulated in microservices aiming to achieve a predefined goal-with varying QoR and resource requirements. Our methodology facilitates dynamic optimization at runtime, which is automatically driven by the MARQ framework. Our evaluations show that MARQ can be used effectively for the dynamic selection of AI service chains in real-time while maintaining the required application constraints of mission-critical AI software. Notably, our approach achieves a 100x acceleration in service chain selection and an average 10% improvement in QoR compared to existing methods. Uwe Gropengießer, Elias Dietz, Florian Brandherm, Achref Doula, Osama Abboud, Xun Xiao, Max Mühlhäuser |
ICSE | 6 |
| 2025 | Backdoor Cleaning without External Guidance in MLLM Fine-tuningabstractMultimodal Large Language Models (MLLMs) are increasingly deployed in fine-tuning-as-a-service (FTaaS) settings, where user-submitted datasets adapt general-purpose models to downstream tasks. This flexibility, however, introduces serious security risks, as malicious fine-tuning can implant backdoors into MLLMs with minimal effort. In this paper, we observe that backdoor triggers systematically disrupt cross-modal processing by causing abnormal attention concentration on non-semantic regions—a phenomenon we term **attention collapse**. Based on this insight, we propose **Believe Your Eyes (BYE)**, a data filtering framework that leverages attention entropy patterns as self-supervised signals to identify and filter backdoor samples. BYE operates via a three-stage pipeline: (1) extracting attention maps using the fine-tuned model, (2) computing entropy scores and profiling sensitive layers via bimodal separation, and (3) performing unsupervised clustering to remove suspicious samples. Unlike prior defenses, BYE equires no clean supervision, auxiliary labels, or model modifications. Extensive experiments across various datasets, models, and diverse trigger types validate BYE's effectiveness: it achieves near-zero attack success rates while maintaining clean-task performance, offering a robust and generalizable solution against backdoor threats in MLLMs. Xuankun Rong, Wenke Huang 0003, Jian Liang 0003, Jinhe Bi, Xun Xiao, Yiming Li 0004, Bo Du 0001, Mang Ye |
NeurIPS | 5 |
| 2025 | X-DINC: Toward Cross-Layer ApproXimation for theDistributed and In-Network ACceleration of Multi-Kernel ApplicationsabstractWith the rapid evolution of programmable network devices and the urge for energy-efficient and sustainable computing, network infrastructures are mutating toward a computing pipeline, providing In-Network Computing (INC) capability. Despite the initial success in offloading single/small kernels to the network devices, deploying multi-kernel applications remains challenging due to limited memory, computing resources, and lack of support for Floating Point (FP) and complex operations. To tackle these challenges, we present a cross-layer approximation and distribution methodology (X-DINC) that exploits the error resilience of applications. X-DINC utilizes a chain of techniques to facilitate kernel deployment and distribution across heterogeneous devices in INC environments. First, we identify approximation and optimization opportunities in data acquisition and computation phases of multi-kernel applications. Second, we simplify complex arithmetic operations to cope with the computation limitations of the programmable network switches. Third, we perform application-level sensitivity analysis to measure the trade-off between performance gain and Quality of Results (QoR) loss when approximating individual kernels via various techniques. Finally, a greedy heuristic swiftly generates Pareto/near-Pareto mixed-precision configurations that maximize the performance gain while maintaining the user-defined QoR. X-DINC is prototyped on a Virtex-7 Field Programmable Gate Array (FPGA) and evaluated using the Blind Source Separation (BSS) application on industrial audio dataset. Results show that X-DINC performs separation up to 35% faster with up to 88% lower Area-Delay Product (ADP) compared to an Accurate-Centralized approach, when distributed across 2 to 7 network nodes, while maintaining audio quality within an acceptable range of 15–20 dB. Zahra Ebrahimi, Maryam Eslami, Xun Xiao, Akash Kumar 0001 |
Future Gener. Comput. Syst. | 3 |
| 2025 | DPReF: Decentralized Key Generation Using Physical-Related FunctionsabstractPhysical Unclonable Functions (PUFs) serve as a lightweight source to generate cryptographic keys utilizing the inherent physical device properties, making them particularly suitable for resource-constrained environments such as Internet of Things (IoT) devices. Recently, Physical-Related Functions (PReFs) extended PUFs to enable multiple devices to generate similar keys without the need to exchange or store them, improving security. However, state-of-the-art PReF implementations rely on a Trusted Third Party (TTP) to identify relative challenges, introducing a potential vulnerability if the TTP is compromised. In this work, we propose the first decentralized PReF protocol, removing reliance on the TTP and mitigating associated security risks. The proposed protocol allows relative challenges to be identified directly between devices in a decentralized manner. Additionally, we formalize a mathematical model to estimate the minimum number of devices required to build a network, based on the sizes of the PUF and the shared Challenge-Response Pair (CRP).. We demonstrate the generality of our model by verifying it across different types of state-of-the-art PUFs (Arbiter-based Non-Volatile Memory PUF (ANV-PUF) and Pseudo Linear Feedback Shift Register PUF (PLPUF).). We establish a 128 bit cryptographic key using the proposed protocol that matches the state-of-the-art but in a decentralized manner. Moreover, we prove that our protocol can be used to construct hardware-assisted attestation networks using ANV-PUF and PLPUF implementations with a shared secret of 16 bit that allows for both integrity and identity verification. Mohamed Alsharkawy, Hassan Nassar, Jeferson González-Gómez, Xun Xiao, Osama Abboud, Jörg Henkel |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2025 | Estimating Mean Time to Failure of Solid-State Drives via Self-Organizing Map and Model AveragingabstractIn this article, a two-step approach is developed to estimate mean time to failure (MTTF) of solid-state drives (SSD) by first formulating a composite health indicator via multichannel signal fusion and further predicting the remaining useful life(RUL) under degradation model misspecification. Specifically, an unsupervised neural network based on self-organizing map is constructed to approximate the highly nonlinear relationship between multivariate monitoring attributes and a univariate SSD health indicator. For each SSD, the composite health indicator over time is further calibrated by smoothing techniques and formulated into a general path degradation model with a uniform failure threshold. By extrapolating each degradation path to hit the failure threshold, the RULs of SSDs are obtained as pseudofailure times, which are fitted by various lifetime distributions. Finally, a novel model averaging strategy is proposed to weigh the MTTFs estimated by multiple combinations of candidate degradation models and lifetime distributions to alleviate the impact of model misspecification. A real-world SSD dataset is used to demonstrate the feasibility of the proposed two-step approach. Numerical results suggest that the proposed approach better characterizes the underlying degradation process under different model assumptions and settings. Peng Li 0038, Xun Xiao, Jiayu Chen 0002 |
IEEE Trans. Reliab. | 2 |
| 2024 | Unjammable: A Practical Approach to Jamming-Safe Wireless ChannelsabstractCommunication systems are delicate to deliberate jamming, which may ruin legitimate communication. Two kinds of jammers exist: partial knowledge and full knowledge jammers. Jammers with partial knowledge only know the encoding and decoding functions, whereas jammers with full knowledge also know the actual message. This paper studies the detectability of Denial of Service (DoS) attacks by jammers with partial knowledge. The theoretical framework that characterizes the detectability of a DoS attack is expressed in the literature for Turing machines with no computing capacity scarcity. Even though there is no computational limitation for Turing machines, they cannot decide if a DoS attack is possible. Conversely, by following the theoretical framework, Turing machines can recognize an Arbitrarily Varying Channel (AVC) which a DoS attack is not possible. In these circumstances, we present an algorithm that detects the scenarios where DoS attack is impossible. We provide a complexity analysis and perform Monte Carlo simulations to investigate the performance of the proposed algorithm in terms of time consumption. We observe that the simulation results are compatible with the complexity analysis. Mehmet Akif Kurt, Prashanth K. H. Sheshagiri, Jennifer Gabriel, Juan Alberto Cabrera Guerrero, Xun Xiao, Frank H. P. Fitzek |
GLOBECOM | 5 |
| 2024 | MOTPE/D: Hardware and Algorithm Co-design for Reconfigurable Neuromorphic ProcessorabstractRecent advances in hardware/algorithm co-design for spiking neural networks have demonstrated its potential for jointly optimizing algorithmic performance while minimizing hardware overhead. However, the gigantic mixed-variable hard-ware/algorithm co-design space and time-consuming hardware verification still pose an intractable challenge for solutions exploration. To tackle these problems, 1) we propose a generic three-phase hardware/algorithm co-design framework. In this framework, 2) we target a reconfigurable neuromorphic processor, and parameterize the hardware and network architecture in a unified design space. 3) We propose a generic analytical model to estimate the parameter size and power consumption, which can support fast candidate evaluation during the exploration. 4) We extend vanilla TPE (a single-objective optimization algorithm) to MOTPE/D, a generic Multi-objective optimization (MOO) algorithm, by introducing a decomposition strategy. Renzhi Chen, Xun Xiao, Jingyue Zhao, Zhenhua Zhu 0002, Huadong Dai, Yuhua Tang |
ICCD | 4 |
| 2024 | Fast and Lightweight Automatic Modulation Recognition using Spiking Neural NetworkabstractAutomatic modulation classification (AMR) is essential for receivers to demodulate signals in communication systems. Currently, various portable devices are capable of receiving a large amount of radio and real-time spectrum data, leading to a growing demand for fast and lightweight modulation recognition solutions. However, most existing AMR schemes emphasize higher recognition accuracy without considering complexity and model size. Therefore, lightweight methods meeting the accuracy requirements are still left to be investigated. In this paper, we propose an efficient AMR model that utilizes a liquid state machine (LSM), a typical spiking neural network (SNN), for the first time. This model is faster and more lightweight than previous solutions, and it utilizes a multilayer perceptron (MLP) classifier and a normalization unit to enhance performance, achieving a recognition accuracy of 96.4%. Compared to the most lightweight method currently, our model has 2× fewer trainable parameters and runs 260 × faster. Canghai Lin, Zhijiao Zhang, Jingyue Zhao, Xun Xiao |
ISCAS | 7 |
| 2024 | Comparison of Neural Network Models for Short-Term Load ForecastingabstractBalancing supply and demand is crucial for efficient energy distribution. To achieve it, accurate short-term electrical load forecasting is essential. This study investigates the applicability of various machine learning architectures for short-term load forecasting, using an NSW load dataset from the Australian Energy Market Operator. The key finding is the superior performance of a hybrid model, which integrates LSTM and GRU layers, on the NSW load dataset. This study demonstrates hybrid neural network models can significantly improve the accuracy and reliability of energy load predictions, thereby suggesting a viable pathway for enhancing future utility management practices. Shinead Surmon, Ahmad Ahmad, Xun Xiao, Huadong Mo |
SMC | 3 |
| 2024 | A Fast and Safe Neuromorphic Approach for Obstacle Avoidance of Unmanned Aerial VehicleabstractObstacle avoidance is a crucial task in unmanned aerial vehicles (UAV) motion planning. The accuracy and consistency of real-time visual information affect the gener-ation of obstacle avoidance commands, raising higher safety demands for obstacle avoidance. The neuromorphic computing-based obstacle avoidance solution can address these challenges. Dynamic vision sensors (DVS) exhibit low latency, low power consumption, and high dynamic range as novel neuromorphic sensors. Spiking neural networks (SNN) also leverage the same mechanism to efficiently process asynchronous and sparse event data generated by DVS, offering latency and energy efficiency advantages. Additionally, the optimal estimation method effectively mitigates the impact of noise and interference within the system, reducing the influence of errors on the algorithm and enhancing safety. Based on these considerations, this paper proposes a fast and safe obstacle avoidance framework. DVS is used to acquire event data from the environment, and a hardware-compatible lightweight SNN is employed to extract dynamic obstacle position information from the data. Compared to baseline methods, this approach reduces latency by 85%. Furthermore, two estimation methods are used to predict the movement of obstacles, ensuring flight safety by generating different UAV obstacle avoidance actions based on confidence intervals, even in the presence of obstacle information errors and omissions. Zhong Wan, Xun Xiao, Jingyue Zhao, Junbo Tie, Renzhi Chen, Guangda Zhang, Huadong Dai |
SMC | 3 |
| 2024 | Multi-Objective Evolutionary Neural Architecture Search for Liquid State MachineabstractLiquid State Machine (LSM) is a brain-inspired computational model that has proven highly effective in various applications, owing to its intrinsic capability to process spatiotemporal information and its minimal training complexity. However, the performance of LSMs significantly depends on the design of their network architecture, which is overly reliant on existing human experience. Furthermore, as the network scale increases, the computing resources required for deployment and operation also increase, so we regarded the network design as a multi-objective problem. To address these challenges, we introduced an effective surrogate-assisted multi-objective evolutionary neural architecture search algorithm that balanced the accuracy and network scale. Our approach utilized parameter sensitivity analysis followed by the upper confidence bound algorithm to reduce the search space. Experimental results demonstrate that we successfully reduced the dimensions of the search space by 11% and the size of the entire search space by 75%. Compared to the state-of-the-art, our approach offered better trade-off solutions, such as a solution that reduced network scale by 32.5% while maintaining the same accuracy, and another that improved accuracy by 1.4% without changing the network scale. Furthermore, the knee point reduced network scale by 25 % and simultaneously increased accuracy by 0.7%. The source code can be accessed at https://github.com/XinSida/MOENAS-PSA. Sida Xin, Renzhi Chen, Xun Xiao |
SMC | 3 |
| 2024 | Hierarchical Mapping of Large-Scale Spiking Convolutional Neural Networks Onto Resource-Constrained Neuromorphic ProcessorabstractNeuromorphic processors have been designed as non-von Neumann systems for energy-efficient spiking neural network (SNN) execution. Spiking convolutional neural networks (SCNNs), combining the advantage of convolutional neural network (CNN) and SNN, have been widely applied to vision tasks. However, as the scale of SCNNs increases, executing large-scale SCNNs on resource-constrained neuromorphic processor faces many challenges, including massive synapse pruning caused by resource competition, execution performance degradation, etc. Addressing these problems, we propose an efficient approach to map large-scale SCNNs onto resource-constrined neuromorphic processor. The approach consists of three steps: splitting, partitioning, and mapping. We explore three acyclic splitting strategies to divide large-scale SCNNs into subnetworks without cyclic dependency. Axon sharing is the guiding principle to partition subnetworks into multiple clusters. To obtain an optimal cluster-to-core mapping scheme, we use Non-dominated Sorting Genetic Algorithm to collaboratively optimize two metrics. We evaluate our approach with eight realistic SCNN applications. The results show that compared with existing state-of-the-art methods, our approach significantly reduces the synapse pruning and accuracy loss, and increases the execution performance. Xun Xiao, Yao Wang 0002, Junbo Tie, Lei Wang 0011, Weixia Xu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2024 | Accelerating Tip Selection in Burst Message Arrivals for DAG-Based Blockchain SystemsabstractBlockchain systems (e.g., IOTA) use Directed Acyclic Graph (DAG) topology to organize its ledger records. A new message is added by attaching to the tips selected from the DAG. For tip selection, a stochastic approach is widely used, where a random walk is simulated until it ends at a tip of the DAG. Prior art intensively focused on the fairness and security issues of this random walk approach. However, its computational efficiency issue was largely neglected. This paper reports that under a burst message arrival condition, the random walk approach, even with parallelization, will suffer from severe delays. To solve this problem, this paper proposes a new approach inspired by Absorbing Markov chain (AMC) theory. Specifically, the new approach first periodically calculates a tip selection probability distribution (TSPD) of the DAG ledger. With this information, a processing node only needs to do sampling from the calculated TSPD, which significantly accelerates the tip selection process. Rigorous theoretical complexity analysis is provided; in addition, the approach is compared with both single- and multi-processing random walk schemes. Evaluation results confirm the key findings and demonstrate the benefits of the solution. Xun Xiao |
IEEE Trans. Serv. Comput. | 1 |
| 2024 | LSM-Based Hotspot Prediction and Hotspot-Aware Routing in NoC-Based Neuromorphic ProcessorabstractThe traffic patterns of spiking neural networks (SNNs) exhibit high variability and stochastic, leading to the emergence of elevated traffic hotspots on the network-on-chip (NoC)-based neuromorphic processors. Predicting the occurrence of hotspots remains one of the most challenging issues in NoC design. This article presents the first attempt toward traffic hotspot prediction by utilizing liquid state machine (HP-LSM). The predictor extracts essential information reflecting the current state of the NoC to predict potential routing hotspots in the subsequent time step. Furthermore, we designed the hardware architecture for HP-LSM, which incorporates leaky-integrate-and-fire (LIF) neurons with configurable biological parameters. Meanwhile, we introduce a novel hotspot-aware path-based multicast (HaPM) routing algorithm that utilizes advanced knowledge acquired from HP-LSM to guide packet routing throughout the network, aiming to improve the performance of NoC. Results indicate that the HP-LSM can forecast hotspot formation with an accuracy up to 89.36% and 90.19% for two spiking-based datasets, respectively. The hardware experiment results demonstrate a 92.03% reduction in the average execution time of zero skipping compared with nonzero skipping. Moreover, the HP-LSM exhibits a reduction of up to 79.30% in the number of neurons compared with other related SNN predictor models. The experiments reveal a reduction of 73.67% and 53.42% in the average length of the multicast path when compared with dual-path (DP) or multipath (MP) multicast routing. The HaPM demonstrates improved performance in terms of average latency and throughput compared with DP, MP, and path-based multicast (PbM) multicast routing. Ziyang Kang, Xun Xiao, Lei Wang 0011, De Ma, Gang Pan 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2023 | F-E Fusion: A Fast Detection Method of Moving UAV Based on Frame and Event Flow
Xun Xiao, Zhong Wan, Shasha Guo 0001, Junbo Tie |
ICANN (8) | 1 |
| 2023 | An Efficient Graph-Based IOTA Tangle Generation AlgorithmabstractIOTA is a recent distributed ledger technology that relies on Directed Acyclic Graph (DAG) for its ledger organization. To improve IOTA mechanisms, the state of the art methodology employs graph analysis and, for that, heavily relies on synthetic graph generation. Herein, the most popular generation method simulates IOTA protocol execution. Although this method produces realistic IOTA ledgers, it requires too much memory and time due to repeated random walks on the DAG. In this paper, we propose an alternative Graph Generation and Refinement (GraGR) algorithm designed to generate realistic IOTA ledgers while strongly relaxing memory and timing constraints. The evaluations show that, compared to the state of the art, GraGR can generate a ledger with the same properties with only half of memory and up to 10 times faster. Fengyang Guo, Xun Xiao, Artur Hecker, Schahram Dustdar |
ICC | 2 |
| 2023 | Brain-Inspired Binaural Sound Source Localization Method Based on Liquid State Machine
Jingyue Zhao, Xun Xiao, Renzhi Chen |
ICONIP (3) | 3 |
| 2023 | Multimodal Video Emotional Analysis of Time Features Alignment and Information Auxiliary LearningabstractWith the rapid growth of various images and video data on the network platform, multimodal emotional analysis and recognition have become an increasingly popular research field. Inspired by the human emotional judgment and cross information between different characteristics, this paper proposed a multimodal video emotional analysis network of time features alignment and information auxiliary learning(METI). This network not only fully consider the time alignment between modals to obtain the preliminary information but also pay attention to the auxiliary information in the video and make emotional judgments. METI network is mainly composed of the following three parts: Time alignment network pays attention to time alignment information between different modes; Prepose feedback module uses the auxiliary information in the video to simulate the emotional judgment process; Multimodal gate control module pays attention to the emotional correlation of the context. The comparison of experimental results shows that the performance of the METI model on the dataset exceeds the current most advanced emotional analysis model. At the same time, it has been proved from ablation experiments that the modules and methods proposed in this paper can improve the accuracy of emotion recognition by nearly 4%, and has been respectively proved in the seven classification experiment of emotions and three classification experiment of sentiment. Yujie Chang, Tianliang Liu, Qinchao Xu, Xun Xiao |
IJCNN | 5 |
| 2023 | A Theoretical Model Characterizing Tangle Evolution in IOTA Blockchain NetworkabstractIOTA blockchain system is lightweight without heavy proof-of-work mining phases, which is considered a promising service platform of Internet of Things applications. IOTA organizes ledger data in a directed acyclic graph (DAG), called Tangle, rather a chain structure as in traditional blockchains. With arriving messages, IOTA tangle grows in a special way, as multiple messages can be attached to the tangle at different locations in parallel. Hence, the network dynamics of an operational IOTA system would justify a thorough study, which is currently unexplored in the literature. In this article, we present the first theoretical modeling for the evolving IOTA tangle based on stochastic analysis. After analyzing snapshots of the real-world IOTA ledger data, our key finding suggests that IOTA tangle follows a rather atypical double Pareto Lognormal (dPLN) degree distribution. In contrast, typical power-law and exponential distributions do not accurately reflect the fact. For model parameter estimation, we further realize that using generic optimization solvers cannot yield quality fitting results. Thus, we design an alternative algorithm based on expectation-maximization (EM) framework. We evaluate the proposed model and fitting algorithm with official data provided by the IOTA Foundation. Quantitative comparisons confirm the fitting quality of our proposed model and algorithm. The whole analysis reveals a deeper understanding of the internal mechanism of the IOTA network. Fengyang Guo, Xun Xiao, Artur Hecker, Schahram Dustdar |
IEEE Internet Things J. | 2 |
| 2023 | Accelerating Industrial IoT Acoustic Data Separation With In-Network ComputingabstractAcoustic data from the Industrial Internet of Things (IIoT) are widely used in anomaly detection because audio information reflects richer internal statuses of monitored working machines than the video does. Since multiple acoustic data sources interfere with each other by nature, source data estimation is a prerequisite of subsequent anomaly detection. Existing schemes often use a centralized manner to separate full data on a remote node in clouds. However, such a centralized manner may delay reactions to anomalies due to data transmission delay and the complexity of solving data separation problems. This article shows that the data separation phase can be substantially accelerated with an in-network computing approach. The key idea is to offload data processing jobs to intermediate network nodes along the forwarding path. We first propose a distributed algorithm so that the data separation jobs can be done in a progressive manner; likewise, we modify the forwarding layer in order to eliminate hop-by-hop data transmission delay that hurts the performance of using in-network computing. We further derive theoretical upper and lower bounds of the required number of intermediate nodes that achieve the maximum acceleration. We also implement our proposed solution in a full-stack network emulator. Based on an open and professional data set, evaluation results justify the feasibility and advantages of our idea with nearly 32.18% acceleration on total processing time. This work exemplifies the convergence of IIoT, edge, and clouds. Huanzhuo Wu, Yunbin Shen, Xun Xiao, Giang T. Nguyen 0002, Artur Hecker, Frank H. P. Fitzek |
IEEE Internet Things J. | 3 |
| 2023 | Back to Homogeneous Computing: A Tightly-Coupled Neuromorphic Processor With Neuromorphic ISAabstractIn recent years, neuromorphic processors are widely used in many scenarios, showing extreme energy efficiency over traditional architectures. However, almost all existing neuromorphic hardware are following the heterogeneous computing methodology without Instruction Set Architecture (ISA), leading to inflexibility in programming. In this paper, we first propose a RISC-V Neuromorphic Extension (RVNE) to enable fine-grained and flexible homogeneous programming for neuromorphic algorithms while utilizing SNN sparsity from different levels of granularity and computing flows. Based on RVNE, we next implement a neuromorphic micro-architecture that is tightly coupled to the CPU pipeline to accelerate neuromorphic computing. To demonstrate the proposed homogeneous neuromorphic architecture, we implement a prototype processor called NeuroRVcore based on RISC-V ISA and an open-source RISC-V core. The evaluation results show that RVNE achieves a 2.8 × −4.3 × reduction in code density compared with the general-purpose ISAs. Compared with the state-of-the-art neuromorphic processor, the proposed homogeneous computing reduces energy consumption by 3.4%−22.5% while enabling fine-grained and flexible homogeneous programming. Lei Wang 0011, Yao Wang 0002, Junbo Tie, Feng Wang 0050, LingHui Peng, Xun Xiao, Gan Zhou, Xuhu Yu, Xia Zhao 0004, Yuhua Tang, Weixia Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 10 |
| 2022 | Unicorn: a multicore neuromorphic processor with flexible fan-in and unconstrained fan-out for neuronsabstractNeuromorphic processor is popular due to its high energy efficiency for spatio-temporal applications. However, when running the spiking neural network (SNN) topologies with the ever-growing scale, existing neuromorphic architectures face challenges due to their restrictions on neuron fan-in and fan-out. This paper proposes Unicorn, a multicore neuromorphic processor with a spike train sliding multicasting mechanism (STSM) and neuron merging mechanism (NMM) to support unconstrained fan-out and flexible fan-in of neurons. Unicorn supports 36K neurons and 45M synapses and thus supports a variety of neuromorphic applications. The peak performance and energy efficiency of Unicorn reach 36TSOPS and 424GSOPS/W respectively. Experimental results show that Unicorn can achieve 2×-5.5× energy reduction over the state-of-the-art neuromorphic processor when running an SNN with a relatively large fan-out and fan-in. Lei Wang 0011, Yao Wang 0002, LingHui Peng, Xun Xiao, Weixia Xu 0001 |
DAC | 6 |
| 2022 | Modeling Ledger Dynamics in IOTA BlockchainabstractIOTA blockchain is a new type of distributed ledger systems that is lightweight without mining and feeless-of-using. Rather than using a chain structure as in traditional blockchains, IOTA organizes ledger records with a directed acyclic graph (DAG), called Tangle. When message entries are committed into the ledger, the ledger tangle grows in a special way where multiple messages could be attached by different processing nodes in parallel. Such a unique evolution process motivates us to study the ledger tangle dynamics, which is unexplored so far. In this paper, we present the first generative modeling for IOTA tangle based on stochastic analysis. A key finding is that IOTA tangle renders a double Pareto Lognormal (dPLN) distribution, rather not typical network models (e.g., Power-Law and Exponential distributions). Quantitative comparisons show that the fitting quality of our model outperforms existing popular models on official real world datasets published by IOTA Foundation. Estimated model parameters are provided, which is immediately instrumental for a more realistic IOTA network generator design. The proposed generative model also provides a deeper understanding of the internal mechanics of IOTA network. Fengyang Guo, Xun Xiao, Artur Hecker, Schahram Dustdar |
GLOBECOM | 2 |
| 2022 | Fast Tip Selection for Burst Message Arrivals on A DAG-based Blockchain Processing Node at EdgeabstractWith the rapid evolution of blockchain technology, a clear trend is that new blockchain systems (e.g., IOTA) tend to use a Directed Acyclic Graph (DAG) rather a chain structure to organize ledger records. Such a DAG-based blockchain system shows higher scalability as multiple locations are available in the ledger for new message attachment. To decide an attachment location, a popular type of tip selection algorithms follow an approach using weighted random walks on the DAG ledger. In a burst message arrival scenario, however, a processing node deployed at edge using such a method may become a bottleneck because sequentially repeating random walks significantly increases processing delay. In this paper, we propose a new tip selection algorithm for the burst message arrival scenario on an edge node. Our solution abandons the weighted random walk approach, instead, with similar efforts we transfer to calculate in advance the tip selection probability distribution of the DAG ledger. Such a new scheme reduces tip selection to a probability distribution sampling task, which can be done extremely fast. We implement our solution and demonstrate the benefits of our approach by comparing with the random walk approach. We believe our attempt can effectively mitigate the congestion at the edge node and inspire tip selection algorithm design with a new vision for DAG-based blockchain systems. Xun Xiao, Fengyang Guo, Artur Hecker, Schahram Dustdar |
GLOBECOM | 1 |
| 2022 | An Event Based Gesture Recognition System Using a Liquid State Machine AcceleratorabstractIn this paper, we design a spiking neural network (SNN) accelerator based on the Liquid State Machine (LSM) which is more lightweight and bionic. In this accelerator, 512 leaky integrate-and-fire (LIF) neurons with configurable biological parameters are integrated. For the sparsity of computation and memory of the LSM, we use zero-skipping and weight compression to maximize the performance. The quantized 4-bit model deployed on the accelerator can achieve a classification accuracy of 97.42% on the DVS128 gesture dataset. We implement the accelerator on FPGA. Results indicate that its end-to-end average inference latency is 3.97 ms, which is 26 times better than the gesture recognition system based on TrueNorth. Xun Xiao, Ziyang Kang, LingHui Peng |
ACM Great Lakes Symposium on VLSI | 3 |
| 2022 | Dynamic Vision Sensor Based Gesture Recognition Using Liquid State Machine
Xun Xiao, Lei Wang 0011, Lianhua Qu, Shasha Guo 0001, Yao Wang 0002, Ziyang Kang |
ICANN (3) | 1 |
| 2022 | A Spatio-Temporal Event Data Augmentation Method for Dynamic Vision Sensor
Xun Xiao, Ziyang Kang, Shasha Guo 0001, Lei Wang 0011 |
ICONIP (6) | 1 |
| 2021 | In-Network Processing Acoustic Data for Anomaly Detection in Smart FactoryabstractModern manufacturing is now deeply integrating new technologies such as 5G, Internet-of-things (IoT), and cloud/edge computing to shape manufacturing to a new level – Smart Factory. Autonomic anomaly detection (e.g., malfunctioning machines and hazard situations) in a factory hall is on the list and expects to be realized with massive IoT sensor deployments. In this paper, we consider acoustic data-based anomaly detection, which is widely used in factories because sound information reflects richer internal states while videos cannot; besides, the capital investment of an audio system is more economically friendly. However, a unique challenge of using audio data is that sounds are mixed when collecting thus source data separation is inevitable. A traditional way transfers audio data all to a centralized point for separation. Nevertheless, such a centralized manner (i.e., data transferring and then analyzing) may delay prompt reactions to critical anomalies. We demonstrate that this job can be transformed into an in-network processing scheme and thus further accelerated. Specifically, we propose a progressive processing scheme where data separation jobs are distributed as microservices on intermediate nodes in parallel with data forwarding. Therefore, collected audio data can be separated 43.75% faster with even less total computing resources. This solution is comprehensively evaluated with numerical simulations, compared with benchmark solutions, and results justify its advantages. Huanzhuo Wu, Yunbin Shen, Xun Xiao, Artur Hecker, Frank H. P. Fitzek |
GLOBECOM | 3 |
| 2021 | Does class size matter? An in-depth assessment of the effect of class size in software defect prediction
Amjed Tahir, Kwabena Ebo Bennin, Xun Xiao, Stephen G. MacDonell |
Empir. Softw. Eng. | 3 |
| 2020 | Characterizing IOTA Tangle with Empirical DataabstractIOTA organizes transactions in the ledger as a Directed Acyclic Graph (DAG) called Tangle, instead of a hash chain of transaction blocks used by most of traditional blockchains. IOTA is considered a promising platform to support Internet-of-Things (IoT) applications with its key features such as micropayment support and absence of transaction fees. While prior art shows extensive analysis based on synthetic data generated through simulations, an analysis based on empirical data from a deployed IOTA network is still missing. In this paper, we provide the first comprehensive analysis by using real transaction data officially published by IOTA Foundation. Our key finding is that neither the tangle's topological features nor the actual observed performance is consistent with the main conclusions from the literature. In particular, most of transactions take roughly 10 minutes to be officially confirmed, which is not exactly instant as commonly assumed; yet, what is arguably worse is that there is a certain amount (5%) of transactions experiencing exceptionally long confirmation time. This shows that IOTA still has gaps to meet the stringent requirements of IoT applications that are delay sensitive. Fengyang Guo, Xun Xiao, Artur Hecker, Schahram Dustdar |
GLOBECOM | 2 |
| 2016 | When group-buying meets cloud computingabstractAs a major driving force for adopting cloud computing, continuous cost reduction has been constantly pursued by cloud users. For a group of users with heterogeneous cloud resource demands, it may be possible for them to buy resources in a collaborative way in order to save the purchase cost, which is known as group-buying in business. While group-buying can benefit cloud users in principle, the question is how to design an implementation scheme to support group-buying on the cloud market. In this paper, we address the question by studying a coalition formation game, aiming to design a way under which the users can form stable coalitions for group-buying. It turns out that group-buying on the cloud market is challenging in that most popular solution concepts may fail to constitute stable coalitions. In order to sustain group-buying for cloud services, we propose a new solution concept, contractually group stable, which is an extension of an existing concept in the literature. We show that this new solution concept can guarantee the existence of stable coalitions, making group-buying always possible on the cloud market. We also develop computing algorithms for solving the coalition formation game under our concept. Computational experiments show that our concept can bring in substantial cost reduction for cloud users. Juntao Wang 0004, Xun Xiao, Jianping Wang 0001, Kejie Lu, Xiaotie Deng, Ashwin Gumaste |
INFOCOM | 2 |
| 2016 | An optimal pricing scheme to improve transmission opportunities for a mobile virtual network operator
Xun Xiao, Rui Zhang 0031, Jianping Wang 0001, Chunming Qiao, Kejie Lu |
Comput. Networks | 1 |
| 2016 | Automatic optimal filament segmentation with sub-pixel accuracy using generalized linear models and B-spline level-setsabstractBiological filaments, such as actin filaments, microtubules, and cilia, are often imaged using different light-microscopy techniques. Reconstructing the filament curve from the acquired images constitutes the filament segmentation problem. Since filaments have lower dimensionality than the image itself, there is an inherent trade-off between tracing the filament with sub-pixel accuracy and avoiding noise artifacts. Here, we present a globally optimal filament segmentation method based on B-spline vector level-sets and a generalized linear model for the pixel intensity statistics. We show that the resulting optimization problem is convex and can hence be solved with global optimality. We introduce a simple and efficient algorithm to compute such optimal filament segmentations, and provide an open-source implementation as an ImageJ/Fiji plugin. We further derive an information-theoretic lower bound on the filament segmentation error, quantifying how well an algorithm could possibly do given the information in the image. We show that our algorithm asymptotically reaches this bound in the spline coefficients. We validate our method in comprehensive benchmarks, compare with other methods, and show applications from fluorescence, phase-contrast, and dark-field microscopy. Xun Xiao, Veikko F. Geyer, Hugo Bowne-Anderson, Jonathon Howard, Ivo F. Sbalzarini |
Medical Image Anal. | 1 |
| 2016 | Optimal Design for Destructive Degradation Tests With Random Initial Degradation Values Using the Wiener ProcessabstractThis study investigates modeling, estimation and optimization of destructive degradation tests (DDTs) for highly reliable products with random initial degradation values. It is common to observe that the degradation paths of distinct products start from different values specified by a random variable. The random initial value introduces additional uncertainties to the degradation of the product. In this study, Wiener-process-based degradation models are developed for products with random initial values. We first consider a DDT without stress acceleration. In a DDT, the measurement of the degradation destroys a test unit and, thus, only one measurement is available for each unit. Closed-form maximum likelihood (ML) estimators are derived. Then, an accelerated DDT (ADDT) is considered. Based on these results, we investigate optimal designs of both DDT and ADDT with the objective of minimizing the asymptotic variance of the estimated p th-quantile of the failure time distribution under use conditions. The optimal test plans have to be obtained through a numerical approach. Optimality of the plans is verified by the general equivalence theorem. An adhesive bond example with real degradation data is analyzed to show the performance of the proposed methods. Xun Xiao, Zhisheng Ye 0001 |
IEEE Trans. Reliab. | 1 |
| 2013 | Coordinated resource provisioning and maintenance scheduling in cloud data centersabstractLack of proper maintenance is the root cause of anywhere from a third to a half of downtime events in a cloud data center. To help safeguard the uptime of data centers, regular preventive maintenance must be conducted. During the maintenance time, some accommodated virtual machines (VMs) may be re-provisioned to the other available (backup) resource through migration, and some VMs may be terminated. One way that can allow a data center to perform all necessary preventive maintenance activities without causing too much disruption to VMs is to design an appropriate maintenance schedule. In this paper, given the available resource in a data center and the required maintenance activities with their deadlines, we consider the joint VM resource provisioning and maintenance scheduling problem to maximize the revenue of the data center. We tackle the problem by firstly proposing a heuristic for the resource provisioning under a given maintenance schedule. Using such a heuristic algorithm as the building block, we then propose another heuristic algorithm to solve the joint resource provisioning and maintenance scheduling problem and also derive its upper bound. Extensive simulations have shown that our proposed heuristic algorithms can effectively maximize the revenue of the data center. Minming Li, Xun Xiao, Jianping Wang 0001 |
INFOCOM | 3 |
| 2012 | Optimal resource allocation to defend against deliberate attacks in networking infrastructuresabstractProtecting networking infrastructures from malicious attacks is important as a successful attack on a high data rate link can cause the loss or delay of large amounts of data. In this paper, we consider a proactive approach where the ISPs are willing to allocate some (limited) resources to defend the networking infrastructures against the attacks. We aim to answer where and how much the defending resource should be placed so that the expected data loss can be minimized no matter where the attacker may launch the attack. We model the problem as a 2-player zero-sum game where the payoffs are measured by the maximum network flow. In order to overcome the unique challenges of such payoffs, we transform the payoffs into explicit piece-wise functions through multi-parametric linear programming (MP-LP) and divide the entire strategy space into a set of critical regions. We prove that a global Nash Equilibrium (NE) exists when there is only one critical region. However, when the number of critical regions is greater than 1, there is no global NE. We also prove that there exists one and only one local NE in each critical region. We then design a mixed-strategy solution. Our results have shown that to dedicate all defending resources to one min-cut set when there are multiple min-cut sets will not be an optimal solution, however, min-cut strategies will have higher probabilities to be selected in the mixed-strategy solution when the defending resource is limited. Xun Xiao, Minming Li, Jianping Wang 0001, Chunming Qiao |
INFOCOM | 1 |