Xiyu Shi

dblp:72/4382 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
9since 2021 · last 2026
0000-0001-6174-3383ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 From Threads to Tiles: T2T, a Compiler for CUDA-to-NPU Translation via 2D Vectorization
abstract
CUDA’s programming model, exposing massive parallelism via fine-grained scalar threads, has become the de facto standard for GPU computing. Concurrently, NPUs are emerging as highly efficient accelerators, but their architecture is fundamentally different, relying on coarse-grained, explicit 2-D tile-based instructions. This creates a critical challenge: bridging the semantic gap "From Threads to Tiles". A direct translation is infeasible, as it requires lifting the implicit parallelism of CUDA’s scalar model into the explicit, multi-dimensional vector space of NPUs, a problem we formalize as a lifting challenge.This paper introduces T2T, a compiler framework that automates this "Threads to Tiles" translation via the 2-D Vectorization technique. T2T first transforms a CUDA kernel’s implicit SIMT parallelism into a structured, explicit loop nest via our Unified Parallelism Abstraction (UPA), making the parallelism analyzable. From this representation, T2T’s core vectorization engine systematically selects optimal pairs of loops and maps them onto the NPU’s 2-D tile instructions to maximize hardware utilization. To ensure correctness and handle performance-critical CUDA features, a final set of semantics-preserving optimizations is applied, including efficient control-flow management and vectorization of warp-level intrinsics.We implement T2T based on Polygeist and evaluate representative NPU architectures. On a diverse set of benchmarks, kernels translated by T2T achieve up to 73% of native CUDA performance on an A100 GPU and outperform baseline translation approaches by up to 6.9×. Our work demonstrates that a systematic, compiler-driven approach to 2-D vectorization is a principled and high-performance path for porting the rich CUDA ecosystem to the evolving landscape of NPU accelerators.
Shuaijiang Li, Ying Liu 0055, Shuoming Zhang, Yijin Li, Yangyu Zhang, Runyu Zhou, Xiyu Shi, Chunwei Xia, Yuan Wen, Xiaobing Feng 0002, Huimin Cui
CGO10
2026 Symbiotic MLLM Serving: Dynamically Balancing Parallelism Across GPUs and Resources Within GPUs
Yangyu Zhang, Zhaolin Duan, Shuoming Zhang, Shuaijiang Li, Donglin Yu, Yuan Wen, Chunwei Xia, Xiyu Shi, Huimin Cui
ISCA12
2026 A Novel Solution for Zero-day Attack Detection in IDS using Self-attention and Jensen-Shannon divergence in WGAN-GP
abstract
The increasing sophistication of cyber threats, especially zero-day attacks, poses a significant challenge to cybersecurity. Zero-day attacks exploit unknown vulnerabilities, making them difficult to detect and defend against. Existing approaches patch flaws and deploy an Intrusion Detection System (IDS). Using advanced Wasserstein GANs with Gradient Penalty (WGAN-GP), this paper makes a novel proposition to synthesize network traffic that mimics zero-day patterns, enriching data diversity and improving IDS generalization. SA-WGAN-GP is first introduced, which adds a Self-Attention (SA) mechanism to capture long-range cross-feature dependencies by reshaping the feature vector into tokens after dense projections. A JS-WGAN-GP is then proposed, which adds a Jensen-Shannon (JS) divergence-based auxiliary discriminator that is trained with Binary Cross-Entropy (BCE), frozen during updates, and used to regularize the generator for smoother gradients and higher sample quality. Third, SA-JS-WGAN-GP is created by combining the SA mechanism with JS divergence, thereby enhancing the data generation ability of WGAN-GP. As data augmentation does not equate with true zero-day attack discovery, we emulate zero-day attacks via the leave-one-attack-type-out method on the NSL-KDD dataset for training all GANs and IDS models in the assessment of the effectiveness of the proposed solution. The evaluation results show that integrating SA and JS divergence into WGAN-GP yields superior IDS performance and more effective zero-day risk detection.
Ziyu Mu, Xiyu Shi, Safak Dogan
Comput. Networks2
2026 A Systematic Review of Virtual Reality Technology and Human Memory Augmentation
abstract
Researchers have long been intrigued in memory augmentation through Virtual Reality (VR) applications. However, inconsistent findings have been reported in literature due to variety of methodologies, VR and memory types adopted in each study. This paper provides a systematic review of the literature in this domain by thoroughly mapping and meta-analysing previous research on VR and human memory augmentation to offer methodological guidance for more rigorous and reliable studies in the future. The review highlights a large proportion of explanatory, quantitative, and experimental research with various demographic samples, indicating the accumulation of rich empirical evidence in literature. Our review also brings to attention that the memory types, except for the spatial memory, have been less frequently and vaguely mentioned in literature, whereas the VR types have often been highlighted. Meta-analysis has revealed a medium effect size (Cohen’s d = 0.54) for the relationship between VR and human memory.
Yunsun A. Hong, Xiyu Shi, Safak Dogan
Int. J. Hum. Comput. Interact.2
2025 ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models
abstract
Activation sparsity refers to the existence of considerable weakly-contributed elements among activation outputs, serving as a promising paradigm for accelerating model inference. Nevertheless, most large language models (LLMs) adopt activation functions without intrinsic activation sparsity (e.g., GELU and Swish). Some recent efforts have explored introducing ReLU or its variants as the substitutive activation function to pursue activation sparsity and acceleration, but few can simultaneously obtain high activation sparsity and comparable model performance. This paper introduces a simple and effective method named “ProSparse” to sparsify LLMs while achieving both targets. Specifically, after introducing ReLU activation, ProSparse adopts progressive sparsity regularization with a factor smoothly increasing for multiple stages. This can enhance activation sparsity and mitigate performance degradation by avoiding radical shifts in activation distributions. With ProSparse, we obtain high sparsity of 89.32% for LLaMA2-7B, 88.80% for LLaMA2-13B, and 87.89% for end-size MiniCPM-1B, respectively, with comparable performance to their original Swish-activated versions. These present the most sparsely activated models among open-source LLaMA versions and competitive end-size models. Inference acceleration experiments further demonstrate the significant practical acceleration potential of LLMs with higher activation sparsity, obtaining up to 4.52x inference speedup.
Xu Han 0007, Zhengyan Zhang, Shengding Hu, Xiyu Shi, Kuai Li, Zhiyuan Liu 0001, Guangli Li, Maosong Sun 0001
COLING5
2025 Investigating Transferability in Multi-Agent Reinforcement Learning
abstract
Effectively transferring knowledge from one task to another in cooperative Multi-Agent Reinforcement Learning (MARL) can be key to accelerate learning in complex scenarios. For instance, in sports it is common to train in simpler situations and then apply what was learned in the real game. The same logic applies to other scenarios that involve increasing levels of difficulty, such as robotics tasks, or healthcare applications. In the realm of MARL, transferring knowledge can become extremely challenging due to factors such as changes in the observation and action spaces, or the underlying dynamics of the environment. Consequently, most current methods still opt to train each new task from scratch. However, as tasks become increasingly complex, there is a growing interest in leveraging behavioural similarities that emerge among semantically similar tasks. In this context, we explore techniques to facilitate the rapid transfer of knowledge from one policy network to another within off-policy value-based methods. Additionally, we introduce a special case that enables function-preserving transfers of centralised functions between tasks. Our work offers a promising strategy to reduce training time, enabling zero-shot task transfers.
Corentin Artaud, Rafael Pina, Varuna De Silva, Xiyu Shi
COMPSAC4
2025 SpaceServe: Spatial Multiplexing of Complementary Encoders and Decoders for Multimodal LLMs
abstract
Recent multimodal large language models (MLLMs) marry modality-specific vision or audio encoders with a shared text decoder. While the encoder is compute- intensive but memory-light, the decoder is the opposite, yet state-of-the-art serving stacks still time-multiplex these complementary kernels, idling SMs or HBM in turn. We introduce SpaceServe, a serving system that space-multiplexes MLLMs: it decouples all modality encoders from the decoder, and co-locates them on the same GPU using fine-grained SM partitioning available in modern runtimes. A cost-model-guided Space-Inference Scheduler (SIS) dynamically assigns SM slices, while a Time-Windowed Shortest-Remaining-First (TWSRFT) policy batches en- coder requests to minimise completion latency and smooth decoder arrivals. Evaluation shows that SpaceServe reduces time-per-output-token by 4.81× on average and up to 28.9× on Nvidia A100 GPUs. SpaceServe is available at https://github.com/gofreelee/SpaceServe
Shuoming Zhang, Xiyu Shi, Yangyu Zhang, Shuaijiang Li, Donglin Yu, Zheming Yang, Yuan Wen, Huimin Cui
NeurIPS5
2025 Advancements in neuromorphic computing for bio-inspired artificial vision: A review
abstract
Neuromorphic computing is revolutionising artificial vision by emulating the human brain’s remarkable efficiency, adaptability, and spatio-temporal processing. This review synthesises recent advances in neuromorphic vision, with a special focus on wave-based dynamics; particularly the role of cortical travelling waves and neural oscillations in coordinating activity across the visual cortex. We examine how these biological mechanisms inspire cutting-edge computational models, including Physics-Informed Neural Networks, reservoir computing, and spiking neural networks, each enabling real-time, energy-efficient visual processing. The review also highlights breakthroughs in hardware, from memristive devices and photonic circuits to optoelectronic polymers, which support in-sensor and event-based processing while dramatically reducing power consumption. By integrating insights from computational neuroscience, materials science, and machine intelligence, we identify persistent challenges; such as scalable training, robust hardware integration, and biologically plausible modelling and outline actionable directions for future research. Our synthesis provides a comprehensive roadmap for the next generation of neuromorphic vision systems, paving the way toward artificial perception that rivals the efficiency and adaptability of biological vision.
Sharmarke A. Gabayre, Mindula Illeperuma, Varuna De Silva, Xiyu Shi, Sergey E. Savel'ev
Neurocomputing4
2021 Scalar Product Lattice Computation for Efficient Privacy-Preserving Systems
abstract
Privacy-preserving (PP) applications allow users to perform online daily actions without leaking sensitive information. The PP scalar product (PPSP) is one of the critical algorithms in many private applications. The state-of-the-art PPSP schemes use either computationally intensive homomorphic (public-key) encryption techniques, such as the Paillier encryption to achieve strong security (i.e., 128 b) or random masking technique to achieve high efficiency for low security. In this article, lattice structures have been exploited to develop an efficient PP system. The proposed scheme is not only efficient in computation as compared to the state-of-the-art but also provides a high degree of security against quantum attacks. Rigorous security and privacy analyses of the proposed scheme have been provided along with a concrete set of parameters to achieve 128-b and 256-b security. Performance analysis shows that the scheme is at least five orders faster than the Paillier schemes and at least twice as faster than the existing randomization technique at 128-b security. Also the proposed scheme requires six-time fewer data compared to the Paillier and randomization-based schemes for communications.
Yo Rahul, Safak Dogan, Xiyu Shi, Rongxing Lu, Muttukrishnan Rajarajan, Ahmet M. Kondoz
IEEE Internet Things J.3
2019 Adaptive blind moving source separation based on intensity vector statistics
Areeb Riaz, Xiyu Shi, Ahmet M. Kondoz
Speech Commun.2
2014 An approach to immersive audio rendering with wave field synthesis for 3D multimedia content
abstract
This paper proposes an immersive audio rendering scheme for networked 3D multimedia systems. The spatial audio rendering method based on wave field synthesis is particularly useful for applications where multiple listeners experience a true spatial soundscape while being free to move without losing spatial sound properties. The proposed approach can be considered as a general solution to the static listening restriction imposed by conventional methods, which rely on an accurate sound reproduction within a sweet spot only. The paper reports on the results of numerical analysis and experimental validation using various sound sources. It is demonstrated and confirmed that while covering the majority of the listening area, the developed approach can create a variety of virtual audio objects at target positions with very high accuracy. Subjective evaluation results show that an accurate spatial impression can be achieved with multiple simultaneous audible depth cues improving localization accuracy over single object rendering.
H. Lim, Erhan Ekmekcioglu, Safak Dogan, Andrew Peter Hill, Ahmet M. Kondoz, Xiyu Shi
ICIP7
2014 Analysis by synthesis spatial audio coding
abstract
This study presents a novel spatial audio coding (SAC) technique, called analysis by synthesis SAC (AbS‐SAC), with a capability of minimising signal distortion introduced during the encoding processes. The reverse one‐to‐two (R‐OTT), a module applied in the MPEG Surround to down‐mix two channels as a single channel, is first configured as a closed‐loop system. This closed‐loop module offers a capability to reduce the quantisation errors of the spatial parameters, leading to an improved quality of the synthesised audio signals. Moreover, a sub‐optimal AbS optimisation, based on the closed‐loop R‐OTT module, is proposed. This algorithm addresses a problem of practicality in implementing an optimal AbS optimisation while it is still capable of improving further the quality of the reconstructed audio signals. In terms of algorithm complexity, the proposed sub‐optimal algorithm provides scalability. The results of objective and subjective tests are presented. It is shown that significant improvement of the objective performance, when compared to the conventional open‐loop approach, is achieved. On the other hand, subjective test show that the proposed technique achieves higher subjective difference grade scores than the tested advanced audio coding multichannel.
Ikhwana Elfitri, Xiyu Shi, Ahmet M. Kondoz
IET Signal Process.2