Xinyu Kang

dblp:217/8318 · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ResKANNet: A residual Kolmogorov-Arnold network with multi-scale attention for brain tumor segmentation
Zhongfeng Kang, Yutong Wang 0004, Xinyu Kang, Shantian Yang
Neurocomputing4
2026 RV-WINO: A RISC-V Neural Network Accelerator Based on Winograd Algorithm Fabricated in 55-nm CMOS Process
abstract
The rapid evolution of artificial intelligence (AI) in IoT applications necessitates the execution of inference tasks on edge devices. However, the deployment of computation-intensive neural networks on resource-constrained edge systems presents a significant challenge. This brief presents the RV-WINO processor, the first silicon implementation of a RISC-V processor based on the Winograd algorithm for convolution and general matrix multiplication (GEMM) acceleration. The processor incorporates a Winograd module, which significantly reduces multiplication operations during convolutions, leading to a substantial decrease in energy consumption. In addition, the processor includes a matrix multiplication module that reuses the multipliers of the Winograd module, accelerating fully connected and dot product operations in neural networks. The RV-WINO processor fabricated in a 55-nm CMOS process achieves the peak computational performance of 0.95 and 2.39 GOPS in INT32 and INT8 modes, with its peak energy efficiency reaching 112 and 237 GOPS/W. In convolutional neural network (CNN) inference tasks, the execution time is reduced by over 80% compared with the baseline processor.
Yucong Huang, Qu Lu, Xinyu Kang, Yuru Li, Qi Wang 0051, Terry Tao Ye
IEEE Trans. Very Large Scale Integr. Syst.4
2025 Logic Gate Network Inference Acceleration with RISC-V Custom Instruction Set
abstract
Logic Gate Networks (LGNs) exploit the similarity between neural networks and logic circuit networks and replace the neurons with logic gates.Consequently, the computation inside the neurons can be replaced by Boolean operations (16 operations for two-input logic).LGNs can be implemented by logic-based instructions in processors and significantly reduce the computation overhead during inference.However, the encoding and decoding processes at the input and output stages of LGNs face efficiency challenges when using traditional RISC-V instruction sets.This limitation arises because these processes rely on one-bit operations, which cannot fully utilize the 32-bit bandwidth of standard instructions.In this work, we proposed four custom RISC-V-based instructions to accelerate the encoding and decoding processes of LGNs.An applicationspecific RISC-V processor, called RV-LGN, has been implemented on FPGA and synthesized using Synopsys® Design Compiler with the CMOS 55nm process.The custom instructions can be called via in-line assembly in C code, making RV-LGN highly promising for implementation in edge devices.Benchmark tests on MIT-BIH, MNIST, and CIFAR-10 classification tasks demonstrate that RV-LGN achieves a runtime reduction of over 87% compared to a generic RISC-V RV32IM ISA processor.Additionally, power consumption during LGN inference is significantly reduced.For the MIT-BIH dataset, the energy consumption is 0.098 µJ/Beat, while MNIST and CIFAR-10 tasks require 0.18 µJ/Image and 0.51 µJ/Image, respectively.These results highlight the superior efficiency of RV-LGN compared to other processors.
Chenxi Feng, Xinyu Kang, Yuru Li, Yucong Huang, Terry Tao Ye
CF3
2025 Design and Evaluation of Farm Farce: AI-Driven Improvisational Comedy in Farming Simulation Games
abstract
Contemporary farming simulations face significant limitations in dynamic humor generation, relying predominantly on static comic scripts that lack personalization and contextual awareness. While procedural content generation has advanced terrain and narrative creation, its application to humor remains underdeveloped-existing systems deploy isolated visual gags without integrating player actions, environmental states, or cultural preferences. Farm Farce addresses this gap through an innovative three-layer AI architecture designed to serve as an improvisational comedy director. The system combines a repository of$\mathbf{2 0 0 +}$humor templates with real-time context sensors and player feedback-driven reinforcement learning, enabling layered absurdity generation such as crops unionizing during harvests or tools developing rebellious personalities in response to player behavior. Experimental validation with 120 players demonstrated clear advantages over alternatives. Participants rated the AI-generated humor at 4.3/5-significantly higher than context-aware systems (3.4/5) and random events (2.1/5). The solution reduced gameplay disruption complaints by 30 % through cultural adaptation, with Asian players showing preference for wordplay-driven events while European users favored anthropomorphic scenarios. Longterm engagement increased substantially, with players exceeding 20 hours of gameplay reporting 58 % higher appreciation for personalized humor. Technically, the architecture maintains 60 FPS performance with 200+ concurrent agents on consumergrade hardware while reducing content production costs by 73% compared to manually scripted alternatives. This research establishes a new paradigm for contextually grounded procedural humor. By resolving critical gaps in crossdomain triggering and cultural intelligence while demonstrating quantifiable improvements in player engagement, Farm Farce transitions game comedy from predetermined routines to dynamic co-creation. The framework's scalability suggests immediate applicability beyond farming simulators to RPGs and life simulations, with future explorations targeting multimodal humor generation through audio and augmented reality interfaces.
Yuantong Yun, Xinyu Kang, Yulin Dong
CoG2
2025 MSA-Former: Multi-scale Adaptive Transformer for Image Snow Removal
Zekun Chen, Shili Liang, Sijia Guo, Xinyu Kang, Huajing Li
MMM (3)6
2025 NNia-8: An 8-Core RISC-V Neural Network Inference Accelerator with Efficient Processing Elements and Memory Utilization
Yucong Huang, Xinyu Kang, Yuru Li, Qi Wang 0051, Terry Tao Ye
NPC (2)3
2025 qLIF: Mitigating the memory and computation overhead to implement spiking convolutional neural networks
Silong Li, Xinyu Kang, Chunlin Yu, Terry Tao Ye
Neural Comput. Appl.3
2025 Spike-Count Reduction Techniques for Low Power Spiking Neural Networks
abstract
Spiking neural network (SNN) has demonstrated its great potential in low-power neuromorphic applications. In SNN, computation activities are associated with the arrival and firing of spikes, its power consumption is directly correlated with the number of spikes propagated in the network. In this paper, we explore two methods to reduce the spike-count in the network, aiming to reduce the power consumption of SNN. We use Poisson distribution function in the input layer and through adjusting the correlation (called the gain in the paper) between the probability of spike generation and input values, the number of spikes in the input layer can be reduced to only 20% of the baseline model with the accuracy degradation of less than 1%. We also exploit the leaky-integrate-and-fire (LIF) mechanism and use the refractory period to reduce the generation of spikes from the neurons in the hidden layers. Through this method, the spike-count is reduced by 20% $$\sim $$ 50% while the performance degradation is still less than 1%. These two spike-count reduction techniques are implemented in Verilog RTL; the power simulation results demonstrate significant power reduction in performing SNN computations. We further discovered that for different network architectures, these two techniques have different trade-offs to achieve optimal spike-count reduction while maintaining satisfactory results. Compared with other spike-count reduction techniques, the proposed scheme is efficient and straightforward for hardware implementation, making it well-suited for edge computing scenarios.
Xinyu Kang, Zhitao Yang, Terry Tao Ye
Neural Process. Lett.1
2025 RV-SCNN: A RISC-V Processor With Customized Instruction Set for SNN and CNN Inference Acceleration on Edge Platforms
abstract
The rapid advancement of artificial intelligence (AI) applications has driven an increasing demand for conducting inference tasks on edge devices. However, implementing computation-intensive neural networks on resource-constrained edge systems remains a significant challenge. In this article, we propose a novel processor architecture called RV-SCNN to address this challenge. The architecture is based on the RISC-V generic instruction set and incorporates various single instruction multiple data (SIMD) custom instruction extensions to accelerate the computation of spike neural networks (SNNs) and convolutional neural networks (CNNs), enabling efficient execution of complex neural network models. The core operators of the processor are shared by both SNN and CNN operations, thus supporting both computation modes. Other acceleration implementations include an internal hardware loop control unit that reduces the instruction overhead, an address calculation unit and an interlayer fusion unit that minimize the memory access overhead, as well as an image to column (IM2COL) unit that improves the computational efficiency of the$3 \times 3$convolutions in SNNs and CNNs. The custom instructions are called through inline assembly in the C program, providing higher flexibility compared to traditional ASICs and supporting custom complex SNN/CNN network structures. Compared to traditional instruction sets, the RV-SCNN processor reduces the execution time of CNNs and SNNs by over 90%. We validate the processor on FPGA platform and evaluate its performance under CMOS 55-nm process. The processor achieves an operational efficiency of 9.88 pJ/SOP in SNN network inference tasks, while the peak energy efficiency reaches 679 GOPS/W in CNN network inference.
Chenxi Feng, Xinyu Kang, Qi Wang 0051, Yucong Huang, Terry Tao Ye
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2025 Joint Retrieval of Ozone Profile in Near Space Based on the Atmospheric and Near Infrared Atmospheric Bands of O2 Airglow
abstract
The vertical distribution of ozone in the near space is a key scientific issue for revealing atmospheric photochemical processes and energy transmission mechanisms. However, given the complex environment in near space and limitations in detection methods, realizing high-precision and wide-range ozone profile retrieval with traditional methods is difficult. This study innovatively proposes a joint retrieval method for near-space ozone profiles based on the theory of O2 molecular airglow spectroscopy. The research establishes an airglow photochemical model based on the generation and loss mechanisms of O2 (a${}^{1} \Delta _{\mathrm { g}}$) and O2 (b${}^{1} \Sigma ^{+}_{\mathrm { g}}$) states and combines Scanning Imaging Absorption Spectrometer for Atmospheric CHartographY limb-viewing data to achieve accurate retrieval of ozone distribution in the near space. The retrieval results are cross-validated using multisource satellite observation data, such as those on Sounding of the Atmosphere using Broadband Emission Radiometry (SABER) and Michelson Interferometer for Passive Atmospheric Sounding (MIPAS). The findings demonstrate that the joint retrieval method successfully retrieves ozone concentrations in the 30–130-km altitude range, with overlapping spectral bands providing mutual verification. The retrieved ozone concentrations exhibit a deviation of less than 1 part per million (ppm) by volume compared with the SABER data and a linear correlation coefficient of 0.96 compared with MIPAS data in the 40–110-km range. These findings fully validate the effectiveness and reliability of the retrieval method and offer a novel technical approach for ozone detection in the near-space region.
Zhongyi Fu, Daoqi Wang, Xinyu Kang, Kuijun Wu
IEEE Trans. Geosci. Remote. Sens.7
2024 Spatiotemporal Meteorological Prediction Based on Fully Convolutional Neural Network
abstract
Accurate prediction of meteorological data is critical for enhancing the capacity to respond to climate change, reducing disaster risks, and ensuring the sustainable development of human society. However, existing data-driven methods fail to meet the forecasting demand regarding accuracy and efficiency. Therefore, this study introduced spatiotemporal meteorological prediction UNet (STMP-UNet), a UNet-based spatiotemporal meteorological prediction network. The network used an attention-optimized spatial UNet to extract global-local spatial connections. Additionally, it utilized a multiscale temporal pyramid gating module (TPGM) to capture the evolution patterns of meteorological spatial features. Meanwhile, we constructed a spatiotemporal composite loss function based on mean square error and temporal differentiation to better fit the complex nonlinear evolutionary patterns in the meteorological domain. Comparison and ablation experiments demonstrated that the proposed model and its modules in this study can better predict meteorological data for the next three and five days. The meteorological model trained on data from Northeast China is applied to predict conditions in the Pacific region, confirming the transportability of STMP-UNet and the robustness of the multiscale prediction approach. Finally, STMP-UNet achieves weather data prediction for the next five days in just 38 ms, demonstrating its superior efficiency compared to other models. Therefore, the STMP-UNet model can be used as an effective tool for predicting meteorological data, aiding decision-making in meteorological-related fields.
Mingyang Hua, Zekun Chen, Shili Liang, Xinyu Kang
IEEE Trans. Geosci. Remote. Sens.6
2021 Model Composition: Can Multiple Neural Networks Be Combined into a Single Network Using Only Unlabeled Data?
Amin Banitalebi-Dehkordi, Xinyu Kang, Yong Zhang 0004
BMVC2
2021 SimROD: A Simple Adaptation Method for Robust Object Detection
abstract
This paper presents a Simple and effective unsupervised adaptation method for Robust Object Detection (SimROD). To overcome the challenging issues of domain shift and pseudo-label noise, our method integrates a novel domain-centric data augmentation, a gradual self-labeling adaptation procedure, and a teacher-guided fine-tuning mechanism. Using our method, target domain samples can be leveraged to adapt object detection models without changing the model architecture or generating synthetic data. When applied to image corruptions and high-level cross-domain adaptation benchmarks, our method outperforms prior baselines on multiple domain adaptation benchmarks. SimROD achieves new state-of-the-art on standard real-to-synthetic and cross-camera setup benchmarks. On the image corruption benchmark, models adapted with our method achieved a relative robustness improvement of 15-25% AP50 on Pascal-C and 5-6% AP on COCO-C and Cityscapes-C. On the cross-domain benchmark, our method outperformed the best baseline performance by up to 8% and 4% AP50 on Comic and Watercolor respectively.1
Rindranirina Ramamonjison, Amin Banitalebi-Dehkordi, Xinyu Kang, Xiaolong Bai, Yong Zhang 0004
ICCV3