Xinglong Ji

dblp:235/3823 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0002-8116-7974ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Optimizing Spatial Data Structure with Near-Cache Acceleration by Exploiting Physical Locality
Haoran Pei, Zijian Pan, Songchen Ma, Leshan Li, Xinglong Ji
ISCA9
2026 Efficient on-device continuous learning enabling evolutionary edge intelligence
Xinglong Ji
Neurocomputing5
2026 Espresso: Exploiting the Sparsity Property in Brain-Inspired Vision Sensors With Spatiotemporal Ordering
abstract
Brain-inspired vision sensors (BVSs), drawing inspiration from the human visual system, produce sparse, high-temporal-resolution data stream capable of capturing rapid object motion. However, real-time processing of such data while preserving its inherent sparsity presents a critical challenge for practical deployment. A key challenge is to leverage the spatiotemporal correlations among event stream. Overlooking these correlations leads to prohibitively high query costs that even negate the benefits of sparsity, while exploiting them introduces complications such as input-output order conflicts and trade-offs between memory usage and latency. To address this, we present Espresso, an efficient hardware architecture that leverages the spatiotemporal order of events while explicitly preserving sparsity, achieving low-latency stream processing of event data. We first formalize a spatiotemporal order representation that identifies key features for stream processing on sparse events. Building on this, Espresso decouples output window address from input event address, resolving order conflicts via a dedicated queue mechanism and minimizing memory overhead through an optimized hash table. This enables immediate window-wise processing with minimal latency. To coordinate the pipeline, we design the Event-Scheduler, a streamlined finite state machine that prunes computations on zero values and aligns input-output stream order discrepancies. Integrated together, these modules deliver scalable, high-throughput processing for event-driven vision tasks. Espresso achieves up to 5000 fps, offering a 5.1× performance improvement over embedded GPUs. With parallel instantiations, it exceeds over 7000 fps in structured scenes and maintains over 2000 fps under complex-environment scenarios with minimal hardware overhead. These results establish Espresso as an efficient and scalable solution for real-time event-based vision processing, demonstrating the importance of spatiotemporal ordering in unlocking the full potential of BVSs.
Leshan Li, Taoyi Wang, Mingtao Ou, Xinglong Ji
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.8
2025 Espresso: Exploiting the Sparsity Property in Event Sensors with Spatiotemporal Ordering
abstract
Event-based vision sensors are novel cameras inspired by human eyes, capable of capturing the rapid motion of objects with the high-speed sparse event stream. However, it is challenging to efficiently stream process event data without demolishing its sparsity. In this paper, we design Espresso, an architecture for event-based vision processing that can efficiently stream spatiotemporal events while preserving sparsity. We implement Espresso on the FPGA platform and design experiments to compare the performance with embedded GPU and line-buffer-based accelerator. In real-world scenarios, Espresso achieves throughput up to 5000 fps, which is $5.1 \times$ higher than the embedded GPU.
Leshan Li, Mingtao Ou, Xinglong Ji
DAC6
2025 RoboSpike: Fully Utilizing the Heterogeneous System With Subcallback Scheduling in ROS 2
abstract
The advancement in artificial intelligence (AI) has greatly propelled the development of robotics, requiring the adoption of heterogeneous computing architectures with multicore CPUs, GPUs, and accelerators to meet the growing computational needs of edge computing. Such heterogeneity, coupled with the inherently IO-intensive nature of robotic applications, poses substantial challenges for task scheduling and resource management. These challenges are particularly acute for systems striving to maximize computational resource utilization, which cannot be effectively addressed through callback-level scheduling. To overcome these obstacles, we developed RoboSpike, a systematic solution built on the Robot Operating System 2 (ROS 2). We first implemented a subcallback scheduling mechanism utilizing coroutines to utilize the blocked CPUs which wait for I/O operations. Building on this mechanism, we extended the design to incorporate the coprocessor and introduced an auto-tuning algorithm to adapt to system performance variations. Finally, we performed the response time analysis to ensure that the RoboSpike is predictable in time. The evaluation results demonstrate that RoboSpike achieves substantial improvements, increasing throughput by 1.65–2.25 times in real-world scenarios. RoboSpike enhances the scheduling capabilities of ROS 2 by refining the granularity from the callback level, thus opening up new opportunities for performance improvement in robotic systems, especially in resource-limited scenarios with complex workloads.
Songchen Ma, Xinglong Ji
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2024 Multiple-local feature and attention fused person re-identification method
abstract
Person re-identification (ReID) is widely used in intelligent security, monitoring, criminal investigation and other fields. Aiming at the problems of local occlusion, scale misalignment and attitude change of pedestrian images in actual scenes, we propose a Multi-local Feature and Attention fused network (MFA) used for person re-identification task. Firstly, Channel Point Affinity Attention module (CPAA) is embedded in the backbone network to enhance the ability of the network for extracting local details. The feature map output from the backbone network is horizontally segmented into four local feature maps, and further four branch networks are concatenated to the feature map of the backbone network. The four local feature maps are used to guide the four branch networks to pay more attention on different areas of pedestrians through Global Local Aligned loss (GLA) function. Finally, the pedestrian feature vector containing multi-local features is obtained. The mAP of the network on Market-1501, DukeMTMC-reID,CUHK03 and MSMT17 datasets were 88.6%, 81.4%, 79.5% and 64.7%, and the Rank-1 was 95.8%, 90.1%, 81.2% and 84.1% respectively. In addition, the model also obtained 73.2% and 68.1% of Rank-1 on partial dataset Patial-REID and Patial-iLIDS, respectively. Recently, The MFA model parameter is 28.3M and the inference efficiency is approximately 32 fps to an image with a resulation of 256 × 128. Compared with other ReID methods, our proposed methods achieved a competitive performance for ReID task. The code was available at github:[email protected]:ISCLab-Bistu/MFA.git.
Mingxin Yu, Rui You, Xinglong Ji, Wenshuai Lu
Intell. Data Anal.4
2024 SongC: A Compiler for Hybrid Near-Memory and In-Memory Many-Core Architecture
abstract
Building hybrid systems that incorporate various processing-in-memory (PIM) devices and processing-near-memory (PNM) technologies can offer complementary advantages in both efficiency and flexibility, while many-core architectures show great potential in deploying data-centric parallel applications with high performance. Compilers for the hybrid PN/IM architecture are critical for enabling such computing systems to be put into practical use. However, most of the existing neural network compilers for PIM or PNM are optimized from the perspective of an operator, and cannot effectively take advantage of a decentralized core-level dataflow with large on-chip memory access bandwidth. Here, we propose a full-stack System-on-graph Compiler (SongC) framework for many-core architecture, which optimizes the efficiency of the PIM devices and leverages the flexibility of the PNM architectures. SongC establishes multi-level graph abstractions to clarify the critical deployment challenges at different levels and generalizes the standard optimizations, decoupling versatile algorithms and diverse types of hardware. To handle the complexity of many-core resource utilization, we also establish a simulation-compilation interaction flow, including a just-in-time evaluator to boost the scheduling search and an extended Roofline model, referred to as the Palace model, to guide the search. Experiments demonstrate the various optimizations and overall performance of SongC and reveal the capability of strategy exploration.
Junfeng Lin, Huanyu Qu, Songchen Ma, Xinglong Ji, Xiaochuan Li 0003, Chenhang Song
IEEE Trans. Computers4
2023 A lightweight vision transformer with symmetric modules for vision tasks
abstract
Transformer-based networks have demonstrated their powerful performance in various vision tasks. However, these transformer-based networks are heavyweight and cannot be applied to edge computing (mobile) devices. Despite that the lightweight transformer network has emerged, several problems remain, i.e., weak feature extraction ability, feature redundancy, and lack of convolutional inductive bias. To address these three problems, we propose a lightweight visual transformer (Symmetric Former, SFormer), which contains two novel modules (Symmetric Block and Symmetric FFN). Specifically, we design Symmetric Block to expand feature capacity inside the module and enhance the long-range modeling capability of attention mechanism. To increase the compactness of the model and introduce inductive bias, we introduce convolutional cheap operations to design Symmetric FFN. We compared the SFormer with existing lightweight transformers on several vision tasks. Remarkably, on the image recognition task of ImageNet [13], SFormer gains 1.2% and 1.6% accuracy improvements compared to PVTv2-b0 and Swin Transformer, respectively. On the semantic segmentation task of ADE20K [64], SFormer delivers performance improvements of 0.2% and 0.7% compared to PVTv2-b0 and Swin Transformer, respectively. On the cityscapes dataset [11], SFormer delivers performance improvements of 2.5% and 4.2% compared to PVTv2-b0 and Swin Transformer, respectively. The code is open-source and available at: https://github.com/ISCLab-Bistu/Symmetric_Former.git.
Shengjun Liang, Mingxin Yu, Wenshuai Lu, Xinglong Ji, Xiongxin Tang, Rui You
Intell. Data Anal.4
2022 In Situ Aging-Aware Error Monitoring Scheme for IMPLY-Based Memristive Computing-in-Memory Systems
abstract
Stateful logic through memristor is a promising technology to build Computing-in-Memory (CIM) systems. However, aging-induced degradation of memristors’ threshold voltage imposes a major challenge to the reliability and guardbands estimation of memristive CIM systems, especially the Material Implication (IMPLY) logic based CIM systems. In this paper, a novel in-situ aging-aware error monitoring scheme for memristor-based IMPLY logic is proposed. The proposed in-situ error monitoring scheme can achieve faster error detection speed and higher detection accuracy than the straightforward program-verify monitoring scheme. Simulation results under Monte-Carlo simulation show that the proposed monitoring scheme can effectively detect the major operation failures existing in IMPLY logic operations with a detection accuracy up to 99.95%. Moreover, a case study of error monitoring design of 4-bit IMPLY-based adder is carried out. The analysis result exhibits that the proposed in-situ monitoring scheme can achieve 75.2% improvement on the detection speed against the program-verify scheme. Further analysis on a convolution filter in VGG-11 based Binarized Neural Network shows that 74% improvement on the detection speed can also be achieved by using the proposed monitoring scheme, which suggests that the proposed in-situ error monitoring scheme is an efficient solution to improve the reliability of IMPLY-based memristive CIM systems.
Jiajun Wu 0006, Xinglong Ji, Guoyi Yu, Chao Wang 0096
IEEE Trans. Circuits Syst. I Regul. Pap.5
2021 Efficient Design of Spiking Neural Network With STDP Learning Based on Fast CORDIC
abstract
In emerging Spiking Neural Network (SNN) based neuromorphic hardware design, energy efficiency and on-line learning are attractive advantages mainly contributed by bio-inspired local learning with nonlinear dynamics and at the cost of associated hardware complexity. This paper presents a novel SNN design employing fast COordinate Rotation DIgital Computer (CORDIC) algorithm to achieve fast spike timing–dependent plasticity (STDP) learning with high hardware efficiency. In this study, a system design and evaluation method of CORDIC-based SNN is proposed for finding optimal CORDIC type and precision, from theoretical CORDIC-level error to application-level learning performance. From the proposed design and evaluation method, a reconfigurable SNN design based on fast-convergence CORDIC is designed to achieve high classification accuracy on MNIST, fast on-line learning and good energy efficiency. By utilizing SNN’s fault tolerance and time-division-multiplexing (TDM) strategy, the reconfigurable SNN design employs 8-bit fast-convergence CORDIC and TDM-based hardware accelerator for high efficiency. FPGA implementation results confirm that the proposed fast-convergence CORDIC SNN design outperforms the state-of-the-art CORDIC method by 38.5%−45.3% in terms of learning speed and energy efficiency, with the STDP learning of 30.2 ns/SOP, energy efficiency of 176.6 pJ/SOP, processing speed of 6.1 ms/image, and on-line learning convergence of 21.4 s (time to reach the final accuracy, on average), on MNIST benchmark.
Jiajun Wu 0006, Zixuan Peng, Xinglong Ji, Guoyi Yu, Chao Wang 0096
IEEE Trans. Circuits Syst. I Regul. Pap.4