Shijie Li 0002

dblp:141/7586-2 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0002-0529-0057ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2026 An Energy-Efficient 0.56-pJ/cycle AVFS System Based on a Fast Transient Response Digital LDO and a Self-Calibrating Elastic Clock
Jiliang Liu, Zhengbin Pang, Fangxu Lv, Shijie Li 0002, Qiang Wang 0006, Lizhou Wu, Chengzhuo Zhao
ISCAS4
2025 FusionSVFilter: A Deep-Learning Based Fast Structural Variation Filtering Tool for Long Reads
abstract
Structural variations (SVs) play a critical role in species diversity, biological evolution, and human diseases. Although third-generation sequencing technology has enhanced the ability to detect long and complex SVs through long reads that can span complex genomic regions, it still suffers from high false-positive (FP) SV calls due to three factors: the complexity of SVs, limitations of detection algorithms, and relatively high single-base error rates in long reads. To address this challenge, we propose FusionSVFilter, a deep learning-based algorithm for SV filtering. It first transforms genomic sequence features into multi-level grayscale images via sequence-to-image encoding, which captures the structural complexity of SVs. Moreover, the encoder is accelerated with parallel computing to improve speed. Afterward, these grayscale images are enhanced to retain key features and converted into RGB format to meet the input requirements of the model. In addition, FusionSVFilter further leverages transfer learning, initializing with a pre-trained ResNet50 model that is fine-tuned on our curated dataset to recognize SV-specific patterns. Experimental results show that FusionSVFilter sub-stantially reduces FP calls while maintaining true-positive (TP) detection at near-constant levels. The code and documentation are available at: https://github.com/nudt-bioinfo/FusionSVFilter.
Xinghai Zeng, Tao Tang 0001, Shijie Li 0002, Yingbo Cui 0001
BIBM3
2025 A Novel High-Speed Adaptive Duobinary Digital Detector Based on the Feed-Forward Equalizer and the Maximum Likelihood Sequence Detector for Wireline Transceivers
abstract
To solve the high bit error rate (BER) problem of conventional 56-Gb/s nonreturn-to-zero (NRZ) transceivers under high-insertion loss (IL) channels, this study proposes a high-speed adaptive duobinary (DB) digital detector based on the feed-forward equalizer (FFE) and the maximum likelihood sequence detector (MLSD). In this detector, adaptive FFE is combined with channel characteristics to generate DB signals and complete equalization, thus extending the transmission bandwidth and eye height and allowing a larger sampling phase offset. The parallel MLSD is used to complete the detection and decoding of DB signals to reduce the BER. An adaptive algorithm is proposed to avoid the long convergence time of the conventional zero-forcing (ZF) algorithm applied to the DB detector, so that it can be applied to various bit rates and IL channels. In this study, the verification of this DB detector is accomplished at 56 Gb/s. The platform based on a 56-Gb/s analog front-end chip (AFEC) and field-programmable gate array (FPGA) proves that the detector can work well in 12–56 Gb/s and multiple IL channels. The BER was less than 2e-8 at 56 Gb/s on −42-dB channel loss at 28 GHz. The structure can be well used for higher rate transceivers, such as 112 Gb/s.
Chaolong Xu, Fangxu Lv, Xingyun Qi, Qiang Wang 0006, Zhang Luo, Shijie Li 0002, Geng Zhang 0001
IEEE Trans. Very Large Scale Integr. Syst.7
2024 TianheStar: Orchestrating SSSP Applications on Tianhe Supercomputer
abstract
Computing single-source shortest paths (SSSP) is one of the fundamental problems in graph theory and is also essential for data-intensive applications. As the potential of artificial intelligence (AI) continues to be explored, and with the advent of exascale supercomputing, there is a growing need for an extremely fast graph engine for SSSP applications. Current distributed SSSP engines for large-scale graph applications, unfortunately, often exhibit poor efficiency when running on supercomputers. In this paper, we introduce TianheStar, an ultra-fast SSSP engine designed specifically for graph search on the Tianhe supercomputer. TianheStar effectively minimizes communication costs and establishes a new balance between computation and communication. The key idea of TianheStar is to leverage network topology information for performing topology-aware message aggregation and architecture-aware group communication. These two techniques effectively reduce the number of messages and the average number of communication hops, respectively. We validate TianheStar using Graph500, a widely adopted benchmark for graph search on supercomputers. Extensive evaluation demonstrates that, compared to the state-of-the-art solutions, TianheStar achieves a remarkable performance improvement. We have deployed TianheStar on the latest Tianhe supercomputer and secured the top position in the latest Graph500. We achieved an outstanding performance of 23,021 GTEPS (Giga Traversed Edges Per Second) for SSSP using 4096 nodes. Furthermore, we have delved into real-world graphs representing the USA road networks and conducted computations to determine the shortest paths between vertices. Our experimental results demonstrate that TianheStar can traverse the USA road network, comprising over 58,333,344 edges, in less than 0.1 second on the Tianhe supercomputer. This performance represents a speedup of over a thousand times compared to parallel shortest-path graph computations on the Aziz supercomputer, a globally renowned high-performance computing system, using the same input data.
Xinbiao Gan, Shijie Li 0002, Bo Yang 0023
CCGrid4
2024 MST: Topology-Aware Message Aggregation for Exascale Graph Processing of Traversal-Centric Algorithms
abstract
This article presents MST, a communication-efficient message library for fast graph traversal on exascale clusters. The key idea is to follow the multi-level network topology to perform topology-aware message aggregation, where small messages are gathered and scattered at each level of domain. To facilitate message aggregation, we equip MST with flexible buffer management including active buffer switching and dynamic buffer expansion. We implement MST on the newest-generation Tianhe supercomputer and evaluated its performance using various traversal-centric algorithms on both synthetic trillion-scale graphs and real-world big graphs. The results show that MST-based graph traversal is orders of magnitude faster than that based on Active Messages Library (AML). For the Graph500-BFS benchmark, MST-based Tianhe (with 77.2 K nodes) outperforms the Fugaku supercomputer (with 148.5 K nodes) by 18.53%, while Fugaku is ranked No. 1 in the latest Graph500-BFS ranking (June 2023). MST also greatly improves graph processing performance on other commercial large-scale computing systems at the National Supercomputing Center in Changsha (NSCC) and WuzhenLight.
Xinbiao Gan, Bo Yang 0023, Xinhai Chen 0001, Chunye Gong, Shijie Li 0002, Kai Lu 0001, Qiao Li 0001, Yiming Zhang 0003
ACM Trans. Archit. Code Optim.7
2023 MTMap: A Long-Read Alignment Tool based on Multi-Core DSPs
abstract
Read alignment is a basic and important task in genomic data analysis. The popularity of the third—generation sequencing technology has brought the need of sequence alignment algorithms to analyze long-read sequences with longer read length and high error rate. Moreover, the rapid growth of sequence data has also presented challenges for read alignment. To improve the ability to process large volume of sequencing reads, we developed a long-read sequence alignment algorithm MTMap on the heterogeneous processor FT-m7032. MTMap utilizes multi-level parallel technologies: firstly, we tailored the data structure for the wide vector processing units of DSP to speedup the score matrix computation. Secondly, we developed multithread parallelization for base-level alignment on each DSP cluster. Finally, we implemented multi-process parallelization between DSP clusters to fully exploit the computing power of FT-m7032. Experiments show that, MTMap achieves up to 16 times of parallel acceleration performance compared with the original algorithm under the condition of ensuring accuracy.
Xinjie An, Shijie Li 0002, Yingbo Cui 0001, Peng Zhang 0061, Biao Long
BIBM3
2019 Exploring frame segmentation networks for temporal action localization
Ke Yang 0004, Xiaolong Shen, Peng Qiao, Shijie Li 0002, Dongsheng Li 0001, Yong Dou
J. Vis. Commun. Image Represent.4
2018 mmCNN: A Novel Method for Large Convolutional Neural Network on Memory-Limited Devices
abstract
Deep learning recently has been widely used in many interactive application fields including but not limited to object recognition, speech recognition, natural language processing and so on. At the same time more and more attractive interactive applications (face recognition and augmented reality) are available on wearable and mobile devices. However, traditional deep learning methods such as CNN cost a lot of memory resources. This challenge makes it difficult to apply the powerful deep learning method on mobile memory limited platforms. In this paper we present a novel memory management strategy called mmCNN to solve this problem. This method helps us deploy a trained large size CNN on an any memory size platform including GPU, FPGA and memory-limited mobile devices. In our experiments, we run a feed-forward CNN process in an extremely small memory size (as low as 5MB) on a GPU platform. The result shows that our method saves more than 98% memory compared to a traditional CNN algorithm and further saves more than 90% compared to the sate-of-the-art related work "vDNN". Our work improve the computing scalability of interaction applications and break the memory bottleneck of using deep learning method on a memory-limited devices.
Shijie Li 0002, Yong Dou, Jinwei Xu, Qiang Wang 0006, Xin Niu 0002
COMPSAC (1)1
2018 Deep Image Clustering Using Convolutional Autoencoder Embedding with Inception-Like Block
abstract
Image clustering is one of the challenging tasks in machine learning, and has been extensively used in various applications. Recently, various deep clustering methods has been proposed. These methods take a two-stage approach, feature learning and clustering, sequentially or jointly. We observe that these works usually focus on the combination of reconstruction loss and clustering loss, relatively little work has focused on improving the learning representation of the neural network for clustering. In this paper, we propose a deep convolutional embedded clustering algorithm with inception-like block (DCECI). Specifically, an inception-like block with different type of convolution filters are introduced in the symmetric deep convolutional network to preserve the local structure of convolution layers. We simultaneously minimize the reconstruction loss of the convolutional autoencoders with inception-like block and the clustering loss. Experimental results on multiple image datasets exhibit the promising performance of our proposed algorithm compared with other competitive methods.
Qiang Wang 0006, Rongchun Li, Peng Qiao, Ke Yang 0004, Shijie Li 0002, Yong Dou
ICIP6
2017 A fast and memory saved GPU acceleration algorithm of convolutional neural networks for target detection
Shijie Li 0002, Yong Dou, Xin Niu 0002, Qiang Wang 0006
Neurocomputing1
2017 Heterogeneous blocked CPU-GPU accelerate scheme for large scale extreme learning machine
Shijie Li 0002, Xin Niu 0002, Yong Dou, Yueqing Wang
Neurocomputing1
2017 Throughput-Optimized FPGA Accelerator for Deep Convolutional Neural Networks
abstract
Deep convolutional neural networks (CNNs) have gained great success in various computer vision applications. State-of-the-art CNN models for large-scale applications are computation intensive and memory expensive and, hence, are mainly processed on high-performance processors like server CPUs and GPUs. However, there is an increasing demand of high-accuracy or real-time object detection tasks in large-scale clusters or embedded systems, which requires energy-efficient accelerators because of the green computation requirement or the limited battery restriction. Due to the advantages of energy efficiency and reconfigurability, Field-Programmable Gate Arrays (FPGAs) have been widely explored as CNN accelerators. In this article, we present an in-depth analysis of computation complexity and the memory footprint of each CNN layer type. Then a scalable parallel framework is proposed that exploits four levels of parallelism in hardware acceleration. We further put forward a systematic design space exploration methodology to search for the optimal solution that maximizes accelerator throughput under the FPGA constraints such as on-chip memory, computational resources, external memory bandwidth, and clock frequency. Finally, we demonstrate the methodology by optimizing three representative CNNs (LeNet, AlexNet, and VGG-S) on a Xilinx VC709 board. The average performance of the three accelerators is 424.7, 445.6, and 473.4GOP/s under 100MHz working frequency, which outperforms the CPU and previous work significantly.
Yong Dou, Jingfei Jiang, Jinwei Xu, Shijie Li 0002, Yongmei Zhou, Yingnan Xu
ACM Trans. Reconfigurable Technol. Syst.5
2016 Multi-view clustering with extreme learning machine
Qiang Wang 0006, Yong Dou, Xinwang Liu 0002, Shijie Li 0002
Neurocomputing5