Guoxiang Li

dblp:181/9553 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RV-CIM: Energy-Delay Optimized Mapping and Architecture Co-Design for a RISC-V Multi-core SoC with Configurable DCIM Cluster
Ninghui Shang, Ying Liu 0069, Zecheng Zhou, Jiyong Hu, Zhiqiang Guo, Zhiyuan Chen 0009, Guoxiang Li, Yufei Ma 0002, Le Ye
APPT8
2026 Adaptive elitism election algorithm for B-type random 2-satisfiability in discrete hopfield neural network
Guoxiang Li, Mohd Shareduwan Mohd Kasihmuddin, Nurul Atiqah Romli, Baorong Yu, Zhaoqin Cao, Yueling Guo, Yujian Gan
Neurocomputing2
2026 Quartet: A Digital Compute-in-Memory Versatile AI Accelerator With Heterogeneous Tensor Engines and Off-Chip-Less Dataflow
abstract
Although most AI core operations can be formulated as matrix multiplications (MMs), their characteristics are quite different. In some cases, one input may be a constant weight matrix while the other is a dynamic feature matrix, i.e., FWMM, or both inputs may be dynamic features as FFMM. Furthermore, the data involved can also be sparse, leading to variations like SpFWMM and SpFFMM. To address this challenge, this paper investigates a versatile accelerator architecture for AI algorithms based on heterogeneous tensor engines. For MM operators with varying characteristics, this paper proposes four-quadrant heterogeneous tensor engines to handle FWMM, SpFWMM, FFMM, and SpFFMM, respectively. These four tensor engines are comprised of SRAM-based single address digital compute-in-memory (CIM) array, SRAM-based multi-address digital CIM array, systolic array, and multi-SIMD array, respectively. In addition, to improve the AI execution efficiency, this paper proposes a dual-level multi-issue mechanism to achieve inter-operator and inter-block parallelization along with an off-chip-less dataflow enabled by on-chip unified memory pool. Thanks to the integration of aforementioned innovations, this paper develops a versatile AI acceleration chip Quartet, which achieves exceptional energy efficiency. Specifically, for graph convolutional network on PubMed, it demonstrates a$19.56\times $and$3.47\times $improvement in energy efficiency compared to similar works, ReDCIM and TensorCIM, respectively.
Yikan Qiu, Guoxiang Li, Meng Wu 0005, Yifan Jia 0009, Le Ye, Yufei Ma 0002
IEEE Trans. Circuits Syst. I Regul. Pap.2
2025 3D-SubG: A 3D Stacked Hybrid Processing Near/In-Memory Accelerator for Subgraph GNNs
abstract
Subgraph Graph Neural Networks (GNNs) are emerging as a promising approach to enhance GNN expressiveness, but their more complex graph structures with numerous independent and irregular subgraphs pose significant hardware deployment challenges. In this work, we propose 3D-SubG, a 3D stacked hybrid processing-near/in-memory accelerator for subgraph GNNs. With hybrid bonding packaging technology, a logic die is 3D stacked with a DRAM die for highly parallel memory accesses. The logic die employs digital SRAM-based processing-in-memory (PIM) macros to boost computation density and minimize data transfer. We further propose a bit-level non-zero gathering method to exploit graph sparsity for PIM, a workloadbalanced mapping strategy for subgraph allocation onto different logic-to-DRAM blocks, and a distributed global pooling approach to reduce inter-block data movements. Experimental results show that 3D-SubG achieves average improvements of $146.11 \times$ in performance, $934.18 \times$ in area efficiency, and $1171.80 \times$ in energy efficiency compared to RTX 3090Ti.
Guoxiang Li, Runnan Xu, Ruohang Xu, Yikan Qiu, Renati Tuerhong, Muhan Zhang, Le Ye, Yufei Ma 0002
DAC1
2025 T-NER: Triple embedding for Chinese Named Entity Recognition
abstract
Named Entity Recognition (NER) aims to automatically extract specific entities from unstructured text. Compared with English NER, Chinese NER faces challenges due to heterophony, where the same Chinese character may have different pronunciations and meanings. Additionally, the lack of clear separators between Chinese characters exacerbates these challenges, leading to difficulties in boundary detection and entity category determination. Inspired by the hieroglyphic and phonetic features of Chinese characters, this study proposes a multi-feature fusion embedding model (T-NER). The model employs CNN for extracting radicals and phonetic features of Chinese characters, combines the encoded information from these features with pre-trained characters vectors to generate fusion embedding vectors, and uses a fully-connected layer for feature transformation. Experiments were conducted on the Chinese benchmark datasets Resume and Weibo. Compared to current mainstream models, the proposed model demonstrates superior performance in terms of F1 score, F1 score stability, and individual entity recognition accuracy. Compared to MFE-NER, T-NER improves F1 score by +0.30% on the Resume dataset and +2.16% on the Weibo dataset. Ablation experiments further validate the effectiveness of the introduced radicals and phonetic features. The experimental results demonstrate that this model effectively captures the semantic information of Chinese characters, addresses the problem of Chinese character heterophony, and improves entity recognition performance.
Guopeng Cheng, Yushan Guo, Guoxiang Li, Shuanghong Qu, Bin Ai
IJCNN5
2025 Secret image restoration with high-bit correction and symbiotic organisms search
Jianzhong Yang, Xianquan Zhang, Chunqiang Yu, Guoxiang Li, Zhenjun Tang
Expert Syst. Appl.4
2025 Secret image restoration with interpolation and social network search
Jianzhong Yang, Xianquan Zhang, Chunqiang Yu, Xuemao Zhang, Guoxiang Li, Zhenjun Tang
Neurocomputing5
2025 Reversible data hiding in encrypted images using prediction error modification and basic block compression
Xuemao Zhang, Xianquan Zhang, Chunqiang Yu, Guoxiang Li, Zhenjun Tang
Signal Process.4
2025 Reversible Data Hiding in Encrypted Images With Secret Sharing and Multivariate Linear Equation
Chunqiang Yu, Xianquan Zhang, Guoxiang Li, Peng Liu 0044, Xinpeng Zhang 0001, Zhenjun Tang
IEEE Trans. Dependable Secur. Comput.3
2024 An In-Memory Computing Accelerator with Reconfigurable Dataflow for Multi-Scale Vision Transformer with Hybrid Topology
abstract
Transformer models equipped with multi-head attention (MHA) mechanism have demonstrated promise in computer vision (CV) tasks, i.e., vision transformers (ViTs). Nevertheless, the lack of inductive bias in ViTs leads to substantial computational and storage requirements, hindering their deployment on resource-constrained edge devices. To this end, multi-scale hybrid models are proposed to take the advantages of both transformers and convolutional neural networks (CNNs). However, existing domain-specific architectures focus on the optimization of either convolution or MHA at the expense of flexibility. In this work, an in-memory computing (IMC) accelerator is proposed to efficiently accelerate ViTs with hybrid MHA and convolution topology by introducing pipeline reordering. SRAM-based digital IMC macro is utilized to mitigate memory access bottleneck, while avoiding analog non-ideality. The reconfigurable processing engines and interconnections are investigated to enable the adaptable mapping of both convolution and MHA. Under typical workloads, experimental results exhibit that our proposed IMC architecture delivers 2.20× to 2.52× speedup and 40.6% to 74.8% energy reduction compared with the baseline design.
Zhiyuan Chen 0009, Yufei Ma 0002, Yifan Jia 0009, Guoxiang Li, Meng Wu 0005, Le Ye, Ru Huang 0001
DAC5
2024 DCIM-GCN: Digital Computing-in-Memory Accelerator for Graph Convolutional Network
abstract
Graph convolutional network (GCN) has gained great success in a diverse range of intelligent tasks. However, the hardware performance of GCNs is often bounded by random and non-continuous memory accesses due to the sparse graph data, which incur high latency and high power consumption. The emerging computing-in-memory (CIM) architecture significantly reduces the overhead of data movements, which is suitable for memory-intensive GCN acceleration. Existing analog-based CIM solutions require a large amount of analog-to-digital (AD) and digital-to-analog (DA) conversions, which dominate the overall area and power consumption. Furthermore, the analog non-ideality can degrade accuracy and reliability of CIM. To address these challenges, this work proposes a digital CIM accelerator based on SRAM, called DCIM-GCN, to accelerate GCN algorithm. DCIM-GCN introduces innovations on three levels: circuit, architecture, and algorithm. At the circuit level, digital CIM is proposed with SRAM sub-arrays to eliminate the power and area expensive AD/DA converters. Furthermore, we have incorporated the multi-address feature into the digital CIM, thereby leveraging its ability to efficiently process sparse matrix multiplication. At the architecture level, the sparsity-aware computation engine takes advantage of sparsity in GCNs and leverages CIM to minimize memory accesses and data movements. Finally, at the algorithm level, the balance mapping algorithm tackles workload imbalance issues, while the vertex reorder algorithm reduces idle states for aggregation engines, resulting in increased hardware utilization. Our DCIM-GCN achieves 1.89$\times$and 2.42$\times$speedup and 4.58$\times$and 9.46$\times$energy efficiency improvement on average over other CIM-based graph accelerators, e.g., PASGCN and PIM-GCN, respectively.
Yufei Ma 0002, Yikan Qiu, Guoxiang Li, Meng Wu 0005, Le Ye, Ru Huang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.4
2022 Reversible data hiding with adaptive difference recovery for encrypted images
Chunqiang Yu, Xianquan Zhang, Guoxiang Li, Shanhua Zhan, Zhenjun Tang
Inf. Sci.3
2022 Reversible Data Hiding With Hierarchical Embedding for Encrypted Images
abstract
Reversible data hiding in encrypted images (RDHEI) is an effective technique of data security. Most state-of-the-art RDHEI methods do not achieve desirable payload yet. To address this problem, we propose a new RDHEI method with hierarchical embedding. Our contributions are twofold. (1) A novel technique of hierarchical label map generation is proposed for the bit-planes of plaintext image. The hierarchical label map is calculated by using prediction technique, and it is compressed and embedded into the encrypted image. (2) Hierarchical embedding is designed to achieve a high embedding payload. This embedding technique hierarchically divides prediction errors into three kinds: small-magnitude, medium-magnitude, and large-magnitude, which are marked by different labels. Different from the conventional techniques, pixels with small-magnitude/large-magnitude prediction errors are both used to accommodate secret bits in the hierarchical embedding technique, and therefore contribute a high embedding payload. Experiments on two standard datasets are discussed to validate the proposed RDHEI method. The results demonstrate that the proposed RDHEI method outperforms some state-of-the-art RDHEI methods in payload. The average payloads of the proposed RDHEI method are 3.4568 bpp and 3.6823 bpp for BOWS-2 dataset and BOSSbase dataset, respectively.
Chunqiang Yu, Xianquan Zhang, Xinpeng Zhang 0001, Guoxiang Li, Zhenjun Tang
IEEE Trans. Circuits Syst. Video Technol.4
2021 Jura: Towards Automatic Compliance Assessment for Annual Reports of Listed Companies
abstract
The initial public offering (IPO) market in Hong Kong is consistently one of the largest in the world. As part of its regulatory responsibilities, Hong Kong Exchanges and Clearing Limited (HKEX) reviews annual reports published by listed companies (issuers). The number of issuers has grown at a fast pace, reaching 2,538 as the end of 2020. This poses a challenge for manually reviewing these annual reports against the many diverse regulatory obligations (listing rules). We propose a system named Jura to improve the efficiency of annual report reviewing with the help of machine learning methods. This system checks the compliance of an issuer's published information against listing rules in four steps: panoptic document recognition, relevant passage location, fine-grained information extraction, and compliance assessment. This paper introduces in detail the passage location step, how it is critical for speeding up compliance assessment, and the various challenges faced. We argue that although a passage is a relatively independent unit, it needs to be combined with document structure and contextual information to accurately locate the relevant passages. With the help of Jura, HKEX reports saving 80% of the time on reviewing issuers' annual reports.
Zhengqi Xu, Yixuan Cao 0001, Rongyu Cao, Guoxiang Li, Xuanqiang Liu, Yangbin Wang, Allie Cheung, Matthew Tam, Lukas Petrikas, Ping Luo 0001
CIKM4