Qimin Zhou

dblp:220/1007 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Riemannian Graph Tokenizer for Structural Knowledge Transfer
Qimin Zhou, Li Sun 0008, Chuan Shi 0001
WWW1
2025 Harnessing Language Model for Cross-Heterogeneity Graph Knowledge Transfer
abstract
Heterogeneous graphs (HGs) that contain various node and edge types are ubiquitous in real-world scenarios. Considering the common label sparsity problem in HGs, some researchers propose to pretrain on source HGs to extract general knowledge and then fine-tune on a target HG for knowledge transfer. However, existing methods often assume that source and target HGs share a single heterogeneity, meaning that they have the same types of nodes and edges, which contradicts the real-world scenarios requiring cross-heterogeneity transfer. Although a recent study has made some preliminary attempts in cross-heterogeneity learning, its definition of general knowledge heavily rely on human knowledge, which lacks flexibility and further leads to a suboptimal transfer. To address the problem, we propose a novel Language Model-enhanced Cross-Heterogeneity learning model, namely LMCH. Specifically, we first design a metapath-based corpus construction method to unify HG representations as languages. The corpora of source HGs are then used to fine-tune a pretrained Language Model (LM), enabling the LM to autonomously extract general knowledge across different HGs. Furthermore, to fully utilize the extensive unlabeled nodes in a few-labeled target HG, we propose an iterative training pipeline with the help of an extra Graph Neural Network (GNN) predictor, enhanced by LM-GNN contrastive alignment at the end of each iteration. Extensive experiments on four real-world datasets have demonstrated the superior performance of LMCH over state-of-the-art methods.
Cheng Yang 0002, Qimin Zhou, Yang Juan, Chuan Shi 0001
AAAI5
2025 A Bit Level Weight Reordering Strategy Based on Column Similarity to Explore Weight Sparsity in RRAM-Based NN Accelerator
abstract
Compute-in-Memory (CIM) and weight sparsity are two effective techniques to reduce data movement during Neural Network (NN) inference. However, they can hardly be employed in the same accelerator simultaneously because CIM requires structural compute patterns which are disrupted in sparse NNs. In this paper, we partially solve this issue by proposing a bit level weight reordering strategy which can realize compact mapping of sparse NN weight matrices onto Resistive Random Access Memory (RRAM) based NN Accelerators (RRAM-Acc). In specific, when weights are mapped to RRAM crossbars in a binary complement manner, we can observe that, which can also be mathematically proven, bit-level sparsity and similarity commonly exist in the crossbars. The bit reordering method treats bit sparsity as a special case of bit similarity, reserve only one column in a pair of columns that have identical bit values, and then map the compressed weight matrices into Operation Units (OU). The performance of our design is evaluated with typical NNs. Simulation results show a 61.24 % average performance improvement and$1.51 \times-2.52 \times$energy savings under different sparsity ratios, with only slight overhead compared to the state-of-the-art design.
Weiping Yang, Shilin Zhou 0001, Yujiao Nie, Qimin Zhou, Changlin Chen
ICPADS4
2025 Pearson correlation coefficient-guided large-scale fuzzy cognitive maps learning algorithm
Qimin Zhou, Yingcang Ma, Zhiwei Xing, Xiaofei Yang 0004
Fuzzy Sets Syst.1
2025 OPASCA: Outer Product-Based Accelerator With Unified Architecture for Sparse Convolution and Attention
abstract
Vision transformer (ViT)-based models have achieved state-of-the-art accuracy in many computer vision tasks, but their attention mechanism is more computation and communication intensive than convolutional neural networks (CNNs). To adapt ViT-based models for resource-constrained edge computing platforms, techniques, such as network sparsity and convolution-attention combination, have been proposed to reduce processing costs without compromising accuracy. Therefore, specific hardware designs able to handle sparsity in both convolution and attention operations are required to accelerate the processing speed. In view of this, this work proposes OPASCA, an accelerator featuring a unified hardware architecture that supports irregular activation and weight sparsity in both operations. Specifically, this target is achieved with the following contributions: 1) we employ outer product dataflow to efficiently handle sparse weights and input neurons, supporting both convolution and attention computing with minimal hardware overhead; 2) we design a hierarchical butterfly network to route the output neurons to the accumulation buffers, minimizing the conflicts among outer product results and reducing the hardware overhead of accumulation banks; and 3) we propose a novel encoding scheme that achieves more compact sparse inputs and enhances multiplier utilization. Evaluations on VGG-16, ViT, BoTNet, and Conformer models show that OPASCA outperforms state-of-the-art accelerators by$1.60 \times -2.08 \times $on sparse convolution tasks and by$1.18 \times -1.70 \times $on sparse attention tasks in term of performance. It also reduces DRAM access to 19%–86% and energy consumption to 36%–92% of those of the counterpart designs.
Qimin Zhou, Tuo Ma, Changlin Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2024 An Integration and Time-Sampling based Readout Circuit with Current Compensation for Parallel MAC operations in RRAM Arrays
abstract
In Resistive Random Access Memory (RRAM) based compute-in-memory designs, the column current readout circuits still consume too much area and power overhead, even if plenty of methods have been proposed to optimize the circuits. To alleviate this problem, this paper presents a novel current readout circuit to sense multiply-and-accumulate (MAC) result of RRAM array. Specifically, the circuit first integrate the stabilized and proportionally mirrored MAC current on a small capacitor until it fires, then sample the integration time with a set of reference signals with different carefully designed delays, and finally code the sampled result into a digital value. The proposed readout circuit has fine stability due to simple and determined relationship among the inputs, the RRAM cells’ states, and the MAC current. Meanwhile, the proposed design can achieve accurate MAC result readout at low resistance switching ratios, for it employs a current compensation circuit to remove background current caused by high resistance state RRAM cells. Our design is implemented using 28nm CMOS technology with a read latency of 2.8ns and an area occupation of 267μm2/channel, which is 12.5% and 89% less than state of the art design. Its power consumption, 0.092mW/channel, is also less than most counterpart designs.
Weiping Yang, Shilin Zhou 0001, Qimin Zhou, Qingjiang Li, Changlin Chen
ISCAS4
2024 Sparse and regression learning of large-scale fuzzy cognitive maps based on adaptive loss function
Qimin Zhou, Yingcang Ma, Zhiwei Xing, Xiaofei Yang 0004
Appl. Intell.1
2022 Cross-Grade Curriculum Group Based Teaching Experiment System for Innovative Design of IoT Intelligent Dynamic Measurement and Control
abstract
One of the key hot spots and difficulties in the construction of “Engineering Education Certification” and “New Engineering” is how to strengthen the cultivation of the college students’ ability to solve complex engineering problems. In this paper, a software and hardware experimental platform has been designed, which organically integrates automatic judgement and sensing communications with electromechanical integrated IoT measurement and control. And a teaching system of cross-grade supporting experimental course group based on this platform has been proposed, which is driven by the application function topics from the shallower to the deeper. It is an engineering ability training mechanism that can continuously carry out experimental teaching such as design, installation and adjustment from the lower grade to the higher grade. More than five years of teaching practice shows that this system has significantly improved the students’ professional research interest and ability to solve complex engineering problems.
Hongqing Ma, Peihong Li, Xihua Li 0004, Xiangdong Jin, Qimin Zhou, Xiaoxing Shi, Huizhong Li
ISCAS6
2020 Effective metric learning with co-occurrence embedding for collaborative recommendations
Hao Wu 0010, Qimin Zhou, Rencan Nie, Jinde Cao
Neural Networks2
2019 Spatio-temporal context-aware collaborative QoS prediction
Qimin Zhou, Hao Wu 0010, Kun Yue, Ching-Hsien Hsu
Future Gener. Comput. Syst.1