EDBT 2026 Demo / reviewers in the wild / expert
Xiaoyang Lin
dblp:29/3714
· DBLP profile ↗
15ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Fully-Parallel Digital MRAM Computing-in-Memory Macro Featuring a High-Efficient Dynamic Adder Tree and Bit-Splitting MAC
Zhongzhen Tong, Jiye Yao, Shaohui Ma, Yulong Qiu, Zhaohao Wang, Amara Amara, Xiaoyang Lin |
ISCAS | 7 |
| 2026 | BaM-CIM: A High Throughput Booth Algorithm-Based In-MRAM Computing Macro Using Hybrid VGSOT-MTJ/GAA-CNTFETabstractAs artificial intelligence (AI) and computational models grow in scale, the demand for computational power and storage has significantly increased. The computing-in-memory (CIM) architecture addresses this challenge by performing computations directly within the memory array, reducing data transfer between the processor and memory. This paper introduces a Booth algorithm-based In-MRAM computing architecture (BaM-CIM) using a hybrid voltage-gated spin-orbit torque MTJ (VGSOT-MTJ) and gate-all-around carbon nanotube field-effect transistors (GAA-CNTFETs) for efficient multiply-and-accumulate (MAC) computing. The key contributions of BaM-CIM are as follows: 1) A Voltage divider reference (VDR) cell is proposed, which enables read operations using only a 2T1M cell structure. Compared to complementary read cells, the VDR reduces the area by half and achieves robust data sensing without requiring precharge/discharge operations. 2) The BaM-CIM circuit is proposed to complete 8b-W/8b-IN/21b-OUT computations in only two cycles (1.6 ns), reducing the number of cycles by 75% compared to single-bit input serial operations and by 50% compared to two-bit serial operations. 3) A three-input 8b Booth computing adder (BCA), along with Modified computing shift adder (MCSA) and Modified computing post adder (MCPA), which can achieve higher energy efficiency. BaM-CIM with 128 Kb is simulated, achieving throughput and energy efficiency of 0.93 TOPS and 258.4 TOPS/W, respectively, at a 0.6 V supply voltage and 1.28 TOPS and 169.5 TOPS/W, respectively, at a 0.8 V supply voltage with 8b-IN, 8b-W, and 21b-OUT. Chenghang Li, Zhongzhen Tong, Yulong Qiu, Jiye Yao, Chao Wang 0094, Zhaohao Wang, Xiaoyang Lin, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2025 | Spectral Subspace Clustering for Attributed GraphsabstractSubspace clustering seeks to identify subspaces that segment a set of n data points into k (k« n) groups, which has emerged as a powerful tool for analyzing data from various domains, especially images and videos. Recently, several studies have demonstrated the great potential of subspace clustering models for partitioning vertices in attributed graphs, referred to as SCAG. However, these works either demand significant computational overhead for constructing the nxn self-expressive matrix, or fail to incorporate graph topology and attribute data into the subspace clustering framework effectively, and thus, compromise result quality. Xiaoyang Lin, Renchi Yang, Xiangyu Ke |
KDD (1) | 1 |
| 2025 | Effective Clustering for Large Multi-Relational GraphsabstractMulti-relational graphs (MRGs) are an expressive data structure for modeling diverse interactions/relations among real objects (i.e., nodes), which pervade extensive applications and scenarios. Given an MRG G with N nodes, partitioning the node set therein into K disjoint clusters (referred to as MRGC) is a fundamental task in analyzing MRGs, which has garnered considerable attention. However, the majority of existing solutions towards MRGC either yield severely compromised result quality by ineffective fusion of heterogeneous graph structures and attributes, or struggle to cope with sizable MRGs with millions of nodes and billions of edges due to the adoption of sophisticated and costly deep learning models. In this paper, we present DEMM and DEMM+, two effective MRGC approaches to address the aforementioned limitations. Specifically, our algorithms are built on novel two-stage optimization objectives, where the former seeks to derive high-caliber node feature vectors by optimizing the multi-relational Dirichlet energy specialized for MRGs, while the latter minimizes the Dirichlet energy of clustering results over the node affinity graph. In particular, DEMM+ achieves significantly higher scalability and efficiency over our based method DEMM through a suite of well-thought-out optimizations. Key technical contributions include (i) a highly efficient approximation solver for constructing node feature vectors, and (ii) a judicious and theoretically-grounded problem transformation together with carefully-crafted techniques that enable the linear-time clustering without explicitly materializing the N x N dense affinity matrix. Further, we extend DEMM to handle attribute-less MRGs through non-trivial adaptations. Extensive experiments, comparing DEMM+ against 20 baselines over 11 real MRGs, exhibit that DEMM+ is consistently superior in terms of clustering quality measured against ground-truth labels, while often being remarkably faster. Xiaoyang Lin, Runhao Jiang, Renchi Yang |
Proc. ACM Manag. Data | 1 |
| 2025 | A Self-Decryption Pass Transistor Logic-Based In-MRAM Computing Macro Using Hybrid VGSOT-MTJ/GAA-CNTFETabstractSpintronic devices and gate-all-around carbon nanotube field-effect-transistors (GAA-CNTFETs)-based computing in-memory architecture are competitive candidates for applications in battery-powered tiny artificial intelligence (AI) edge devices. Meanwhile, data encryption and decryption are also necessary to protect AI model weights and the customized data used to guarantee neural network (NN) inference accuracy. In this study, we propose a self-decryption pass transistor logic (PTL)-based in-MRAM computing macro (SP-CIM) that utilizes hybrid voltage-gated spin-orbit torque magnetic tunnel junctions (VGSOT-MTJ)/GAA-CNTFET. The proposed SP-CIM macro enables simultaneous data access, decryption, and full-accuracy multiply-and-accumulate (MAC) operations using the newly introduced voltage-divider self-decryption cell, without the need for additional decryption logic. Compared to existing in-memory decryption strategies, this design reduces energy consumption by 45.7% and decreases decryption delay by 87.2%. To enhance area efficiency and reduce computing latency, we propose a PTL-based multiplication cell that achieves full-accuracy local 2b-IN TEXPRESERVE0 2b-W operations with only 20 transistors (20T). Additionally, novel PTL-based full-swing output half adders (10T-HA) and full adders (14T-FA) are proposed to construct the local adder tree, achieving reductions of 31.8%, 76.4%, and 41.4% in energy, delay, and area, respectively, compared to conventional adder trees in CIM macros. Simulations of the 288 kb SP-CIM macro demonstrated throughput and energy efficiency of 2.25 TOPS and 226.6 TOPS/W, respectively, at a 0.6 V supply voltage, and 2.97 TOPS and 154.1 TOPS/W, respectively, at a 0.8 V supply voltage, with 8b-IN, 8b-W, and 24b-OUT. Zhongzhen Tong, Sifan Sun, Chenghang Li, Jiye Yao, Yulong Qiu, Chao Wang 0094, Zhaohao Wang, Amara Amara, Xiaoyang Lin, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 11 |
| 2024 | Efficient Topology-aware Data Augmentation for High-Degree Graph Neural NetworksabstractIn recent years, graph neural networks (GNNs) have emerged as a potent tool for learning on graph-structured data and won fruitful successes in varied fields. The majority of GNNs follow the message-passing paradigm, where representations of each node are learned by recursively aggregating features of its neighbors. However, this mechanism brings severe over-smoothing and efficiency issues over high-degree graphs (HDGs), wherein most nodes have dozens (or even hundreds) of neighbors, such as social networks, transaction graphs, power grids, etc. Additionally, such graphs usually encompass rich and complex structure semantics, which are hard to capture merely by feature aggregations in GNNs.Motivated by the above limitations, we propose TADA, an efficient and effective front-mounted data augmentation framework for GNNs on HDGs. Under the hood, TADA includes two key modules: (i) feature expansion with structure embeddings, and (ii) topology- and attribute-aware graph sparsification. The former obtains augmented node features and enhanced model capacity by encoding the graph structure into high-quality structure embeddings with our highly-efficient sketching method. Further, by exploiting task-relevant features extracted from graph structures and attributes, the second module enables the accurate identification and reduction of numerous redundant/noisy edges from the input graph, thereby alleviating over-smoothing and facilitating faster feature aggregations over HDGs. Empirically, \algo considerably improves the predictive performance of mainstream GNN models on 8 real homophilic/heterophilic HDGs in terms of node classification, while achieving efficient training and inference processes. Yurui Lai, Xiaoyang Lin, Renchi Yang |
KDD | 2 |
| 2024 | Effective Clustering on Large Attributed Bipartite GraphsabstractAttributed bipartite graphs (ABGs) are an expressive data model for describing the interactions between two sets of heterogeneous nodes that are associated with rich attributes, such as customer-product purchase networks and author-paper authorship graphs. Partitioning the target node set in such graphs into k disjoint clusters (referred to as k-ABGC) finds widespread use in various domains, including social network analysis, recommendation systems, information retrieval, and bioinformatics. However, the majority of existing solutions towards k-ABGC either overlook attribute information or fail to capture bipartite graph structures accurately, engendering severely compromised result quality. The severity of these issues is accentuated in real ABGs, which often encompass millions of nodes and a sheer volume of attribute data, rendering effective k-ABGC over such graphs highly challenging. Renchi Yang, Yidu Wu, Xiaoyang Lin, Qichen Wang 0001, Tsz Nam Chan, Jieming Shi 0001 |
KDD | 3 |
| 2024 | BSTCIM: A Balanced Symmetry Ternary Fully Digital In-MRAM Computing Macro for Energy Efficiency Neural NetworkabstractSilicon-based traditional binary computing in-memory (TBCIM) architectures are approaching their energy efficiency and throughput limits owing to challenges facing Moore’s Law. Thus, it is essential to explore architecture based on novel devices and computing paradigms to fulfill data-centric applications, such as artificial intelligence. In this paper, we propose a balanced symmetry ternary (BST) fully digital in-MRAM computing macro (BSTCIM) using hybrid voltage-gated spin-orbit torque magnetic tunnel junctions (VGSOT-MTJ) and gate-all-around carbon nanotube field-effect-transistors (GAA-CNTFET) technology. The overall computing is based on the highest efficiency multi-bit ternary system. BSTCIM includes a ternary dot product (TDP) unit with 4 GAA-CNTFETs and 2 VGSOT-MTJs achieving TDP operation without complex logic circuits. The multi-bit ternary multiply-and-accumulate (MAC) operation is realized through the proposed ternary adder tree and ternary post adder which accumulate TDP results within the digital domain enabling high accuracy neural network inference. Furthermore, due to the advantages of BST, ternary signed MAC is more easily performed compared to TBCIM macros that adapt 2’s complement or separate signed bit calculations. BSTCIM with 288 kb is simulated, achieving throughput and energy efficiency of 0.72 TOPS and 54.5 TOPS/W, respectively, at a 0.6 V supply voltage and 1.15 TOPS and 33.7 TOPS/W, respectively at a 0.8 V supply voltage with 8b-IN, 8b-W, and 20b-OUT. Moreover, the figure-of-merit for BSTCIM is 1.13–33.6 times higher than that of existing CIM macros. Zhongzhen Tong, Chenghang Li, Chao Wang 0094, Suteng Zhao, Qianyong Peng, Daming Zhou, Zhaohao Wang, Xiaoyang Lin, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 10 |
| 2024 | A High Throughput In-MRAM-Computing Scheme Using Hybrid p-SOT-MTJ/GAA-CNTFETabstractSilicon-based semiconductor transistors are approaching their physical limits due to shrinking feature sizes. Simultaneously, traditional silicon-based von Neumann architectures exhibit significant latency and power consumption issues in data-centric applications, such as the Internet of Things and artificial intelligence. To tackle these challenges, this study introduces a novel approach: Magnetoresistance Random Access Memory (MRAM) computing in-memory (CIM) using gate-all-around carbon nanotube field-effect transistors (GAA-CNTFET). The proposed MRAM array comprised three transistors and one perpendicular magnetic anisotropy spin-orbit torque magnetic tunnel junction (p-SOT-MTJ) (3T1M) cell and achieves full-array Boolean logic operations and half/full-adder operations. The calculated results can be stored in-situ during the computing phase without requiring additional peripheral circuits. A 16 Kb MRAM was simulated in both GAA-CNTFET/p-SOT-MTJ and 14-nm FinFET/p-SOT-MTJ technologies to examine the effectiveness of the proposed design. Compared to its 14-nm FinFET/p-SOT-MTJ counterparts, the write and computing latencies of the GAA-CNTFET/p-SOT-MTJ CIM macro were reduced by approximately 21% and 20.6%, respectively, while the read and computing energy consumption by approximately 45.3% and 24.7%, respectively. Moreover, the proposed in-memory Boolean logic throughput was 8192 GOPS, which was approximately 160–250 times higher than that of existing CIM solutions, in which only two rows of word lines can be activated. Zhongzhen Tong, Yunlong Liu 0006, Xinrui Duan, Suteng Zhao, Chenghang Li, Zhi-Ting Lin, Xiulong Wu, Zhaohao Wang, Xiaoyang Lin |
IEEE Trans. Circuits Syst. I Regul. Pap. | 11 |
| 2023 | Neuromorphic terahertz imaging based on carbon nanotube circuits
Zhizhong Si, Xiaoyang Lin |
Sci. China Inf. Sci. | 4 |
| 2023 | In-Memory Transposable Multibit Multiplication Based on Diagonal Symmetry Weight BlockabstractA possible approach to overcome the von Neumann bottleneck and meet the increasing demand for better computing performance is to computing in-memory (CIM). The results of the in-memory calculations are primarily reflected in the vertical bitline (BL) analog voltage. However, the nonlinearity of the BL discharge deteriorates with the increase in discharge voltage. In this study, we propose a diagonal symmetry weight block (DSWB) based on an eight-transistor (8T) static random access memory (SRAM) that can achieve multibit transposable operations. In addition, to guarantee linearity and complete multibit multiplication operations, we propose a cascode current mirror (CCM)-based multiplier. To achieve low-overhead and more efficient quantification, our proposed CIM macro uses a counter-type quantization circuit to read out the analog calculation results. We simulated the performance of the proposed 8T SRAM in a 28-nm complementary metal–oxide–semiconductor process. The integral nonlinearity (INL) of the proposed CCM-based CIM decreased by approximately 54.4% compared with the traditional CIM. Furthermore, the proposed in-memory multibit multiplication throughput density was 6.74 GOPS/kb; this throughput density improvement is approximately 3.3–10.5 times higher than the existing CIM works. Zhongzhen Tong, Yue Zhao 0029, Jin Zhang 0036, Zhi-Ting Lin, Xiaoyang Lin, Xiulong Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2015 | Modeling the Task of Google MapReduce WorkloadabstractIn order to better understand and describe tasks and improve the ability of Cloud, the analyzing of tasks inessential. A coarse-grained analysis, cluster analysis, anointer-cluster analysis are used to model tasks for the analysis of a one-month trace of a Google MapReduce cluster across about 12,000 machines. In this paper, we consider the k value which is central to the performance of k-means algorithm can effect on modelling. Besides, we also take the selection of attributes into account which are used as the dimension when tasks are classified. Experiment results by using different type of task attributes and k value show the well performance odour approach. Xiaoyang Lin, Piyuan Lin, Peijie Huang, Linxiao Chen, Ziwei Fan 0001, Peisen Huang |
CCGRID | 1 |
| 2006 | A control-theoretic approach to improving fairness in DCF based WLANsabstractAchieving fair bandwidth distribution among uplink and downlink flows in the infrastructure based wireless local area networks (WLAN) which the distributed coordination function (DCF) mode is difficult. In this paper we present a new control theoretic approach to achieve a fair bandwidth distribution among the flows regardless of the transport protocol used by the flows. In addition, we explore methods to improve the channel bandwidth utilization and reduce the delay while improving the fair distribution. Our approach combines an AQM scheme for IFQ queue and the MAC layer design to achieve this goal. The effectiveness of our approach is demonstrated through extensive simulations over a wide range of network scenarios. Xiaolin Chang, Xiaoyang Lin, Jogesh K. Muppala |
IPCCC | 2 |
| 2005 | VQ-RED: An Efficient Virtual Queue Management Approach to Improve Fairness in Infrastructure WLANabstractIn this paper, we consider two fairness problems (downlink/uplink fairness and fairness among flows in the same direction) that arise in the infrastructure WLAN. We propose a virtual queue management approach, named VQ-RED to address the fairness problems. We demonstrate the effectiveness of our approach by conducting a series of simulations. The results show that compared with standard DCF, VQRED not only greatly improves the fairness, but also reduces packet delays. Xiaoyang Lin, Xiaolin Chang, Jogesh K. Muppala |
LCN | 1 |
| 2005 | An adaptive queue management mechanism for improving TCP fairness in the infrastructure WLANabstractResearch studies have uncovered the unfair WLAN bandwidth distribution between uplink and downlink TCP flows in the infrastructure WLAN using the distributed coordination function mode. Different mechanisms at or above the MAC layer have been proposed for handling this problem. This paper presents a simple but effective approach in order to achieve the fair distribution by using an adaptive queue management mechanism to manage the queue between the logic link and MAC layers at the access point. Extensive simulation results show that this approach outperforms other mechanisms discussed in the literature over a wide range of network scenarios in terms of fair distribution, small delay, and high WLAN bandwidth utilization Xiaolin Chang, Xiaoyang Lin, Jogesh K. Muppala |
PIMRC | 2 |