Lingyi Kong

dblp:266/4563 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Reinforcement Learning for Enhanced Advanced QEC Architecture Decoding
Lingyi Kong, Yifeng Peng, Zhiding Liang
ASP-DAC2
2026 Group-wise attentive enhancements for unsupervised feature selection
Jie Wang 0164, Yongjin Yuan, Lingyi Kong, Zheng Wang 0037, Rong Wang 0001, Feiping Nie 0001
Knowl. Based Syst.4
2026 An LSGQ-FFS Framework for Adaptive Optimization of Hybrid INT-CIM Architecture
abstract
Hybrid computing-in-memory (CIM) has recently gained significant attention due to its ability to leverage the strengths of both digital CIM (DCIM) and analog CIM (ACIM). The multibit fusion (MF) scheme enhances energy efficiency by fusing low-bit results, which typically require multiple read-out cycles, into a single-cycle read out. However, the relationship between hybrid INT-CIM circuit design and network performance based on the MF scheme has not yet been systematically explored. In addition, we investigate how different MF configurations affect the performance of various neural networks. To address this gap, we first propose a less-significant group quantization (LSGQ) model, which defines and explores the design space of hybrid INT-CIM. Second, we develop a FastFuse-Search (FFS) algorithm, which optimizes configurations for different networks to strike a better balance between model accuracy and energy efficiency. Based on the experimental results, some key considerations on hybrid CIM design are derived. FFS yields a$1.72\times $energy-efficiency boost with negligible accuracy loss. Finally, we fabricate a 28-nm hybrid INT-CIM test chip, achieving 59.74 TOPS/W and 0.96 TOPS/mm2, with performance metrics of 23.21 perplexity for GPT-2, 68.69% accuracy for ResNet18, and 80.53% accuracy for ViT.
Shaochen Li, Xi Chen 0107, Yujia Xiong, Lingyi Kong, He Wang 0028, Tianhui Jiao, Yan Yan 0030, Xin Si
IEEE Trans. Very Large Scale Integr. Syst.5
2025 Modeling of Less-Significant Group Quantization for Hybrid CIM Architecture
abstract
Hybrid computing-in-memory (CIM) has gained growing interest in recent times due to its ability to combine the strengths of both digital CIM (DCIM) and analog CIM (ACIM). The Multi-Bit Fusion (MF) scheme enhances energy efficiency by fusing low-bit results, which would typically require multiple readout cycles, into a single-cycle readout. However, the relationship between hybrid INT-CIM circuit design and network performance based on the MF scheme has yet to be systematically explored. To fill this gap, a less-significant group quantization (LSGQ) model is proposed, defining and exploring the hybrid INT-CIM design space. Experimental results demonstrate up to 1.83x improvement in energy efficiency with minimal impact on performance. A 28nm hybrid INT-CIM test chip is fabricated, achieving 52.16 TOPS/W, 0.96 TOPS/mm2, with respective performance metrics of 48.58 perplexity for GPT-2, 75.51% accuracy for ResNet18, and 81.5% accuracy for ViT.
Shaochen Li, Xi Chen 0107, Lingyi Kong, He Wang 0028, Yi Yang 0001, Xin Si
ISCAS3
2025 Direct Spectral Clustering With New Graph Learning for Better Fitting
abstract
Traditional spectral clustering methods struggle with scalability and robustness in large datasets due to their reliance on similarity matrices and EigenValue Decomposition. We introduce two innovative models: Rcut-based Coordinate Descent Clustering (R-CDC) and Ncut-based Doubly Stochastic Clustering (N-DSC). These models integrate graph construction and segmentation into a unified process optimized through the coordinate descent method, significantly enhancing clustering efficacy. A novel graph structure enhances robustness against noise and outliers, simplifying the clustering process and improving outcomes across diverse datasets. Our extensive experiments show that these models surpass existing spectral clustering techniques in managing large-scale data and complex structures.
Lingyi Kong, Feiping Nie 0001, Xuelong Li 0001
IEEE Trans. Knowl. Data Eng.1
2024 Towards Fault-tolerant Design of Quaternary Quantum Arithmetic
abstract
Qudit has emerged as a promising system for next-generation quantum computers because of its significant advantages in multi-phase problems and quantum error correction, while multi-qudit operations are yet to be implemented. As one of the basic operations, quantum addition has been widely applied in number factorization and discrete logarithms. Meanwhile, the physical constraints of noisy intermediate-scale quantum (NISQ) devices necessitate the development of fault-tolerant quantum adders and the exploration of quaternary operations on binary devices is still in its infancy. In this work, we propose the first approach for implementing quaternary quantum addition algorithms by employing primitive quantum gates. A library of quaternary quantum gates and quaternary quantum full adders (Q2FA) able to produce carry-first results, along with lower depth and fewer T-gates optimizations are proposed and evaluated, where all circuits are implemented on IBM Qiskit SDK. Extensive experiments show that our proposed Q2FA design, together with the optimization techniques, reduces T-depth by up to 1.4× and T-count by 1.7× compared with baseline quantum circuits without depth and T-gate optimizations. Meanwhile, the scalability of the proposed Q2FA is demonstrated by constructing quantum carry-ripple adders. Under noisy conditions, our proposed design can achieve an overall fidelity increase by 1.4×.
Yunchen Zhu, Ruixuan Yang, Yuhang Gu, Fangtian Gu, Lingyi Kong, He Li 0008
ITC-Asia5
2024 Security and Reliability Tradeoff of UAV Relays Assisted Cognitive Transmissions With Hardware Impairments
abstract
In this article, we analyze the security and reliability tradeoff (SRT) performance of a unmanned aerial vehicle (UAV) relays assisted cognitive network consisting of a cognitive source (S), multiple cognitive UAV relays (Rs), a cognitive destination (D) in the face of an eavesdropper (E). We consider a practical scenario where mutual interference exists between the primary users and cognitive users. In order to protect the wireless transmission against the eavesdropping attack, we propose two UAV relay selection schemes, namely, noninterference aware relay selection (NIARS) scheme and interference aware relay selection (IARS) scheme. Specifically, in the NIARS scheme, a relay having the best instantaneous (R-D) link (spanning from Rs to D) is selected to forward message. By contrast, a relay maximizing the main channel capacity is selected, which relies not only on the channel state informations (CSIs) of main links but also on the CSIs of the interference links. For the purpose of comparison, we present the round robin (RR) scheme as a baseline. To evaluate the SRT performance, we derive the closed-form intercept probability (IP) and outage probability (OP) expressions for RR, NIARS, and IARS schemes under the impact of hardware impairments (HIs) over Nakagami-$m$fading channels. Numerical results show that the proposed IARS and NIARS schemes perform better than RR scheme in terms of SRT performance. Additionally, the HI at D has a negative impact on SRT performance, while HI at E has a positive impact on SRT performance.
Lingyi Kong, YuLong Zou, Bin Li 0022
IEEE Internet Things J.1
2021 A progressive CNN in-loop filtering approach for inter frame coding
Dandan Ding, Lingyi Kong, Fengqing Zhu 0001
Signal Process. Image Commun.2
2020 Guided CNN Restoration with Explicitly Signaled Linear Combination
abstract
State-of-the-art Convolutional Neural Network (CNN) based loop restoration generally involves a CNN structure with a large number of parameters and applies the CNN model to those degraded frames uniformly to generate their restored version, even though the contents within these frames are different. By contrast, in this paper, we propose a Guided CNN Restoration (GNR) scheme, where a CNN is used in conjunction with explicitly signaled guide parameters, with an aim to adapt the CNN model to different input contents. Specifically, the CNN architecture is designed such that the final restoration is constrained within the subspace generated by various output channels of the CNN, and meanwhile the weighting parameters for a linear combination of the output channels to obtain the final restoration are explicitly signaled by the encoders. The proposed GNR is incorporated into an AV1 encoder to replace the anchor in-loop filters and the weighting parameters are written into the encoded bitstream. Experimental results show that given a small CNN with 3,312 parameters, the proposed approach achieves a BD-rate reduction of 3.06% over the AV1 anchor, while the traditional CNN-based method only achieves 1.39%.
Lingyi Kong, Dandan Ding, Fuchang Liu, Debargha Mukherjee, Urvang Joshi, Yue Chen 0040
ICIP1
2020 A Switchable Deep Learning Approach for In-Loop Filtering in Video Coding
abstract
Deep learning provides a great potential for in-loop filtering to improve both coding efficiency and subjective quality in video coding. State-of-the-art work focuses on network structure design and employs a single powerful network to solve all problems. In contrast, this paper proposes a deep learning based systematic approach that includes an effective Convolutional Neural Network (CNN) structure, a hierarchical training strategy, and a video codec oriented switchable mechanism. First, we propose a novel CNN structure, i.e., Squeeze-and-Excitation Filtering CNN (SEFCNN), as an optional in-loop filter. To capture the non-linear interaction between channels, the SEFCNN is comprised of two subnets, i.e., Feature EXtracting (FEX) subnet and Feature ENhancing (FEN) subnet. Then, we develop a hierarchical model training strategy to adapt the two subnets to different coding scenarios. For high-rate videos with small artifacts, we train a single global model using the FEX for all types of frames, whereas for low-rate videos with large artifacts, different models are trained using both FEX and FEN for different types of frames. Finally, we propose an adaptive enhancing mechanism which is switchable between the CNN-based and the conventional methods. We selectively apply the CNN model to some frames or some regions in a frame. Experimental results show that the proposed scheme outperforms state-of-the-art work in coding efficiency, while the computational complexity is acceptable after GPU acceleration.
Dandan Ding, Lingyi Kong, Zoe Liu, Yong Fang 0001
IEEE Trans. Circuits Syst. Video Technol.2