Xiaolin He

dblp:255/8320 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 MIREDO: MIP-Driven Resource-Efficient Dataflow Optimization for Computing-in-Memory Accelerator
abstract
Computing-in-Memory (CIM) architectures have emerged as a promising solution for accelerating Deep Neural Networks (DNNs) by mitigating data movement bottlenecks. However, realizing the potential of CIM requires specialized dataflow optimizations, which are challenged by an expansive design space and strict architectural constraints. Existing optimization approaches often fail to fully exploit CIM accelerators, leading to noticeable gaps between theoretical and actual system-level efficiency. To address these limitations, we propose the MIREDO framework, which formulates dataflow optimization as a MixedInteger Programming (MIP) problem. MIREDO introduces a hierarchical hardware abstraction coupled with an analytical latency model designed to accurately reflect the complex data transfer behaviors within CIM systems. By jointly modeling workload characteristics, dataflow strategies, and CIM-specific constraints, MIREDO systematically navigates the vast design space to determine the optimal dataflow configurations. Evaluation results demonstrate that MIREDO significantly enhances performance, achieving up to $3.2 \times$ improvement across various DNN models and hardware setups.
Xiaolin He, Cenlin Duan, Yingjie Qi, Xiao May, Jianlei Yang 0001
ASP-DAC1
2026 CIMinus: Empowering Sparse DNN Workloads Modeling and Exploration on SRAM-Based CIM Architectures
abstract
Compute-in-memory (CIM) has emerged as a pivotal direction for accelerating workloads in the field of machine learning, such as Deep Neural Networks (DNNs). However, the effectively exploitation of sparsity in CIM systems presents numerous challenges, due to the inherent limitations in their rigid array structures. Designing sparse DNN dataflows and developing efficient mapping strategies also become more complex when accounting for diverse sparsity patterns and the flexibility of a multi-macro CIM structure. Despite these complexities, there is still an absence of a unified systematic view and modeling approach for diverse sparse DNN workloads in CIM systems. In this paper, we propose CIMinus, a framework dedicated to cost modeling for sparse DNN workloads on CIM architectures. It provides an in-depth energy consumption analysis at the level of individual components and an assessment of the overall workload latency. We validate CIMinus against contemporary CIM architectures and demonstrate its applicability in two use-cases. These cases provide valuable insights into both the impact of sparsity patterns and the effectiveness of mapping strategies, bridging the gap between theoretical design and practical implementation.
Yingjie Qi, Jianlei Yang 0001, Rubing Yang, Cenlin Duan, Xiaolin He, Ziyan He, Weitao Pan, Weisheng Zhao 0001
IEEE Trans. Computers5
2026 Efficient SRAM-PIM Co-Design by Joint Exploration of Value-Level and Bit-Level Sparsity
abstract
Processing-in-memory (PIM) architectures mitigate the Von Neumann bottleneck by integrating computation units into memory arrays. Among PIM architectures, digital SRAMPIM has become a prominent approach, directly integrating digital logic within the SRAM array. However, the rigid crossbar architecture and full array activation pose challenges in efficiently utilizing value-level sparsity. Moreover, neural network models exhibit a high proportion of zero bits within non-zero values, which remain underutilized due to architectural constraints. To overcome these limitations, we present Dyadic Block PIM (DB-PIM), a groundbreaking algorithm-architecture co-design framework to harness both value-level and bit-level sparsity. At the algorithm level, our hybrid-grained pruning technique, combined with a novel sparsity pattern, enables effective sparsity management. Architecturally, DB-PIM incorporates a sparse network and customized digital SRAM-PIM macros, including input pre-processing unit (IPU), dyadic block multiply units (DBMUs), and Canonical Signed Digit (CSD)-based adder trees. It circumvents structured zero values in weights and bypasses unstructured zero bits within non-zero weights and block-wise all-zero bit columns in input features. As a result, the DBPIM framework skips a majority of unnecessary computations, thereby driving significant gains in computational efficiency. Experimental results demonstrate that our DB-PIM framework achieves up to 8.01× speedup and 85.28% energy savings, significantly boosting computational efficiency in digital SRAMPIM systems.
Cenlin Duan, Jianlei Yang 0001, Yiou Wang, Yingjie Qi, Xiaolin He, Bonan Yan, Xiaotao Jia, Weisheng Zhao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2026 ISDNet: High-Fidelity Single-View Reconstruction of Indoor Scenes via Instance Separation and Deformation
abstract
In this work, we aim to reconstruct the 3D shape of an indoor scene from a single view, which includes multiple objects and the background. This task is challenging for existing methods since those instances of indoor scenes regularly occlude each other and contain diverse topologies. To address this, we propose a novel framework, ISDNet, to adaptively separate mixed instances and perform topology-aware reconstruction. Specifically, Specifically, ISDNet consists of two cascaded subnetworks: an instance separation module (ISM) and an instance deformation module (IDM). The ISM learns to separate occluded objects through stepwise sampling, inferring clean features for each instance. On the basis of these features, IDM generates an instance-topology-aware template and deforms it with learned offsets to reconstruct detailed geometry. Quantitative and qualitative experiments on the SUNRGB-D and 3D-FRONT datasets demonstrate that ISDNet outperforms the state-of-theart methods in terms of local details and overall shapes.
Xiaolin He, Ying Liu 0027, Yiming Han, Junxian Chen, Ruihui Li
IEEE Trans. Multim.1
2025 CIMFlow: An Integrated Framework for Systematic Design and Evaluation of Digital CIM Architectures
abstract
Digital Compute-in-Memory (CIM) architectures have shown great promise in Deep Neural Network (DNN) acceleration by effectively addressing the “memory wall” bottleneck. However, the development and optimization of digital CIM accelerators are hindered by the lack of comprehensive tools that encompass both software and hardware design spaces. Moreover, existing design and evaluation frameworks often lack support for the capacity constraints inherent in digital CIM architectures. In this paper, we present CIMFlow, an integrated framework that provides an out-of-the-box workflow for implementing and evaluating DNN workloads on digital CIM architectures. CIMFlow bridges the compilation and simulation infrastructures with a flexible instruction set architecture (ISA) design, and addresses the constraints of digital CIM through advanced partitioning and parallelism strategies in the compilation flow. Our evaluation demonstrates that CIMFlow enables systematic prototyping and optimization of digital CIM architectures across diverse configurations, providing researchers and designers with an accessible platform for extensive design space exploration.
Yingjie Qi, Jianlei Yang 0001, Yiou Wang, Dayu Wang, Cenlin Duan, Xiaolin He, Weisheng Zhao 0001
DAC8
2025 High-Fidelity Single-View Reconstruction of Indoor Scenes using 3D Shape Prior Template and Pixel-Aligned Deformation
abstract
This paper presents a novel pipeline for estimating room layouts and reconstructing the 3D shapes of indoor objects. This task remains challenging due to occlusions of indoor scenes, which lead to incomplete shape and poor geometric quality manifested as non-smooth meshes. Our key insight is that occlusions of indoor objects inherently lead to insufficient information in images. Pixel-level features alone are inadequate to recover the complete structure of objects; Therefore, additional prior information is required to supplement the missing occluded details. To address this, we propose a two-stage training strategy. First, a VQ-VAE encodes 3D shapes into a latent space, with a decoder leveraging priors to predict occluded regions. Second, pixel-aligned features are used for deformation, ensuring consistency between the reconstructed shape and the image. Quantitative and qualitative evaluations on the 3D-FRONT and SUNRGB-D datasets demonstrate that our approach surpasses state-of-the-art methods in reconstructing more complete geometric topologies and smoother meshes.
Xiaohao Zhang, Xiaolin He, Zhuo Tang, Ruihui Li
ICASSP2
2024 Towards Efficient SRAM-PIM Architecture Design by Exploiting Unstructured Bit-Level Sparsity
abstract
Bit-level sparsity in neural network models harbors immense untapped potential. Eliminating redundant calculations of randomly distributed zero-bits significantly boosts computational efficiency. Yet, traditional digital SRAM-PIM architecture, limited by rigid crossbar architecture, struggles to effectively exploit this unstructured sparsity. To address this challenge, we propose Dyadic Block PIM (DB-PIM), a groundbreaking algorithm-architecture co-design framework. First, we propose an algorithm coupled with a distinctive sparsity pattern, termed a dyadic block (DB), that preserves the random distribution of non-zero bits to maintain accuracy while restricting the number of these bits in each weight to improve regularity. Architecturally, we develop a custom PIM macro that includes dyadic block multiplication units (DBMUs) and Canonical Signed Digit (CSD)-based adder trees, specifically tailored for Multiply-Accumulate (MAC) operations. An input pre-processing unit (IPU) further refines performance and efficiency by capitalizing on block-wise input sparsity. Results show that our proposed co-design framework achieves a remarkable speedup of up to 7.69× and energy savings of 83.43%.
Cenlin Duan, Jianlei Yang 0001, Yiou Wang, Yingjie Qi, Xiaolin He, Bonan Yan, Xiaotao Jia, Weisheng Zhao 0001
DAC6
2024 DDC-PIM: Efficient Algorithm/Architecture Co-Design for Doubling Data Capacity of SRAM-Based Processing-in-Memory
abstract
Processing-in-memory (PIM), as a novel computing paradigm, provides significant performance benefits from the aspect of effective data movement reduction. SRAM-based PIM has been demonstrated as one of the most promising candidates due to its endurance and compatibility. However, the integration density of SRAM-based PIM is much lower than other nonvolatile memory-based ones, due to its inherent 6T structure for storing a single bit. Within comparable area constraints, SRAM-based PIM exhibits notably lower capacity. Thus, aiming to unleash its capacity potential, we propose DDC-PIM, an efficient algorithm/architecture co-design methodology that effectively doubles the equivalent data capacity. At the algorithmic level, we propose a filter-wise complementary correlation (FCC) algorithm to obtain a bitwise complementary pair. At the architecture level, we exploit the intrinsic cross-coupled structure of 6T SRAM to store the bitwise complementary pair in their complementary states$(Q/\overline {Q})$, thereby maximizing the data capacity of each SRAM cell. The dual-broadcast input structure and reconfigurable unit support both depthwise and pointwise convolution, adhering to the requirements of various neural networks. Evaluation results show that DDC-PIM yields about$2.84\times $speedup on MobileNetV2 and$2.69\times $on EfficientNet-B0 with negligible accuracy loss compared with PIM baseline implementation. Compared with state-of-the-art SRAM-based PIM macros, DDC-PIM achieves up to$8.41\times $and$2.75\times $improvement in weight density and area efficiency, respectively.
Cenlin Duan, Jianlei Yang 0001, Xiaolin He, Yingjie Qi, Yiou Wang, Ziyan He, Bonan Yan, Xiaotao Jia, Weitao Pan, Weisheng Zhao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 On the Pilot-Aided Channel Estimation for Windowed OTFS With Data Interference in Rapidly Time-Varying Channels
abstract
The orthogonal time-frequency space (OTFS) modulation is an effective technique to deal with the high-mobility challenge in vehicular wireless communications, whose data detection depends heavily on accurate channel estimation (CE). In pilot-aided CE, reducing guard symbols can achieve higher spectral efficiency. However, data interference is inevitable, especially in fractional Doppler channels. Therefore, this paper investigates the impact of data interference on CE, and aims to mitigate such data interference by adding a non-rectangular window in the time-frequency (TF) domain. To fulfill this goal, a Cramer-Rao lower bound (CRLB) is derived by considering the presence of data interference and a non-rectangular window for CE in OTFS. Different from the existing CRLB analysis for embedded-pilot OTFS, data interference is considered and treated as noise interference, resulting in a tighter derived CRLB that effectively reflects the impact of data interference on CE. By minimizing the derived CRLB with the rectangular window, a new pilot sequence is obtained, which can achieve lower CRLB and better BER performance compared to the Zadoff-Chu (ZC) sequence. On the other hand, windowing can loosen the CRLB due to its ability to suppress data interference. Under the same main lobe width, it is found that the Kaiser window and Slepian window are more effective in reducing data interference than Dolph-Chebyshev (DC) window and rectangular window. Our simulation results indicate that the interference from data to the pilot can be significantly reduced by employing appropriate windowing techniques, and the windowing functions employed play an important role in pilot-aided CE.
Xiaolin He, Weijie Yuan 0001, Pingzhi Fan
IEEE Trans. Wirel. Commun.1
2024 Two-Dimensional Delay-Doppler Pilots and Channel Estimation for Multi-Antenna OTFS in Doubly Dispersive Channels
abstract
Orthogonal time frequency space (OTFS) is a novel modulation scheme to handle the high Doppler effect under time-varying channels. In this paper, in order to improve channel estimation accuracy and to reduce pilot overhead, two types of two-dimensional (2D) pilots and the corresponding matched filters are designed for multi-antenna OTFS systems. Our 2D pilots are formed using perfect arrays (such as Frank array, Chu array, etc) or perfect-sequence based Kronecker array (PKA). Different from the previous multi-antenna OTFS pilots, these 2D pilots are placed on the same area in the delay-Doppler domain, using code division multiplexing to deal with the interference between pilots of different antennas, known as pilot pollution. To improve the channel estimation performance, matched filters are also designed to better compensate the phase shift of the 2D pilot response in the delay-Doppler domain. Compared with the conventional multi-antenna OTFS schemes, the proposed scheme achieves significantly better NMSE performance, while having lower pilot overhead under multi-antenna scenarios.
Pingzhi Fan, Xiaolin He
IEEE Trans. Wirel. Commun.4
2023 SD-Net: Spatially-Disentangled Point Cloud Completion Network
abstract
Point clouds obtained from 3D scanning are typically incomplete, noisy, and sparse. Previous completion methods aim to generate complete point clouds, while taking into account the densification of point clouds, filling small holes, and proximity-to-surface, all through a single network. After revisiting the task, we propose SDNet, which disentangles the task based on the spatial characteristics of point clouds and formulates two sub-networks, a Dense Refiner and a Missing Generator. Given a partial input, the Dense Refiner produces a dense and clean point cloud, as a more reliable partial surface, which assists the Missing Generator to better infer the remaining point cloud structure. To promote the alignment and interaction across these two modules, we propose a Cross Fusion Unit with designed Non-Symmetrical Cross Transformers to capture geometric relationships between partial and missing regions, contributing to a complete, dense and well-aligned output. Extensive quantitative and qualitative results demonstrate that our method outperforms the state-of-the-art methods.
Junxian Chen, Ying Liu 0027, Yiqi Liang, Dandan Long, Xiaolin He, Ruihui Li
ACM Multimedia5