EDBT 2026 Demo / reviewers in the wild / expert
Huan Cheng
dblp:225/8073
· DBLP profile ↗
10ranked-venue papers
0as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 6 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An IR drop-robust Mapping Method for Reliable Memristive AcceleratorsabstractMemristive accelerators (MAs) facilitate efficient matrix-vector multiplication (MVM) by performing in situ computation within memory crossbar arrays, thereby ensuring a fast and energy-efficient application acceleration. A significant challenge associated with the MA lies in the limited computing accuracy caused by the IR drop effect. However, existing IR drop mitigation works provide an approximate compensation, resulting in less accurate results. In this paper, we propose an IR drop-robust mapping method for reliable memristive accelerators. Firstly, the IR drop-robust mapping (IRM) method exploits the residuals between the equivalent matrix after the IR drop effect and the original matrix, and iteratively maps them to the crossbars for IR drop compensation. Based on the IRM method, a novel mechanism of the matrix-vector multiplication (MVM) operation is derived, ensuring that MVM is computed correctly. Secondly, the Calibrate-Shift-Reflect (CSR) strategy is developed to significantly reduce the number of arrays required by the IRM method to map the residuals. Thirdly, the hardware support for the IRM method is designed, and the overhead is reduced by sharing drivers/selectors between neighboring arrays. The experimental results indicate that the IRM-CSR method can effectively mitigate the IR drop effect, restoring inference accuracy by at most 80% (for neural network applications), and achieving a reduction in the relative root-mean-squared error by 103×~1010× (for scientific computing), compared with the state-of-the-art methods. Shiyi Song, Bing Wu 0001, Huan Cheng, Xueliang Wei, Wei Tong 0001, Dan Feng 0001 |
DATE | 4 |
| 2026 | pTree: Building Efficient B${}^{+}$+-Tree on Non-Volatile Memory With Processing-in-MemoryabstractB+-Trees are widely used in storage systems and diverse applications. However, their performance is hindered by frequent memory accesses required for key comparisons and structural maintenance. The limited parallelism of CPUs, along with their sensitivity to data order and volume, further restricts the efficiency of B+-Trees. Processing-in-memory (PIM) offers a promising alternative with its in-situ parallel computing capability. In this paper, we propose pTree, a novel set of PIM techniques and architecture specifically designed to accelerate B+-Tree operations. pTree introduces in-situ parallel comparison mechanisms that significantly reduce costly memory accesses and inefficient CPU-side comparisons during tree traversal. This parallelism eliminates the need to maintain intra-node key order and, together with our in-situ parallel node bisecting techniques, greatly minimizes tree structure maintenance overhead during insertions and deletions. Additionally, pTree decouples both comparison and modification efficiency from node size, enabling the use of larger nodes to reduce tree height and traversal complexity. Evaluation shows that pTree reduces the average latency to 35%, 26%, 52%, and 32% compared to state-of-theart B+-Trees forInsert, Search, Update,andDeleteoperations. Bing Wu 0001, Shiyi Song, Xueliang Wei, Huan Cheng, Wei Tong 0001, Dan Feng 0001 |
IEEE Trans. Computers | 5 |
| 2026 | Integrated Photoacoustic-Metabolomic Platform for Multimodal Assessment of Sepsis-Induced Brain DysfunctionabstractSepsis-induced brain dysfunction (SIBD), a critical determinant of mortality and long-term neurological sequelae in sepsis patients. The mechanistic understanding of SIBD has been limited by conventional single-modality approaches, which fail to capture the complex oxygenation-metabolism interplay. Here, we present an integrated photoacoustic-metabolomic platform that combines high-resolution photoacoustic imaging with targeted metabolomics to comprehensively assess cerebral oxygenation and metabolic alterations during sepsis. Our imaging system provides high spatiotemporal resolution, enabling mapping of key parameters, including cerebral oxygen saturation (sO2), oxygen extraction fraction (OEF), and the spatial heterogeneity of oxygen metabolism. By integrating these imaging capabilities with region-specific metabolomic profiling, we uncover a dynamic relationship between sepsis-driven oxygenation disruptions and metabolic reprogramming. Specifically, we demonstrate that sepsis induces heterogeneous cortical hypoxia, dysregulated OEF dynamics, and a metabolic shift marked by enhanced glycolysis and suppressed pentose phosphate pathway activity in high-OEF regions. This multimodal platform not only advances our understanding of the pathophysiology of SIBD but also offers a powerful tool for early diagnosis, personalized therapeutic strategies, highlighting the promise in bedside monitoring and precision medicine applications in sepsis management. Cong Mai, Xiaoxuan Zhong, Xiaoran Huang, Huan Cheng, Yongji Cui, Yukai Xu, Haoying Lan, Yi-Zhi Liang, Long Jin 0002, Bai-ou Guan |
IEEE Trans. Medical Imaging | 4 |
| 2025 | NeRF dynamic scene reconstruction based on motion, semantic information and inpaintingabstractIn this work, we address the inherent limitations of Neural Radiance Field (NeRF) in synthesizing novel viewpoints within dynamic environments, particularly those compromised by moving objects. Such scenarios frequently yield reconstructions of suboptimal quality, characterized by blurriness and the presence of artifacts, which significantly undermines the fidelity of synthetic scenes. This limitation significantly restricts the potential applications of NeRF in autonomous driving contexts, such as scene editing, high-precision map construction, and related functionalities. To overcome these challenges, we propose a novel NeRF-based approach tailored to address the complexities associated with moving objects in monocular driving scenarios. The proposed approach combines optical flow analysis and semantic information to precisely detect and localize moving objects. This was then followed by an inpainting technique that guides the NeRF reconstruction process, effectively mitigating the adverse impacts of dynamic elements within the scene. Our model is further enhanced by incorporating depth and semantic data to refine the training process. We validate the efficacy of our approach through comprehensive experimentation on both synthetic and real-world driving datasets, as well as on challenging self-recorded realistic driving scenes. Our method achieves a performance improvement of up to 13% compared to previous state-of-the-art methods. Additionally, we verify the efficacy of our approach through comprehensive ablation analyses. Both the quantitative and qualitative results demonstrate the superiority especially in dynamic driving scenes, advancing the potential applications in autonomous driving contexts. Our code and self-collected data are available at https://github.com/GandalfTGrey/Nerf-KBS.git . Huan Cheng, Shuo Wang 0030, Meng Li 0048 |
Neurocomputing | 2 |
| 2025 | Envelope rotation forest: A novel ensemble learning method for classification
Huan Cheng, Yongming Li 0003, Yinghua Shen |
Neurocomputing | 2 |
| 2025 | FADESIM: Enable Fast and Accurate Design Exploration for Memristive Accelerators Considering NonidealitiesabstractMemristive accelerators (MAs), built with memristive crossbar arrays (MCAs), have gained significant attention for their ability to efficiently perform matrix-vector multiplication in diverse applications. Modeling and simulation are indispensable tools for exploring and evaluating architectural design. Specifically, for MAs, maintaining high computation accuracy has been challenging because of realistic nonidealities like wire resistance (i.e., IR drop), I–V nonlinearity, program variation, and so on. Thus, for the system architects, fast and exact analysis of the effects of nonidealities is highly desirable, especially during the vast early design space exploration of the architecture. However, the SPICE model and existing MCA compact models (CMs) do not offer acceptable speeds for accurate simulation purposes, making them less practical. Additionally, the existing MCA simplified circuit models and predictive models fail to produce effective results due to the complexity of IR drop, let alone the coupling of multiple nonidealities. To enable fast and accurate design exploration for MAs with nonidealities considered, in this article, we propose FADESIM which includes fast IR drop simulation methods and a processing chain for the joint simulation of multiple nonidealities. Starting with the analysis of the accurate CM for IR drop, we explore the special properties of the model to enable fast iterative methods as well as a method to skip invalid calculations. This significantly reduces the time complexity of the simulation from naive$O(n^{6})$to$O(n^{3})$. For less severe IR drop cases, a custom iterative update algorithm is presented for faster simulations as a supplement, specifically with a time complexity of near$O(n^{2})$and proven applicable conditions. To simulate multiple nonidealities, we introduce a processing chain to inject corresponding processing functions before, during, and after the proposed fast IR drop simulation process according to the stages at which nonidealities take effect. The array-level experimental results show that our method achieves accurate simulations, with$19.8 \times - 884.9 \times $faster and$276.7 \times - 8018.8 \times $reduced memory usage compared to SPICE simulation. Further experiments at the algorithm-level demonstrate the effectiveness of our method in assisting architects with evaluating their designs. Bing Wu 0001, Huan Cheng, Xueliang Wei, Wei Tong 0001, Dan Feng 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2024 | DRCTL: A Disorder-Resistant Computation Translation Layer Enhancing the Lifetime and Performance of Memristive CIM ArchitectureabstractThe memristive Computing-in-Memory (CIM) sys-tem can efficiently accelerate matrix-vector multiplication (MVM) operations through in-situ computing. The data layout has a significant impact on the communication performance of CIM systems. Existing software-level communication opti-mizations aim to reduce communication distance by carefully designing static data layouts, while wear-leveling (WL) and error mitigation methods use dynamic scheduling to enhance system reliability, resulting in randomized data layouts and increased communication overhead. Besides, existing CIM compilers di-rectly map data to physical crossbars and generate instructions, which causes inconvenience for dynamic scheduling. To address these challenges of balancing communication performance and reliability while coordinating existing CIM compilers and dy-namic scheduling, we propose a disorder-resistant computation translation layer (DRCTL), which improves system lifetime and communication performance through co-optimization of data layout and dynamic scheduling. It consists of three parts: (1) We propose an address conversion method for dynamic scheduling, which updates the addresses in the instruction stream after dynamic scheduling, thereby avoiding recompilation. (2) Dynamic scheduling strategy for reliability improvement. We propose a hierarchical wear-leveling (HWL) strategy, which reduces communication by increasing scheduling granularity. (3) Communication optimization for dynamic scheduling. We propose data layout-aware selective remapping (LASR), which helps dynamic scheduling methods improve communication lo-cality and reduce latency by exploiting data dependencies. The experiments demonstrate that HWL extends lifetime by 100.3-205.9 x compared to not using WL. Even with a slight lifetime decrease compared to the state-of-the-art WL (TIWL), it still supports continuous neural network training for 7 years. After applying LASR to HWL, the number of execution cycles, energy consumption, on-chip and off-chip NoC accesses decrease by an average of 26.91 %, 26.88%, 36.41 %, and 80.62%, respectively. Bing Wu 0001, Huan Cheng, Taoming Lei, Dan Feng 0001, Wei Tong 0001 |
MICRO | 3 |
| 2024 | Deep Fuzzy Envelope Sample Generation Mechanism for Imbalanced Ensemble ClassificationabstractEnsemble methods are widely used to tackle class imbalance problem. However, for existing imbalanced ensemble (IE) methods, the samples in each subset are resampled from the same dataset, and are directly input to the classifier for training, so the quality (diversity and separability) of the subsets is unsatisfactory usually. To solve the problem, a deep fuzzy envelope sample generation mechanism is proposed. First, the fuzzy C-means clustering based deep sample envelope prenetwork (DSEN) is designed to mine correlation information among samples, thereby increasing the quality of the subsets. Second, the local manifold structure metric and global structure distribution metric are designed to construct local-global structure consistency mechanism (LGSCM) to enhance distribution consistency of interlayer samples of DSEN. Third, the DSEN and LGSCM are combined to form the final deep sample envelope network–DSENLG to refresh the existing subsets. Finally, base classifiers are applied on the new subsets generated by the DSENLG and then fused, thereby realizing a new IE algorithm. The experimental results show that the proposed algorithm is significantly better than existing representative IE algorithms and it achieves the highest improvement of 10.64%, 19.5%, 18.67% and 22.33% on four criteria over the state-of-the-art methods. The originality of the article is threefold: proposing the concept of “deep fuzzy samples” or “envelope samples”, which comprehensively considers the correlation information among original samples; proposing the LGSCM to resolve the distribution inconsistency of interlayer samples; and forming an fuzzy envelope sample based IE algorithm. Fan Li 0024, Yongming Li 0003, Yinghua Shen, Witold Pedrycz, Pufei Li, Chuanyan Zhou, Huan Cheng |
IEEE Trans. Fuzzy Syst. | 9 |
| 2023 | ODLPIM: A Write-Optimized and Long-Lifetime ReRAM-Based Accelerator for Online Deep LearningabstractReRAM-based Processing-In-Memory (PIM) architectures have demonstrated high energy efficiency and performance in deep neural network (DNN) acceleration. Most of the existing PIM accelerators for DNN focus on offline batch learning (OBL) which requires the whole dataset to be available before training. However, in the real world, data instances arrive in sequential settings, and even the data pattern may change, which calls concept drift. OBL requires expensive retraining to solve concept drift, whereas online deep learning (ODL) is evidenced to be a better solution to keep the model evolving over streaming data. Unfortunately, when ODL optimizes models over a large-scale data stream in the PIM system, unbalanced writes are more severe than OBL, due to the heavier weight updates, resulting in the amplification of unbalanced writes and lifetime deterioration. In this work, we propose ODLPIM, an online deep learning PIM accelerator that extends the system lifetime through algorithm-hardware co-optimization. ODLPIM adopts a novel write-optimized parameter update (WARP) scheme that reduces the non-critical weight updates in hidden layers. Besides, a table-based inter-crossbar wear-leveling (TIWL) scheme is proposed and applied to the hardware controller to achieve wear-leveling between crossbars for lifetime improvement. Experiments show that WARP reduces weight updates on average to 15.25% and up to 24% compared to that without WARP, and eventually prolongs system lifetime on average to 9.65% and up to 26.81%, with a negligible rise in cumulative error rate (up to 0.31%). By combining WARP with TIWL, the lifetime of ODLPIM is improved by an average of$\mathbf{12}.\mathbf{59}\times$and up to$\mathbf{17}.\mathbf{73}\times$. Bing Wu 0001, Huan Cheng, Wei Zhao 0034, Xueliang Wei, Dan Feng 0001, Wei Tong 0001 |
DATE | 3 |
| 2023 | ICON: An IR Drop Compensation Method at OU Granularity with Low Overhead for eNVM-based AcceleratorsabstractProcessing at operating unit (OU) granularity can alleviate the program variation effect and ADC conversion overhead in emerging non-volatile memory (eNVM) based accelerators. However, our experiments show that the IR drop effect can severely decrease computing accuracy when processing at OU granularity. Moreover, the IR drop effect on the entire array differs from the IR drop effect on OUs, meaning compensating at OU granularity is necessary. We also notice that the IR drop effect differs among OUs, and previous IR drop mitigation methods introduce more latency, area, and power overhead to adapt to these differences. Compensation modules from their methods calibrated for one OU do not apply to other OUs and need to be configured for each OU compensation using configuration modules. This paper proposes ICON, an IR drop compensation method at OU granularity with low overhead for eNVM-based accelerators. In order to decrease compensation latency, area, and power overhead, we perform several optimizations. First, the designed compensation circuit is simplified and does not compensate for the IR drop effect caused by parasitic resistances inside an OU. This simplification is based on our observation that the parasitic wire resistances inside an OU can be ignored using the OU size mentioned in previous works. Second, the compensation circuit is designed without the help of configuration circuits. We take the IR drop differences among OUs as input parameters of the compensation circuit so that it can apply to all OUs. Furthermore, the compensation circuit is pipelined into six stages to increase throughput. Experiments show that our compensation method can overcome the IR drop problem when processing in the eNVM-based crossbar array at OU granularity, with 1.3× ~ 13× lower latency, 1.5× ~ 33.1× lower area, and 1.4× ~ 8.4× lower power overhead compared with state-of-the-art methods. Wei Tong 0001, Bing Wu 0001, Huan Cheng, Chengning Wang |
ICCD | 4 |