EDBT 2026 Demo / reviewers in the wild / expert
Lixia Han
dblp:78/1605
· DBLP profile ↗
16ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0002-4721-1564ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Re-RIS: A Reconfigurable 3D RRAM In-Sensor Architecture for Low-Latency Machine VisionabstractCutting-edge machine vision applications impose stringent latency and energy efficiency demands on edge devices. To address these demands, In-Sensor Computing (ISC) architectures aim to eliminate data movement overhead, while 3D RRAM technology provides the hardware foundation of high memory density and massive computing parallelism. However, existing ISC architectures rely on static resource allocation, failing to address the dynamic "shifting bottleneck" in CNNs— where early layers are compute-bound and later layers are readout-bound. To address this, we propose Re-RIS, a Reconfigurable 3D RRAM In-Sensor architecture. By dynamically switching hardware granularity between high-parallelism and high-throughput modes, Re-RIS optimizes resource utilization for varying layer characteristics. Experimental results on VGG-16 demonstrate an end-to-end latency of 0.93 ms, achieving a 75% reduction compared to static baselines, with an energy efficiency of 244.6 TOPS/W and an area efficiency of 1.85 TOPS/mm2. Lixia Han, Lifeng Liu, Peng Huang 0004 |
DATE | 2 |
| 2026 | PIM-NoC: A NoC Architecture with In-Router Processing-in-Memory for DNN Acceleration
Haixiang Ren, Lixia Han, Cuiyu Qi, Hui Chen 0015 |
ISCAS | 3 |
| 2026 | C2NoC: A Communication-Computation Coupled NoC-based Neural Network Accelerator
Cuiyu Qi, Hui Chen 0015, Lixia Han, Chenkai Cao, Boxiang Zhang, Weiqiang Liu 0001 |
ISCAS | 4 |
| 2026 | An 8T-SRAM Near-Memory Architecture for Multiplierless Approximate DCT
Ke Chen 0018, Bi Wu 0002, Chenggang Yan 0002, Lixia Han, Chenghua Wang, Weiqiang Liu 0001 |
ISCAS | 5 |
| 2026 | Instant-CIM: An Instant Neural Radiance Field Computing-In-Memory Architecture for Low-Power and Real-Time AR/VR RenderingabstractNovel View Synthesis is a foundational technique for creating immersive Augmented and Virtual Reality (AR/VR) experiences, aiming to generate photorealistic images of a scene from arbitrary camera viewpoints using only a limited set of source images, with Neural Radiance Fields (NeRF) emerging as the state-of-the-art solution. However, real-time NeRF rendering on low-power devices remains challenging due to its memory-intensive hash encoding and compute-intensive Multilayer Perception (MLP). In this work, we propose Instant-CIM, the fully on-chip Computing-in-Memory (CIM) architecture for efficient NeRF rendering. At the algorithm level, Instant-CIM proposes a spatially-adaptive framework that dynamically selects the number of active hash encoding levels per spatial region based on a composite importance score derived from density and gradient. The approach replaces uniform level allocation with a threshold-based strategy that activates finer encoding levels only in regions with high representation complexity. At the hardware level, Instant-CIM proposes an in-situ hash engine that implements in-memory hash query and interpolation through 3D scene grid decomposition and Z-order based mapping schemes. Meanwhile, Instant-CIM proposes a sparse MLP engine that leverages differential-based input complemented by a precision-adjustable skipping mechanism to fully exploit spatial similarities. Comprehensive evaluation across synthetic datasets demonstrates that Instant-CIM achieves 3.0×~4.9× improvement in rendering speed and 8.6×~33× enhancement in energy efficiency compared to state-of-the-art NeRF architecture. Lixia Han, Hui Chen 0015, Xueming Fu, Ke Chen 0018, Peng Huang 0004, Yijun Cui, Weiqiang Liu 0001 |
IEEE Trans. Computers | 1 |
| 2025 | Mitigating methodology of hardware non-ideal characteristics for non-volatile memory based neural networks
Lixia Han, Peng Huang 0004, Yijiao Wang, Haozhang Yang, Jinfeng Kang |
Sci. China Inf. Sci. | 1 |
| 2025 | CIMUS: 3D-Stacked Computing-in-Memory Under Image Sensor Architecture for Efficient Machine VisionabstractComputational image sensors with CNN processing capabilities are emerging to alleviate the energy-intensive and time-consuming data movement between sensors and external processors. However, deploying CNN models onto these computational image sensors faces challenges from the limited on-chip memory resources and insufficient image processing throughput. This work proposes a 3D-stacked NAND flash-based computing-in-memory under image sensor architecture (CIMUS) to facilitate the complete deployment of CNN model. To fully leverage the potential of high bandwidth from the 3D-stacked integration, we design a novel distributed CNN mapping and dataflow to process the full focal plane image in parallel, which senses and recognizes ImageNet tasks with >1000fps. To tackle the computational error of inputs “0” in 3D NAND flash-based CIM, we propose an input-independent offset compensation method, which reduces the average vector-matrix multiplication (VMM) error by 48%. Evaluation results indicate that CIMUS architecture achieves a 9.8× improvement in CNN inference speed and a 33× boost in energy efficiency compared to the state-of-the-art computational image sensor in the ImageNet recognition task. Lixia Han, Haozhang Yang, Ao Shi, Guihai Yu, Yijiao Wang, Yanzhi Wang 0001, Jinfeng Kang, Peng Huang 0004 |
IEEE Trans. Computers | 1 |
| 2024 | Pipeline Design of Nonvolatile-based Computing in Memory for Convolutional Neural Networks Inference AcceleratorsabstractNonvolatile-based computing-in-memory inference chips show great potential to accelerate convolutional neural networks. The intrinsic weight stationary characteristic makes pipeline design a crucial solution to further enhance throughput. In this work, we propose a balanced pipeline design and establish performance/area evaluation models for the optimal pipeline solution. The evaluation results indicate that our pipeline design achieves$30\times$computational efficiency improvement. Lixia Han, Peng Huang 0004, Haozhang Yang, Jinfeng Kang |
DATE | 1 |
| 2024 | Low Quantization Error Readout Circuit with Fully Charge-Domain Calculation for Computation-in-Memory Deep Neural NetworkabstractThis work presents a low quantization error readout circuit with fully-charge-domain calculation for quantization and post-process of computation-in-memory (CIM)-based neural network. The contributions include: (1) A novel residual charge accumulation function is designed to achieve charge-domain summation of quantized partial sum, and reduces 38% quantization error; (2) Charge reset is introduced in the integrate & fire circuit to realize <1 LSB INL at ±7 bits and speed of 285MHz/LSB; (3) Sample & hold, current subtraction and bidirectional counter are designed to improve 3.95× energy efficiency and 2.48× area efficiency. Ao Shi, Lixia Han, Lifeng Liu, Linxiao Shen, Peng Huang 0004, Jinfeng Kang |
ISCAS | 3 |
| 2024 | CoMN: Algorithm-Hardware Co-Design Platform for Nonvolatile Memory-Based Convolutional Neural Network AcceleratorsabstractComputing in memory (CIM) convolutional neural network (CNN) accelerators based on nonvolatile memory (NVM) show great potential to improve energy efficiency and throughput, while the multiple design levels and huge design space of CIM-based CNN acceleration system make cross-level co-design methodology and platforms extremely desired. In this work, an algorithm-hardware co-design platform CoMN with the graphic user interface is proposed for designers to fast verify and further optimize the designments. In the platform, 1) a mapper is developed to automatically map CNN models to CIM chips through optimizing pipeline, weight transformation, partition, and placement; 2) accuracy evaluator and performance evaluator are built to jointly estimate accuracy, energy, latency, and area overheads considering the design dependencies across multiple levels; 3) algorithm adapter is exploited to retrain CNN weights for higher hardware accuracy within limited energy budget through nonidealities aware training and energy aware training; 4) hardware optimizer is developed to search hardware microarchitecture and circuit design space in the early design stage. We conduct several case studies to verify the effectiveness of the CoMN platform. Results indicate that CoMN platform can enable algorithm-hardware mapping, hardware-aware algorithm adaption, hardware configuration exploration, and overall algorithm-hardware co-design efficiently. The CoMN platform can be accessed online at http://101.42.97.22:8081/index.html with username “tcad” and password “comnuser”. Lixia Han, Renjie Pan 0003, Hairuo Lu, Haozhang Yang, Peng Huang 0004, Guangyu Sun 0003, Jinfeng Kang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | Specific ADC of NVM-Based Computation-in-Memory for Deep Neural NetworksabstractNon-volatile memory (NVM)-based Computation-in-memory has demonstrated a significant advantage in high-efficiency neural networks. However, the requirement of analog-to-digital converter (ADC) and post-processing circuits not only cost high energy and area but also results in high computation errors, which tradeoffs the performance boost brought by CIM. Here, we present a specific ADC and post-processing circuit of the NVM-based CIM neural network to address these issues. The main contributions include: (1) A novel residual charge accumulation function (RCA) is designed to achieve charge-domain summation of quantized partial sum and reduces 38% quantization error; (2) Charge reset is introduced in the integrate & fire circuit to realize$3.95\times $energy efficiency and$2.48\times $area efficiency. Evaluation based on the measured results of the fabricated chip shows that the VGG-11 neural network with the proposed ADC circuit can achieve a 3.28-time improvement in energy efficiency while maintaining the same network recognition rate. Ao Shi, Lixia Han, Haozhang Yang, Lifeng Liu, Linxiao Shen, Jinfeng Kang, Peng Huang 0004 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | A Convolution Neural Network Accelerator Design with Weight Mapping and Pipeline OptimizationabstractThe pipeline is an efficient solution to boost performance in non-volatile memory based computing in memory (nvCIM) convolution neural network (CNN) accelerators. However, the previous works seldom focus on pipeline optimization from the perspective of the whole system, especially overlooking the effect of buffer access. In this work, we propose a high-performance NVM-based CNN accelerator with a balanced pipeline design, which takes account of both the macro computing and the buffer access. At the operator level, a matrix-based weight mapping method is proposed to reduce buffer access delay. At the macro level, decoupled access and execution design is introduced to shorten the single-layer latency. At the system level, a hybrid inter/intra-tile design is presented to balance the overall latency across CNN layers. With the collaboration among three methods, we construct a well-balanced pipeline for the nvCIM accelerator at a smaller hardware cost. Experiments show that our pipeline design can achieve 3.7х, 7.5х, and 3.5х throughput improvement for recognition of ImageNet with ResNet18, VGG19, and ResNet34 models, respectively. Lixia Han, Peng Huang 0004, Jinfeng Kang |
DAC | 1 |
| 2023 | Co-optimization strategy between array operation and weight mapping for flash computing arrays to achieve high computing efficiency and accuracy
Guihai Yu, Peng Huang 0004, Runze Han, Lixia Han, Jinfeng Kang |
Sci. China Inf. Sci. | 4 |
| 2022 | Efficient Discrete Temporal Coding Spike-Driven In-Memory Computing Macro for Deep Neural Network Based on Nonvolatile MemoryabstractNonvolatile memory (NVM) based neural network can directly perform in situ computation in memory to significantly reduce energy consumption resulting from the data movement. However, the energy consumption by the analog-to-digital converter (ADC) restricts the efficiency of the mixed-signal in-memory computing macro. The rate coding spike-driven in-memory computing macro can increase the energy efficiency via eliminating the ADC, but the improvement is limited because substantial energy is consumed for the coding of multiple spikes. In this work, we propose a discrete temporal coding spike-driven in-memory computing macro, including input coding scheme, weight mapping method, and improved leaky integrate-and-fire (LIF) neuron circuit, to perform the efficient forward inference of deep neural networks based on NVM array. We then optimize the designment of the proposed in-memory computing macro to mitigate the neural network accuracy loss due to the nonlinearity of the LIF neuron and voltage drop caused by interconnect resistance. Because the temporal coding scheme reduces spike numbers and the improved-LIF circuit simultaneously integrates two bit-lines current corresponding to positive and negative weight, the proposed macro achieves 46.63TOPS/W energy efficiency and 1.92TOPS throughput for 3bit temporal coding precision. Lixia Han, Peng Huang 0004, Yijiao Wang, Jinfeng Kang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2010 | Decentralized adaptive practical tracking of nonlinear interconnected systems with dynamic input and output interactionsabstractIn this paper, the adaptive practical output tracking problem on a class of nonlinear interconnected systems is considered, in which the dynamic interactions and unmodelled dynamics depend on both the subsystem inputs and outputs. Under the assumption that both the references and their first-order derivatives are bounded, decentralized adaptive backstepping controllers have been proposed to make sure that all the signals in the interconnected systems are globally uniformly bounded, with the tracking errors converging to a small set around the origin. Lixia Han, Huijin Fan |
ICARCV | 1 |
| 2009 | A clustering multi-objective evolutionary algorithm based on orthogonal and uniform designabstractDesigning efficient algorithms for difficult multi-objective optimization problems is a very challenging problem. In this paper a new clustering multi-objective evolutionary algorithm based on orthogonal and uniform design is proposed. First, the orthogonal design is used to generate initial population of points that are scattered uniformly over the feasible solution space, so that the algorithm can evenly scan the feasible solution space once to locate good points for further exploration in subsequent iterations. Second, to explore the search space efficiently and get uniformly distributed and widely spread solutions in objective space, a new crossover operator is designed. Its exploration focus is mainly put on the sparse part and the boundary part of the obtained non-dominated solutions in objective space. Third, to get desired number of well distributed solutions in objective space, a new clustering method is proposed to select the non-dominated solutions. Finally, experiments on thirteen very difficult benchmark problems were made, and the results indicate the proposed algorithm is efficient. Yuping Wang 0003, Chuangyin Dang, Hecheng Li, Lixia Han, Jingxuan Wei |
IEEE Congress on Evolutionary Computation | 4 |