EDBT 2026 Demo / reviewers in the wild / expert
Haitong Li
dblp:154/1110
· DBLP profile ↗
11ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0003-3393-9252ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dual-stream frequency-domain framework with contextual graph enhancer and consensus-difference fusion for cross-view geo-localization
Haitong Li, Chaoyi Ma, Yuehuan Wang, Ruonan Wei |
J. Vis. Commun. Image Represent. | 2 |
| 2026 | Cross-Layer Design of Vector-Symbolic Computing: Bridging Cognition and Brain-Inspired Hardware AccelerationabstractVector Symbolic Architectures (VSAs), also known as hyperdimensional (HD) computing, are increasingly deployed in cognitive applications due to their simple and efficient operations. The widespread adoption has, in turn, spurred the development of a diverse set of hardware solutions that optimize VSA performance for embedded and edge AI systems. Despite these advances, there remains a lack of comprehensive, unified discussion on the co-design and co-evolution of VSA algorithms and hardware. This survey aims at bridging that gap by linking theoretical, software-level explorations with efficient hardware architectures and emerging technology fabrics for VSAs, providing co-design insights that are accessible to both algorithm and hardware communities. First, we introduce the principles of vector-symbolic computing, including its core mathematical operations and learning paradigms. Second, we provide an in-depth discussion on hardware technologies for VSAs, analyzing analog, mixed-signal, and digital circuit design styles. We compare hardware implementations of VSAs by carrying out detailed analysis of their performance characteristics and tradeoffs, from which we distill design guidelines that are applicable across arbitrary VSA formulations. Third, we discuss a methodology for cross-layer design of VSAs that identifies synergies across layers and explores key ingredients for hardware/software co-design of VSAs. Finally, as a concrete case study of this methodology, we present an in-memory computing hardware design for VSA-based hierarchical cognition, illustrating how the proposed co-design principles translate into efficient architectures. The article concludes with a discussion of open research challenges and opportunities for future explorations. Shuting Du, Mohamed Ibrahim 0002, Zishen Wan, Luqi Zheng, Boheng Zhao, Zhenkun Fan, Che-Kai Liu, Tushar Krishna, Arijit Raychowdhury, Haitong Li |
ACM Trans. Embed. Comput. Syst. | 10 |
| 2025 | 3D-CIMlet: A Chiplet Co-Design Framework for Heterogeneous In-Memory Acceleration of Edge LLM Inference and Continual LearningabstractThe design space for edge AI hardware supporting large language model (LLM) inference and continual learning is underexplored. We present 3D-CIMlet, a thermal-aware modeling and co-design framework for 2.5D/3D edge-LLM engines exploiting heterogeneous computing-in-memory (CIM) chiplets, adaptable for both inference and continual learning. We develop memory-reliability-aware chiplet mapping strategies for a case study of edge LLM system integrating RRAM, capacitor-less eDRAM, and hybrid chiplets in mixed technology nodes. Compared to 2 D baselines, $2.5 \mathrm{D} / 3 \mathrm{D}$ designs improve energy efficiency by up to 9.3 x and 12 x, with up to 90.2% and 92.5% energy-delay product (EDP) reduction respectively, on edge LLM continual learning. Shuting Du, Luqi Zheng, Aradhana Mohan Parvathy, Feifan Xie, Tiwei Wei, Anand Raghunathan, Haitong Li |
DAC | 7 |
| 2024 | Special Session: Neuro-Symbolic Architecture Meets Large Language Models: A Memory-Centric PerspectiveabstractLarge language models (LLMs) have significantly transformed the landscape of artificial intelligence, demonstrating exceptional capabilities in natural language understanding and generation. Recently, the integration of LLMs with neurosymbolic architectures has gained traction to enhance contextual awareness and planning capabilities. However, this integration faces computational challenges that hinder scalability and efficiency, especially in edge computing environments. This paper provides an in-depth analysis of these challenges and explores state-of-the-art solutions, focusing on memory-centric computing principles at both algorithmic and hardware levels. Our exploration is centered around the key computational elements of the Transformer, the foundation of all LLMs, and vector-symbolic architecture, the leading neuro-symbolic model for edge applications. Additionally, we propose potential research directions for further investigation. By examining these aspects, this paper aims to bridge critical gaps in the path toward effective artificial general intelligence at the edge. Mohamed Ibrahim 0002, Zishen Wan, Haitong Li, Priyadarshini Panda, Tushar Krishna, Pentti Kanerva, Yiran Chen 0001, Arijit Raychowdhury |
CODES+ISSS | 3 |
| 2024 | Co-designing 2.5D Silicon Photonic Accelerators for Distributed Transformer at the EdgeabstractThe efficient execution of attention-based transformers and large language models on traditional CPUs and GPUs presents significant challenges related to performance and energy efficiency. While innovative solutions like ASICs, FPGAs, and ReRAMs have been explored, the field of silicon photonics has emerged as a promising avenue for developing energy-efficient accelerators for deep AI models. Notably, existing endeavors in silicon photonics have predominantly concentrated on inference for deep AI algorithms, leaving a limited number of initiatives focused on creating comprehensive deep learning accelerators capable of real-time training for transformer-like algorithms. This paper utilizes the superior merits of silicon photonics to realize a full-fledged transformer accelerator equipped for both inference and training. Introducing PHOTRAN, an AI analog photonics accelerator, we harness silicon microdisk-based convolution, photonic phase-change memory-based cache, and dense-wavelength-division-multiplexing to achieve energy-efficient and ultrafast transformer acceleration. Through evaluations using a commercial CAD framework on benchmark models, including Vision Transformers and Large Language models, our results showcase the superior performance of PHOTRAN. This work underscores the significant potential of photonic computing for on-chip training of large deep AI models. Dharanidhar Dang, Priyabrata Dash, Luqi Zheng, Haitong Li |
ICCAD | 4 |
| 2024 | Small Fixed-Wing Unmanned Aerial Vehicle Path Following Under Low Altitude Wind Shear DisturbanceabstractThe wind shear in low altitude can bring about time-varying and unknown disturbances, which makes small fixed-wing Unmanned Aerial Vehicle (UAV) hard to accurately follow desired curved paths. In this paper, a Vector Field (VF) based curved path following algorithm is designed for UAV to overcome the above difficulties. Firstly, the path following problem is mathematically formulated to give the kinematics model of UAV with constant airspeed. Secondly, the curved path following control laws under both constant and time-varying unknown wind disturbances are designed with VF. A ground velocity estimator is additionally designed to overcome the difficulties of unmeasurable ground velocity, and the Lyapunov stability of curved path following is analyzed. Finally, both simulations and field tests on a real small fixed-wing UAV are carried out to evaluate the algorithm performance in the presence of time-varying and unknown wind disturbances. Verification results demonstrate that the algorithm proposed in this paper outperforms other widely used path following algorithms and can effectively follow arbitrary curved paths. Zhouyu Zhang 0002, Chenyuan He, Hongshen Chen, Hai Wang 0003, Yingfeng Cai, Long Chen 0003, Haitong Li, Tongwei Lu |
IEEE Trans. Intell. Transp. Syst. | 8 |
| 2023 | Emerging Hardware Technologies and 3D System Integration for Ubiquitous Machine IntelligenceabstractNext-generation semiconductor hardware technologies and system integration serve as the physical foundation in the pursuit of ubiquitous machine intelligence, with unprecedented requirements in energy efficiency, performance, cost effectiveness, and security. Here, we provide an overview of emerging technologies with an emphasis on 3D system integration, and discuss on cross-layer designs for memory-centric computing in the 3D era. Haitong Li |
DAC | 1 |
| 2020 | Timely: Pushing Data Movements And Interfaces In Pim Accelerators Towards Local And In Time DomainabstractResistive-random-access-memory (ReRAM) based processing-in-memory (R2PIM) accelerators show promise in bridging the gap between Internet of Thing devices' constrained resources and Convolutional/Deep Neural Networks' (CNNs/DNNs') prohibitive energy cost. Specifically, R2PIM accelerators enhance energy efficiency by eliminating the cost of weight movements and improving the computational density through ReRAM's high density. However, the energy efficiency is still limited by the dominant energy cost of input and partial sum (Psum) movements and the cost of digital-to-analog (D/A) and analog-to-digital (A/D) interfaces. In this work, we identify three energy-saving opportunities in R2PIM accelerators: analog data locality, time-domain interfacing, and input access reduction, and propose an innovative R2PIM accelerator called TIMELY, with three key contributions: (1) TIMELY adopts analog local buffers (ALBs) within ReRAM crossbars to greatly enhance the data locality, minimizing the energy overheads of both input and Psum movements; (2) TIMELY largely reduces the energy of each single D/A (and A/D) conversion and the total number of conversions by using time-domain interfaces (TDIs) and the employed ALBs, respectively; (3) we develop an only-once input read (O2IR) mapping method to further decrease the energy of input accesses and the number of D/A conversions. The evaluation with more than 10 CNN/DNN models and various chip configurations shows that, TIMELY outperforms the baseline R2PIM accelerator, PRIME, by one order of magnitude in energy efficiency while maintaining better computational density (up to 31.2×) and throughput (up to 736.6×). Furthermore, comprehensive studies are performed to evaluate the effectiveness of the proposed ALB, TDI, and O2IR in terms of energy savings and area reduction. Pengfei Xu 0011, Yang Zhao 0013, Haitong Li, Yuan Xie 0001, Yingyan (Celine) Lin |
ISCA | 4 |
| 2019 | On-Chip Memory Technology Design Space Explorations for Mobile Deep Neural Network AcceleratorsabstractDeep neural network (DNN) inference tasks have become ubiquitous workloads on mobile SoCs and demand energy-efficient hardware accelerators. Mobile DNN accelerators are heavily area-constrained, with only minimal on-chip SRAM, which results in heavy use of inefficient off-chip DRAM. With diminishing returns from conventional silicon technology scaling, emerging memory technologies that offer better area density than SRAM can boost accelerator efficiency by minimizing costly off-chip DRAM accesses. This paper presents a detailed design space exploration (DSE) of technology-system co-design for systolic-array accelerators. We focus on practical/mature on-chip memory technologies, including SRAM, eDRAM, MRAM, and 3D vertical RRAM (VRRAM). The DSE employs state-of-the-art optimizations (e.g., model compression and optimized buffer scheduling), and evaluates results on important models including ResNet-50, MobileNet, and Faster-RCNN. Compared to an SRAM/DRAM baseline, MRAM-based accelerators show up to 4.68× energy benefits (57% area overhead), while a 3D VRRAM-based design achieves 2.22× energy benefits (33% area reduction). Haitong Li, Mudit Bhargava, Paul N. Whatmough, H.-S. Philip Wong |
DAC | 1 |
| 2015 | Modeling and design optimization of ReRAMabstractResistive switching memories (ReRAM) have been widely studied for applications in next-generation data storage and neurormorphic computing systems. To enable device-circuit-system co-design and optimization, a SPICE model of ReRAM that can reproduce the device characteristics in circuit simulations is needed. In this paper, we present a novel tool for ReRAM design including a physics-based SPICE model, the model parameters extraction strategy, as well as the system assessment method. This physics-based SPICE model can capture all the essential features of HfOx-based ReRAM including the DC/AC and multi-level switching behaviors, switching reliability, and intrinsic device variations. A strategy is developed to extract the critical model parameters from the fabricated ReRAM devices. A variety of electrical measurements on various ReRAMs are performed to verify and calibrate the model. The assessment method based on the experimentally verified SPICE model can be applied to explore a wide range of applications including: 1) variation-aware and reliability-emphasized system design; 2) system performance evaluation; 3) array architecture optimization. This verified design tool not only enables system design but also enables system optimization that capitalizes on device/circuit interaction for both data storage and neuromorphic computing applications. Jinfeng Kang, Haitong Li, Peng Huang 0004, Bin Gao 0006, Zizhen Jiang, H.-S. Philip Wong |
ASP-DAC | 2 |
| 2015 | Variation-aware, reliability-emphasized design and optimization of RRAM using SPICE model
Haitong Li, Zizhen Jiang, Peng Huang 0004, Hong-Yu Chen, Bin Gao 0006, Jinfeng Kang, H.-S. Philip Wong |
DATE | 1 |