EDBT 2026 Demo / reviewers in the wild / expert
Jintao Li 0002
dblp:414/6882-2
· DBLP profile ↗
16ranked-venue papers
8as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Priority-Based Graph-Enhanced Reinforcement Learning for Robust Analog Circuit OptimizationabstractA primary motivation for analog integrated circuit (IC) design automation is the inefficiency of manual design in meeting increasingly stringent specifications, which often involve over 10 objectives. Recent advances in reinforcement learning (RL) emerge as a promising method, yet gaps remain when considering full design specifications, especially under process-voltage-temperature (PVT) variations. Excessive objectives lead to diminished reward signals, while varying PVT conditions result in conflicting gradients, both of which result in inefficient exploration. To address these, we propose a priority-based graph-enhanced RL framework. Specifically, using fuzzy logic converts quantitative rewards into qualitative priority signals, mitigating reward deterioration and enhancing exploration via entropy regularization. Furthermore, a graph-based representation compresses high-dimensional objective spaces under PVT variations into low-dimensional manifolds, enabling dynamic resource allocation to variation-sensitive regions and resolving gradient conflicts. Empirical results on various real-world analog ICs demonstrate that our method significantly outperforms existing RL algorithms, achieving superior solution quality and reducing simulation overhead. Jintao Li 0002, Zhenxin Chen, Aojin Li, Shui Yu 0002 |
AAAI | 1 |
| 2026 | Multitask evolution with problem reformulation for global exploration in analog circuit design
Jintao Li 0002, Aojin Li, Shui Yu 0002, Yun Li 0002 |
Adv. Eng. Informatics | 1 |
| 2025 | Balancing Objective Optimization and Constraint Satisfaction for Robust Analog Circuit OptimizationabstractAutomated design of analog integrated circuits (ICs) involves balancing multiple objectives under process, voltage, and temperature (PVT) variations. An excess of constraints can ensnare algorithms in local optima, while the variations elevate the costs of simulation. To address this challenge, we propose a two-search mode multi-task evolutionary framework to balance objective optimization and constraint satisfaction under variations. Specifically, considering the inherent relationships between objective optimizations and constraint violations, our method adaptively switches between unconstrained surrogate-assisted and constrained simulation-driven search modes. Furthermore, our framework treats PVT variations as a multi-task challenge, facilitating inter-corner knowledge transfer via multi-task evolution, substantially lowering simulation costs. Our framework has been evaluated using two different sensing elements and an amplifier within a 22 nm process. Based on Monte-Carlo simulations, compared to multi-task reinforcement learning, this method attains a 60% to 80% reduction in the relative inaccuracy of sensing elements and accomplishes a 60% decrease in total runtime. Jintao Li 0002, Haochang Zhi, Jiang Xiao 0002, Yanhan Zeng, Weiwei Shan, Yun Li 0002 |
ASP-DAC | 1 |
| 2025 | Analog Circuit Transfer Method Across Technology Nodes via Transistor BehaviorabstractIn the post-Moore era, chips integrate multiple technology node chiplets, necessitating repeated implementations of the same circuit topology across nodes, highlighting the need for technology-independent circuit representation. We use a four-parameter vector---gm, ft, VDS, and ΔVGS-to represent the behavior of each transistor, called the transistor behavioral vector (TBV). The TBVs are vertically concatenated to form the transistor behavioral circuit representation (TBCR) matrix, which precisely reflects the circuit's performance and provides a technology-independent representation. Furthermore, we propose a transistor behavioral model (TBM) to convert the TBV into the corresponding sizing. Finally, we propose a method to transfer analog circuits between different technology nodes using TBM (TNT), translating the modifications in the process parameters into the corresponding adjustments in ΔVGS. The experimental results show that for a single transistor, the mapping accuracy from TBV to simulation result was reached 99%. Multiple amplifiers were transferred from 180nm to 22nm technology, compared to the conventional transfer method based on gm/id, our transfer method based on TBCR achieved a success rate of up to 5× higher, along with additional performance improvements from the scaling down. Haochang Zhi, Jintao Li 0002, Yun Li 0002, Weiwei Shan |
ASP-DAC | 2 |
| 2025 | Decoupling Analog Circuit Representation from Technology for Behavior-Centric OptimizationabstractAnalog IC design is mainly manual and implemented at the device level. A major reason is circuit behavior-extraction. Unlike its digital counterpart, analog IC design is strongly coupled with technology nodes and is difficult to represent by an abstract behavioral model. The lack of accurate and efficient analog modeling has become a bottleneck in analog design automation. This paper proposes a behavior-centric optimization framework for analog circuits that represents circuit behavior using transistor electrical properties instead of sizes, improving model generalization and reducing optimization complexity. To characterize the process, we propose a method for mapping transistor electrical properties to sizes. Moreover, we developed a radial basis functions-based Kolmogorov-Arnold network (RBF-KAN) to accurately approximate circuit nonlinear behavior with limited simulations. Compared to blackbox modeling, our approach enables constructing surrogate models via KAN under a set specification with just a few hundred simulations. Experiments on the testing suite showed our framework achieved a $1.76 \times$ to $2.64 \times$ improvement in large signal figure of merit (FOM) and $1.73 \times$ to $2.48 \times$ in small signal FOM over state-of-the-art methods, while also enabling $3.5 \times$ to $6.2 \times$ acceleration in design porting. Jintao Li 0002, Haochang Zhi, Jiang Xiao 0002, Keren Zhu 0001, Yun Li 0002 |
DAC | 1 |
| 2025 | FlightVGM: Efficient Video Generation Model Inference with Online Sparsification and Hybrid Precision on FPGAsabstractVideo Generation Model (VGM), as a representative of multi-modal large models, has revolutionized the productivity of video content creation. VGMs are compute-bound due to adopting the Diffusion Transformer (i.e., DiT) structure. Sparsification is a common method for accelerating compute-intensive models. Still, sparse VGMs cannot fully exploit the effective throughput (i.e., TOPS) of GPUs. FPGAs are good candidates for accelerating sparse deep learning models. However, existing FPGA accelerators still face low throughput ( < 2TOPS) on VGMs due to the significant gap in peak computing performance (PCP) with GPUs ( > 21× ). To achieve a higher throughput than GPUs, FPGA-based acceleration of sparse VGMs still faces the following challenges: large redundancy in activations, low performance of DSPs under hybrid precision, and under-utilization using static compilation for online compression. Jun Liu 0117, Shulin Zeng, Li Ding 0012, Widyadewi Soedarmadji, Hao Zhou 0008, Jinhao Li 0006, Jintao Li 0002, Yadong Dai, Kairui Wen, Yaqi Sun, Yu Wang 0002, Guohao Dai 0001 |
FPGA | 8 |
| 2025 | FMC-LLM: Enabling FPGAs for Efficient Batched Decoding of 70B+ LLMs with a Memory-Centric Streaming ArchitectureabstractFor large language model (LLM) acceleration, FPGAs face two challenges: insufficient peak computing performance and unacceptable accuracy loss of model compression. This paper proposes FMC-LLM to enable FPGAs for efficient batched decoding of 70B+ LLMs. Wenheng Ma, Shulin Zeng, Tengxuan Liu, Libo Shen, Jiewen Wang, Jintao Li 0002, Zhenhua Zhu 0002, Xuefei Ning, Tsung-Yi Ho, Guohao Dai 0001, Yu Wang 0002 |
FPGA | 11 |
| 2025 | Dynamically Reconfigurable NPU Acceleration for Knowledge Loading in LLM Retrieval-Augmented GenerationabstractRetrieval-Augmented Generation (RAG) provides large language models (LLMs) a means of retrieving relevant external knowledge, but its document parsing leads to increased latency and energy consumption. To address this issue, we propose a dynamically reconfigurable Neural Processing Unit (NPU) that accelerates both RAG document parsing and inference. By leveraging compute-in-memory fusion, dynamic convolution, and multi-level parallelism, our approach reduces memory transfer overhead and optimizes hardware resource allocation. Experimental results show that our design achieves a 1.8x speedup in document parsing and a 2.83x improvement in energy efficiency. Additionally, it achieves an 11.71% reduction in inference time and a 35.59% boost in energy efficiency over traditional CPU/GPU methods, offering a scalable solution for large-scale RAG tasks. Peidong Lin, Jintao Li 0002, Shihong Li, Shui Yu 0002, Yun Li 0002 |
SMC | 2 |
| 2025 | Enhancing Small Object Detection in Aerial Images via Transformer Scaling and Dynamic FusionabstractAt present, accurate detection of small objects in an aerial imagery remains a challenge in remote sensing due to limited pixel resolution, background clutter, and scale variations. To address these issues for high-precision detection in a complex remote sensing scene, we propose a novel detection framework based on RepViT Dynamic Fusion and YOLOv11, termed RDF-YOLO. The RDF-YOLO brings in two core innovations:(1) a Dynamic Scale RepViT module that integrates lightweight Transformer operations into the backbone to enhance global context modeling and semantic discrimination under noisy conditions, and (2) a dynamic fusion module that incorporates spatially aware dilated convolutions and channel-adaptive fusion strategies to enable flexible, scale-aware feature interaction. Extensive experiments on the challenging AI-TOD dataset show that the RDF-YOLO outperforms state-of-the-art methods by substantial margins. In particular, the RDF-YOLO improves AP50:95 by 6.9% and AP50by 8.6% over the YOLOv11 baseline and on small-object metrics, including APvt, APt, and APs. These results verify the effectiveness of the RDF-YOLO architecture for robust and efficient detection of small objects in remote sensing imagery. The source code is available at https://github.com/AssiiKk/RDF-YOLO. Jintao Li 0002, Weixuan Liu, Shui Yu 0002, Yun Li 0002 |
SMC | 2 |
| 2025 | Hierarchical multi-task circuit modeling for PVT robustness via KAN-CNN integration
Hanjie Cai, Jintao Li 0002, Tongyu Luo, Wenyue Cai, Chaoying Tang, Yanhan Zeng |
Expert Syst. Appl. | 2 |
| 2024 | Performance-Driven Analog Layout Automation: Current Status and Future Directions (Invited Paper)abstractOptimizing circuit performance presents a pivotal challenge in the realm of automatic analog physical design. The intricacy of analog performance arises from its sensitivity to layout implementation, frequently lacking a viable approach for direct optimization. This talk initiates with a comprehensive overview of the present challenges and the techniques currently in use. The emphasis will be laid on the recent advancements in employing black-box optimization for enhancing analog performance. Subsequently, we will delve into a detailed case study and analysis of post-layout performance distribution for a typical analog circuit. This study will showcase various layout implementations generated by the open-source analog layout generator, MAGICAL. Future directions will be discussed based on the case study. Peng Xu 0052, Jintao Li 0002, Tsung-Yi Ho, Bei Yu 0001, Keren Zhu 0001 |
ASPDAC | 2 |
| 2024 | FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGAsabstractTransformer-based Large Language Models (LLMs) have made a significant impact on various domains. However, LLMs' efficiency suffers from both heavy computation and memory overheads. Compression techniques like sparsification and quantization are commonly used to mitigate the gap between LLM's computation/memory overheads and hardware capacity. However, existing GPU and transformer-based accelerators cannot efficiently process compressed LLMs, due to the following unresolved challenges: low computational efficiency, underutilized memory bandwidth, and large compilation overheads. This paper proposes FlightLLM, enabling efficient LLMs inference with a complete mapping flow on FPGAs. In FlightLLM, we highlight an innovative solution that the computation and memory overhead of LLMs can be solved by utilizing FPGA-specific resources (e.g., DSP48 and heterogeneous memory hierarchy). We propose a configurable sparse DSP chain to support different sparsity patterns with high computation efficiency. Second, we propose an always-on-chip decode scheme to boost memory bandwidth with mixed-precision support. Finally, to make FlightLLM available for real-world LLMs, we propose a length adaptive compilation method to reduce the compilation overhead. Implemented on the Xilinx Alveo U280 FPGA, FlightLLM achieves 6.0× higher energy efficiency and 1.8× better cost efficiency against commercial GPUs (e.g., NVIDIA V100S) on modern LLMs (e.g., LLaMA2-7B) using vLLM and SmoothQuant under the batch size of one. FlightLLM beats NVIDIA A100 GPU with 1.2× higher throughput using the latest Versal VHK158 FPGA. Shulin Zeng, Jun Liu 0117, Guohao Dai 0001, Tianyu Fu 0004, Wenheng Ma, Hanbo Sun, Zixiao Huang 0001, Yadong Dai, Jintao Li 0002, Kairui Wen, Xuefei Ning, Yu Wang 0002 |
FPGA | 12 |
| 2024 | AnalogGym: An Open and Practical Testing Suite for Analog Circuit SynthesisabstractRecent advances in machine learning (ML) for automating analog circuit synthesis have been significant, yet challenges remain. A critical gap is the lack of a standardized evaluation framework, compounded by various process design kits (PDKs), simulation tools, and a limited variety of circuit topologies. These factors hinder direct comparisons and the validation of algorithms. To address these shortcomings, we introduced AnalogGym, an open-source testing suite designed to provide fair and comprehensive evaluations. AnalogGym includes 30 circuit topologies in five categories: sensing front ends, voltage references, low dropout regulators, amplifiers, and phase-locked loops. It supports several technology nodes for academic and commercial applications and is compatible with commercial simulators such as Cadence Spectre, Synopsys HSPICE, and the open-source simulator Ngspice. AnalogGym standardizes the assessment of ML algorithms in analog circuit synthesis and promotes reproducibility with its open datasets and detailed benchmark specifications. AnalogGym's user-friendly design allows researchers to easily adapt it for robust, transparent comparisons of state-of-the-art methods, while also exposing them to real-world industrial design challenges, enhancing the practical relevance of their work. Additionally, we have conducted a comprehensive comparison study of various analog sizing methods on AnalogGym, highlighting the capabilities and advantages of different approaches. AnalogGym is available in the GitHub repository1. The documentations are also available at2. Jintao Li 0002, Haochang Zhi, Ruiyu Lyu, Wangzhen Li, Zhaori Bi, Keren Zhu 0001, Yanhan Zeng, Weiwei Shan, Changhao Yan, Fan Yang 0001, Yun Li 0002, Xuan Zeng 0001 |
ICCAD | 1 |
| 2024 | Robust circuit optimization under PVT variations via weight optimization problem reformulation
Jintao Li 0002, Yongfu Li 0002, Yanhan Zeng |
Expert Syst. Appl. | 1 |
| 2024 | Knowledge Transfer Framework for PVT Robustness in Analog Integrated CircuitsabstractProcess, voltage, and temperature (PVT) variations in chip fabrication or operation pose a significant challenge to the robustness of analog integrated circuits. Existing design techniques for mitigating PVT variations involve analyzing offsets of DC operating points, but this approach often leads to compromises in circuit performance. To address this challenge, we developed a ‘PVT-Transfer’ framework to facilitate knowledge transfer with evolutionary design. Specifically, by cross-operating the circuit parameters under variations, design knowledge is transferred through parameter migration, thus enhancing the robustness of the resultant circuit. In addition, we leverage data-driven learning to discover potential similarities among PVT variations, thereby mitigating negative knowledge transfer. The PVT-Transfer Framework is evaluated on three integrated voltage references and compared with four state-of-the-art circuit sizing methods. Based on post-layout Monte-Carlo simulations, this framework is verified to offer superior performance to existing methods, yielding a 60% reduction in power consumption, an 80% increase in temperature resilience, and up to 70$\times$enhancement in the figure of merit. Further, it leads to a 60% reduction in the number of required circuit simulations and is suitable for parallel computation. Jintao Li 0002, Yanhan Zeng, Haochang Zhi, Jingci Yang, Weiwei Shan, Yongfu Li 0002, Yun Li 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2023 | Multi-Task Evolutionary to PVT Knowledge Transfer for Analog Integrated Circuit OptimizationabstractDesigning analog integrated circuits (ICs), particularly sensors and reference circuits, requires a significant amount of human expertise and time, largely due to the requirement of maintaining process, voltage, and temperature (PVT) consistency. So far, there has been plenty of work on tuning the circuit to meet the PVT consistency requirements by comparing the offset of the DC operating point, but this inevitably leads to circuit performance degradation. To improve, we propose a ‘PVT-Transfer’ framework that utilizes knowledge transfer among PVT corners through evolutionary multitasking. Specifically, via cross-operating the circuit parameters under different PVT corners, knowledge is transferred through parameter migration to improve the robustness of the circuit. Further, PVT-Transfer employs data-driven learning to identify potential similarities among PVT variations, thereby leading to more cost-effective optimization. This framework is evaluated on two voltage references and compared with four state-of-the-art circuit sizing methods. The post-layout Monte-Carlo simulation results verify that PVT-Transfer outperforms the existing methods. It reduces the number of simulations required by 60% compared to the GCN-RL method. Besides, PVT-transfer achieves up to 10× improvement in the figure of merit over the human design. Jintao Li 0002, Haochang Zhi, Weiwei Shan, Yongfu Li 0002, Yanhan Zeng, Yun Li 0002 |
ICCAD | 1 |