EDBT 2026 Demo / reviewers in the wild / expert
Li Shang 0002
dblp:32/4492-2
· DBLP profile ↗
14ranked-venue papers
0as first author
14since 2021 · last 2026
0009-0005-9216-1921ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 8 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Atelier: An Automated Analog Circuit Design Framework via Multiple Large Language Model-Based AgentsabstractThis paper introduces Atelier, a large language model (LLM)-based framework for analog circuit design to address the issues of data scarcity and the substantial domain-specific knowledge required in this field. Atelier integrates general-purpose LLMs with a high-quality, compact knowledge base to fulfill the considerable knowledge requirements of analog circuit design, obviating the need for extensive domain-specific training or fine-tuning. The knowledge base is meticulously curated to be task-oriented and encapsulates critical information from pertinent literature within user-defined templates, leveraging the LLMs’ capabilities in text comprehension and summarization. The framework comprises several LLM agents, structured in a graph-of-thoughts architecture, with each agent specialized in a distinct task in analog circuit design, including circuit analysis, topology selection, topology modification, parameter tuning, and design decision. This collaborative multi-agent system, enriched with access to the compact knowledge base and advanced mechanisms such as self-reflection, backtracking, and tool integration, automates the analog circuit design process. It significantly enhances design quality and efficiency while ensuring interpretability. Experimental results highlight Atelier’s superiority over state-of-the-art black-box methods, general-purpose LLMs, and LLM-based methods, demonstrating notable improvements in success rates, design quality, and runtime. Jinyi Shen, Ji Zhuang, Jiangli Huang, Fan Yang 0001, Li Shang 0002, Zhaori Bi, Changhao Yan, Dian Zhou, Xuan Zeng 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2026 | Variation-aware Analog Circuit Design via Contextual Modeling and Robust OptimizationabstractRobust analog circuit design is becoming increasingly challenging due to process, voltage, and temperature (PVT) variations at advanced technology nodes. In this article, we formulate analog circuit synthesis as a robust optimization problem, and propose a Contextual Robust OptimiZAtion (CROZA) method for variation-aware analog circuit design. The proposed method uses Contextual Gaussian process to model both the design parameters and perturbation parameters, and a hybrid strategy of adversarially robust optimization and stochastically perturbed robust optimization to find robust solutions. Compared to state-of-the-art methods, our proposed approach achieves significant simulation and runtime speedups while delivering superior optimization results. Jiangli Huang, Jinyi Shen, Fan Yang 0001, Li Shang 0002, Zhaori Bi, Changhao Yan, Wenchuang Walter Hu, Dian Zhou, Xuan Zeng 0001 |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2025 | INTO-OA: Interpretable Topology Optimization for Operational AmplifiersabstractThis paper presents INTO-OA, an interpretable topology optimization method for operational amplifiers (op-amps). We propose a Bayesian optimization-based approach to effectively explore the high-dimensional, discrete topology design space of op-amps. Our method integrates a Gaussian process surrogate model with the Weisfeiler-Lehman graph kernel to extract structural features from a dedicated circuit graph representation. It also employs a candidate generation strategy that combines random sampling with mutation to balance global exploration and local exploitation. Additionally, INTO-OA enhances interpretability by assessing the impact of circuit structures on performance, providing designers with valuable insights into generated topologies and enabling the interpretable refinement of existing designs. Experimental results demonstrate that INTO-OA achieves higher success rates, a 1.84× to 19.10x improvement in op-amp performance, and a 3.20x to 14.33× increase in topology optimization efficiency compared to state-of-the-art methods. Jinyi Shen, Fan Yang 0001, Li Shang 0002, Zhaori Bi, Changhao Yan, Dian Zhou, Xuan Zeng 0001 |
DATE | 3 |
| 2025 | Oracle-MoE: Locality-preserving Routing in the Oracle Space for Memory-constrained Large Language Model InferenceabstractMixture-of-Experts (MoE) is widely adopted to deploy Large Language Models (LLMs) on edge devices with limited memory budgets. Although MoE is, in theory, an inborn memory-friendly architecture requiring only a few activated experts to reside in the memory for inference, current MoE architectures cannot effectively fulfill this advantage and will yield intolerable inference latencies of LLMs on memory-constrained devices. Our investigation pinpoints the essential cause as the remarkable temporal inconsistencies of inter-token expert activations, which generate overly frequent expert swapping demands dominating the latencies. To this end, we propose a novel MoE architecture, Oracle-MoE, to fulfill the real on-device potential of MoE-based LLMs. Oracle-MoE route tokens in a highly compact space suggested by attention scores, termed the oracle space, to effectively maintain the semantic locality across consecutive tokens to reduce expert activation variations, eliminating massive swapping demands. Theoretical analysis proves that Oracle-MoE is bound to provide routing decisions with better semantic locality and, therefore, better expert activation consistencies. Experiments on the pretrained GPT-2 architectures of different sizes (200M, 350M, 790M, and 2B) and downstream tasks demonstrate that without compromising task performance, our Oracle-MoE has achieved state-of-the-art inference speeds across varying memory budgets, revealing its substantial potential for LLM deployments in industry. Jixian Zhou, Ruijun Huang, Hengjie Cao, Mengyi Chen, Anrui Chen, Mingzhi Dong, Yujiang Wang 0001, Dongsheng Li 0002, David A. Clifton, Qin Lv, Rui Zhu 0006, Fan Yang 0001, Tun Lu, Ning Gu 0001, Li Shang 0002 |
ICML | 18 |
| 2025 | An Efficient Placement Speedup Technique Based on Graph Signal ProcessingabstractPlacement is a critical task with high computation complexity in VLSI physical design. Modern analytical placers formulate the placement objective as a nonlinear optimization task, which suffers a long iteration time. To accelerate and enhance the placement process, recent studies have turned to deep learning-based approaches, particularly leveraging graph convolution networks (GCNs). However, learning-based placers require time- and data-consuming model training due to the complexity of circuit placement that involves large-scale cells and design-specific graph statistics. This article proposes GiFt, a parameter-free initialization technique for accelerating placement, rooted in graph signal processing. GiFt excels at capturing multiresolution smooth signals of circuit graphs to generate optimized initial placement solutions without the need for time-consuming model training, and meanwhile significantly reduces the number of iterations required by analytical placers. Moreover, we present GiFtPlus, an enhanced version of GiFt, which is more efficient in handling large-scale circuit placement and can accommodate location constraints. Experimental results on public benchmarks show that GiFt and GiFtPlus significantly improve placement efficiency, while achieving competitive or superior performance compared to state-of-the-art placers. In particular, the recently proposed GPU-accelerated analytical placer DREAMPlace uses up to 50% more total runtime than GiFtPlus-DREAMPlace. Yiting Liu 0002, Hai Zhou 0001, Jia Wang 0003, Fan Yang 0001, Xuan Zeng 0001, Li Shang 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2025 | ATOM: An Automatic Topology Synthesis Framework for Operational AmplifiersabstractBayesian optimization (BO) is more efficient in automatically synthesizing operational amplifier (opamp) topologies compared to conventional methods. However, the design space for behavior-level opamp topologies involves numerous connections that are difficult to comprehend, and evaluating each topology incurs substantial computational costs. To tackle these challenges, this brief introduces ATOM, an automatic opamp topology synthesis framework. We construct a concise design space for behavior-level opamp topologies, consisting of topologies that designers can easily understand. We propose an opamp topology optimization method that incorporates freeze-thaw BO. This method efficiently explores the design space and expedites the evaluation process. Experimental studies demonstrate that ATOM outperforms state-of-the-art topology synthesis methods in terms of success rate and optimization results while reducing the number of required simulations by up to 8.15 times. The source code for ATOM is available athttps://github.com/Jinyi-Shen/ATOM. Jinyi Shen, Fan Yang 0001, Li Shang 0002, Changhao Yan, Zhaori Bi, Dian Zhou, Xuan Zeng 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | Adaptive ILT via Multi-Level Lithography Simulation
Shuyuan Sun, Fan Yang 0001, Bei Yu 0001, Li Shang 0002, Xuan Zeng 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | Denoising Reuse: Exploiting Inter-Frame Motion Consistency for Efficient Video GenerationabstractDenoising-based diffusion models have attained impressive image synthesis; however, their applications on videos can lead to unaffordable computational costs due to the per-frame denoising operations. In pursuit of efficient video generation, we present a Diffusion Reuse MOtion (Dr. Mo) network to accelerate the video-based denoising process. Our crucial observation is that the latent representations in early denoising steps between adjacent video frames exhibit high consistencies with motion clues. Inspired by the discovery, we propose to accelerate the video denoising process by incorporating lightweight, learnable motion features. Specifically, Dr. Mo will only compute all denoising steps for base frames. For a non-based frame, Dr. Mo will propagate the pre-computed based latents of a particular step with inter-frame motions to obtain a fast estimation of its coarse-grained latent representation, from which the denoising will continue to obtain more sensitive and fine-grained representations. On top of this, Dr. Mo employs a meta-network named Denoising Step Selector (DSS) to dynamically determine the step to perform motion-based propagations for each frame, ensuring the correct transformation of multi-granularity visual features. Extensive evaluations on video generation and editing tasks indicate that Dr. Mo delivers widely applicable acceleration for diffusion-based video generations while effectively retaining the visual quality and style. Video generation and visualization results can be found athttps://drmo-denoising-reuse.github.io. Yixuan Chen 0003, Yujiang Wang 0001, Mingzhi Dong, Dongsheng Li 0002, Rui Zhu 0006, David A. Clifton, Robert P. Dick, Qin Lv, Fan Yang 0001, Tun Lu, Ning Gu 0001, Li Shang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 16 |
| 2025 | Introduction to Special Issue on Large Language Models for Electronic System Design AutomationabstractLarge Language Models are having a substantial impact on electronic design automation in areas ranging from hardware architecture to verification and optimization. The special issue provides a snapshot of work on this topic. This introduction describes and provides context for the research area, describes the organization of the special issue, and provides terse summaries of each of its papers. Robert P. Dick, Hammond A. Pearce, Li Shang 0002, Fan Yang 0001 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2024 | Once Read is Enough: Domain-specific Pretraining-free Language Models with Cluster-guided Sparse Experts for Long-tail Domain KnowledgeabstractLanguage models (LMs) only pretrained on a general and massive corpus usually cannot attain satisfying performance on domain-specific downstream tasks, and hence, applying domain-specific pretraining to LMs is a common and indispensable practice.
However, domain-specific pretraining can be costly and time-consuming, hindering LMs' deployment in real-world applications.
In this work, we consider the incapability to memorize domain-specific knowledge embedded in the general corpus with rare occurrences and long-tail distributions as the leading cause for pretrained LMs' inferior downstream performance.
Analysis of Neural Tangent Kernels (NTKs) reveals that those long-tail data are commonly overlooked in the model's gradient updates and, consequently, are not effectively memorized, leading to poor domain-specific downstream performance.
Based on the intuition that data with similar semantic meaning are closer in the embedding space, we devise a Cluster-guided Sparse Expert (CSE) layer to actively learn long-tail domain knowledge typically neglected in previous pretrained LMs.
During pretraining, a CSE layer efficiently clusters domain knowledge together and assigns long-tail knowledge to designate extra experts. CSE is also a lightweight structure that only needs to be incorporated in several deep layers.
With our training strategy, we found that during pretraining, data of long-tail knowledge gradually formulate isolated, outlier clusters in an LM's representation spaces, especially in deeper layers. Our experimental results show that only pretraining CSE-based LMs is enough to achieve superior performance than regularly pretrained-finetuned LMs on various downstream tasks, implying the prospects of domain-specific-pretraining-free language models. Mengyi Chen, Jixian Zhou, Yubin Shi, Yixuan Chen 0003, Mingzhi Dong, Yujiang Wang 0001, Dongsheng Li 0002, Rui Zhu 0006, Robert P. Dick, Qin Lv, Fan Yang 0001, Tun Lu, Ning Gu 0001, Li Shang 0002 |
NeurIPS | 16 |
| 2023 | cVTS: A Constrained Voronoi Tree Search Method for High Dimensional Analog Circuit SynthesisabstractA constrained Voronoi tree-based domain decomposition method for high-dimensional Bayesian optimization is proposed to solve large scale analog circuit synthesis problems, which can be formulated as high-dimensional heterogeneous black-box optimization. Hierarchical Voronoi tree progressively breaks down the design space into partitions with implicit performance boundaries such that promising regions are efficiently explored. Fast exploitation is ensured in Voronoi nest via local Bayesian optimization with a few observations. A slice-enhanced Gibbs sampling method is proposed to sample acquisition function cMES in irregular polyhedrons with design constraints. Compared with state-of-the-art methods, cVTS achieves significant speed up without loss of accuracy. Aidong Zhao, Xianan Wang, Zixiao Lin, Zhaori Bi, Changhao Yan, Fan Yang 0001, Li Shang 0002, Dian Zhou, Xuan Zeng 0001 |
DAC | 8 |
| 2023 | Over-parameterized Model Optimization with Polyak-Łojasiewicz Condition
Yixuan Chen 0003, Yubin Shi, Mingzhi Dong, Dongsheng Li 0002, Yujiang Wang 0001, Robert P. Dick, Qin Lv, Fan Yang 0001, Ning Gu 0001, Li Shang 0002 |
ICLR | 12 |
| 2023 | Train Faster, Perform Better: Modular Adaptive Training in Over-Parameterized ModelsabstractDespite their prevalence in deep-learning communities, over-parameterized models convey high demands of computational costs for proper training. This work studies the fine-grained, modular-level learning dynamics of over-parameterized models to attain a more efficient and fruitful training strategy. Empirical evidence reveals that when scaling down into network modules, such as heads in self-attention models, we can observe varying learning patterns implicitly associated with each module's trainability. To describe such modular-level learning capabilities, we introduce a novel concept dubbed modular neural tangent kernel (mNTK), and we demonstrate that the quality of a module's learning is tightly associated with its mNTK's principal eigenvalue $\lambda_{\max}$. A large $\lambda_{\max}$ indicates that the module learns features with better convergence, while those miniature ones may impact generalization negatively. Inspired by the discovery, we propose a novel training strategy termed Modular Adaptive Training (MAT) to update those modules with their $\lambda_{\max}$ exceeding a dynamic threshold selectively, concentrating the model on learning common features and ignoring those inconsistent ones. Unlike most existing training schemes with a complete BP cycle across all network modules, MAT can significantly save computations by its partially-updating strategy and can further improve performance. Experiments show that MAT nearly halves the computational cost of model training and outperforms the accuracy of baselines. Yubin Shi, Yixuan Chen 0003, Mingzhi Dong, Dongsheng Li 0002, Yujiang Wang 0001, Robert P. Dick, Qin Lv, Fan Yang 0001, Tun Lu, Ning Gu 0001, Li Shang 0002 |
NeurIPS | 13 |
| 2022 | Recursive Disentanglement Network
Yixuan Chen 0003, Yubin Shi, Dongsheng Li 0002, Yujiang Wang 0001, Mingzhi Dong, Robert P. Dick, Qin Lv, Fan Yang 0001, Li Shang 0002 |
ICLR | 10 |