Ting-Jung Lin

dblp:06/7249 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0002-5208-482XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Mixture-of-Trees: Learning to Select and Weigh Reasoning Paths for Efficient LLM Inference
abstract
We introduce Mixture-of-Trees (MoT), a novel framework that integrates sparse expert activation with structured tree-based reasoning for efficient LLM inference. MoT employs a learned gating mechanism to selectively activate only the most relevant expert reasoning trees for each problem, where experts use models of varying capacities based on task complexity. The framework features three key innovations: (1) sparse expert activation through unified gating networks, (2) specialized expert trees that leverage domain-specific expertise while optimizing the quality-efficiency trade-off, and (3) collaborative debate mechanisms for conflicting solutions. Additionally, MoT includes a shared baseline tree with early stopping—activated experts perform lightweight validation and terminate early when confidence is high. Experiments across five benchmarks (GSM8K, MATH, AIME 2024, MMLU, HotpotQA) show that MoT achieves 2-7 percentage point accuracy improvements while reducing LLM calls by 37-40% compared to existing multi-path methods.
Yangbo Wei, Zhen Huang 0007, Shaoqiang Lu, Junhong Qian, Dongge Qin, Ting-Jung Lin, Wei W. Xing, Lei He 0001
AAAI6
2026 VFlow: Discovering Optimal Agentic Workflows for Verilog Generation
Yangbo Wei, Zhen Huang 0007, Lei He 0001, Ting-Jung Lin, Wei W. Xing
ASP-DAC5
2025 Self-Attention to Operator Learning-based 3D-IC Thermal Simulation
abstract
Thermal management in 3D ICs is increasingly challenging due to higher power densities. Traditional PDESolving based methods, while accurate, are too slow for iterative design. Machine learning approaches like FNO provide faster alternatives but suffer from high-frequency information loss and high-fidelity data dependency. We introduce Self-Attention UNet Fourier Neural Operator (SAU-FNO), a novel framework combining self-attention and U-Net with FNO to capture longrange dependencies and model local high-frequency features effectively. Transfer learning is employed to fine-tune low-fidelity data, minimizing the need for extensive high-fidelity datasets and speeding up training. Experiments demonstrate that SAUFNO achieves state-of-the-art thermal prediction accuracy and provides an $842 \times$ speedup over traditional FEM methods, making it an efficient tool for advanced 3D IC thermal simulations.
Zhen Huang 0007, Wenkai Yang, Muxi Tang, Depeng Xie, Ting-Jung Lin, Yu Zhang 0086, Wei W. Xing, Lei He 0001
DAC6
2025 MambaOPU: An FPGA Overlay Processor for State-space-duality-based Mamba Models
abstract
State-space models (SSMs), such as Mamba, have emerged as a promising alternative to Transformers. However, the recently developed Mamba2, based on state space duality (SSD), is highly memorybound and suffers from limited computation efficiency. This inefficiency arises from its irregular broadcast element-wise multiplications and structured sparse computations. In this work, we propose MambaOPU, an FPGA overlay processor, to accelerate SSD. First, to reduce memory overhead, we introduce a software-hardware co-optimized operator fusion framework. Specifically, operator merging combines adjacent broadcast multiplication and summation operations into a single descriptor, while operator backward shifting embeds segment multiplication into subsequent operations. Both techniques shorten the computation path and improve computation efficiency. Second, to enhance sparse computation efficiency, we skip zero-region computations using a tensor-reorder-and-group algorithm combined with a sparse-predefined data fetcher. Additionally, since Mamba integrates linear operations with SSD, we develop a reconfigurable systolic array to improve data reuse across different computation modes. Extensive experiment results demonstrate that MambaOPU achieves up to $1812 \times$ and $880.79 \times$ higher normalized throughput and up to $12908 \times$ and $24.27 \times$ higher energy efficiency over Intel Xeon Gold 6348 CPU and NVIDIA A100 GPU, respectively.
Shaoqiang Lu, Xuliang Yu, Tiandong Zhao, Siyuan Miao, Xinsong Sheng, Ting-Jung Lin, Lei He 0001
DAC8
2025 SetupKit: Efficient Multi-Corner Setup/Hold Time Characterization Using Bias-Enhanced Interpolation and Active Learning
abstract
Accurate setup/hold time characterization is crucial for modern chip timing closure, but its reliance on potentially millions of SPICE simulations across diverse process-voltage-temperature (PVT) corners creates a major bottleneck, often lasting weeks or months. Existing methods suffer from slow search convergence and inefficient exploration, especially in the multi-corner setting. We introduce SetupKit, a novel framework designed to break this bottleneck using statistical intelligence, circuit analysis and active learning (AL). SetupKit integrates three key innovations: BEIRA, a bias-enhanced interpolation search derived from statistical error modeling to accelerate convergence by overcoming stagnation issues, initial search interval estimation by circuit analysis and AL strategy using Gaussian Process. This AL component intelligently learns PVT-timing correlations, actively guiding the expensive simulations to the most informative corners, thus minimizing redundancy in multi-corner characterization. Evaluated on industrial 22nm standard cells across 16 PVT corners, SetupKit demonstrates a significant 2.4× overall CPU time reduction (from 720 to 290 days on a single core) compared to standard practices, drastically cutting characterization time. SetupKit offers a principled, learning-based approach to library characterization, addressing a critical EDA challenge and paving the way for more intelligent simulation management.
Junzhuo Zhou, Haoxuan Xia, Yuxin Yan, Chengyu Zhu, Ting-Jung Lin, Wei W. Xing, Lei He 0001
ICCAD6
2025 Abuttable Analog Cell Library and Automatic AMS Layout
abstract
The state of the art analog circuit design applies mainly a full-custom layout methodology. This demands high expertise and heavy manual workload. Additionally, neither can the resulting layout be re-used easily across different designs or different PDKs. Learning from digital standard cells, existing work has proposed stem cells that are abuttable. But stem cells have a fixed area ratio of 2 over same-sized Pcells, limiting its wide application. In this paper we develop a new type of abuttable analog cells (called Acells) for transistors and passive elements. Acells are compatible with digital standard cells and can be abutted in all directions, enabling the use of automatic digital place and route (PnR) engines. We automate Acell generation and show that the average area ratio over same-sized Pcell is 1.49 for 65nm technology and 1.3 for 28nm technology, and is expected to decrease for more advanced technologies. We then use digital PnR to automatically layout a number of analog and mixed-signal (AMS) circuits mainly in 28nm, and show that compared to Pcell-based manual layout, Acell-based layout obtains similar performance and its circuit level layout area is about 2% higher for large scale AMS circuits in our experiments.
Tianjia Zhou, Jingyun Gu, Zexin Ji, Hailang Liang, Zhanfei Chen, Ting-Jung Lin, Na Bai, Zhengping Li, Lei He 0001
ISPD9
2025 LVFGen: Efficient Liberty Variation Format (LVF) Generation Using Variational Analysis and Active Learning
abstract
As transistor dimensions shrink, process variations significantly impact circuit performance, signifying the need for accurate statistical circuit analysis. In digital circuit timing analysis, the Liberty Variation Format (LVF) has emerged as an industrial leading representation of timing distributions in cell libraries at 22 nm and below. However, LVF characterization relies on the Monte Carlo (MC) method, which requires excessive SPICE simulations of cells with process variations. Similar challenges also exist for uncertainty propagation and quantification in chip manufacturing and the broader scientific communities. To resolve this foundational challenge, this paper presents LVFGen, a novel method that reduces the simulation costs of MC while generate high-accuracy LVF library. LVFGen utilizes an active learning strategy based on variational analysis to identify process variation samples that impact timing distributions more significantly. Compared to the state-of-the-art Quasi-MC method, LVFGen demonstrates an overall 2.27× speedup in LVF library generation within an accuracy level of 5k-sample MC and a 4.06× speedup within a 100k-sample MC accuracy.
Junzhuo Zhou, Haoxuan Xia, Wei W. Xing, Ting-Jung Lin, Lei He 0001
ISPD4
2025 Symbol and Footprint Database for Electronic Components by Agentic Recognition and Generation
Zhuofu Tao, Yuhao Gao, Ting-Jung Lin, Lei He 0001
PRCV (7)6
2025 AMSnet-KG: A Netlist Dataset for LLM-based AMS Circuit Auto-design Using Knowledge Graph RAG
abstract
High-performance analog and mixed-signal (AMS) circuits are mainly full-custom designed, which is time-consuming and labor-intensive. A significant portion of the effort is experience-driven, which makes the automation of AMS circuit design a formidable challenge. Large language models (LLMs) have emerged as powerful tools for electronic design automation (EDA) applications, fostering advancements in the automatic design process for large-scale AMS circuits. However, the absence of high-quality datasets has led to issues such as model hallucination, which undermines the robustness of automatically generated circuit designs. To address this issue, this article introduces AMSnet-KG, a dataset encompassing various AMS circuit schematics and netlists. We construct a knowledge graph with annotations on detailed functional and performance characteristics. Facilitated by AMSnet-KG, we propose an automated AMS circuit generation framework that utilizes the comprehensive knowledge embedded in LLMs. The flow first formulate a design strategy (e.g., circuit architecture using a number of circuit components) based on required specifications. Next, matched subcircuits are retrieved and assembled into a complete topology, and transistor sizing is obtained through Bayesian optimization. Simulation results of the netlist are automatically fed back to the LLM for further topology refinement, ensuring the circuit design specifications are met. We perform case studies of operational amplifier and comparator design to verify the automatic design flow from specifications to netlists with minimal human effort. The dataset used in this article is available at https://ams-net.github.io/ .
Zhuofu Tao, Yuhao Gao, Tianjia Zhou, Bingyu Chen 0007, Genhao Zhang, Alvin Liu, Zhiping Yu, Ting-Jung Lin, Lei He 0001
ACM Trans. Design Autom. Electr. Syst.11
2025 ModelGen: Automating Semiconductor Parameter Extraction with Large Language Model Agents
abstract
Device models require large numbers of parameters to characterize complex physical effects. Although the latest advancements in machine learning and automated tools have drastically improved efficiency over the classic methods, they still demand a considerable amount of human intervention in the loop to gain accuracy. This drastically limits further automation. Inspired by the success of Multimodal Large Language Models (MLLMs) in addressing tasks across diverse fields, we propose ModelGen, the first in-depth study to leverage MLLMs with RAG (Retrieval-Augmented Generation) to significantly reduce human effort in parameter extraction for compact model. Our contributions include (1) Automated Agentic Workflow Construction that learns to build and refine extraction workflows through iterative optimization, (2) MLLM Judge, a visual scoring mechanism that evaluates fitting quality using actual device characteristic plots rather than simple numerical metrics, and (3) Model-specific RAG for providing relevant domain knowledge during the extraction process. Experimental results demonstrate that ModelGen achieves a 26.8%–33.1% improvement in pass@1,3,5 compared to base LLM methods. The system completes complex model extractions for BSIMs and ASM-HEMT in hours (up to 168× faster) rather than days or weeks, making parameter extraction more accessible to non-experts while maintaining professional engineer-level accuracy.
Yangbo Wei, Zhanfei Chen, Jinlong Yan, Ting-Jung Lin, Zhen Huang 0007, Wei W. Xing, Lei He 0001
ACM Trans. Design Autom. Electr. Syst.6
2025 MCoreOPU: An FPGA-based Multi-Core Overlay Processor for Transformer-based Models
abstract
Transformer-based models have achieved extensive success with increasingly large numbers of parameters and computations, for which many multi-core accelerators have been developed. Nevertheless, they suffer from limited throughput due to either low operating frequency or high communication overhead between cores. This article proposes an FPGA-based multi-core overlay processor, named MCoreOPU, to optimize intra-core computation and inter-core communication. First, we boost the operating frequency of the processing element (PE) array to double the rest of the processor to improve the intra-core throughput. Second, we develop on-chip synchronization routers to reduce off-chip memory traffic, where only the partial sum and maximum are communicated between cores rather than entire vectors for layer normalization and softmax. Moreover, we pipeline synchronization to reduce synchronization latency and develop a bypass of the interconnect bus to reduce the off-chip memory access latency. Finally, we optimize the multi-core model allocation and scheduling to minimize the inter-core communications and maximize the intra-core computation efficiency. The MCoreOPU is implemented in 8-bit fixed-point precision with four cores and four DDRs on the Xilinx U200 FPGA, where the PE array runs at 600 MHz while the rest runs at 300 MHz. Experimental results show that the throughput per MAC of MCoreOPU for BERT, ViT, GPT-2, and LLaMA inference is 1.31 \(\times\) –7.18 \(\times\) higher than other FPGA-based accelerators. Compared with the A100 GPU, the throughput per equivalent MAC efficiency is improved by 22.52 \(\times\) –27.12 \(\times\) .
Shaoqiang Lu, Tiandong Zhao, Ting-Jung Lin, Rumin Zhang, Lei He 0001
ACM Trans. Reconfigurable Technol. Syst.3
2024 LVF2: A Statistical Timing Model based on Gaussian Mixture for Yield Estimation and Speed Binning
abstract
As transistor size continues to scale down, process variation has become an essential factor determining semiconductor yield and economic return. The Liberty Variation Format (LVF) is the current industrial standard that expresses statistical timing behaviors based on single Gaussian model. However, it loses accuracy when the timing distribution is non-Gaussian due to growing process variations. This paper proposes a novel LVF2 distribution model that combines two weighted skewed-normal (SN) distributions, which better captures the multi-Gaussian timing distribution while maintaining backward compatibility with LVF. Experiments using TSMC 22nm standard cells show that, compared to LVF, LVF2 reduces binning error by 7.74X in delay and 9.56X in transition time, and reduces 3σ-yield error by 4.79X and 7.18X in delay and transition time, respectively. The error reduction for path delay is diminished due to Central Limit Theorem (CLT). But it is still 2X for a typical circuit path with 8 Fanout-of-4 (FO4) inverter delays.
Junzhuo Zhou, Haoxuan Xia, Leilei Jin, Xiao Shi 0001, Wei W. Xing, Ting-Jung Lin, Lei He 0001
DAC8
2015 FDR 2.0: A Low-Power Dynamically Reconfigurable Architecture and Its FinFET Implementation
abstract
Large area/delay/power overheads are required to support the reconfigurability of field-programmable gate arrays (FPGAs). We proposed a hybrid CMOS/nanotechnology dynamically reconfigurable architecture, called NATURE, earlier to address this challenge. It uses the concept of temporal logic folding and fine-grain (i.e., cycle-level) dynamic reconfiguration to increase logic density and save area. Because logic folding reduces area significantly, most of the on-chip communications become localized. To take full advantage of localized communications, we then presented a new CMOS-based fine-grain dynamically reconfigurable (FDR) architecture. It consists of an array of homogeneous logic elements (LEs), which can be configured into logic or interconnect or a combination of both. FDR eliminates most of the long-distance and global wires, which occupy a large amount of area in conventional FPGAs. FDR improves the area-delay product by an order of magnitude relative to conventional architectures. In this paper, we present an augmented FDR 2.0 architecture, where: 1) the LE is augmented with dedicated carry logic to facilitate arithmetic operations; 2) diagonal direct links are incorporated to improve the flexibility of local communication; and 3) coarse-grain blocks, including embedded memories and digital signal processing (DSP) blocks, are added to support fast data-intensive computations. Experimental results show that the coarse-grain design can improve circuit performance by 3.6× compared with the fine-grain FDR architecture. Incorporation of the DSP blocks in FDR 2.0 also enables more effective area-delay and power-delay tradeoffs, allowing the users to trade performance for smaller area or power consumption. We have implemented the design in the 22-nm FinFET technology, which enables more flexible and effective power management. Finally, different types of FinFETs and power management techniques have been explored in FDR 2.0 to optimize power.
Ting-Jung Lin, Wei Zhang 0012, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.1
2014 A Fine-Grain Dynamically Reconfigurable Architecture Aimed at Reducing the FPGA-ASIC Gaps
abstract
Prior work has shown that due to the overhead incurred in enabling reconfigurability, field-programmable gate arrays (FPGAs) require 21× more silicon area, 3× larger delay, and 10× more dynamic power consumption compared with application-specific integrated circuits (ASICs). We have earlier presented a hybrid CMOS/nanotechnology reconfigurable architecture (NATURE). It uses the concept of temporal logic folding and fine-grain (i.e., cycle-level) dynamic reconfiguration to increase logic density by an order of magnitude. Since logic folding reduces area usage significantly, on-chip communications tend to become localized. To take full advantage of this fact, we propose a new architecture, called fine-grain dynamically reconfigurable (FDR), that consists of an array of homogeneous reconfigurable logic elements (LEs). Each LE can be arbitrarily configured into a lookup table (LUT) or interconnect or a combination of both. This significantly enhances the flexibility of allocating hardware resources between LUTs and interconnects based on application needs. The proposed FDR architecture eliminates most of the long-distance and global wires, which occupy most of the area in conventional FPGAs. Fine-grain dynamic reconfiguration is enabled by local embedded static RAM blocks. The experiments show that, on an average, area, delay, and power are improved by 9.14×, 1.11×, and 1.45×, compared with a conventional FPGA architecture that does not use the concept of logic folding. Compared with NATURE with deep logic folding, area, delay, and power are improved by 2.12×, 3.28×, and 1.74×, respectively. Although this does not eliminate the FPGA-ASIC area/delay/power gaps, it makes progress toward bridging these gaps.
Ting-Jung Lin, Wei Zhang 0012, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.1
2012 SRAM-Based NATURE: A Dynamically Reconfigurable FPGA Based on 10T Low-Power SRAMs
abstract
We presented a hybrid CMOS/nanotechnology reconfigurable architecture (NATURE), earlier. It was based on CMOS logic and nano RAMs. It used the concept of temporal logic folding and fine-grain (e.g., cycle-level) dynamic reconfiguration to increase logic density by an order of magnitude. This dynamic reconfiguration is done intra-circuit rather than inter-circuit. However, the previous design of NATURE required fine-grained distribution of nano RAMs throughout the field-programmable gate array (FPGA) architecture. Since the fabrication process of nano RAMs is not mature yet, this prevents immediate exploitation of NATURE. In this paper, we present a NATURE architecture that is based on CMOS logic and CMOS SRAMs that are used for on-chip dynamic reconfiguration. We use fast and low-power SRAM blocks that are based on 10T SRAM cells. We have also laid out the various FPGA components in a 65-nm technology to evaluate the FPGA performance. We hide the dynamic reconfiguration delay behind the computation delay through the use of shadow SRAM cells. Experimental results show more than an order of magnitude improvement in logic density and improvement in the area-delay product relative to a traditional baseline FPGA architecture that does not use the concept of logic folding.
Ting-Jung Lin, Wei Zhang 0012, Niraj K. Jha
IEEE Trans. Very Large Scale Integr. Syst.1
2008 Modeling Orbit Dynamics of FORMOSAT-3/COSMIC Satellites for Recovery of Temporal Gravity Variations
abstract
The precise GPS high-low tracking data from the joint Taiwan-USA mission FORMOSAT-3/COSMIC (COSMIC) can be used for gravity recovery. The current orbital accuracy of COSMIC kinematic orbit is 2 cm and is better than 1 cm for 60-s normal points. We model the perturbing forces acting on the COSMIC spacecraft based on standard models of orbit dynamics. The major tool for the numerical work of force modeling is NASA Goddard's GEODYN II software. Considering that COSMIC spacecraft are not equipped with accelerometers, the accelerations due to atmospheric drag, solar radiation pressure, and other minor surface forces are modeled by estimating relevant parameters over one orbital period from COSMIC's kinematic and reduced dynamic orbits. We carry out experimental solutions of time-varying geopotential coefficients using one month of COSMIC kinematic orbits (August 2006). With the nongravity origin forces properly modeled by GEODYN II, residual orbital perturbations (difference between kinematic and reference orbits) are assumed to be linear functions of time-varying geopotential coefficients and are used as observations to estimate the latter. Both COSMIC and combined COSMIC and GRACE gravity solutions are computed. The COSMIC solution shows some well-known temporal gravity signatures but contains artifacts. The combined COSMIC and GRACE solution enhances some local temporal gravity signatures in the GRACE solution.
Cheinway Hwang, Ting-Jung Lin, Tzu-Pang Tseng, Benjamin Fong Chao
IEEE Trans. Geosci. Remote. Sens.2