EDBT 2026 Demo / reviewers in the wild / expert
Gengsheng Chen
dblp:94/4080
· DBLP profile ↗
15ranked-venue papers
0as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoEA: A Mixture of Experts Accelerator with Direct Token Access and Dynamic Expert SchedulingabstractTransformer-based large language models (LLMs) have achieved widespread adoption, but their growing model sizes impose substantial computational costs. Mixture-of-Experts (MoE) mechanism alleviates this by sparsely activating a subset of experts for each token. However, it still faces two critical challenges: (1) token rearrangement incurs non-trivial overhead of off-chip data movement, and (2) static execution flow fails to exploit inter-expert token reuse. In this paper, we present MoEA, a specialized accelerator designed to address the inefficiencies in MoE inference. MoEA introduces two key innovations: a Direct Token Access mechanism that leverages a hardware-managed metadata queue to eliminate off-chip token rearrangement, and a Dynamic Expert Scheduler that captures inter-expert token reuse patterns and optimizes expert execution order to maximize token reuse. Evaluations on representative MoE benchmarks show that MoEA reduces off-chip memory access by 12.06% and achieves $259.24 \times, 9.63 \times$ and $1.16 \times$ speedups over CPU, GPU and EdgeMoE accelerator, respectively. Zifeng Zhao, Jiewen Zheng, Tianxing Xie, Xinghao Zhu, Gengsheng Chen |
ASP-DAC | 5 |
| 2026 | FlexEdge: A Hardware-Software Co-Design for Flexible Edge Transformer Inference
Jiewen Zheng, Zifeng Zhao, Gengsheng Chen, Wenbo Yin |
ISCAS | 4 |
| 2026 | SCASN: Sparse Cross-Attention Self-supervised Network for Endoscopic Surgical Video Desmoking
Wanyi Zhou, Yinna Zhu, Gengsheng Chen, Meng Shi |
ISCAS | 4 |
| 2025 | RUNoC: Re-inject into the Underground Network to Alleviate Congestion in Large-Scale NoCabstractIn modem high-performance systems, the demand for large-scale Networks-on-Chip (NoC) has grown rapidly. As NoCs are scaled up, the issue of network congestion becomes increasingly critical and complex to address. Existing researches utilize partition-based NoC and Two-Level Network (TLN) techniques to alleviate congestion in large-scale NoCs. However, most of these solutions have limitations concerning universality, hardware overhead and workload balancing due to their inherent complexity and design constraints. In this article, we propose RUNoC, a new partition-based TLN architecture consisting of a Main Network for normal transmission and a sparse Underground Network enabling fast transmission. A special hardware unit is designed and integrated into each Main Network Router to decide when a packet needs to be routed to Underground Network based on network congestion information and the distance to the packet's destination, thereby alleviating congestion in Main Network and ensuring a subtle load balance between the two networks. Additionally, RUNoC is further enhanced by Shared Row Buffers (SRBs) and specialized network interface to guarantee deadlock and livelock freedom. Evaluation results indicate that RUNoC improves up to 60% in performance compared to XY routing scheme and up to 34% in performance-area ratio compared to existing researches. Xinghao Zhu, Jiyuan Bai, Zifeng Zhao, Qirong Yu, Gengsheng Chen |
ASP-DAC | 6 |
| 2024 | PAIR: Periodically Alternate the Identity of Routers to Ensure Deadlock Freedom in NoCabstractNetwork-on-Chip (NoC) has become widely adopted in multi/many-core systems for on-chip communication. Avoiding deadlock is a critical issue in NoC design. Recently proposed techniques partition the network resources into one or more express paths. Each path allows a blocked packet to move forward to its destination, thereby breaking deadlocks. However, as the scaling up of NoC, the existing methods are experiencing a decline in efficiency due to the coarse-grained partitioning strategy. In this work, we present PAIR, a novel scheme that guarantees deadlock freedom. PAIR adopts a fine-grained resource partitioning approach, significantly increasing the number of express paths available. The express paths in PAIR not only allow blocked packets but also permit non-blocked packets to be forwarded to the next hop, minimizing the performance degradation for the normal flow control. We implemented and evaluated PAIR on the mesh network using classic synthetic traffic patterns. Our experiments show that PAIR significantly improves throughput performance, achieving an average increase of 80.36% compared with state-of-the-art deadlock-free solutions. Furthermore, PAIR outperforms the latest deadlock-free scheme by 37.78% in terms of throughput. Zifeng Zhao, Xinghao Zhu, Jiyuan Bai, Gengsheng Chen |
ASPDAC | 4 |
| 2022 | Characterization of Single Event Upsets of Nanoscale FDSOI Circuits Based on the Simulation and Irradiation ResultsabstractThe advanced FDSOI technology has improved performance and inherent SEU resistance of integrated circuits, which is beneficial to the space applications. This paper provides the comprehensive characterization of SEU sensitivities based on the 3D-TCAD and SPICE simulations, as well as the irradiation results. We concentrate on the transient pulse, charge sharing, and collection effects of FDSOI circuits. The impact of strike location on transient features is evaluated in simulation, and the influence of charge sharing effects on SEU thresholds of SRAM is also analyzed. The SEU sensitive regions are characterized, which are closely related to the internal bipolar amplification effect and affected by the strike location. Additionally, the charge sharing effects are analyzed and verified by the combination of our circuit-level simulations and irradiation experiments based on the 256 Kbit pulse-mitigated SRAM. The split charge injection simulations show that the SEU threshold reduces about ~97% for the worst condition. Whereas, the actual SEU threshold of the pulse-mitigated FDSOI SRAM is not very small due to the limited charges shared by adjacent cells. The results provide a meaningful guidance for the radiation hardening design of FDSOI integrated circuits. Luchang Ding, Gengsheng Chen, Jun Yu 0010 |
ISCAS | 3 |
| 2022 | Real-Time Image Inpainting using PatchMatch Based Two-Generator Adversarial Networks with Optimized Edge Loss FunctionabstractImage inpainting algorithms based on the deep learning are rapidly developed in the past few years, as they can synthesize textures and structures while maintaining semantic continuity. However, due to the limitations of existing loss functions and convolution, the generated images are usually blurry and have artificial textures. In this paper, a real-time image inpainting system using PatchMatch based two-generator adversarial network (PatchMatch-GAN), is proposed to improve the clarity of generated images. The first generator is committed to continuous semantic textures, and the second one focuses on the image sharpness. Parallel-dilated convolution is used to enlarge the receptive field of filters. For the first generator, a pre-trained VGG16 is used as an encoder to extract features. For the second generator, a DCGAN-like network is designed to generate images. And a feature replacement method is utilized to improve the resolution. The edge loss has also been proposed as a part of the loss function to emphasize the role of the shape of objects. Experiments are carried out on place2 and irregular mask datasets. Compared with GLGAN and CAGAN, in PSNR, the improvements of our system are 3.33% and 2.7%, respectively; in FID, the improvements of our system are 59.54% and 35.59%, respectively. Moreover, the average processing time of the whole system is 29.96 ms, which is 6.56 times faster than GLGAN, and 21.33 times faster than CAGAN. These show that our system can achieve better image inpainting results while meeting the requirements of real-time processing speed. Luchang Ding, Gengsheng Chen |
ISCAS | 5 |
| 2022 | SEU sensitivity and large spacing TMR efficiency of Kintex-7 and Virtex-7 FPGAs
Bingxu Ning, Lingyun Ke, Gengsheng Chen, Ze He, Liewei Xu, Jie Liu 0032 |
Sci. China Inf. Sci. | 6 |
| 2020 | HybridDNN: A Framework for High-Performance Hybrid DNN Accelerator Design and ImplementationabstractTo speedup Deep Neural Networks (DNN) accelerator design and enable effective implementation, we propose HybridDNN, a framework for building high-performance hybrid DNN accelerators and delivering FPGA-based hardware implementations. Novel techniques include a highly flexible and scalable architecture with a hybrid Spatial/Winograd convolution (CONV) Processing Engine (PE), a comprehensive design space exploration tool, and a complete design flow to fully support accelerator design and implementation. Experimental results show that the accelerators generated by HybridDNN can deliver 3375.7 and 83.3 GOPS on a high-end FPGA (VU9P) and an embedded FPGA (PYNQ-Z1), respectively, which achieve a 1.8x higher performance improvement compared to the state-of-art accelerator designs. This demonstrates that HybridDNN is flexible and scalable and can target both cloud and embedded hardware platforms with vastly different resource constraints. Hanchen Ye, Xiaofan Zhang 0001, Zhize Huang, Gengsheng Chen, Deming Chen |
DAC | 4 |
| 2010 | Wideband reduced modeling of interconnect circuits by adaptive complex-valued sampling methodabstractIn this paper, we propose a new wideband model order reduction method for interconnect circuits by using a novel adaptive sampling and error estimation scheme. We try to address the outstanding error control problems in the existing sampling-based reduction framework. In the new method, called WBMOR, we explicitly compute the exact residual errors to guide the sampling process. We show that by sampling along the imaginary axis and performing a new complex-valued reduction, the reduced model will match exactly with the original model at the sample points. We show theoretically that the proposed method can achieve the error bound over a given frequency range. Practically the new algorithm can help designers choose the best order of the reduced model for the given frequency range and error bound via adaptive sampling scheme. As a result, it can perform wideband accurate reductions of interconnect circuits for analog and RF applications. We compare several sampling schemes such as linear, logarithmic, and recently proposed re-sampling methods. Experimental results on a number of RLC circuits show that WBMOR is much more accurate than all the other simple sampling methods and the recently proposed re-sampling scheme with the same reduction orders. Compared with the real-valued sampling methods, the complex-valued sampling method is more accurate for the same computational costs. Hai Wang 0002, Sheldon X.-D. Tan, Gengsheng Chen |
ASP-DAC | 3 |
| 2010 | Efficient model reduction of interconnects via double gramians approximationabstractThe gramian approximation methods have been proposed recently to overcome the high computing costs of classical balanced truncation based reduction methods. But those methods typically gain efficiency by projecting the original system only onto one dominant subspace of the approximate system gramian (for instance using only controllability gramian). This single gramian reduction method can lead to large errors as the subspaces of controllability and observability can be quite different for general interconnects with unsymmetric system matrices. In this paper, we propose a fast balanced truncation method where the system is balanced in terms of two approximate gramians as achieved in the classical balanced truncation method. The novelty of the new method is that we can keep the similar computing costs of the single gramian method. The proposed algorithm is based on a generalized SVD-based balancing scheme such that the dominant subspace of the approximate gramian product can be obtained in a very efficient way without explicitly forming the gramians. Experimental results on a number of published benchmarks show that the proposed method is much more accurate than the single gramian method with similar computing costs. Boyuan Yan, Sheldon X.-D. Tan, Gengsheng Chen, Yici Cai |
ASP-DAC | 3 |
| 2010 | Variational Capacitance Extraction and Modeling Based on Orthogonal Polynomial MethodabstractIn this paper, we propose a novel statistical capacitance extraction method for interconnect conductors considering process variations. The new method is called statCap, where orthogonal polynomials are used to represent the statistical processes in a deterministic way. We first show how the variational potential coefficient matrix is represented in a first-order form using Taylor expansion and orthogonal decomposition. Then, an augmented potential coefficient matrix, which consists of the coefficients of the polynomials, is derived. After this, corresponding augmented system is solved to obtain the variational capacitance values in the orthogonal polynomial form. Finally, we present a method to extend statCap to the second-order form to give more accurate results without loss of efficiency compared to the linear models. We show the derivation of the analytic second-order orthogonal polynomials for the variational capacitance integral equations. Experimental results show that statCap is two orders of magnitude faster than the recently proposed statistical capacitance extraction method based on the spectral stochastic collocation approach and many orders of magnitude faster than the Monte Carlo method for several practical conductor structures. Ruijing Shen, Sheldon X.-D. Tan, Wenjian Yu, Yici Cai, Gengsheng Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2009 | Statistical analysis of on-chip power grid networks by variational extended truncated balanced realization methodabstractIn this paper, we present a novel statistical analysis approach for large power grid network analysis under process variations. The new algorithm is very efficient and scalable for huge networks with a large number of variational variables. This approach, called varETBR for variational extended truncated balanced realization, is based on model order reduction techniques to reduce the circuit matrices before the variational simulation. It performs the parameterized reduction on the original system using variation-bearing subspaces. varETBR calculates variational response Gramians by Monte-Carlo based numerical integration considering both system and input source variations for generating the projection subspace. varETBR is very scalable for the number of variables and is flexible for different variational distributions and ranges as demonstrated in experimental results. After the reduction, Monte-Carlo based statistical simulation is performed on the reduced system and the statistical responses of the original system are obtained thereafter. Experimental results, on a number of IBM benchmark circuits [15] up to 1.6 million nodes, show that the varETBR can be 4500X faster than the Monte-Carlo method and is much more scalable than one of the recently proposed approaches. Sheldon X.-D. Tan, Gengsheng Chen, Xuan Zeng 0001 |
ASP-DAC | 3 |
| 2008 | Variational capacitance modeling using orthogonal polynomial methodabstractIn this paper, we propose a novel statistical capacitance extraction method for interconnects considering process variations. The new method, called statCap, is based on the spectral stochastic method where orthogonal polynomials are used to represent the statistical processes in a deterministic way. We first show how the variational potential coefficient matrix is represented in a first-order form using Taylor expansion and orthogonal decomposition. Then an augmented potential coefficient matrix, which consists of the coefficients of the polynomials, is derived. After that, corresponding augmented system is solved to obtain the variational capacitance values in the orthogonal polynomial form. Experimental results show that our method is two orders of magnitude faster than the recently proposed statistical capacitance extraction method based on the spectral stochastic collocation approach and many orders of magnitude faster than the Monte Carlo method for several practical interconnect structures. Gengsheng Chen, Ruijing Shen, Sheldon X.-D. Tan, Wenjian Yu, Jiarong Tong |
ACM Great Lakes Symposium on VLSI | 2 |
| 2008 | Modeling and simulation for on-chip power grid networks by locally dominant Krylov subspace methodabstractFast analysis of power grid networks has been a challenging problem for many years. The huge size renders circuit simulation inefficient and the large number of inputs further limits the application of existing Krylov-subspace macromodeling algorithms. However, strong locality has been observed that two nodes geometrically far have very small electrical impact on each other because of the exponential attenuation. However, no systematic approaches have been proposed to exploit such locality. In this paper, we propose a novel modeling and simulation scheme, which can automatically identify the dominant inputs for a given observed node in a power grid network. This enables us to build extremely compact models by projecting the system onto the locally dominant Krylov subspace corresponding to those dominant inputs only. The resulting simulation can be very fast with the compact models if we only need to view the responses of a few nodes under many different inputs. Experimental results show that the proposed method can have at least 100X speedup over SPICE-like simulations on a number of large power grid networks up to 1M nodes. Boyuan Yan, Sheldon X.-D. Tan, Gengsheng Chen, Lifeng Wu 0002 |
ICCAD | 3 |