Li-De Chen

dblp:54/9911 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
3since 2021 · last 2025
0009-0006-4206-1022ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 2 since 2021Computer networks · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Computational photography and imaging · 50% Rendering · 25% Image and video processing · 25%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 73% Electronic design automation · 27%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computational photography and imaging
light field display
0.912025
Temporal Fusion: Continuous-Time Light Field Video Factorization · IEEE Trans. Image Process. 2025
Computational photography and imaging › light field imaging
light field factorization
0.912025
Temporal Fusion: Continuous-Time Light Field Video Factorization · IEEE Trans. Image Process. 2025
Rendering › image-based rendering
light field rendering
0.912025
Temporal Fusion: Continuous-Time Light Field Video Factorization · IEEE Trans. Image Process. 2025
Image and video processing › image fusion
temporal fusion
0.912025
Temporal Fusion: Continuous-Time Light Field Video Factorization · IEEE Trans. Image Process. 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator
CNN accelerator
0.412019
eCNN: A Block-Based and Highly-Parallel CNN Accelerator for Edge Inference · MICRO 2019
Hardware accelerators and domain-specific architectures › machine learning accelerator › DNN inference accelerator
edge inference accelerator
0.412019
eCNN: A Block-Based and Highly-Parallel CNN Accelerator for Edge Inference · MICRO 2019
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.412019
eCNN: A Block-Based and Highly-Parallel CNN Accelerator for Edge Inference · MICRO 2019
Electronic design automation
physical design
0.312017
Generating Routing-Driven Power Distribution Networks With Machine-Learning Technique · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Hardware accelerators and domain-specific architectures › accelerator architecture
memory-efficient accelerator design
0.112019
eCNN: A Block-Based and Highly-Parallel CNN Accelerator for Edge Inference · MICRO 2019
Electronic design automation
design flow
0.112017
Generating Routing-Driven Power Distribution Networks With Machine-Learning Technique · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Electronic design automation
machine learning for EDA
0.112017
Generating Routing-Driven Power Distribution Networks With Machine-Learning Technique · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017

Methods — techniques the papers use, named apart from their topics

persistence-of-vision modeling · 0.9low-rank factorization · 0.9GPU implementation · 0.9block-based parallel processing · 0.4DRAM bandwidth reduction · 0.4wire length prediction · 0.3machine learning · 0.3
YearPublicationVenuePosition
2025 Temporal Fusion: Continuous-Time Light Field Video Factorization
abstract
A factored display emits full-parallax dense-view light fields for a glasses-free 3D experience without sacrificing the spatial resolution of a liquid-crystal display (LCD). For static light fields, it achieves high-quality reconstruction by applying frame-based low-rank factorization to time-multiplexed sub-frame contents of stacked LCDs. However, for light field videos such frame-based factorization could introduce reconstruction artifacts and visual flickers and further cause human discomfort. The artifacts mainly come from incomplete constraints for the emitted light fields that are actually perceived in continuous time, instead of discrete frames. In particular, the perceived light fields are related to the persistence-of-vision (POV) effect of human eyes and the refresh rates of LCD displays, which is not well explored in previous work. In this work, we introduce a light-field video factorization framework-temporal fusion (TF)-to resolve these issues. To begin with, we explicitly formulate the continuous-time POV effect into a global factorization objective functional to eliminate visual flickers and enhance image quality. We further show that this optimization problem can be solved by sequence-level iterative updates on LCD sub-frames. Then, to tackle the enormous requirement of memory access for the sequence-level processing flow, we devise an efficient cuboid-wise factorization algorithm which enables practical GPU implementation. We also devise another lightweight causal framework, TF-C, for supporting low-latency applications. Finally, extensive experiments are performed to verify the effectiveness. Compared to the plain frame-based factorization, TF/TF-C can improve temporal consistency by reducing flicker values by 85%/91% and enhance reconstruction quality by increasing PSNR values by 5.0dB/3.7dB. In addition, we present a prototype dual-layer factored display, which was built with two 240-Hz high-refresh-rate LCDs, to demonstrate the visual quality for real-life applications.
Li-De Chen, Li-Qun Weng, Hao-Chien Cheng, An-Yu Cheng, Chao-Tsung Huang
IEEE Trans. Image Process.1
2024 VLSI Design of Light-Field Factorization for Dual-Layer Factored Display
abstract
This article introduces a VLSI design for light-field factorization, aimed at enhancing immersive 3-D visual experiences for computational light-field factored displays. The main design challenges are intensive memory-access demands and high computational complexity. Accordingly, we first propose half-block-based factorization (HBBF) and sparse ray sampling (SRS) to reduce DRAM bandwidth by 99% and SRAM size by 74%. Then, we devise integer hybrid quantization (INTH) to cut down computational logic by 41%, leading to improvements in die area and power efficiency. Finally, we fabricated a processor chip that incorporates 75.1 kB of SRAM and 5.9M logic gates using 40-nm CMOS technology. It can operate with three different performance modes: high quality (56.9 MPixel/s at 971 mW), balanced (62.5 MPixel/s at 442 mW), and low power (61.7 MPixel/s at 283 mW). Across these modes, its normalized energy ranges between 4.4 and 16.2 nJ/pixel. This implementation surpasses existing GPU platforms and offers an$85\times $increase in processing speed and a$311\times $reduction in power consumption. We also showcase a real-time computational 3-D display system with this chip, demonstrating its practical efficacy in computational 3-D display technology.
Li-De Chen, Li-Qun Weng, Hao-Chien Cheng, An-Yu Cheng, Kai-Ping Lin, Chao-Tsung Huang
IEEE Trans. Very Large Scale Integr. Syst.1
2021 Flow Scheduling in a Heterogeneous NFV Environment using Reinforcement Learning
abstract
Network function virtualization (NFV) allows net-work functions executed on general-purpose servers or virtual machines (VMs) instead of proprietary hardware, greatly improving the flexibility and scalability of network services. Recent trends in using programmable accelerators to speed up NFV performance introduce challenges in flow scheduling in a dynamic NFV environment. Reinforcement learning (RL) trains machine learning models for decision making to maximize returns in uncertain environments such as NFV. In this paper, we study the allocation of heterogeneous processors (CPUs and FPGAs) to minimize the delays of flows in the system. We conduct extensive simulations to evaluate the performance of reinforcement learning based scheduling algorithms such as Advantage Actor Critic (A2C), Trust Region Policy Optimization (TRPO) and Proximal Policy Optimization (PPO), and compare with greedy policies. The results show that RL based schedulers can effectively learn from past experiences and converge to the optimal greedy policy. We also analyze in-depth how the policies lead to different processor utilization and flow processing time, and provide insights into these policies.
Chun Jen Lin, Yan Luo 0001, Liang-Min Wang 0002, Li-De Chen
NAS4
2019 Build an SR-IOV Hypervisor
Liang-Min Wang 0002, Alex Zelezniak, E. Scott Daniels, Timothy Miskell, Li-De Chen
IM5
2019 Edison: Event-driven Distributed System of Network Measurement
Xiaoban Wu, Timothy Miskell, Yan Luo 0001, Liang-Min Wang 0002, Li-De Chen
IM5
2019 eCNN: A Block-Based and Highly-Parallel CNN Accelerator for Edge Inference
abstract
Convolutional neural networks (CNNs) have recently demonstrated superior quality for computational imaging applications. Therefore, they have great potential to revolutionize the image pipelines on cameras and displays. However, it is difficult for conventional CNN accelerators to support ultra-high-resolution videos at the edge due to their considerable DRAM bandwidth and power consumption. Therefore, finding a further memory- and computation-efficient microarchitecture is crucial to speed up this coming revolution.
Chao-Tsung Huang, Yu-Chun Ding, Huan-Ching Wang, Chi-Wen Weng, Kai-Ping Lin, Li-Wei Wang 0013, Li-De Chen
MICRO7
2018 A 320M Pixel/S Vlsi Architecture Design of Weighted Mode Filter for 4K Ultra-Hd Depth Upsampling
abstract
High-quality and high-resolution depth maps have opened tremendous possibilities for various applications, such as ARlVR display, 3D reconstruction, image refocusing, and view synthesis. But high-resolution depth estimation requires heavy hardware resources. Depth upsampling with weighted mode filtering is an efficient way to overcome this challenge. However, its hardware implementation has two major design issues: large on-chip memory for storing high-precision depth labels and high logic cost for computing adaptive range weight. In this work, we present two techniques, histogram candidate mapping and binary range weight kernel, which can reduce on-chip memory size and logic gate count by 46.9 % and 64.3 % respectively. Furthermore, we also implement a VLSI circuit for 4K Ultra-HD depth video upsampling using TSMC 40nm technology. It has 25.5-KB SRAM and 420K-gate logic, and the core area is 1.1 ×1.1 mm2. When operating at 200 MHz and 0.9V, it delivers 320M pixel/s to support 4K Ultra-HD depth video at 40 fps, and consumes 104 mW based on post-layout simulation.
Bo-Hsiang Yang, Li-De Chen, Chao-Tsung Huang
ICASSP2
2017 Generating Routing-Driven Power Distribution Networks With Machine-Learning Technique
abstract
As technology node keeps scaling and design complexity keeps increasing, power distribution networks (PDNs) require more routing resource to meet IR-drop and electro-migration (EM) constraints. This paper presents a design flow to generate a PDN that can result in near-minimal overhead for the routing of the underlying standard cells while satisfying both IR-drop and EM constraints based on a given cell placement. The design flow relies on a machine-learning model to quickly predict the total wire length of global route associated with a given PDN configuration in order to speed up the search process. The experimental results based on various 28 nm industrial block designs have demonstrated the accuracy of the learned model for predicting the routing cost and the effectiveness of the proposed framework for reducing the routing cost of the final PDN.
Wen-Hsiang Chang, Chien-Hsueh Lin, Szu-Pang Mu, Li-De Chen, Cheng-Hong Tsai, Yen-Chih Chiu, Mango Chia-Tso Chao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2016 VLSI architecture design of weighted mode filter for Full-HD depth map upsampling at 30fps
abstract
High-resolution depth maps are necessary for advanced computer vision applications but difficult to generate on portable devices. In this paper, we aim to provide a realtime depth upsampling engine using weighted mode filtering to alleviate the hardware requirement in such scenarios. An appropriate filter window size is essential for both performance and complexity, and extensive experiments are conducted to choose an 8×8 window. The design bottlenecks are then mainly twofold: high SRAM bandwidth due to large-window source data access and complex histogram processing for a high depth-label count up to 128. These issues were addressed by two proposed techniques accordingly: source spreading and one-the-fly maximum finding. Based on the synthesis results using TSMC 40nm technology, the proposed architecture can provide Full-HD depth upsampling at 43 fps with 247k logic gates and 5.4 kbytes of SRAM.
Li-De Chen, Yu-Ling Hsiao, Chao-Tsung Huang
ISCAS1
2016 Generating Routing-Driven Power Distribution Networks with Machine-Learning Technique
abstract
As technology node keeps scaling and design complexity keeps increasing, power distribution networks (PDNs) require more routing resource to meet IR-drop and EM constraints. This paper presents a design flow to generate a PDN that can result in minimal overhead for the routing of the underlying standard cells while satisfying both IR-drop and EM constraints based on a given cell placement. The design flow relies on a machine-learning model to quickly predict the total wire length of global route associated with a given PDN configuration in order to speed up the search process. The experimental results based on various 28nm industrial block designs have demonstrated the accuracy of the learned model for predicting the routing cost and the effectiveness of the proposed framework for reducing the routing cost of the final PDN.
Wen-Hsiang Chang, Li-De Chen, Chien-Hsueh Lin, Szu-Pang Mu, Mango Chia-Tso Chao, Cheng-Hong Tsai, Yen-Chih Chiu
ISPD2
2007 A Parallel Positive Boolean Function approach to supervised multispectral image classification
abstract
In this paper, we present a parallel computing technique, referred to as parallel positive Boolean function (PPBF), for supervised classification of multispectral images. The approach is based on the generalized positive Boolean function (GPBF) scheme, which has been successfully applied in multispectral image classification. The GPBF classifier is developed from a stack filter. The stack filter is defined as the class of all nonlinear digital filters. Each stack filter corresponding to a GPBF possesses the weak superposition property and the ordering property. In order for the GPBF to be effective, the proposed PPBF is performed to improve the computational speed by using parallel cluster computing techniques. It creates a set of stack filters in each parallel node implemented by message passing interface (MPI). The proposed PPBF technique reduces the structure complexity of original GPBF. The effectiveness of the proposed PPBF is evaluated by fusing Systeme Pour l’Observation de la Terre (SPOT) images and digital elevation model (DEM) information for land cover classification during the post 921 Earthquake period in Taiwan. The experimental results demonstrated that PPBF not only significantly improves the computational loads of GPBF classification, but also substantially improves the precision of classification compared to conventional classification.
Yang-Lang Chang, Jyh-Perng Fang, Li-De Chen, Long-Shin Liang, Kun-Shan Chen
IGARSS3