EDBT 2026 Demo / reviewers in the wild / expert
Chun-Ming Huang
dblp:93/6150
· DBLP profile ↗
23ranked-venue papers
7as first author
4since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-authorTheory of computation · 2Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | LSvT: EEG Channel Localization and Selection via Training for BCI ApplicationsabstractElectroencephalography (EEG) is widely utilized in neuroscience and clinical applications, serving as a noninvasive method for monitoring brain activity. However, despite its extensive use, the multitude of channels recorded by scalp electrodes presents challenges, including impractical usage, heightened model complexity, and potential overfitting issues during training. This paper addresses the challenges of high dimensionality in EEG data and introduces two innovative EEG channel selection algorithms, achieving significant reductions in channels, model size, and complexity while maintaining high classification accuracy. Validated through experiments on EEGNet and the MNE/BCI Competition IV 2a datasets, these algorithms prove valuable for practical and cost-efficient scenarios. LSvT-S emphasizes high channel reduction and low complexity, while LSvT-G targets faster channel selection, offering users choices. Experiments on MNE and BCI Competition IV 2a datasets show that LSvT-S achieves a remarkable 96.7% and 81.8% reduction in channels, along with 44% and 12.4% reductions in model size, and 96.2% and 80.8% in computation complexity, respectively. Meanwhile, LSvT-G achieves an additional speedup ranging from 32.8x to 7.6x. Chun-Ming Huang, Wei-Lin Lai, Chih-Chyau Yang, Yi-Jie Hsieh, Chien-Ming Wu, Chu-Hui Lee |
IJCNN | 1 |
| 2023 | DLA-SP: A System Platform for Deep-Learning Accelerator IntegrationabstractMany deep-learning Accelerators (DLAs) are presented to meet high-performance needs in versatile deep-learning applications. However, they usually lack a flexible platform and flow for system integration. For assisting the professors and students in the academia of Taiwan to speed up their deep-learning accelerator implementation and verification of innovative system designs, the Taiwan Semiconductor Research Institute provides a new service platform named DLA-SP. The developed DLA-SP platform is highly flexible in contrast to existing DLA prototyping systems by providing three design resources: A parameterized DLA wrapper, an integration flow, and a hardware/software board support package. The parameterized DLA wrapper is capable of rate adaptation and format conversion between the DLA and the system. The number of DLA I/Os can be reduced to 29.7% to avoid the pad-limit DLA chip problem; The integration flow and methodology enable the DLA hardware system integration and DLA patterns reuse; The hardware/software board support package is capable of rapid deep-learning application development. The DLA-SP helps professors and students to concentrate their efforts on their deep-learning accelerators, and easily reuse system platforms, which greatly reduces the developing cycle of an embedded system. In this paper, a flexible system platform for deep-learning accelerators is presented, the accelerators can be easily integrated into our proposed DLA-SP system platform in both FPGA and chip ways for system verification and demonstration. A case study of a keyword-spotting accelerator is adopted as an example to illustrate how a DLA is integrated and works properly with our DLA-SP platform. Chih-Chyau Yang, Fu-Chen Cheng 0001, Tsung-Jen Hsieh, Chien-Ming Wu, Chun-Ming Huang |
IECON | 6 |
| 2022 | A 40nm CMOS SoC for Real-Time Dysarthric Voice Conversion of Stroke PatientsabstractThis paper presents the first dysarthric voice conversion SoC, which can translate stroke patients' voice into more intelligible and clearer speech in real time. The SoC is composed of a RISC-V MPU and a compact DNN engine with a single 16-bit multiply-accumulator, which improves 12x performance and > 100x energy efficiency, and has been implemented in 40nm CMOS. The silicon area is 0.68×0.79mm2, and the measured power is 18.4mW for converting 3-sec dysarthric voice within 0.5 sec (at 200MHz and 0.8V) and 4.8mW for conversion < 1 sec (at 100MHz and 0.6V). Tay-Jyi Lin, Chen-Zong Liao, You-Jia Hu, Wei-Cheng Hsu, Zheng-Xian Wu, Shao-Yu Wang, Chun-Ming Huang, Ying-Hui Lai, Chingwei Yeh, Jinn-Shyan Wang |
ASP-DAC | 7 |
| 2022 | A Real-Time 1280 × 720 Object Detection Chip With 585 MB/s Memory TrafficabstractMemory bandwidth has become the real-time bottleneck of current deep learning accelerators (DLAs), particularly for high definition (HD) object detection. Under resource constraints, this article proposes a low memory traffic DLA chip with joint hardware and software optimization. To maximize hardware utilization under memory bandwidth, we morph and fuse the object detection model into a group fusion-ready model to reduce intermediate data access. This reduces the YOLOv2’s feature memory traffic from 2.9 to 0.15 GB/s. To support group fusion, our previous DLA-based hardware employees a unified buffer with write-masking for simple layer-by-layer processing in a fusion group. When compared to our previous DLA with the same processing element (PE) numbers, the chip implemented in a 40-nm process supports$1280\times 720$at 30 frames per second (FPS) object detection and consumes$7.9\times $less external dynamic random access memory (DRAM) access energy, from 2607 to 327.6 mJ. Kuo-Wei Chang, Hsu-Tung Shih, Tian-Sheuan Chang, Shang-Hong Tsai, Chih-Chyau Yang, Chien-Ming Wu, Chun-Ming Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2019 | A Smart Sensor Development Platform and Its System DemonstrationabstractThe smart sensor plays an important role for Internet of Things (IoT) development in the past few years. Due to the fast advance of IC fabrication and electronic design automation technologies, integrating a smart sensor design into a single chip has become practical. To assist the MEMS sensor teams in Taiwan academia to accelerate their smart sensor development, this paper presents a smart sensor development platform which consists of a common platform unit and a sensor unit. Our proposed smart sensor development platform provides the solutions of FPGA-based design, Smart Sensor on Chip (SSoC) design, and Smart Sensor in Package (SSiP) design for smart sensor development. The design flow and deliverables for the SSoC and SSiP smart sensor implementations are also presented in this paper. The proposed smart sensor common platform unit was taped out with UMC 0.18um process to perform the silicon proof. Moreover, this paper also presents a modularized wireless sensor system, MorSensor, to facilitate the smart sensor system demonstration. A case study for the FPGA-based smart sensor is also given in this paper. The experiment results show that the presented smart sensor platform is very suitable for smart sensor development, while MorSensor is suitable for the smart sensor system demonstration. Chun-Ming Huang, Chih-Chyau Yang, Yi-Jie Hsieh, Chun-Wen Cheng, Yi-Jun Liu, Jia-Rong Chang, Yu-Tsang Chang, Chien-Ming Wu |
ISCAS | 1 |
| 2017 | A modular wireless sensor platform and its applicationsabstractIn this paper, we propose a modular wireless sensor platform that consists of sensor modules. Each sensor module is a part of sensor system and in charge of one job in the system, such as computation, communication, output or sensing. Users can stack multiple modules together to build a unique sensor platform. Since users are able to easily replace one module with others, the proposed platform is highly extendable and reusable. Besides, we also design different kinds of mounts, so that sensors can be mounted on objects. This is especially helpful for some moving sensing applications such as attitude monitoring. To demonstrate the proposed platform, we show an alcohol detection application in the paper. The results show that the proposed platform is suitable for academic researches and industrial prototype verification. Chun-Ming Huang, Yi-Jie Hsieh, Wei-Lin Lai, Yi-Jun Liu, Chun-Ying Juan, Ssu-Ying Chen, Jin-Ju Chue, Chih-Chyau Yang, Chien-Ming Wu |
ISCAS | 1 |
| 2017 | IoTtalk: A Management Platform for Reconfigurable Sensor DevicesabstractIoTtalk is a platform for Internet of Things (IoT) device interaction, which nicely integrates a reconfigurable multi-sensor device called MorSensor with the proposed IoT management platform in the network domain. The sensors can be dynamically plugged in/out of a MorSensor device without being turned off, and IoTtalk automatically generates/reuses the application software for these sensors. We propose a dynamic ranging concept that automatically specifies the value range of a sensor so that it can send “meaningful” data to any connected output IoT device. Therefore, a MorSensor device can “talk” to other IoT devices, such as a light bulb or an electric fan. We also show that IoTtalk is a simple yet almost free solution for automatic sensor calibration. Finally, we illustrate that IoTtalk is a powerful tool for developing interactive science experiments. Yi-Bing Lin, Yun-Wei Lin, Chun-Ming Huang, Chang-Yen Chih, Phone Lin |
IEEE Internet Things J. | 3 |
| 2012 | Construction of One-Coincidence Sequence Quasi-Cyclic LDPC Codes of Large GirthabstractOne approach for designing the one-coincidence sequence (OCS) low-density parity-check (LDPC) codes of large girth is investigated. These OCS-LDPC codes are quasi-cyclic, and their parity-check matrices are composed of circulant permutation matrices. Generally, the cycle structures in these codes are determined by the shift values of circulant permutation matrices, and the existence of cycles in the corresponding Tanner graph is governed by certain cycle-governing equations (CGEs). Therefore, finding the proper shift values is the key point to increase the girth of these codes. In this paper, we provide an effective method to systematically find out the CGEs for these codes of girth 6, 8, and 10, respectively. Then, one less computation-intensive algorithm is used to generate the proper shift values for constructing the OCS-LDPC codes of large girth. Simulation results show that significant gains in signal-to-noise ratio over an additive white-Gaussian noise channel can be achieved by increasing the girth of the OCS-LDPC codes. Jen-Fa Huang, Chun-Ming Huang, Chao-Chin Yang |
IEEE Trans. Inf. Theory | 2 |
| 2010 | High-speed and low-power programmable frequency dividerabstractThis paper presents a novel 2/3 divider cell circuit design for a truly modular programmable frequency divider with high-speed, low-power, and high input-sensitivity features. In this paper, the proposed flip-flop based 2/3 divider cell adopts dynamic E-TSPC circuit that not only reduces power consumption, but also improves operation speed and input sensitivity. The whole design was implemented using the TSMC 0.18 μm 1P6M CMOS process. With an 8-stage 2/3 divider cell, the measurement results indicate that the proposed circuit operates up to 5.8 GHz with the power-consumption less than 3.24 mW. Ting-Hsu Chien, Chi-Sheng Lin, Chin-Long Wey, Ying-Zong Juang, Chun-Ming Huang |
ISCAS | 5 |
| 2010 | A packet-based emulating platform with serializer/deserializer interface for heterogeneous IP verificationabstractThis paper proposes a packet-based verification platform with serial link interface for emulating the hardware of the heterogeneous IPs before tape out. With the serial link interface Serializer/Deserializer (SerDes) added between IPs, significant amount of pin counts can be reduced in the platform. An adapter is inserted between IP and SerDes to convert parallel bus into packets and handle the handshaking. Under our proposed adapter architecture and handshaking scheme, the limitation on the number of the master adapter is eliminated compared with Bus-based Advanced High-performance Bus (AHB) architecture. Simulation results show the data transfer through our proposed architecture works correctly without the limitation on the number of masters. With the proposed adapter and SerDes architecture, the number of required signals in the interconnect is reduced from 79 to two for the AHB bus. Chih-Hsing Lin, Yung-Chang Chang, Wen-Chih Huang, Wei-Chih Lai, Ching-Te Chiu, Jen-Ming Wu, Shuo-Hung Hsu, Chun-Ming Huang, Chih-Chyau Yang, Shih-Lun Chen |
ISCAS | 8 |
| 2009 | Implementation and Prototyping of a Complex Multi-project System-on-a-chipabstractA silicon prototyping methodology is presented for Multi-Project System-on-a-Chip (MP-SoC) implementation. A multi-projects platform was created for integrating heterogeneous SoC projects into a single chip. The total silicon prototyping cost of these projects can be greatly reduced by sharing a common platform. To demonstrate the effectiveness of the proposed methodology, a MP-SoC chip was implemented with eleven SoC projects sharing the common platform. The total silicon area is about 37.97 mm2in the TSMC 0.13 um CMOS generic logic process technology. Compared with the total chip area 129.39 mm2by implementing these projects separately, the results show that there are 91.42 mm2silicon areas reduced by the MP-SoC platform. In order to verify MP-SoC through silicon prototyping, a system modeling and hardware/ software co-design virtual platform were implemented. A configurable SoC prototyping system, namely CONCORD, is also created as a verification platform for emulating the hardware of MP-SoC before chip being taped out. The CONCORD system provides higher connection flexibility, modularization, and architecture consistence than conventional FPGA systems. Chun-Ming Huang, Chien-Ming Wu, Chih-Chyau Yang, Wei-De Chien, Shih-Lun Chen, Chi-Shi Chen, Jiann-Jenn Wang, Chin-Long Wey |
ISCAS | 1 |
| 2008 | PrSoC: Programmable System-on-chip (SoC) for silicon prototypingabstractThis paper presents a Programmable SoC (System- on-chip) design methodology which integrates multiple heterogeneous SoC design projects into a single chip such that the total silicon prototyping cost for these projects can be greatly reduced by sharing the common SoC platform. Results show that an integrated SoC platform is comprised of eight SoC projects. When these eight SoC projects are designed separately, the total area is approximately 143.03mm , while the area of the integrated platform is about 24.43mm . The area reduction is significant, so is the fabrication cost. Once the integrated platform chip is fabricated, three programming schemes are carried out to allow the integrated chip to act as the individual SoC design projects. A test chip is designed and implemented using the TSMC 0.13um CMOS generic logic process technology. Chun-Ming Huang, Chien-Ming Wu, Chih-Chyau Yang, Chin-Long Wey |
ISCAS | 1 |
| 2008 | The strong distance problem on the Cartesian product of graphs
Justie Su-tzu Juan, Chun-Ming Huang, I-Fan Sun |
Inf. Process. Lett. | 2 |
| 2007 | Error resilient GOP structures on video streaming
Kai-Chao Yang, Chun-Ming Huang, Jia-Shung Wang |
J. Vis. Commun. Image Represent. | 2 |
| 2007 | Design of frame dependency for VCR streaming videos
Kai-Chao Yang, Chun-Ming Huang, Jia-Shung Wang |
Signal Process. Image Commun. | 2 |
| 2006 | Efficient path metric access for reducing interconnect overhead in Viterbi decodersabstractEfficient management of the path metric memory and minimization of interconnection networks between the memory and addcompareselect unit (ACSU) are always the key concerns on the design and implementation of Viterbi decoders. In this paper, we derive a set of simple equations to partition the memory into P banks such that the equivalent memory bandwidth can be increased with very simple interconnection networks. Compared with the previous work, our proposed approach reveals the following superiority: (1) Each memory bank can be treated as a local memory of a specific ACS; thus, the interconnection network is simplified. (2) The P memory banks can be merged into only two pseudo-banks regardless of the number of ACS operations. This not only further reduces the hardware requirements of address generation, but also makes smaller the required memory space. Ming-Der Shieh, Tai-Ping Wang, Chien-Ming Wu, Chun-Ming Huang |
ISCAS | 4 |
| 2006 | A cost-effective reconfigurable accelerator for platform-based SOC designabstractIn this paper, we propose a cost-effective reconfigurable accelerator for the platform-based system-on-a-chip (SoC) design. Based on the proposed design methodology, the reconfigurable computation array (RCA) can be landed with the features of high usage rate and low hardware cost without sacrificing multimedia computation performance. The RCA consisting of 8 type 1 grouped processing elements (GPE1s), 3 GPE2s and 1 GPE3 is capable of configuring two 16times16-bit multiplication, eight 8times8 multiplication, and sixteen 8-bit absolute operations in different connection topologies. Via the cost-effective RCA, the number of GPEs can be saved up to 25% and the usage rates of the RCA compared with that of for motion estimation (ME), RGB2YUV and DCT/IDCT can be improved by 25%, 18.7%, and 23.9%, respectively Lan-Da Van, Hsin-Fu Luo, Nien-Hsiang Chang, Chun-Ming Huang |
ISCAS | 4 |
| 2005 | High-performance low-complexity bit-plane coding scheme for MPEG-4 FGSabstractMPEG-4 FGS (Fine Granularity Scalability) has received tremendous attentions because it has ability to adapt to the network bandwidth variation. In this paper, we present a novel and effective bit-plane coding technique to further improve the coding efficiency of MPEG-4 FGS. The proposed approach reveals three superiorities to the MPEG-4 FGS based bit-plane variable-length coding (VLC): (1) better video quality in about 1.2 dB in terms of average PSNR, (2) less memory requirement, and (3) lower implementation complexity and power dissipation. Thus, it is well suitable for efficient hardware or software implementations. Hong-Yu Chao, Jia-Shung Wang, Juin-Long Lin, Kai-Chao Yang, Chien-Ming Wu, Chun-Ming Huang, Lan-Da Van |
ICME | 6 |
| 2004 | Error resilience supporting bi-directional frame recovery for video streamingabstractIn this paper, we propose a novel coding dependency among video frames which supports efficient error resilience without adding any redundancy. Our approach is based upon reorganizing the regular GOP structure. The proposed scheme can effectively recover a lost frame and prevent the error propagation phenomenon. For a single frame loss, we guarantee that both of its next and previous frames still can be successfully decoded, thus we have sufficient temporal and spatial information to reconstruct the damaged frame. Our new coding structure even can recover successive lost frames. The experimental results showed that the video quality has a graceful degradation when the loss rate increases rapidly. Comparing with the conventional GOP, PSNR values can improve from 0.5 dB to 3 dB. Chun-Ming Huang, Kai-Chao Yang, Jia-Shung Wang |
ICIP | 1 |
| 2004 | Support fast scan operations with video streaming technologyabstractIn this article, we present a novel approach to support fast scan operations for video streaming applications. Our approach is based upon reshaping the ordinary linear encoded GOP (group of pictures) structure into a hierarchically encoded binary tree, along with a proximity-based approximation scheme. Our scheme can support forward and backward fast playback operations at any speed-up factor with few bandwidth requirements even for the normal speed playback. Chun-Ming Huang, Kai-Chao Yang, Jia-Shung Wang |
ICME | 1 |
| 2004 | A high-performance area-aware DSP processor architecture for video codecsabstractIn this paper, we propose a high-performance and area-aware very long instruction word (VLIW) DSP architecture using a flexible single instruction multiple data (SIMD) approach and a grouped permutation (GP) structure register file, respectively. Via the proposed data path architecture, the reduction of the execution cycles for digital filter and RGB2YUV benchmarks can be improved up to 50% compared with that of Hinrichs et al. (2000) and Lin et al. (2003). For motion estimation, the number of pixels per cycle applying the proposed architecture can be four times than that of Hinrichs and Lin. For the register file, using the proposed GP structure, the saving of switching network overhead can be anticipated compared with the work in Lin Lan-Da Van, Hsin-Fu Luo, Chien-Ming Wu, Wen-Hsiang Hu, Chun-Ming Huang, Wei-Chang Tsai |
ICME | 5 |
| 2003 | Restructuring GOP Algorithm to Reduce Video Server Load on VCR FunctionalityabstractWe address serious video server load resulting from the VCR functionality, and the restructuring algorithm of frame dependencies is presented to reduce the server load and minimize the requirements of the network bandwidth. The requested frames are transmitted and decoded without sending any redundant data. Thus, even for the VCR operations with high speed factors, the server still sends data at normal transmission rate. The video server cost can significantly reduce as well. Kai-Chao Yang, Chun-Ming Huang, Jia-Shung Wang |
ICPP | 2 |
| 1994 | Performance of a global circuit-switched satellite communication networkabstractA circuit-switched network consisting of multiple low Earth orbiting satellites and three geostationary satellites is considered. Crosslinks among low-orbit satellites are used as the major communication channels, while calls are routed via geostationary satellites only when low-orbit routes exceed or equal to a hop-count threshold. Simulation results are used to illustrate the network throughput under various traffic conditions and different link capacities. The impact of satellite failures on the network performance is also investigated.> Zsehong Tsai, Chuan-Chen Chuang, Jin-Fu Chang, Chun-Ming Huang |
VTC | 4 |