Po-Yuan Chen

dblp:61/5790 · DBLP profile ↗
← Back
22ranked-venue papers
14as first author
3since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 10 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 A K-Band CMOS High-gain Power Amplifier Using Transformer-feedback Cascode Topology
abstract
The paper presents a K-band high-gain power amplifier (PA) using transformer-feedback cascode topology, implemented in a 0.18-μm CMOS process. Due to the tradeoff between high gain and wide operational bandwidth, different coupling coefficients of the two transformers are addressed. The proposed high-gain PA demonstrated a measured maximum small-signal gain of 18.6 dB, with a 3-dB bandwidth from 20 to 26.1 GHz, corresponding to a fractional bandwidth of 26.5%. The measured output 1-dB compression point (P1dB) is 11.7 dBm. Furthermore, the proposed PA also achieves a measured output third-intercept point (OIP3) of 26.7 dBm. The measured dc power consumption is 216 mW with a supply voltage of 3.6 V. The chip size of the proposed high-gain PA is 0.53 mm2. This work is highly suitable for certain millimeter-wave applications due to its high-gain characteristic.
Yi-Fu Chen, Guan-Han Lin, Po-Yuan Chen, Hong-Yeh Chang
ISCAS3
2023 Multi-Scale Dynamic Fixed-Point Quantization and Training for Deep Neural Networks
abstract
State-of-the-art deep neural networks often require extremely high computational power which results in the deployment of deep neural networks on embedded devices being impractical. Therefore, model quantization is important for the deployment of deep neural networks on edge devices. The purpose of this paper is to quantize the deep neural networks from high-precision to low-precision (e.g. INT8) dynamic fixed-point format at the layer-by-layer level quantization. In addition, we further improve the uniform dynamic fixed-point quantization to multi-scale dynamic fixed-point quantization for lower quantization loss. The proposed multi-scale dynamic fixed-point quantization scheme divides the quantization ranges into two regions, and each region is assigned different quantization levels and quantization parameters to better approximate the bell-shaped distributions. The proposed quantization pipeline is composed of post-training quantization followed by model fine-tuning which can keep the accuracy drop of the quantized model within 1% mean average precision (mAP). Furthermore, the proposed quantization and fine-tuning method can be combined with model pruning to obtain a compact and accurate deep neural network with low bit-width.
Po-Yuan Chen, Hung-Che Lin, Jiun-In Guo
ISCAS1
2022 WRAP: Weight RemApping and Processing in RRAM-based Neural Network Accelerators Considering Thermal Effect
abstract
Resistive random-access memory (RRAM) has shown great potential for computing in memory (CIM) to support the requirements of high memory bandwidth and low power in neuromorphic computing systems. However, the accuracy of RRAM-based neural network (NN) accelerators can degrade significantly due to the intrinsic statistical variations of the resistance of RRAM cells, as well as the negative effects of high temperatures. In this paper, we propose a subarray-based thermal-aware weight remapping and processing framework (WRAP) to map the weights of a neural network model into RRAM subarrays. Instead of dealing with each weight individually, this framework maps weights into subarrays and performs subarray-based algorithms to reduce computational complexity while maintaining accuracy under thermal impact. Experimental results demonstrate that using our framework, inference accuracy losses of four DNN models are less than 2% compared to the ideal results and 1% with compensation applied even when the surrounding temperature is around 360K.
Po-Yuan Chen, Fang-Yi Gu, Yu-Hong Huang, Ing-Chao Lin
DATE1
2013 Generalization of an Enhanced ECC Methodology for Low Power PSRAM
abstract
Error control codes (ECCs) have been widely used to maintain the reliability of memories, but ordinary ECC codes are not suitable for memories with long codewords. For portable products, power reduction in memories with DRAM-like cells can be done by reducing the refresh frequency, but the loss of data integrity should be taken care of seriously. To solve these issues, we have proposed a parallel encoding and decoding ECC scheme to reduce refresh power for an industrial pseudo-SRAM (PSRAM) with long codewords. In this paper, we briefly review the scheme and propose a systematic way to generate the parity check matrix for the new ECC scheme. We also modify the parity correction mechanism to reduce the operating power of the scheme. As for the 70 ns access time of the 256-MB PSRAM with 64-bit codewords and 16-bit I/O, experimental results show that the new ECC scheme can be integrated with the READ/WRITE operations with about 0.2 percent circuit area overhead and less than 3.5 ns encoding/decoding time. The new ECC architecture provides a flexible solution for memories with different widths of ECC codewords and I/O ports, without the error masking effect or reduction in reliability.
Po-Yuan Chen, Chin-Lung Su, Chao-Hsun Chen, Cheng-Wen Wu
IEEE Trans. Computers1
2012 A memory yield improvement scheme combining built-in self-repair and error correction codes
abstract
Error correction code (ECC) and built-in self-repair (BISR) schemes have been wildly used for improving the yield and reliability of memories. Many built-in redundancy-analysis (BIRA) algorithms and ECC schemes have been reported before. However, most of them focus on either BIRA algorithms or ECC schemes. In this paper, we propose an ECC-Enhanced Memory Repair (EEMR) scheme for yield improvement. Many modern memories are equipped with ECC in addition to BISR. We evaluate the back-end flow that combines both ECC and BIRA to determine whether yield can be improved by proper sequencing of the two steps. We also collect and identify important failure patterns and their distributions from over 100,000 sample memory instances, which are used to enhance the EEMR scheme that incorporates ECC. As ECC is failure pattern sensitive, careful evaluation from realistic failure bitmaps is necessary. We also verify the feasibility of implementing the proposed EEMR scheme by real test data. Experimental results from industrial 4Mb memory instances show that the proposed EEMR scheme gains over 2% instance yield on average, as compared with the traditional scheme. We also investigate the reliability of the EEMR scheme with different ECC specifications and BIRA algorithms.
Tze-Hsin Wu, Po-Yuan Chen, Mincent Lee, Bin-Yen Lin, Cheng-Wen Wu, Chen-Hung Tien, Hung-Chih Lin, Hao Chen 0053, Ching-Nen Peng, Min-Jer Wang
ITC2
2012 Cost modeling and analysis for interposer-based three-dimensional IC
abstract
Three-dimensional (3D) integration has recently become a popular technology for integrated circuits (IC). 3D IC with the passive silicon interposer is currently the main trend in the industry, especially for processor-memory integration. Evaluating the economic efficiency of test operations in the interposer-based 3D IC thus is important. We propose a cost model for the Die-to-Wafer (D2W) and Die-to-Die (D2D) stacking, including manufacturing cost and test cost. A tool which is based on the proposed cost model is developed. We use this tool for cost analysis and for finding the most cost effective test flow. The results show that, in some applications, test flows including the iterative known-good stack (KGS) test and the pre-bond interposer test significantly reduce the cost, when the KGS test yield is lower than 98.2% and the pre-bond interposer test yield is lower than 99.38%. A Shmoo plot is depicted to show the lower bound of the yield of the final package level test, given the number of stacked dies and the final yield. For different applications, the proposed model evaluates the critical yield or cost values, which helps the designers to determine the most cost effective test flow and the system architecture.
Ying-Wen Chou, Po-Yuan Chen, Mincent Lee, Cheng-Wen Wu
VTS2
2012 Decision support for foreign investment strategy under hybrid uncertainty
Po-Yuan Chen, Horng-Jinh Chang
Expert Syst. Appl.1
2011 Predication the Suppressing Human Oral Cancer Cell Line by Curcumin through the Research of Fas Receptor
abstract
Curcumin is commonly applied as food coloring and seasoning but recent research show that curcumin is able to enhance the activation of Fas receptor to the cancer cell apoptosis through extrinsic pathway. Oral cancer results from the particular habit of chewing betel nuts that Taiwanese have. The additive in betel nuts (such as lime) is the main reason that increase the cancer rate of people chewing betel nuts and further increase the death rate gradually. The major treatments for oral cancer are still chemotherapy and surgery. This research expects to perform the computer simulation to calculate the activity of curcumin to oral cancer as medicine. The pharmaceutical activity are evaluate by the score from docking procedure perform by simulation program. We hope this experiment can be the evidence that curcumin can practically be used as medicine for oral cancer therapy in future.
Po-Yuan Chen, Yu-Chi Wu, Tzu-Hurng Cheng, Tzu-Ching Shih, Cing-Tsan Tsai, Chieh-Hsi Wu, Tzu-Yu Hua, Yen-Yu Huang, Ming-Jen Fan
BIBE1
2010 On-chip testing of blind and open-sleeve TSVs for 3D IC before bonding
abstract
Pre-bond test is preferred for a three-dimensional integrated circuit (3D IC), since it reduces stacking yield loss and thus saves cost. In this paper, we present two schemes for testing through-silicon vias (TSVs) by performing on-chip screening before wafer thinning and bonding. The first scheme is for blind TSVs, which have one end floating, using a charge-sharing technique commonly seen in DRAM. The second scheme is for open-sleeve TSVs, which have one end shorted to the substrate, using a voltage-dividing technique commonly seen in ROM. By virtue of the inherent capacitive and resistive characteristics, we detect the TSVs out of a specified range as anomalies, taking into account the effects of process variations in the detection circuitry. The statistical design by Monte Carlo simulation using TSMC 65nm low-power process shows that for blind TSVs, the best overkill ratio is below 6%. For open-sleeve TSVs, inherent limitations restrict the applicability, so more work needs to be done in the future. Our implementation enjoys little area overhead, requiring only a simple sense amplifier and a write buffer that are shared among a number of TSVs. Reducing the number of TSVs that share a test module will reduce the test time, but increase the area overhead. For blind TSVs, the parallelism also affects the overkill and escape rates.
Po-Yuan Chen, Cheng-Wen Wu, Ding-Ming Kwai
VTS1
2009 On-Chip TSV Testing for 3D IC before Bonding Using Sense Amplification
abstract
We present a novel testing scheme for TSVs in a 3D IC by performing on-chip TSV monitoring before bonding, using a sense amplification technique that is commonly seen on a DRAM. By virtue of the inherent capacitive characteristics, we can detect the faulty TSVs with little area overhead for the circuit under test.
Po-Yuan Chen, Cheng-Wen Wu, Ding-Ming Kwai
Asian Test Symposium1
2009 Exploring the Receptor-Based Virtual Screening Study of the Aromatic Substituents for the Hepatitis C Virus NS5B Polymerase by Docking and Scoring
abstract
HCV (Hepatitis C virus) that the NS3 protease and NS5B RNA-dependent RNA polymerase (RbRp) which the enzymes for virtual replication. HCV plays an important role that to cause the chronic and liver diseases. The computer aided drug design (CADD) that is the new method to design the new molecules as like the drugs from the potent compounds. We took the program of Discovery Studio 2.0 and the scoring function in this that the Dock Score, -PLP1, -PLP2, -PMF, and the Jain score in the program. The compound 26a that have the highest biological activity (IC50), but the compound didnpsilat have the highest score in the study. In the scoring function of Dock Score, the compound 13a has the highest score value. The compound 19a has the highest scoring value in the both score functions of -PMF and Jain. The compound 20 showed the highest value in the scoring functions that the -PLP1 and -PLP2. So the de novo evolution in the program that we want to design the HCV NS5B inhibitor that has higher docking scoring value than the potent compound in the future.
Po-Yuan Chen, Wei-Tse Hsu, Mien-De Jhuo, Tzu-Hurng Cheng
BIBE1
2009 The MAPK Signal Pathway Research and New Drug Discovery
abstract
MAPK cell signal transduction pathways determine the survival of the cells. If one can control this pathways, and then they will prohibit the proliferation of the cancer cells. Furthermore, they will heal the cancer smoothly. In order to attain this goal, we use lots of drugs to interact with MEK1 in MAPK, by using computer aided drug design to analyze the ligand activity of proteins in MEK1.
Po-Yuan Chen, Mien-De Jhuo, Wei-Tse Hsu, Tzu-Ching Shih, Tzu-Hurng Cheng
BIBE1
2009 Leakage reduction, delay compensation using partition-based tunable body-biasing techniques
abstract
In recent years, fabrication technology of CMOS has scaled to nanometer dimensions. As scaling progresses, several new challenges follow. Among them, the most noticeable two are process variations and leakage current of the circuit. To tackle the problems of process variations and leakage current, an effective way is to use a body-biasing technique. In substance, using the RBB technique can minimize leakage current but increase the delay of a gate. Contrary to RBB, the FBB technique decreases the delay but increases leakage current of a gate. In the previous work, a single body-biasing is applied to the whole circuit. In a slow circuit, since the FBB is applied to the whole circuit, the leakage current of all gates in the circuit increases dramatically. On the other hand, in a fast circuit, RBB is applied to decrease the leakage current. However, without violating the timing specification, the value of body-biasing is restricted by the critical paths, and the saving of leakage current is limited. In this article, we propose a design flow to partition the circuit into subcircuits so that each subcircuit can be applied its individual RBB or FBB. Experiments show that our method is able to save leakage current from 42% to 47% as compared to designs not using a body-biasing technique. Under process variations, our method can save 42% to 49% leakage on fast circuits and 20% to 35% on slow circuits.
Po-Yuan Chen, Chiao-Chen Fang, TingTing Hwang, Hsi-Pin Ma
ACM Trans. Design Autom. Electr. Syst.1
2009 Skew-aware polarity assignment in clock tree
abstract
In modern sequential VLSI designs, clock tree plays an important role in synchronizing different components in a chip. To reduce peak current and power/ground noises caused by clock network, assigning different signal polarities to clock buffers is proposed in previous work. Although peak current and power/ground noises are minimized by signal polarities assignment, an assignment without timing information may increase the clock skew significantly. As a result, a timing-aware signal polarities assigning technique is necessary. In this article, we propose a novel signal polarities assigning technique which can not only reduce peak current and power/ground noises simultaneously but also render the clock skew in control. The experimental result shows that the clock skew produced by our algorithm is 94% of original clock skew in average while the clock skews produced by three algorithms (Partition, MST, Matching) in the absence of post clock tuning steps in the previous work are 235%, 272%, and 283%, respectively. Moreover, our algorithm is as efficient as the three algorithms of the previous work in reducing peak current and power/ground noises.
Po-Yuan Chen, Kuan-Hsien Ho, TingTing Hwang
ACM Trans. Design Autom. Electr. Syst.1
2008 Exploring 3D-QSAR pharmacophore mapping of azaphenanthrenone derivatives for mPGES-1 inhibition Using HypoGen technique
abstract
Microsomal prostablandin E synthase-1 (mPGES-1) has been recently investigated to be a novel and promising target for inflammation-related diseases. The quantitative structure-activity relationship (QSAR) study was used to explore the critical pharmacophore features of mPGES-1 by using a set of 35 azaphenanthrenone derivatives. Twenty four selected pharmacophore models derived from 240 hypotheses were employed to identify the critical features. The best two pharmacophore hypotheses exhibited the residuals of approximately 150 and the high correlation coefficient of 0.92. The selected four hypotheses all showed a confidence level of 95 % in the Fischerpsilas randomization test. The final four pharmacophore model showed that the four dominant features (hydrogen bond donor and 3 hydrophobic features, occasionally replaced by ring aromatic feature) had significant impact on activity of mPGES-1 inhibitors. The database virtual screening and drug design can be further implement to searching the novel mPGES-1 inhibitors.
Winston Yu-Chen Chen, Po-Yuan Chen, Calvin Yu-Chian Chen, Jing-Gung Chung
CIBCB2
2008 Transition-aware decoupling-capacitor allocation in power noise reduction
abstract
Dynamic power noises may not only degrade the circuit performance but also reduce the noise margin which may result in the functional errors in integrated circuit. Decoupling capacitor (decap) allocation is one of the most effective way in reducing serious dynamic power noises (hotspots). To allocate decap before placement, we observed that not only locations but also rising time of functional cells are required to accurately predict power noises. Compared to a previous work which only takes neighborhood relation into consideration, our method is more efficient in reducing hotspots. Furthermore, to reduce the hotspots after placement, instead of only using the empty space as proposed in the previous work, we move out cells in the area with serious power noise area (hot area). The obtained empty space can be used to accommodate decaps to further reduce the hotspots. The experimental result shows, compared to the previous work [1], our estimation function to allocate decap before placement is 23% better in reducing power noises. Moreover, compared to a method which fills decaps to all remaining empty space, our cell move algorithm can almost eliminate all the remaining hot grid nodes and hot cells. In summary, compared to the original circuits (without decap), about 60% of hotspots can be removed using our prediction function before placement, and most of the remaining hotspots are removed by our cell moving step after placement.
Po-Yuan Chen, Che-Yu Liu, TingTing Hwang
ICCAD1
2007 Skew aware polarity assignment in clock tree
abstract
In modern sequential VLSI designs, clock tree plays an important role in synchronizing different components in a chip. To reduce peak current and power/ground noises caused by clock network, assigning different signal polarities to clock buffers is proposed in previous work. Although peak current and power/ground noises are minimized by signal polarities assignment, an assignment without timing information may increase the clock skew significantly. As a result, a timing-aware signal polarities assigning technique is necessary. In this paper, we propose a novel signal polarities assigning technique which can not only reduce peak current and power/ground noises simultaneously but also render the clock skew in control. The experimental result shows that the clock skew produced by our algorithm is 94% of original clock skew in average while the dock skews produced by three algorithms (Partition. MST, Matching) [5] are 235%, 272%, and 283%, respectively. Moreover, our algorithm is as efficient as the three algorithms of [5] in reducing peak current and power/ground noises.
Po-Yuan Chen, Kuan-Hsien Ho, TingTing Hwang
ICCAD1
2007 Reliability Analysis of Physiological Phenomena by Cardiac Action Potential Model
abstract
To simulate ventricular cardiac action potentials used Luo-Rudy phase I cell model with hybrid method, the computation time can be 100 times faster than RK4 (fixed step 0.01). However, the hybrid method of time step 0.8-2.0 causes the irregular S:R ratio by Wenckebach periodicity as stimulation 30 beats; other hybrid methods can adequately simulate cardiac physiology phenomena and prevent from phantom simulation. In this study, we analyse the simulation reliability and accuracy for three different speeds of hybrid methods that compare to RK4 (0.01), thus a suggestion can be understood how to select the right methods to simulate the physiological phenomena under computing speed fast enough and maintain the accuracy, saving huge amount of the computing time in the cardiac simulations.
Ching-Hsing Luo, Po-Yuan Chen, Chun-Hao Teng, Sheng-Nan Wu, Ruey-Jen Sung
ISCAS2
2007 A Bus-Encoding Scheme for Crosstalk Elimination in High-Performance Processor Design
abstract
A crosstalk effect leads to increases in delay and power consumption and, in the worst-case scenario, to inaccurate results. With the scale down of technology to deep-submicrometer level, the crosstalk effect between adjacent wires becomes more and more serious, particularly between long on-chip buses. In this paper, we propose a deassembler/assembler technique to eliminate undesirable crosstalk effects on bus transmission. By taking advantage of the prefetch process, where the instruction/data fetch rate is always higher than the instruction/data commit rate, the proposed method incurs almost no penalty in terms of dynamic instruction count. In addition, when the bus width is 128 b, the required number of extra bus wires is only 7 as compared to the 85 extra bus wires needed in the work of Victor and Keutzer.
Wen-Wen Hsieh, Po-Yuan Chen, Chun-Yao Wang, TingTing Hwang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2006 Switching-activity driven gate sizing and Vth assignment for low power design
abstract
Power consumption has gained much saliency in circuit design recently. One design problem is modeled as "under a timing constraint, to minimize power as much as possible". Previous research regarding this problem focused on either minimizing dynamic power by gate sizing, or reducing leakage power by dual threshold voltage assignment on non-critical path. However, given a timing constraint, an optimization algorithm must be able to utilize gate sizing and threshold-voltage assignment interchangeably, in order to minimize total power consumption including dynamic and leakage power in active mode and leakage power in idle mode. We find that switching-activity of a gate plays an important role in making decision as to choosing gate sizing or threshold-voltage assignment for performance improvement. For high switching-activity gates, threshold-voltage assignment should be used while for low switching-activity gates, gate sizing should be utilized. We develop an algorithm to perform gate sizing and threshold-voltage assignment simultaneously taking switching activity into consideration. The results show that under the same timing constraint, our circuits have 16.26%, and 18.53%, improvement of total power as compared to the original circuits for the cases where the percentage of active time are 100%, and 50%, respectively.
Yu-Hui Huang, Po-Yuan Chen, TingTing Hwang
ASP-DAC2
2006 An Enhanced EDAC Methodology for Low Power PSRAM
abstract
As feature size keeps shrinking, how to maintain the reliability becomes an important issue in IC production, especially for high density memory circuits. Error detection and correction (EDAC) schemes have been widely used for memory circuits for this purpose, but ordinary EDAC schemes are not suitable for memories with long codewords. The demand for low-power memory is increasing due to the growth in portable electronics markets. Power reduction in memories with DRAM-like cells can be done by reducing the refresh frequency, but the loss of data integrity should be taken care of seriously. To solve the above two issues, we propose a parallel encoding and decoding EDAC scheme, which can be used on memories with long codewords. Targeting refresh power reduction, we have implemented our scheme on an industrial pseudo SRAM (PSRAM), and have completed experiments. The major hardware penalty is the parity overhead that is 1/9, and the longest delay of our circuit is 3.6ns for the PSRAM fabricated by a 0.11mum CMOS technology. With respect to the 70ns access time of the PSRAM, the proposed EDAC scheme can be integrated with the read/write operations without increasing the latency. Experimental results show that the refresh time can be extended greatly, without sacrificing reliability
Po-Yuan Chen, Chao-Hsun Chen, Jen-Chieh Yeh, Cheng-Wen Wu, Jeng-Shen Lee, Yu-Chang Lin
ITC1
2003 Bounding the Execution Times of DMA I/O Tasks on Hard-Real-Time Embedded Systems
Chih-Chieh Chou, Po-Yuan Chen
RTCSA3