EDBT 2026 Demo / reviewers in the wild / expert
Pei-Yin Chen
dblp:44/5992
· DBLP profile ↗
40ranked-venue papers
12as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 6 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-authorComputer networks · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorSecurity and privacy · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Retraining: Source-Free Adaptation for Generalizable Intrusion DetectionabstractMachine learning (ML)–based intrusion detection systems (IDS) often degrade when deployed across heterogeneous networks due to domain shifts in traffic and configuration. To mitigate this degradation, conventional domain adaptation (DA) methods aim to align source and target data distributions; however, they require access to source data during deployment—an impractical constraint that undermines scalability and reusability. To overcome this limitation, we propose TRANSFA-IDS (Transformer Source-Free Adaptation for IDS), which removes the need for source data during adaptation while preserving the knowledge encoded in the source-trained model. TRANSFA-IDS transforms tabular flow records into structured color image embeddings and employs a compact Vision Transformer with a Deep Support Vector Data Description (Deep-SVDD) head to learn domain-invariant representations of benign behavior. During deployment, it adapts to new environments using only a small portion of unlabeled target traffic by fine-tuning the last Transformer block, efficiently realigning feature distributions without retraining. Experiments across cross-dataset settings (CICIDS2018↔UNSW-NB15) show that TRANSFA-IDS achieves AUROC scores up to 0.908 and 0.873, outperforming traditional non-adaptive unsupervised baselines by over 40% while adapting more than twice as fast as conventional adaptive unsupervised. These results demonstrate that source-free adaptation can deliver both high accuracy and deployment practicality for scalable IDS across diverse network environments. Didik Sudyana, Wong Yu Xuan, Laurens D'hooge, Ren-Hung Hwang, Narn-Yih Lee, Pei-Yin Chen, Tim Wauters, Bruno Volckaert, Filip De Turck |
ICC | 6 |
| 2026 | High-Performance AES-GCM Hardware via Circuit and Architecture Co-Design of AES and GHASH
Chia-Chou Chuang, Chuan-Kai Su, Ying-Sheng Huang, Da-Huei Lee, Pei-Yin Chen |
ISCAS | 5 |
| 2026 | CRC-Based Error Detection Mechanism for Multiplicative Inverse in Redundant Basis S -Box
Ying-Sheng Huang, Chia-Chou Chuang, Chien-Chung Ho, Pei-Yin Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | Live Demonstration: SoC Design of Lightweight Cryptography for Real-Time Monitoring SystemabstractThis live demonstration shows a low-cost but high-performance hardware IP for video stream encryption and authentication on real-time monitoring system. The proposed IP is essentially based on the established standard of lightweight cryptography (LWC) by NIST in 2023, which is called Ascon cipher. For maintaining the frame rate of the video, a high-speed and low-cost hardware IP is proposed with a step-by-step setup procedure introduced in this brief. Given that the recursive permutation of Ascon algorithm, the cost of the proposed Ascon cipher IP can be sustainably reduced by round-based design. The video stream encryption system is realized on Xilinx PYNQ-Z2 board with some peripheral IPs embedded in that FPGA board. Implementation results show the designed IP and the SoC system can perform better than the previous work. Yi-Hsuan Lee, Po-Chun Chen, Yu-Cheng Huang, Chia-Chou Chuang, Narn-Yih Lee, Pei-Yin Chen |
ISCAS | 6 |
| 2024 | Cryptosystem for IoT Devices With Feedback Operation Modes Based on Shared Buffer and Unrolled-Pipeline TechniquesabstractThis article presents a multichannel cryptosystem based on the proposed shared buffer techniques which enhance security for several Internet of Things (IoT) sensors and devices by using block cipher with feedback operation modes. Information security and its performance are both important characteristics for modern networking devices. In cryptography, the block cipher, for example, advanced encryption standard (AES), using feedback operation modes can enhance the security level but, unfortunately, the encryption cannot be performed in parallel owing to data dependency. Therefore, the pipeline and unrolling techniques are not applicable to increase the throughput of hardware designs. Under this situation, the round-based architecture is popular and regarded as an area-efficient solution. However, this method inherently limits its throughput. The proposed cryptosystem aims to provide many resource-constrained IoT devices with high-speed centralized encryption service to enhance their security levels, which are applicable for various scenarios, such as vehicular network, home network, and network function virtualization (NFV)/software-defined networking (SDN) IoT. Beyond that, the proposed design introduces the shared buffer technique based on linked lists and presents a novel queuing structure to enhance the memory utilization so that it can reduce 72.9% memory requirement of the naïve implementation while achieving the same speedup. According to the implementation result, an aggregate throughput of 130.91 Gb/s for encrypting ten IoT devices in cipher block chaining (CBC), cipher feedback (CFB), and output feedback (OFB) modes can be achieved on TSMC 40 nm. The area efficiency of this work significantly outperforms the state-of-the-art works. Wen-Long Chin, Pin-Wei Chen, Shih-Hsiang Chou, Yu-Hua Yang, Pei-Yin Chen, Tao Jiang 0002 |
IEEE Internet Things J. | 5 |
| 2024 | A Low-Cost Pipelined Architecture Based on a Hybrid Sorting AlgorithmabstractIn this paper, a low-cost pipelined architecture based on a hybrid sorting algorithm is proposed. The proposed architecture is constructed with a bitonic sorter and several cascaded bidirectional insertion sorting units. The bidirectional insertion sorting unit uses the segmented sorted subsequence generated by the bitonic sorter as input, and records the maximum and minimum values of the subsequence. After all segmented subsequences are processed through the cascaded bidirectional insertion sorting units, a sorted sequence is obtained. The proposed architecture is implemented using the Verilog hardware description language (HDL) and synthesized using the Synopsys Design Compiler with a TSMC 90-nm cell library. The experimental results indicate that the proposed architecture can not only shorten sorting cycles but also reduce hardware area costs. Moreover, sorting cycles can be further shortened by increasing the parallelism of the proposed architecture. Under the configuration that 2048 32-bit data to be sorted and 16 data have to be processed simultaneously, the proposed architecture can improve the throughput-to-gate-count ratio by 16%, and throughput-to-power-consumption-ratio by 25% compared to the existing sorting design. The proposed architecture makes the most efficient use of hardware resources. You-Rong Chen, Chien-Chia Ho, Pei-Yin Chen |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2021 | A High-Performance Bidirectional Architecture for the Quasi-Comparison-Free Sorting AlgorithmabstractThis paper proposes a high-performance bidirectional architecture for the quasi-comparison-free sorting algorithm. Our architecture improves the performance of the conventional unidirectional architecture by reducing the total number of sorting cycles via bidirectional sorting along with two auxiliary methods. Bidirectional sorting allows the sorting tasks to be conducted concurrently in the high- and low-index parts of our architecture. The first auxiliary method is boundary finding, which shortens the range for index searching by finding the boundaries of the range. The second auxiliary method is queue storing, which stores each useful index in a queue in advance to reduce the number of miss cycles during index searching. The performance of our architecture highly depends on the distribution of input data. For each set of input data to be sorted, five Gaussian distributions of the input data and four standard derivations for each distribution were adopted in our experiments. The results show that at the expense of some additional area cost, the number of sorting cycles and the energy consumption are significantly reduced by our method. Ren-Der Chen, Pei-Yin Chen, Yu-Che Hsiao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2019 | Architecture-aware Memory Access Scheduling for High-throughput Cascaded ClassifiersabstractCascaded classifier based object detectors are popular for many applications because of their high efficiency. Many researches have been devoted to developing the corresponding hardware accelerators. To reduce the circuit complexity while maintaining sufficient throughput, on-chip memories are commonly partitioned into several banks for parallel data access. However, since the coefficients of feature extraction are irregular, memory access conflict would frequently occur without proper scheduling. The proposed scheme explicitly schedules the access sequence as a post-processing for managing the coefficient memory. By formulating the desired sequence as a graph model, the classical graph coloring theory can then be adopted to solve the scheduling problem. In addition, the proposed graph model also considers the resource constraint on intermediate storage. Experimental results show that the throughput and area-efficiency of the target cascaded classifier can be greatly improved by adopting the proposed scheme as compared to the related work. Hsiang-Chih Hsiao, Chun-Wei Chen, Jonas Wang, Ming-Der Shieh, Pei-Yin Chen |
DDECS | 5 |
| 2019 | VLSI Design of an Efficient Flicker-Free Video Defogging Method for Real-Time ApplicationsabstractDefogging is an essential preprocessing technique for object detection in computer vision-based systems and has been widely used in outdoor surveillance system applications. This paper proposes an efficient defogging algorithm for both static images and videos. Considering real-time applications, the proposed defogging algorithm involves a low-cost hardware oriented design that is based on an atmospheric scattering model and dark channel prior. Compared with previous low-complexity techniques, simulation results indicated that the proposed design demonstrated superior performance in terms of quantitative and qualitative evaluations. A weighting technique and a contour preserving estimation approach are adopted alternately to refine the factors in the defogging process. Furthermore, in the atmospheric light estimation, an adjuster is applied to the video for preventing “flicker” which means that the brightness changes dramatically between two neighbor frames in the video. Hence, the proposed algorithm is suitable for video defogging applications, which have not been dealt with in previous approaches. To achieve the requirement of real-time applications for both static and dynamic images, an implementation of seven-stage very-largescale integration architecture for the proposed algorithm is presented. By using TSMC 0.13-um technology, the design yielded a processing rate of approximately 200 Mpixels/s. Yeu-Horng Shiau, Yao-Tsung Kuo, Pei-Yin Chen, Feng-Yuan Hsu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Live Demonstration: Hardware Design of Video Defogging Method for Real-Time ApplicationsabstractAn efficient flicker-free video defogging method is implemented on FPGA and presented in this live demonstration. Considering real-time applications, the proposed design involves a low-cost hardware oriented architecture on an atmospheric scattering model and dark channel prior. The synthesis results show that the design can achieve a processing rate of 200 Mpixels/s by using TSMC 0.13-um technology. Our design is demonstrated on an embedded platform by using Xilinx Spartan-6 FPGA, and the system input could be either static images or videos stream from USB camera. Pei-Yin Chen, Yao-Tsung Kuo, Shih-Hsiang Lin, Yeu-Horng Shiau |
ISCAS | 1 |
| 2018 | Efficient VLSI Architecture for Edge-Oriented DemosaickingabstractColor filter array interpolation, also known as demosaicking and “debayering,” is a crucial process for image reconstruction in digital still cameras. This paper presents an edge-oriented demosaicking method and an efficient very-large-scale integration (VLSI) architecture for color interpolation. The design uses simple operations (addition, subtraction, shift, and comparator) and nearest neighboring pixels to catch the color difference and edges. The required line buffering of the proposed design is four lines; therefore, its hardware cost is low. Our extensive experiments revealed that the proposed technique preserved edge features and exhibited excellent quantitative evaluation and visual quality performances. Compared with the previous VLSI implementations, the proposed design achieved superior image qualities. The synthesis results revealed that by using Taiwan Semiconductor Manufacturing Company 0.18-$\mu \text{m}$ technology, the proposed design yields a processing rate of approximately 200M samples per second. Chih-Yuan Lien, Fu-Jhong Yang, Pei-Yin Chen, Yi-Wen Fang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Local Binary Pattern Circuit Generator With Adjustable Parameters for Feature ExtractionabstractIn the field of computer vision, local binary pattern (LBP) is one of the most popular feature extraction method and has been used in many object detection frameworks. To efficiently extract LBP features in high-resolution images, hardware architecture is needed to disperse CPU burden and to improve the entire object detection performance. In this paper, a hardware implementation of an approximated LBP method with adjustable parameters is introduced. For simulation, Taiwan Semiconductor Manufacturing Company$0.18~\mu \text{m}$technology is used to implement the LBP hardware, and the hardware can achieve 500 MHz with lower gate count than previous study. The proposed LBP circuit is applied to the pedestrian classification application and the evaluation results show that the approximated LBP values generated by our circuit can achieve comparable classification accuracy with the primitive LBP method. Additionally, the proposed LBP hardware provides adjustable parameters to fit different applications while requires fewer hardware costs as compared with the existing work. Min-Chun Hu 0001, Kiat Siong Ng, Pei-Yin Chen, Yu-Jung Hsiao, Cheng-Hsien Li |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2017 | Smart Cage Implementation with Dependable Safety Agent for DogsabstractThis paper presents a smart cage which consists of an automated excrement detection module, a cleaning module, a feeding module, and multiple additional sensors for a dog. In the smart cage, excrement is detected by using image processing methods and is automatically cleaned up by the conveyance. Moreover, the conveyance can also be a treadmill for the dog to exercise, with its built-in and adjustable speed function which is designed for different levels of training. However, since the smart cage is automatic, it thus operates with minimal involvement from the pet owner. Therefore, the safety of the dog becomes an important issue due to no supervision. In this paper, a dependable safety agent will be proposed to solve the safety issue while the conveyance is still in use. The dependable safety agent will monitor two kinds of information and determine whether or not to stop the conveyance rotation. The first kind is the image processing technique which is used to detect the movement of the dog in the cage, with the purpose to signal for the conveyance to stop immediately if the dog moves abnormally. The second kind of information that is used to stop the conveyance, is by detecting unusual current peaks of the conveyance with the use of a current sensor. The smart cage may become safer and more reliable through using these methods to avoid any harm to the dog. Kiat Siong Ng, Pei-Yin Chen, Pi-Hui Ting |
PRDC | 2 |
| 2017 | Hardware Design of Low-Power High-Throughput Sorting UnitabstractSorting is one of the most fundamental topics in computer science. Since partial sorting with lower costs would be much more feasible than the complete sorting method for some applications, this paper focuses on the architecture sorting N values from M inputs. To meet the real-time processing requirement, hardware acceleration is commonly employed to enhance the performance. However, high-speed computation usually raises power consumption drastically.To overcome that problem, a low-power, high-throughput, and modular hardware design of partial sorting network is presented. By applying a pointer-like design, the comparing modules move the indexes of samples instead of moving the input data directly. Power dissipation is reduced by minimizing switching activities and signal transitions. To prevent unnecessary comparing of the large data set, an iterative architecture is proposed which uses both the low-power sorting module and a novel clipping mechanism. The proposed design was simulated with a 90-nm cell library and the results are compared to those of other works on hardware sorting. Experiment results show that power consumption is reduced by 64.9 percent and high-throughput performance can also be achieved. Shih-Hsiang Lin, Pei-Yin Chen, Yu-Ning Lin |
IEEE Trans. Computers | 2 |
| 2015 | A Low-Cost Hardware Architecture for Illumination Adjustment in Real-Time ApplicationsabstractFor real-time surveillance and safety applications in intelligent transportation systems, high-speed processing for image enhancement is necessary and must be considered. In this paper, we propose a fast and efficient illumination adjustment algorithm that is suitable for low-cost very large scale integration implementation. Experimental results show that the proposed method requires the least number of operations and achieves comparable visual quality as compared with previous techniques. To further meet the requirement of real-time image/video applications, the 16-stage pipelined hardware architecture of our method is implemented as an intellectual property core. Our design yields a processing rate of about 200 MHz by using TSMC 0.13-μm technology. Since it can process one pixel per clock cycle, for an image with a resolution of QSXGA (2560 × 2048) , it requires about 27 ms to process one frame that is suitable for real-time applications. In some low-cost intelligent imaging systems, the processing rate can be slowed down, and our hardware core can run at very low power consumption. Yeu-Horng Shiau, Pei-Yin Chen, Hung-Yu Yang, Shang-Yuan Li |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2014 | Weighted haze removal method with halo prevention
Yeu-Horng Shiau, Pei-Yin Chen, Hung-Yu Yang, Chao-Ho Chen, S.-S. Wang |
J. Vis. Commun. Image Represent. | 2 |
| 2014 | An Efficient Hardware Implementation of HOG Feature Extraction for Human DetectionabstractIn intelligent transportation systems, human detection is an important issue and has been widely used in many applications. Histograms of oriented gradients (HOG) are proven to be able to significantly outperform existing feature sets for human detection. In this paper, we present a low-cost high-speed hardware implementation for HOG feature extraction. The simulation shows that the proposed circuit can achieve 167 MHz with 153-K gate counts by using Taiwan Semiconductor Manufacturing Company 0.13-μm technology. Compared with the previous hardware architectures for HOG feature extraction, our circuit requires fewer hardware costs and achieves faster working speed. Pei-Yin Chen, Chien-Chuan Huang, Chih-Yuan Lien, Yu-Hsien Tsai |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2013 | An Efficient Denoising Architecture for Removal of Impulse Noise in ImagesabstractImages are often corrupted by impulse noise in the procedures of image acquisition and transmission. In this paper, we propose an efficient denoising scheme and its VLSI architecture for the removal of random-valued impulse noise. To achieve the goal of low cost, a low-complexity VLSI architecture is proposed. We employ a decision-tree-based impulse noise detector to detect the noisy pixels, and an edge-preserving filter to reconstruct the intensity values of noisy pixels. Furthermore, an adaptive technology is used to enhance the effects of removal of impulse noise. Our extensive experimental results demonstrate that the proposed technique can obtain better performances in terms of both quantitative evaluation and visual quality than the previous lower complexity methods. Moreover, the performance can be comparable to the higher,- complexity methods. The VLSI architecture of our design yields a processing rate of about 200 MHz by using TSMC 0.18 μm technology. Compared with the state-of-the-art techniques, this work can reduce memory storage by more than 99 percent. The design requires only low computational complexity and two line memory buffers. Its hardware cost is low and suitable to be applied to many real-time applications. Chih-Yuan Lien, Chien-Chuan Huang, Pei-Yin Chen, Yi-Fan Lin |
IEEE Trans. Computers | 3 |
| 2013 | Hardware Implementation of a Fast and Efficient Haze Removal MethodabstractIn this letter, a fast and efficient haze removal method is presented. We employ an extremum approximate method to extract the atmospheric light and propose a contour preserving estimation to obtain the transmission by using edge-preserving and mean filters alternately. Our method can efficiently avoid the halo artifact generated in the recovered image. To meet the requirement of real-time applications, an 11-stage pipelined hardware architecture for our haze removal method is presented. It can achieve 200 MHz with 12.8 K gate counts by using TSMC 0.13- μm technology. Simulation results indicate that our design can obtain comparable results with the least execution time compared to previous algorithms and is suitable for low-cost high-performance hardware implementation for haze removal. Yeu-Horng Shiau, Hung-Yu Yang, Pei-Yin Chen, Ya-Zhu Chuang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2012 | A Novel Interpolation Chip for Real-Time Multimedia ApplicationsabstractImage scaling is an important technique that is widely used in many image processing applications. This paper presents a novel scaling algorithm for the implementation of 2-D image scalar. The proposed interpolation method is based on the interpolation error theorem often mentioned in numerical analysis. A bilateral error-amender is used to make the interpolation more precise, and an edge-weighted scheme enhances the edge features of the scaled images. Extensive experimental results demonstrate that the proposed method can obtain better performance than previous methods in both quantitative evaluation and visual quality. This paper also presents an efficient very large-scale integrated architecture for the proposed method. The cooperation and hardware sharing techniques greatly reduce hardware cost requirements. Using a nine-stage pipeline, the proposed scaling circuit contains 13 k gate counts and yields a processing rate of approximately 278 MHz using TSMC 0.13- μm technology. The hardware cost of the proposed circuit is low, making it a good candidate for high-quality image scaling applications. Chien-Chuan Huang, Pei-Yin Chen, Ching-Hsuan Ma |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | An efficient denoising chip for the removal of impulse noiseabstractIn this paper, we present an efficient chip for the removal of impulse noise in images. Our design uses a mask on each pixel in the image to determine whether it is likely being corrupted by random-valued impulse noise or not. After noise detection, we reconstruct the noisy pixel by considering the possible edges existed in the mask and applying a median filter on it. The memory requirement of the proposed method is quite small. The experimental results demonstrate that our method achieves excellent performance in quantitative evaluation and visual quality. An efficient VLSI architecture for this scheme is developed and it yields a processing rate of about 200 MHz by using TSMC 0.18 μm technology. Chih-Yuan Lien, Pei-Yin Chen, Li-Yuan Chang, Yi-Ming Lin, Po-Kai Chang |
ISCAS | 2 |
| 2010 | A low-power IP design of Viterbi decoder with dynamic threshold settingabstractIn this paper, a low-power design of Viterbi decoder is presented. Based on the adaptive Viterbi algorithm, we use a dynamic setting method to set various threshold values for different decoding stages under a particular SNR and efficient reduce the average number of survivor paths. Furthermore, a flexible soft intellectual property core and an auxiliary software system for low-power Viterbi decoder are proposed. In the VLSI realization, we apply the clock-gating technique to disable the activation of registers for nonsurvivor paths. Hence, the power consumption can be reduced. Compared with others, our design requires the lower power consumption for the same SNR condition. Yi-Ming Lin, Wan-Ching Liu, Li-Yuan Chang, Chih-Yuan Lien, Pei-Yin Chen, Shung-Chih Chen |
ISCAS | 5 |
| 2010 | A Low-Cost VLSI Implementation for Efficient Removal of Impulse NoiseabstractImage and video signals might be corrupted by impulse noise in the process of signal acquisition and transmission. In this paper, an efficient VLSI implementation for removing impulse noise is presented. Our extensive experimental results show that the proposed technique preserves the edge features and obtains excellent performances in terms of quantitative evaluation and visual quality. The design requires only low computational complexity and two line memory buffers. Its hardware cost is quite low. Compared with previous VLSI implementations, our design achieves better image quality with less hardware cost. Synthesis results show that the proposed design yields a processing rate of about 167 M samples/second by using TSMC 0.18 ¿m technology. Pei-Yin Chen, Chih-Yuan Lien, Hsu-Ming Chuang |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2009 | A Hybrid Image Restoring Algorithm for Interlaced VideoabstractIn this letter, a novel image restoring technique for interlaced video is presented. We employ a motion-degree detector to determine the type of motion in the adjacent pictures efficiently and a hybrid interpolation based on spatial and/or temporal information to restore the missing pixels for each motion type. Extensive experimental results demonstrate that our method can obtain better performances in terms of both quantitative evaluation and visual quality than the state-of-the-art image restoring techniques. Since the proposed method requires low computational complexity, it is very suitable for real-time hardware implementation and can be applied to current television systems. Chih-Yuan Lien, Chung-Ping Young, Pei-Yin Chen |
IEEE Signal Process. Lett. | 3 |
| 2009 | A Collaborative sensor-fault detection scheme for robust distributed estimation in sensor networksabstractThis work addresses the problem of robust distributed estimation in the presence of sensor faults when the fusion center sequentially receives quantized messages from local sensors. The mean square error (MSE) of distributed estimation schemes increases dramatically if the information received from the faulty sensors within the network is not excluded from the estimation process. Accordingly, an efficient collaborative sensor-fault detection (CSFD) scheme is proposed in which the results of a homogeneity test are used to identify the faulty nodes within the network such that their quantized messages can be filtered out when estimating the parameter of interest. Utilizing an asymptotic analytical technique, a lower bound is derived for the MSE of the proposed distributed estimation scheme. A good agreement is observed between the simulated MSE results and the lower bound values, and thus it is inferred that the lower bound provides a convenient and reliable means of predicting the performance of the proposed estimation scheme in real-world sensor networks. In addition, a low-complexity CSFD (LC-CSFD) scheme is proposed to identify faulty sensors in WSNs with a very large number of nodes. The simulation results confirm that the accuracy of the estimates obtained from the CSFD and LCCSFD schemes is significantly better than that obtained from a conventional estimation scheme when applied in sensor networks characterized by an unknown number of sensor faults of various types. Tsang-Yi Wang, Li-Yuan Chang, Pei-Yin Chen |
IEEE Trans. Commun. | 3 |
| 2009 | VLSI Implementation of an Edge-Oriented Image Scaling ProcessorabstractImage scaling is a very important technique and has been widely used in many image processing applications. In this paper, we present an edge-oriented area-pixel scaling processor. To achieve the goal of low cost, the area-pixel scaling technique is implemented with a low-complexity VLSI architecture in our design. A simple edge catching technique is adopted to preserve the image edge features effectively so as to achieve better image quality. Compared with the previous low-complexity techniques, our method performs better in terms of both quantitative evaluation and visual quality. The seven-stage VLSI architecture of our image scaling processor contains 10.4-K gate counts and yields a processing rate of about 200 MHz by using TSMC 0.18-mum technology. Pei-Yin Chen, Chih-Yuan Lien, Chi-Pin Lu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2008 | A real-time image denoising chipabstractIn this paper, we present a real-time image denoising chip. For each pixel of the image under processing, our design uses a mask on it to determine whether it is likely being corrupted by impulse noise or not. After the noise detection, we reconstruct the noisy pixel by considering the possible edges existed in the mask. Particularly, our design removes the noise from corrupted images efficiently and requires no previous training. Extensive simulations demonstrate that the proposed method achieves excellent performance in quantitative evaluation and visual quality. Furthermore, the computational complexity of the proposed method is low and its memory requirement is small. An efficient VLSI architecture for this scheme is developed and it yields a processing rate of about 150 MHz by using TSMC 0.18mum technology. Pei-Yin Chen, Chih-Yuan Lien, Yi-Ming Lin |
ISCAS | 1 |
| 2008 | An Efficient Edge-Preserving Algorithm for Removal of Salt-and-Pepper NoiseabstractIn this letter, a novel algorithm for removing salt-and-pepper noise from corrupted images is presented. We employ an efficient impulse noise detector to detect the noisy pixels, and an edge-preserving filter to reconstruct the intensity values of noisy pixels. Extensive experimental results demonstrate that our method can obtain better performances in terms of both subjective and objective evaluations than those state-of-the-art impulse denoising techniques. Especially, the proposed method can preserve edges very well while removing impulse noise. Since our algorithm is algorithmically simple, it is very suitable to be applied to many real-time applications. Pei-Yin Chen, Chih-Yuan Lien |
IEEE Signal Process. Lett. | 1 |
| 2008 | An Efficient Design of Variable Length Decoder for MPEG-1/2/4abstractIn this paper, a novel and area-efficient variable length decoder (VLD) for MPEG-1/2/4 is presented. Instead of carrying out every variable length coding table with one dedicated lookup table (LUT) directly, we employ an efficient clustering-merging technique to reduce both the size of a single LUT and the total number of LUTs required for MPEG-1/2/4. Synthesis results show that our VLD occupies 10666 gate counts and operates at 125 MHz by using the standard cell from Artisan TSMC's 0.18 mum process. As demonstrated, the proposed design outperforms other VLDs with less hardware cost. It can decode a symbol of different standards in every cycle and support video resolution of HD1080 at 30 frames/s for MPEG-1/2/4 real-time decoding. Pei-Yin Chen, Yi-Ming Lin, Min-Yi Cho |
IEEE Trans. Multim. | 1 |
| 2004 | VLSI Implementation for One-Dimensional Multilevel Lifting-Based Wavelet TransformabstractThe lifting scheme has been developed as a flexible tool suitable for constructing biorthogonal wavelets recently. We present an efficient VLSI architecture for the implementation of 1D lifting discrete wavelet transform. The architecture folds the computations of all resolution levels into the same low-pass and high-pass units to achieve higher hardware utilization. Because of its modular, regular, and flexible structure, the design is scalable for different resolution levels. In addition, its area is independent of the length of the 1D input sequence and its latency is independent of the number of resolution levels. Since the architecture has a similar topology to a scan chain, we can modify it easily to become a testable scan-based design by adding very few hardware resources. For the computations of N-sample 1D k-level analysis (5, 3) lifting wavelet transform, the design takes N+1 clock cycles, and requires two multipliers, four adders, and (3 + 2.25 /spl times/ 2/sup k/) registers. In the simulation, it works with a clock period of 10 ns and achieves a processing rate of about 100 /spl times/ 10/sup 6/ samples/sec for k-level lifting wavelet transform. Pei-Yin Chen |
IEEE Trans. Computers | 1 |
| 2004 | An efficient prediction algorithm for image vector quantizationabstractIn this paper, we propose a novel efficient fuzzy prediction algorithm for image vector quantization by using the fact that the pixel values of a block is highly related to those of its adjacent blocks in the same image frame because of its spatial correlation. The experimental results show that our method performs better than other VQ algorithms, such as SOC-VQ, LCIC-VQ, STC-VQ, and DST-VQ, in terms of computational complexity and coding efficiency. Pei-Yin Chen |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2002 | A fuzzy search block-matching chip for motion estimation
Pei-Yin Chen |
Integr. | 1 |
| 2001 | An efficient gray search algorithm for the estimation of motion vectorsabstractMotion vector estimation plays an important role in motion-compensated video coding. An efficient and fast search algorithm is proposed for the estimation of motion vectors. With the help of gray prediction, the algorithm can determine the motion vectors of image blocks quickly and correctly. Since the proposed algorithm performs better than other search algorithms [e.g. the three-step search (TSS), cross-search (CS), new three-step search (NTSS), four-step search (FSS), block-based gradient descent search (BBGDS), simple-and-efficient search (SES), prediction search (PS) and gray prediction search (GPS)], it is very beneficial in applications where the video coding speed is important. Jau-Ling Chen, Pei-Yin Chen |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2001 | An efficient blocking-matching algorithm based on fuzzy reasoningabstractDue to the temporal and spatial correlation of image sequence, the motion vector of a reference block is highly related to the motion vectors of its adjacent blocks in the same image frame. By using that idea, we propose a novel efficient fuzzy search (EFS) algorithm for block motion estimation. The experimental results show that the EFS performs better than other fast search algorithms, such as TSS, CS, NTSS, FSS, BBGDS, SES, and PSA in terms of picture quality, accuracy, computational complexity, and coding efficiency. Pei-Yin Chen, Jer-Min Jou |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2000 | A fast-search motion estimation methodabstractAn efficient fast-search algorithm is proposed. With the help of fuzzy reasoning, the algorithm can determine the motion vectors of image blocks quickly and correctly. According to the experimental results, our algorithm performs better than other search algorithms, such as 3SS, NTSS, 4SS, SES, PS, and GPS in terms of five different measures. Yun-Teng Roan, Pei-Yin Chen |
SMC | 2 |
| 2000 | Adaptive arithmetic coding using fuzzy reasoning and grey prediction
Pei-Yin Chen, Jer-Min Jou |
Fuzzy Sets Syst. | 1 |
| 2000 | An adaptive fuzzy logic controller: its VLSI architecture and applicationsabstractMost previous work about the hardware design of a fuzzy logic controller (FLC) intended to either improve its inference performance for real-time applications or to reduce its hardware cost. To our knowledge, there has been no attempt to design a hardware FLC that can perform an adaptive fuzzy inference for the applications of on-line adaptation. The purpose of this paper is to present such an adaptive memory-efficient FLC and its applications. Taking advantage of the adaptability provided by a symbolic fuzzy rule format and the dynamic membership function generator, as well as the high-speed integration capability afforded by VLSI, the proposed adaptive fuzzy logic controller (AFLC) can perform an adaptive fuzzy inference process using various inference parameters, such as the shape and location of a membership function, dynamically and quickly. Three examples are used to illustrate its applications, and the experimental results show the excellent adaptability provided by AFLC. Jer-Min Jou, Pei-Yin Chen, Sheng-Fu Yang |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1999 | A Scalable Pipelined Architecture for Separable 2-D Discrete Wavelet TransformabstractThis paper presents a highly scalable efficient architecture for separable 2-D Discrete Wavelet Transform (DWT) which is simple, regular, modular and pipelined for the computation of 2-D DWT. With these properties, it is easily scalable for different filter lengths and different octave levels. In addition, the architecture has the characteristics of lower hardware cost, shorter latency, and higher throughput rate. Jer-Min Jou, Pei-Yin Chen, Yeu-Horng Shiau, Ming-Shiang Liang |
ASP-DAC | 2 |
| 1999 | A fast and efficient lossless data-compression methodabstractThis paper describes an online lossless data-compression method using adaptive arithmetic coding. To achieve good compression efficiency, we employ an adaptive fuzzy-tuning modeler that applies fuzzy inference to deal efficiently with the problem of conditional probability estimation. In comparison with other lossless coding schemes, the compression results of the proposed method are good and satisfactory for various types of source data, Since we adopt the table-lookup approach for the fuzzy-tuning modeler, the design is simple, fast, and suitable for VLSI implementation. Jer-Min Jou, Pei-Yin Chen |
IEEE Trans. Commun. | 2 |
| 1999 | The gray prediction search algorithm for block motion estimationabstractDue to the temporal and spatial correlation of the image sequence, the motion vector of a block is highly related to the motion vectors of its adjacent blocks in the same image frame. If we can obtain useful and enough information from the adjacent motion vectors, the total number of search points used to find the motion vector of the block may be reduced significantly. Using that idea, an efficient gray prediction search (GPS) algorithm for block motion estimation is proposed in this paper. Based on the gray system theory, the GPS can determine the motion vectors of image blocks quickly and correctly. The experimental results show that the proposed algorithm performs better than other search algorithms, such as 3SS, CS, PHODS, 4SS, BBGDS, SES and PSA, in terms of six different measures: (1) average mean square error per pixel; (2) average peak signal-to-noise ratio; (3) average prediction errors per pixel; (4) average entropy of prediction errors; (5) average percentage of unpredictable pels per frame; and (6) average search points per block. Jer-Min Jou, Pei-Yin Chen, Jian-Ming Sun |
IEEE Trans. Circuits Syst. Video Technol. | 2 |