VLDB 2026 Research / reviewers in the wild / expert
Jiun-In Guo
dblp:33/3059
· DBLP profile ↗
65ranked-venue papers
8as first author
6since 2021 · last 2025
0000-0003-0402-2621ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 32 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 31 · 5 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Methodology and Results of MIC-OPCC: Multi-Indexed Convolution Model for Octree Point Cloud Compression
Gerald Baulig, Jiun-In Guo |
CGI (2) | 2 |
| 2025 | MIC-OPCC: Multi-Indexed Convolution Model for Octree Point Cloud CompressionabstractMulti-Indexed Convolution introduces an alternative approach to spatial feature extraction, which we use in our entropy model to compress the occupation symbols of octree-encoded point clouds. This method offers reduced time and memory usage per point compared to related work. https://github.com/bugerry87/mic-opcc Gerald Baulig, Jiun-In Guo |
DCC | 2 |
| 2023 | Multi-Scale Dynamic Fixed-Point Quantization and Training for Deep Neural NetworksabstractState-of-the-art deep neural networks often require extremely high computational power which results in the deployment of deep neural networks on embedded devices being impractical. Therefore, model quantization is important for the deployment of deep neural networks on edge devices. The purpose of this paper is to quantize the deep neural networks from high-precision to low-precision (e.g. INT8) dynamic fixed-point format at the layer-by-layer level quantization. In addition, we further improve the uniform dynamic fixed-point quantization to multi-scale dynamic fixed-point quantization for lower quantization loss. The proposed multi-scale dynamic fixed-point quantization scheme divides the quantization ranges into two regions, and each region is assigned different quantization levels and quantization parameters to better approximate the bell-shaped distributions. The proposed quantization pipeline is composed of post-training quantization followed by model fine-tuning which can keep the accuracy drop of the quantized model within 1% mean average precision (mAP). Furthermore, the proposed quantization and fine-tuning method can be combined with model pruning to obtain a compact and accurate deep neural network with low bit-width. Po-Yuan Chen, Hung-Che Lin, Jiun-In Guo |
ISCAS | 3 |
| 2023 | Summary of the 2023 PAIR-LITEON Competition: Embedded AI Object Detection Model Design Contest on Fish-eye Around-view CamerasabstractThis competition is dedicated to achieving fisheye object detection in Asia, particularly in countries like Taiwan, while emphasizing low power consumption and simultaneously achieving a high mean average precision (mAP). This task is notably challenging as it must be accomplished in adverse driving conditions. The objects targeted for detection include cars, pedestrians, motorcycles, and bicycles. To train their models, participants utilized 89,002 annotated training images from the iVS-Dataset [1] and conducted testing on the MemryX platform [2]. To excel in this competition, participants had to master the art of transforming standard images into fisheye images. The judging process involved 6,500 test images, with 1,500 used in the preliminary competition stage, and the rest reserved for the final competition stage. A total of 129 teams registered for this competition, and those with mAP scores exceeding 20% advanced to the final competition stage, where 16 teams are qualified. Out of these, 11 teams submitted their works based on the final competition accuracy, which could not be lower than 5% of the preliminary competition accuracy. Ultimately, five teams attained their final scores and competed for rankings based on paper reviews. Champion is chici_lab, securing the top position in this demanding competition. NCKU_ACVLab, the 1st Runner-up, demonstrated outstanding skills. The 2nd Runner-up, yuhsi44165, also showcased commendable performance. Special Awards recognized excellence in specific categories, with chici_lab sweeping all three accolades. They were bestowed the best pedestrian detection award, the best bicycle detection award, and the best motorbike detection award for their remarkable achievements. Yu-Shu Ni, Chia-Chi Tsai, Jyun-Syu Lin, Hsien-Po Meng, Po-Chi Hu, Jiun-Shiung Chen, Kun-Hung Lin, Chih-Yuan Chuang, Jiun-In Guo |
MMAsia | 9 |
| 2022 | IVS-Caffe - Hardware-Oriented Neural Network Model DevelopmentabstractThis article proposes a hardware-oriented neural network development tool, called Intelligent Vision System Lab (IVS)-Caffe. IVS-Caffe can simulate the hardware behavior of convolution neural network inference calculation. It can quantize weights, input, and output features of convolutional neural network (CNN) and simulate the behavior of multipliers and accumulators calculation to achieve the bit-accurate result. Furthermore, it can test the accuracy of the chosen CNN hardware accelerator. Besides, this article proposes an algorithm to solve the deviation of gradient backpropagation in the bit-accurate quantized multipliers and accumulators. This allows the training of a bit-accurate model and further increases the accuracy of the CNN model at user-designed bit width. The proposed tool takes Faster region based CNN (R-CNN) + Matthew D. Zeiler and Rob Fergus (ZF)-Net, Single Shot MultiBox Detector (SSD) + VGG, SSD + MobileNet, and Tiny you only look once (YOLO) v2 as the experimental models. These models include both one-stage object detection and two-stage object detection models, and base networks include the convolution layer, the fully connected layer, and the modern advanced layers, such as the inception module and depthwise separable convolution. In these experiments, direct quantization of layer-I/O fixed-point models to bit-accurate models will have a 2% mean average precision (mAP) drop of accuracy in the constraint that all layers' accumulators and multipliers are quantized to less or equal to 14 and 12 bit, respectively. After retraining of these quantized models with the proposed IVS-Caffe, we can achieve less than 1% mAP drop in accuracy in the constraint that all layers' accumulators and multipliers are quantized to less or equal to 14 and 11 bit, respectively. With the proposed IVS-Caffe, we can analyze the accuracy of the target model when it is running at hardware accelerators with different bit widths, which is beneficial to fine-tune the target model or customize the hardware accelerators with lower power consumption. Code is available at https://github.com/apple35932003/IVS-Caffe. Chia-Chi Tsai, Jiun-In Guo |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Summary of the 2021 Embedded Deep Learning Object Detection Model Compression Competition for Traffic in Asian CountriesabstractThe 2021 embedded deep learning object detection model compression competition for traffic in Asian countries held in IEEE ICMR2021 Grand Challenges focuses on the object detection technologies in autonomous driving scenarios. The competition aims to detect objects in traffic with low complexity and small model size in the Asia countries (e.g., Taiwan), which contains several harsh driving environments. The target detected objects include vehicles, pedestrians, bicycles and crowded scooters. There are 89,002 annotated images provided for model training and 1,000 images for validation. Additional 5,400 testing images are used in the contest evaluation process, in which 2,700 of them are used in the qualification stage competition, and the rest are used in the final stage competition. There are in total 308 registered teams joining this competition this year, and the top 15 teams with the highest detection accuracy entering the final stage competition, from which 9 teams submitted the final results. The overall best model belongs to team "as798792", followed by team "Deep Learner" and team "UCBH." Two special awards of best accuracy award best and bicycle detections go to the same team "as798792," and the other special award of scooter detection goes to team "abcda." Yu-Shu Ni, Chia-Chi Tsai, Jiun-In Guo, Jenq-Neng Hwang, Bo-Xun Wu, Po-Chi Hu, Ted T. Kuo, Hsien-Kai Kuo |
ICMR | 3 |
| 2019 | Recognizing Chinese Texts with 3D Convolutional Neural NetworkabstractIn this paper, we propose a deep learning system to localize and recognize Chinese texts in scenes with signage and road marks through 3D convolutional neural network. The proposed system adopts YOLO for detecting target location and exploits 3D convolutional neural network for recognizing the contents. The proposed design outperforms the existing designs based on LSTM and achieves real-time processing performance, which is feasible to be implemented on embedded platforms. The proposed system reaches over 90% accuracy in recognizing Chinese texts on bird's-eye viewing road marks in a self-driving vehicle equipped with a fisheye camera. In addition, this system can achieve 20 fps execution speed with NVIDIA DIGITS DevBox with 1080Ti GPU, which is fast enough for autonomous driving applications. Kuan-Chou Chen, Guan-Ting Lin, Che-Tsung Lin, Jiun-In Guo |
ICIP | 4 |
| 2019 | Using C3D to Detect Rear Overtaking BehaviorabstractAvoiding traffic accidents is critical since the death of traffic accidents is the eighth among the top ten leading causes of death in 2018. This paper proposes a light-weight convolutional 3D (C3D) network with five 3D convolution layers and two fully-connected layers to predict overtaking behavior. This network utilizes the last layer of convolution layer to learn the overtaking object location in the final frame. Based on NVIDIA Jetson TX2, the proposed C3D network achieves 91.46% accuracy to detect overtaking behavior on rainy days. To generate this excellent deep learning model, we use an efficient labeling tool, called ezLabel, which is a free SaaS for academia group with 96,000 opened image data samples for deep learning. ezLabel owns outstanding route prediction and fitting functions, which speeds up with the factor of ten compared to traditional tools. Users only label the object in its first frame and in its final frame, and then ezLabel labels the object in all frames in between and fits the bounding box to the object. The ezLabel can be used to label objects captured with any moving or static cameras efficiently. Ching-Kai Tseng, Chien-Chih Liao, Po-Chun Shen, Jiun-In Guo |
ICIP | 4 |
| 2019 | Summary Embedded Deep Learning Object Detection Model CompetitionabstractThe embedded deep learning object detection model competition in IEEE MMSP2019 focuses on the object detection for sensing technology in autonomous driving vehicles, which aims at detecting small objects in worse conditions through embedded systems. We provide a dataset with 89,002 annotated images for training and 1,500 annotated images for validation. We test participants' models through 6,000 testing images, which are separated into 3,000 for qualification and 3,000 for finals. There are 87 teams of participants registered this competition and 14 teams submitted the team composition. At last there are nine teams entering the final competition and five teams submitting their final models that can be realized in NVIDIA Jetson TX-2. At the end, only one team's model passed the target accuracy requirement for grading and became the champion of the contest, which the winner is team R.JD. Jiun-In Guo, Chia-Chi Tsai, Yong-Hsiang Yang, Hung-Wei Lin, Bo-Xun Wu, Ted T. Kuo, Li-Jen Wang |
MMSP | 1 |
| 2018 | A Blind Spot Detection Warning System based on Gabor Filtering and Optical Flow for E-mirror ApplicationsabstractBlind Spot Detection (BSD) is an important technique for ADAS. We propose a BSD algorithm using Gabor filtering and optical flow to detect vehicles in the blind spot region for both day-time and night-time applications. For the day-time scene, the Gabor filtering is used to detect the vehicles, inside lane line, and outside lane line. After detection, the optical flow information calculated according to Horn-Schunck method is used to judge the motion of the vehicle candidates and filter the mistake-judgement. For the night-time scene, we try to find the head-light of the approaching cars. First, we perform binarization on the image first, find the center of gravity of the light-area, classify the light-area into 2 groups and judge it as a vehicle or not. The proposed BSD system achieves 93.58% recall and 95.83% precision in day time scene and 90.22% recall and 92.76% precision in night time scene. The algorithm can achieve performance of 89 fps on Intel Core I7 and 50 fps on Renesas R-Car M2 under 640×480 resolution. Shun-Min Chang, Chia-Chi Tsai, Jiun-In Guo |
ISCAS | 3 |
| 2016 | A variable-voltage low-power technique for digital circuit systemabstractA swing variable voltage technique (CK-Vdd) is proposed to reduce power consume for generic digital circuit system. The proposed CK-Vdd generates a swing variable voltage, which is different from the conventional constant voltage (Vdd) to the digital circuit. The swing voltage is produced from using Voltage Frequency Adjustor (VFA) and Frequency Duty-Cycle Adjustor (FDCA) circuits. The clock rising and falling signals fanin FDCA to generate an adjustable high-low signal to control VFA generates high-low cycling swing voltage. When the clock is at positive-level, a generic positive-edge digital circuit will need large operation current. CK-Vdd supply high-voltage to the digital circuit at this time. On the other hand, when the clock signal transfers to the low-level, CK-Vdd can supply low-voltage to reduce power consumption. From reducing the supply current to the digital circuit at low-level clock, the digital circuit power consumption can be reduced. We implement the CK-Vdd technique in a H.264 video decoder test chip based on TSMC 90 nm CMOS process. The result shows that when CK-Vdd voltage is 0.7v ~ 0.9v it can save average 32% power consumption. To the maximum, decoder chip can save as high as 45% power consumption. An-Tai Xiao, Yung-Siang Miao, Ching-Hwa Cheng, Jiun-In Guo |
ASP-DAC | 4 |
| 2016 | Algorithm derivation and its embedded system realization of speed limit detection for multiple countriesabstractThis paper proposes a low-complexity speed limit detection which can not only support different types of speed limit signs for multiple countries, but maintain good detection rate under inclement weathers. The proposed algorithm steps include shape detection to locate speed limit signs, adaptive threshold and digit recognition algorithm which can adapt to different digit fonts to support different types of speed limit signs for multiple countries. We implement the proposed algorithm on both the desktop and the automotive-grade i.MX 6 embedded system. Under D1 resolution (720×480), the proposed system can achieve 150 fps on the desktop and 30 fps on the i.MX 6 embedded system. Ting Chou, M. S. Vinay, Jiun-In Guo |
ISCAS | 4 |
| 2015 | A wireless panoramic endoscope system design and implementation for minimally invasive surgeryabstractMinimally Invasive Surgery (MIS) is a current major surgery technique. A chief problem with MIS is its narrow field of vision. A wireless MIS Panoramic Endoscope (WMISPE) is developed and implemented to provide doctors with broad fields of view. A WMISPE features a combination of video overlapping and video stitching. The panoramic video apparatus has two side-by-side endoscopic lenses that provide wide-angle inputs for video stitching. A WMISPE can provide doctors with panoramic videos, so that they can easily discriminate an organ's position between surgical operations. Experimental results show that WMISPE can enhance the video size to 158%. A low-power multi-mode video decoder (MMVD) is used to validate the stitching videos. The whole system is validated by personal computer, an embedded system, and an H.264 decoder chip. The test chip is successfully validated after system integration, and obtains about a 24% reduction in power consumption, which is better than that of the same design using a single supply voltage. The animal vivo experimental videos have been successfully validated. Ching-Hwa Cheng, Sheng-Ping Hung, Jiun-In Guo, Kai-Che Liu, Jungle Chi-Hsiang Wu |
ISCAS | 3 |
| 2013 | A low-cost scalable Voltage-Frequency Adjustor for implementing low-power systemsabstractThe proposed low-cost, scalable, synthesizable Voltage-Frequency Adjustor (VFA) design provides voltage and frequency automatic adjustments for implementing low-power systems. This design is effective in reducing power consumption from managing the supplied current and clock frequency for the structurally dynamic varying systems. The VFA supplied electrical current amount and voltage level are automatically adjusted from sensing the voltage drop of the fanout load. This adjustment mechanism controls the activated power switch quantity to allow a structure varied design that can obtain a stable fanin voltage level, regardless of the designed module that is dynamically added or removed from the system. Compared to the conventional dynamic voltage frequency scaling technique, the VFA is cost effective and can be easily integrated with the target design as a single chip. The experimental results show that the combination of VFA chips and an H.264 video decoder is successfully validated with 49%~68% power reduction compared to the case that the H.264 video decoder directly uses an outer power supply. Ching-Hwa Cheng, Sheng-Wei Hsu, Jiun-In Guo |
ISCAS | 3 |
| 2012 | A two level mode decision algorithm for H.264 high profile intra encodingabstractA two level mode decision algorithm considering edge information for H.264 high profile intra encoding is proposed to reduce the computational complexity. In level one mode decision, the edge information is used to decide the appropriate intra coding block size according to the observations that a smooth macroblock usually chooses the large intra coding block size, and that a complex macroblock usually chooses the small one. In level two mode decision, we use the correlation between adjacent blocks to decide the appropriate prediction modes. The proposed method achieves 34%~59% reduction of computational complexity when compared to JM14 and the PSNR only drops 0.09 db in average. Cheng-Yen Chang, Cheng-An Chien, Hsiu-Cheng Chang, Jiun-In Guo |
ISCAS | 5 |
| 2012 | 3D depth map generation for embedded stereo applicationsabstractThis paper presents a low complexity 3D image depth map generation algorithm for embedded stereo applications. The proposed algorithm generates depth information based on a single view 2D image automatically. Owing to different scene characteristics of image, we propose a mechanism to classify images to “Scenery”, “Normal” and “Close-up” types first and generate the associated depth map according to the proposed techniques. In addition, we propose a human detection method for strengthening the depth information in images with humans and post-processing for refining depth map. With good quality in the generated depth map, the proposed algorithm achieves about 93% in complexity reduction as compared to the traditional algorithm, which is suitable for realization in both the hardware and embedded systems for portable stereo applications. Jui-Sheng Lee, Guo-An Jian, Cheng-An Chien, Peng-Sheng Chen, Jiun-In Guo |
VCIP | 5 |
| 2011 | A H.264/MPEG-2 dual mode video decoder chip supporting temporal/spatial scalable videoabstractThis paper proposes a dual mode video decoder with 4-level temporal/spatial scalability and 32/64-bit adjustable memory bus width. A design automation environment for simulation and verification is established to automatically verify the correctness and completeness of the proposed design. Using a 0.13 um CMOS technology, it comprises 439Kgates/10.9KB SRAM and consumes 2~328mW in decoding CIF~HD1080 videos at 3.75~30fps when operating at 1~150MHz, respectively. Cheng-An Chien, Yao-Chang Yang, Hsiu-Cheng Chang, Cheng-Yen Chang, Jiun-In Guo, Jinn-Shyan Wang, Ching-Hwa Cheng |
ASP-DAC | 6 |
| 2011 | Dual-phase pipeline circuit design automation with a built-in performance adjusting mechanismabstractThe high speed dual phase operation domino circuit, which includes high-performance and reliable characteristics is proposed, and the circuit design technique with practical implementation is presented. The cell-based automatic synthesis flow supports the quick design of high performance chips. The test chip of a dual-phase 64 bit high-speed multiplier with a built-in performance adjustment mechanism is successfully validated using TSMC 0.18 technology. The test chip shows ×2.7 performance improvement compared to the conventional static CMOS logic design. Yu-Tzu Tsai, Cheng-Chih Tsai, Cheng-An Chien, Ching-Hwa Cheng, Jiun-In Guo |
ASP-DAC | 5 |
| 2011 | A low-power management technique for high-performance domino circuitsabstractExploiting a charge sharing method enables a performance power management design for domino circuits. The domino circuits have both high performance and low power consumption. A test chip has been successfully validated using TSMC 0.13um CMOS technology. Reductions in dynamic power consumption of 68% and static power consumption of 15% are achieved. Yu-Tzu Tsai, Cheng-Chih Tsai, Cheng-An Chien, Ching-Hwa Cheng, Jiun-In Guo |
ASP-DAC | 5 |
| 2010 | Dynamic voltage domain assignment technique for low power performance manageable cell based designabstractMulti-voltage technique is an effective way to reduce power consumption. In the proposed voltage domain programmable (VDP) technique, high and low voltage domains applied to logic gates are programmable. The different voltage domains allow the chip performance and power consumption to be flexibly adjusted during circuit operation. In this proposed internal of the chip technique, the power switches possess the feature of flexible programming after chip manufacturing. The video decoder test chip proof of this novel methodology has 55% power reduction with good power-performance management mechanism. Elone Lee, Feng-Tso Chien, Ching-Hwa Cheng, Jiun-In Guo |
ASP-DAC | 4 |
| 2010 | A remote thin client system for real time multimedia streaming over VNCabstractThis paper proposes a remote thin client system for real time multimedia streaming over VNC. A remote frame can be split as two parts, i.e. high motion part and low motion part, and transmitted through the Internet from servers to clients according to the proposed hybrid RTP protocol. A Dynamic Image Detection Scheme (DIDS) is proposed to automatically detecting the high motion part of a frame with only 1% of extra CPU loading. In addition, an Error Detection Scheme (EDS) and a Dynamic Bit-rate Control Scheme (DBCS) are also proposed to ensure good video streaming quality under bandwidth limited applications. This paper also proposes a linear time BU-level rate control algorithm to ensure the proposed DBCS can be finished in real time. The proposed algorithm reduces the computational complexity from O(n2)to be O(n). By using the proposed thin client system, we can achieve about 22 fps of real-time SIF video streaming with good video quality under 32 KByte/s of bandwidth limitation, which speeds up about 172 times in remote frame display when compared to pure VNC. Kheng-Joo Tan, Jia-Wei Gong, Bing-Tsung Wu, Dou-Cheng Chang, Hsin-Yi Li, Yi-Mao Hsiao, Yung-Chung Chen, Shi-Wu Lo, Yuan-Sun Chu, Jiun-In Guo |
ICME | 10 |
| 2010 | Low complexity fractional motion estimation with adaptive mode selection for H.264/AVCabstractFractional motion estimation (FME) searches subpixels for blocks of various sizes to find out the best matching candidate, which improves the compression efficiency but leads to high computational complexity. In this paper, we propose a low complexity FME design supporting adaptive mode selection (AMS). Exploiting both the mode correlation among neighboring macroblocks and integer motion estimation result, we can select FME modes adaptively to reduce complexity with good video quality. The simulation result shows that the average number of processing modes in FME is reduced to be 2 for HD1080 video instead of processing 7 modes in JM reference software. With tiny PSNR drop (~0.026dB), the proposed design achieves 71% reduction in computational complexity when compared to JM14.2. At operating frequency of 300MHz, the proposed design could support the real-time processing for H.264 videos with resolution up to 4K × 2K. Chih-Chuan Yang, Kheng-Joo Tan, Yao-Chang Yang, Jiun-In Guo |
ICME | 4 |
| 2010 | A group of macroblock based motion estimation algorithm supporting adaptive search range for H.264 video codingabstractThis paper proposes a group of macroblock (GOMB) based motion estimation (ME) algorithm supporting adaptive search range (ASR) for H.264 video coding. Adopting the concept of GOMB to generate the Predicted Motion Vector (PMV) for doing H.264 ME achieves good results in balancing the computational complexity, video quality and memory bandwidth. Moreover, it supports adaptive search range (ASR) in doing ME to greatly reduce both the computational complexity and memory bandwidth while maintaining good video quality. In HD1080 resolution, compared to the JM full search block matching algorithm (FSBMA) centered at (0, 0) and JM PMV-based block matching algorithm (BMA), the proposed algorithm respectively reduces about 93% and 81% memory bandwidth, as well as 99% and 76% computational complexity with negligible PSNR drop. Chang-Hung Tsai, Kheng-Joo Tan, Ching-Lung Su, Jiun-In Guo |
ISCAS | 4 |
| 2009 | A dynamic quality-scalable H.264 video encoder chipabstractThis paper proposes a dynamic quality-scalable H.264 video encoder that comprises 470Kgates and 13.3Kbytes SRAM using 1P8M 0.13μm CMOS technology. Exploiting parameterized algorithms for motion estimation and intra prediction, the proposed design can dynamically configure the encoding modes with the design trade-off between power consumption and video quality for various video encoding applications. It achieves real-time H.264 video encoding on CIF, D1, and HD720@30fps with 7mW-25mW, 27mW-162mW, and 122mW–183mW power dissipation in different quality modes. Hsiu-Cheng Chang, Yao-Chang Yang, Ching-Lung Su, Cheng-An Chien, Jiun-In Guo, Jinn-Shyan Wang |
ASP-DAC | 6 |
| 2009 | CKVdd: a self-stabilization ramp-vdd technique for dynamic power reductionabstractWe propose a self-stabilized ramp voltage technique, CKVdd, to reduce power dissipation in conventional CMOS circuit. Normal CMOS circuits show a power increase proportional to clock frequency. CKVdd results in a lower-than-usual power increase. This technique is easily implemented in CMOS circuits. CKVdd technique possesses several characteristics that differ from of the current circuits using Vdd power source. First, CKVdd circuits have less average current and peak current consumption, such that it can be a low power design technique applied to generic digital circuits. Second, CKVdd technique combines the power source and clock signal, and can easily implement the power management mechanism. Compared to constant Vdd for multimedia decoders, the proposed technique has 45% of the usual power dissipation and 88% of the usual peak current reduction at the cost of small delay penalty. Chin-Hsien Wang, Ching-Hwa Cheng, Jiun-In Guo |
ASP-DAC | 3 |
| 2009 | A Dynamic Quality-scalable H.264 Video EncoderabstractThis demo proposes a dynamic quality-scalable H.264 video encoder that has been published in ISSCC2007 [8]. Exploiting parameterized algorithms for motion estimation and intra prediction, the proposed design can dynamically configure the encoding modes with the design trade-off between power consumption and video quality for various video applications. Hsiu-Cheng Chang, Yao-Chang Yang, Cheng-An Chien, Tzu-Chun Chang, Jinn-Shyan Wang, Jiun-In Guo |
ISCAS | 7 |
| 2009 | A Multi-standard Video Decoder for High Definition Video ApplicationsabstractThrough reducing 70% of external memory bandwidth and 60% of computational complexity, the proposed 252 Kgates/71 mW/0.13 um multi-standard (JPEG/MPEG-1/2/4/H.264) video decoder reduces 72% in gate count and 87% in power consumption as compared to the state-of-the-art design, when operating at 120 MHz for real-time HD1080 video decoding with single AHB-based SDR memory. Cheng-An Chien, Chih-Da Chien, Jui-Chin Chu, Jiun-In Guo, Ching-Hwa Cheng |
ISCAS | 4 |
| 2009 | A High Throughput Deblocking Filter Design Supporting Multiple Video Coding StandardsabstractThis paper presents a high throughput, VLSI architecture for multi-standard in-loop deblocking filter (ILF) supporting H.264 BP/MP/HP, AVS, and VC-1 video decoding. It comprises 38.4 Kgates and 672 bytes of local memory using TSMC 0.13 mum CMOS technology when operating at 225 MHz which meets the real-time processing requirement for high-resolution video decoding. We develop a PDB scheme and an integrated 1-D filter to realize various coding tools of the deblocking filter supporting multiple video coding standards. Cheng-An Chien, Hsiu-Cheng Chang, Jiun-In Guo |
ISCAS | 3 |
| 2009 | A System Architecture Exploration on the Configurable HW/SW Co-design for H.264 Video DecoderabstractIn this paper we focus on the design methodology to propose a design that is more flexible than ASIC solution and more efficient than the processor-based solution for H.264 video decoder. We explore the memory access bandwidth requirement and different software/hardware partitions so as to propose a configurable architecture adopting a DEM (Data Exchange Mechanism) controller to fit the best tradeoff between performance and cost when realizing H.264 video decoder for different applications. The proposed architecture can achieve more than three times acceleration in performance. Guo-An Jian, Jui-Chin Chu, Ting-Yu Huang, Tao-Cheng Chang, Jiun-In Guo |
ISCAS | 5 |
| 2009 | A Dynamic Quality-Adjustable H.264 Video Encoder for Power-Aware Video ApplicationsabstractThis paper proposes a dynamic quality-adjustable H.264 baseline profile (BP) video encoder that comprises 470 Kgates and 13.3 kB SRAM in a core size of 4.3 × 4.3 mm2using TSMC 0.13 ¿m 1P8M CMOS technology. Exploiting parameterized algorithms for motion estimation and intra prediction, the proposed design can dynamically configure the encoding modes with the design trade-off between power consumption and video quality for various video encoding applications. In addition, the proposed basic unit (BU)-based rate control hardware can maintain a constant and stable bit rate for network video transmission. It achieves real-time H.264 video encoding on CIF, D1, and HD720@30 frames/s with 7 mW to 25 mW, 27 mW to 162 mW, and 122 mW to 183 mW power dissipation in different quality modes. Hsiu-Cheng Chang, Bing-Tsung Wu, Ching-Lung Su, Jinn-Shyan Wang, Jiun-In Guo |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2009 | VisoMT: A Collaborative Multithreading Multicore Processor for Multimedia Applications With a Fast Data Switching MechanismabstractMultithreading and multicore processing are powerful ways to take advantage of parallelism in applications in order to boost a system's performance. However, exploring sufficient parallelism and achieving data locality with low communication overhead are still important research issues in embedded multithreading/multicore design. This paper introduces the design of a fast data switching mechanism between multilevel storage structures in a new multicore architecture. This paper makes several contributions to the development of contemporary sophisticated multimedia applications with advanced standards such as H.264. The first contribution,collaborative-multithreading, tightly unifies reduced instruction set computer and collaborative multithreading digital signal processing (DSP) in order to exploit high parallelism to provide sufficient computing power to applications. Each collaborative thread of our DSP is constructed by a heterogeneous-simultaneously multithreading single instruction, multiple data structure, and four media processing cores, which is connected by a fast switch for providing a fast data exchange mechanism among correlative streams on a thread-level basis. Our second contribution isone-stop streaming processing, which aims to keep data in the system for as long as possible until it is no longer needed, thus making data more efficient to access. Our third contribution is achunk threading programming model, including a thread management library and threading communication directives for reducing data communication and synchronization overhead. By a combination of coarse-grained and fine-grained threading, programmers can choose various threading levels based on the amount of data exchange in a program. With our proposed techniques and an appropriate programming model, we can reduce processing time by 54.9% in H.264 video encoding (common intermediate format video at 16.574 f/s) with the 1-virtual independent and streaming processing by open collaborative multithreading configuration, compared to the Texas Instruments C62 core that owns 8 function units. We realize our design as a prototype by chip implementation, and fabricate it as a chip based on the Taiwan Semiconductor Manufacturing Company Ltd. 0.13$\mu {\rm m}$process. The die size of the processor core is 16.12${\rm mm}^{2}$, including 414 k logic transistors and 34.4 kB of on-chip static random access memory. The processor runs at 180 MH0z/1.2-V and consumes 245 mW by postsimulation results. Wei-Chun Ku, Shu-Hsuan Chou, Jui-Chin Chu, Chi-Lin Liu, Tien-Fu Chen, Jiun-In Guo, Jinn-Shyan Wang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2009 | High-Throughput H.264/AVC High-Profile CABAC Decoder for HDTV ApplicationsabstractIn this letter we propose a high-throughput VLSI architecture design for H.264 high-profile context-based adaptive binary arithmatic coding (HP CABAC) decoding for HDTV applications. To speed up the inherent sequential CABAC decoding, we eliminate the bottleneck by proposing a look-ahead decision parsing technique on the grouped context table with cache registers, which reduces 62% of cycle count on average as compared with the original CABAC decoding. In addition, the proposed design supports the macroblock adaptive frame field coding tools in H.264 main profile coding and 8 times 8 transform in H.264 high-profile coding. It achieves the real-time processing for H.264 CABAC decoding up to L4.1@30 frames/s with maximum 60 Mbits/s when operating at 105 MHz. Yao-Chang Yang, Jiun-In Guo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | A 252Kgates/4.9Kbytes SRAM/71mW multistandard video decoder for high definition video applicationsabstractThis article proposes a low-cost, low-power multistandard video decoder for high definition (HD) video applications. The proposed design supports multiple-standard (JPEG baseline, MPEG-1/2/4 Simple Profile (SP), and H.264 Baseline Profile (BP)) video decoding through interactive parsing control and common parameter bus interface. In order to reduce hardware cost, the shared adder-based structure and reusable data management are proposed to achieve hardware sharing and reduce internal memory size, respectively. In addition, the proposed design is optimized through reducing memory bandwidth by increasing both data reuse amount and burst length of memory access as well as eliminating cycle overhead in data access for supporting HD video decoding with single AHB-based SDR memory. The proposed 252Kgates/4.9kB/71mW/0.13μm multi-standard video decoder reduces 72% in gate count and 87% in power consumption as compared to the state-of-the-art design, when operating at 120MHz for real-time HD1080 video decoding with single AHB-based SDR memory. Chih-Da Chien, Cheng-An Chien, Jui-Chin Chu, Jiun-In Guo, Ching-Hwa Cheng |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2008 | Joint algorithm/code-level optimization of H.264 video decoder for mobile multimedia applicationsabstractIn this paper, we propose a joint algorithm/code-level optimization scheme to make it feasible to perform real-time H.264/AVC video decoding software on ARM-based platform for mobile multimedia applications. In the algorithm-level optimization, we propose various techniques like fast interpolation scheme, zero-skipping technique for texture decoding, fast boundary strength decision for in-loop filtering, and pattern matching algorithm for CAVLD. In the code-level optimization, we propose the design techniques on minimizing memory access and branch times. The experimental result shows that we have reduced the complexity of H.264 video decoder up to 93% as compared to the reference software JM9.7. The optimized H.264 video decoder can achieve the QCIF@30Hz video decoding on an ARM9 processor when operating at 120MHz clock. Ting-Yu Huang, Guo-An Jian, Jui-Chin Chu, Ching-Lung Su, Jiun-In Guo |
ICASSP | 5 |
| 2008 | A H.264 basic-unit level rate control algorithm facilitating hardware realizationabstractRate Control plays an important role for video coding especially in video streaming applications with bandwidth constraints. The inherent sequential processing in H.264 basic unit (BU) level rate control algorithm makes it hard to be realized in a pipelined H.264 hardware encoder without increasing the processing latency. In this paper we propose a new H.264 BU-level rate control algorithm facilitating hardware realization. The proposed algorithm breaks down the sequential processing dependence in the original rate control algorithm in JM and reduces 28% for QCIF, 66% for CIF, 87% for D1 of hardware cycles while maintaining good video quality. Simulation results shows that the proposed algorithm reduces MAD’s memory buffer size to be Nunit* 14bits, which amounts to 26% for QCIF, 59% for CIF, 83% for D1 reduction as compared to JM rate control. Moreover, the proposed algorithm possesses high feasibility for hardware realization. Ping-Tsung Wu, Tzu-Chun Chang, Ching-Lung Su, Jiun-In Guo |
ICASSP | 4 |
| 2007 | An Embedded Coherent-Multithreading Multimedia Processor and Its Programming ModelabstractMultithreading and multi-core processing have been shown to be powerful approaches for boosting a system performance by taking advantage of parallelism in applications. This paper presents a processor design by unifying RISC and multithreading DSP for the sophisticated multimedia applications with advanced standards such as H.264. The proposed design not only minimizes integration costs for embedded multithreading/multi-core design by independent coherent threads, but also reduces the memory bandwidth requirements by one-stop streaming buffer and a very fast data exchange mechanism. With the proposed techniques and appropriate programming model, we can achieve 78% reduction of memory bandwidth and 89% reduction of processing time in H.264 video encoding, compared to traditional single stream micro-processor. Jui-Chin Chu, Wei-Chun Ku, Shu-Hsuan Chou, Tien-Fu Chen, Jiun-In Guo |
DAC | 5 |
| 2007 | A Low Latency Memory Controller for Video Coding SystemsabstractThe dynamic memory controller plays an important role in system-on-a-chip (SoC) designs to provide enough memory bandwidth through external memory for DSP and multimedia processing. However, the overhead cycles in accessing the data located in external memory have much influence on the SoC performance. In this paper, we propose a low latency memory controller with AHB interface to reduce the overhead cycles for the SDR memory access in the SoC designs. Through the pre-calculated addresses of impending transfers, two memory control schemes, i.e. Burst terminates Burst (BTB) and Anticipative Row Activation (ARA), are used to reduce the latency of SDR memory access. The experimental results show that the proposed memory controller reduces the memory bandwidth by 33% in a typical MPEG-4 video decoding system. Chih-Da Chien, Chih-Wei Wang, Chiun-Chau Lin, Tien-Wei Hsieh, Yuan-Hwa Chu, Jiun-In Guo |
ICME | 6 |
| 2007 | Low Complexity Multi-Standard Video Player for Portable Multimedia ApplicationsabstractThe title of this demo is "Low Complexity Multi-standard Video Player for Portable Multimedia Applications". This system can support three modern video coding standards that are MPEG-4, H.264, and VC-1. The proposed multi-standard video player is our research fruitage due to endless efforts so it is not a commercial product. In the past, we devoted ourselves to performance improvement for all kinds of video coding standards. Therefore, we integrate these optimized video decoders into a multi-standard video player for sake of exhibiting a prototype of portable multimedia applications. Guo-An Jian, Jiun-In Guo |
ICME | 2 |
| 2007 | A High-Speed/Low-Power Multiplier Using an Advanced Spurious Power Suppression TechniqueabstractThis study provides the experience of applying an advanced version of our former Spurious Power Suppression Technique (SPST) on multipliers for high-speed and low-power purposes. To filter out the useless switching power, there are two approaches, i.e. using registers and usingANDgates, to assert the data signals of multipliers after the data transition. The simulation results show that the SPST implementation withANDgates owns an extremely high flexibility on adjusting the data asserting time which not only facilitates the robustness of SPST but also leads to a 40% speed improvement. By adopting a 0.18-μm CMOS technology, the proposed SPST-equipped multiplier dissipates only 0.0121 mW per MHz in H.264texture coding applications, and obtains a 40% power reduction. Kuan-Hung Chen, Yuan-Sun Chu, Yu-Min Chen, Jiun-In Guo |
ISCAS | 4 |
| 2007 | A Memory-Based Hardware Accelerator for Real-Time MPEG-4 Audio Coding and ReverberationabstractIn this paper we propose a memory-based hardware accelerator for MPEG-4 audio coding and reverberation to achieve both high quality and reality of audio. The proposed design can realize both the computation-intensive component of 256/2048-point IMDCT and the 1024-point FFT-based reverberation through the same hardware engine by adopting a unified IMDCT/FFT/IFFT algorithm, which greatly reduces the hardware cost. The proposed design can achieve both the real-time 5.1 channel audio decoding at the sampling rate of 44.1 KHz and audio reverberation with the hardware cost of 26,633 gates and 4.6K words of local memory for storing transform coefficients and temporary results. The maximum working frequency achieves 220 MHz when implemented by UMC 0.18mum CMOS technology, which can fit the real-time processing requirement of many high quality MPEG-4 audio coding applications Guo-An Jian, Chih-Da Chien, Jiun-In Guo |
ISCAS | 3 |
| 2006 | Low Complexity Architecture Design of H.264 Predictive Pixel Compensator for HDTV ApplicationabstractIn this paper, we propose a low-complexity architecture design of H.264 predictive pixel compensator (PPC) for HDTV application. In intra prediction, we propose a shared adder-based architecture style that supports all of the 17 intra prediction modes, and reduce computational complexity in the I4MB prediction mode 3~8 up to 50% computation. Besides, we have also proposed the distributed memory access to improve the HW usage. As well as, it can used to reduce the memory size for buffering the neighboring pixels. In inter prediction, we can save about 48% of external memory bandwidth by the data reused through the hybrid block size memory access. Adopting the mixed six-tap FIR filter architecture to design luma interpolation, we can efficiently reduce the hardware cost up to 27%. The implemental result shows the hardware cost of the proposed design is about 60854 gates under a TSMC 0.18 μ m CMOS technology, which achieves the real-time processing requirement for HD-1080 format video@30Hz at the working frequency of 87 MHz. Chien-Chang Lin, Jiun-In Guo, Jinn-Shyan Wang |
ICASSP (3) | 3 |
| 2006 | A Condition-based Intra Prediction Algorithm for H.264/AVCabstractThis paper proposes a condition-based algorithm for H.264/AVC 4times4 intra prediction. Exploiting high correlation existed in neighboring intra prediction modes, we propose the three conditions to skip the less possible candidates in doing intra4times4 block mode decision. When compared to the 9 prediction modes in the full search algorithm, the proposed algorithm can complete a 4times4 intra prediction using 4.4 prediction modes operation in average. The simulation result shows that the proposed algorithm can reduce computational complexity up to 44% at the cost of less than 0.1 dB PSNR loss in average Chun-Hao Chang, Chien-Chang Lin, Yi-Huan Yang, Jiun-In Guo, Jinn-Shyan Wang |
ICME | 5 |
| 2006 | Collaborative Multithreading: An Open Scalable Processor Architecture for Embedded Multimedia ApplicationsabstractNumerous approaches can be employed in exploiting computation power in processors such as superscalar, VLIW, SMT and multi-core on chip. In this paper, a UniCore VisoMT processor is proposed, which unifies VLIW and multithreading by providing an efficient control and data communication model, while offering explicit parallelisms for embedded applications. The architecture concurrently executes a main thread and several accelerative threads, coordinated by the main thread. A switch-based register-file is provided for fast data exchange between these accelerative threads. Moreover, a SMT helper function unit is employed for controlling and resource-sharing between accelerative threads, and an event-driven mechanism is introduced for synchronization between the main thread and these accelerative threads. Our results show that the proposed architecture provides area and performance advantages for embedded multimedia applications Wei-Chun Ku, Shu-Hsuan Chou, Jui-Chin Chu, Chih-Heng Kang, Tien-Fu Chen, Jiun-In Guo |
ICME | 6 |
| 2006 | A High Throughput VLSI Architecture Design for H.264 Context-Based Adaptive Binary Arithmetic Decoding with Look Ahead ParsingabstractIn this paper we present a high throughput VLSI architecture design for context-based adaptive binary arithmetic decoding (CABAD) in MPEG-4 AVC/H.264. To speed-up the inherent sequential operations in CABAD, we break down the processing bottleneck by proposing a look-ahead codeword parsing technique on the segmenting context tables with cache registers, which averagely reduces up to 53% of cycle count. Based on a 0.18 mum CMOS technology, the proposed design outperforms the existing design by both reducing 40% of hardware cost and achieving about 1.6 times data throughput at the same time Yao-Chang Yang, Chien-Chang Lin, Hsui-Cheng Chang, Ching-Lung Su, Jiun-In Guo |
ICME | 5 |
| 2006 | Low-power mechanism with power block managementabstractIn this paper, a low power mechanism with power block management is proposed to reduce the power consumption in DSP chips. Because the digital signal processing (DSP) chips use many functional units in the data-path to achieve parallel processing, the unnecessary functional units are also executed simultaneously, and they dissipate the power. In the paper, we classify the types of instruction sets based on their data flow in the DSP to generate the control signals which active the necessary units in their data-path. The mechanism is called power block management (PBM). We employ particular guarded circuits in the different situations to avoid the actions of unnecessary functional units so as to reduce the power dissipation. The experimental results show that the power consumption can be saved about 14% at the cost of less than 2.7% area increment. Kuo-Chuan Chao, Kuan-Hung Chen, Yuan-Sun Chu, Jiun-In Guo |
ISCAS | 4 |
| 2006 | A performance-aware IP core design for multimode transform coding using scalable-DA algorithmabstractThis paper proposes a performance-aware transform IP design which can be configured to appropriate hardware for different performance requirements on demand without requiring additional data bandwidth in multimode video coding (JPEG/MPEG-1/2/4/H.261/H.263/H.264). Based on the scalable-DA approach, three schemes of hardware configurations which are respectively composed of 3, 6, and 12 data-paths are illustrated. The three schemes of the proposed performance-aware DCT/IDCT can achieve CIF, 720HD, and digital cinema video formats when operated at 9.13 MHz, 41.48 MHz, and 188.75 MHz, respectively. Kuan-Hung Chen, Jinn-Shyan Wang, Jiun-In Guo |
ISCAS | 4 |
| 2006 | A high performance CAVLC encoder design for MPEG-4 AVC/H.264 video coding applicationsabstractThis paper presents a high performance VLSI architecture design for MPEG-4 AVC/H.264 CAVLC encoding. In the proposed design, we propose a forward-based parallel coding (FPC) technique to increase the data throughput rate. Moreover, two approaches called arithmetic table elimination (ATE) and fast look-up table matching (FLM) are exploited to reduce the hardware cost. With the synthesis constraint of 125 MHz clock, the hardware cost of the proposed design is 9724 gates based on a 0.18/spl mu/m CMOS technology, which achieves the real-time processing requiremenwat for H.264 video encoding on HD1080 format video. Chih-Da Chien, Keng-Po Lu, Yi-Hung Shih, Jiun-In Guo |
ISCAS | 4 |
| 2006 | Design of customized functional units for the VLIW-based multi-threading processor core targeted at multimedia applicationsabstractIn this paper, we propose the customized functional units (CFUs) for the UniCore which is a VLIW-based multi-threading processor core. The CFUs work as the hardware accelerators and play a key component to increase the performance for the multimedia application. Compared to ARM9TDMI, the number of execution cycles is reduced a factor of 21.7 in running the operations in MPEG video coding. Besides, the proposed design owns about twice data throughput rate compared to other SIMD-based architectures in average. Jui-Chin Chu, Chih-Wen Huang, He-Chun Chen, Keng-Po Lu, Ming-Shuan Lee, Jiun-In Guo, Tien-Fu Chen |
ISCAS | 6 |
| 2006 | A high-performance direct 2-D transform coding IP design for MPEG-4AVC/H.264abstractThis paper proposes a high-performance direct two-dimensional transform coding IP design for MPEG-4 AVC/H.264 video coding standard. Because four kinds of 4 /spl times/ 4 transforms, i.e., forward, inverse, forward-Hadamard, and inverse-Hadamard transforms are required in a H.264 encoding system, a high-performance multitransform accelerator is inevitable to compute these transforms simultaneously for fitting real-time processing requirement. Accordingly, this paper proposes a direct 2-D transform algorithm which suitably arranges the data processing sequences adopted in row and column transforms of H.264 CODEC systems to finish the data transposition on-the-fly. The induced new transform architecture greatly increases the data processing rate up to 8 pixels/cycle. In addition, an interlaced I/O schedule is disclosed to balance the data I/O rate and the data processing rate of the proposed multitransform design when integrated with H.264 systems. Using a 0.18-/spl mu/m CMOS technology, the optimum operating clock frequency of the proposed multitransform design is 100 MHz which achieves 800 Mpixels/s data throughput rate with the cost of 6482 gates. This performance can achieve the real-time multitransform processing of digital cinema video (4096 /spl times/ 4 2048@30 Hz). When the data throughput rate per unit area is adopted as the comparison index in hardware efficiency, the proposed design is at least 1.94 times more efficient than the existing designs. Moreover, the proposed multitransform design can achieve HDTV 720p, 1080i, digital cinema video processing requirements by consuming only 0.58, 2.91, and 24.18 mW when operated at 22, 50, and 100 MHz with 0.7, 1.0, and 1.8 V power supplies, respectively. Kuan-Hung Chen, Jiun-In Guo, Jinn-Shyan Wang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2006 | An Area-Efficient Variable Length Decoder IP Core Design for MPEG-1/2/4 Video Coding ApplicationsabstractThis paper proposes an area-efficient variable length decoder (VLD) IP core design for MPEG-$hbox 1/2/4$video coding applications. The proposed IP core exploits the parallel numerical matching in the MPEG-$hbox 1/2/4$entropy decoding to achieve high data throughput rate in terms of limited hardware cost. This feature not only improves the performance of VLD, but also facilitates reducing the power consumption through lowering down the supply voltage while maintaining enough data throughput rate. Moreover, we propose a partial combinational component enabling approach for minimizing the power consumption of the proposed design. Based on 0.18-$mu$m CMOS technology, the implementation results show that the proposed IP core operates at 125-MHz clock frequency with the cost of 13 105 gates. In addition, the power consumption of the proposed design reaches 163.4$mu$W operated at 12.5 MHz with 0.9-V supply voltage, which is fast enough for MPEG-1/2/4 real-time decoding on 4CIF video@30 Hz. Compared to the existing designs, the proposed IP core possesses both higher data throughput and less hardware cost. Chih-Da Chien, Keng-Po Lu, Yu-Min Chen, Jiun-In Guo, Yuan-Sun Chu, Ching-Lung Su |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2005 | An efficient spurious power suppression technique (SPST) and its applications on MPEG-4 AVC/H.264 transform coding designabstractThis paper proposes an efficient Spurious Power Suppression Technique (SPST) and its applications on an MPEG-4 AVC/H.264 transform coding design. There are three techniques addressed in this paper, which are (1) the SPST, (2) the direct 2-D algorithm, and (3) the interlaced I/O schedule to solve the design challenges induced by both the real-time processing and low-power requirements. The major novelty of this paper is implementing the SPST concept on the transform architecture for H.264, which save 31.9% power consumption at the cost of 20.9% area price. Moreover, the proposed transform design also possesses 60.05% higher hardware efficiency through the TPUA index than the existing designs Kuan-Hung Chen, Kuo-Chuan Chao, Jinn-Shyan Wang, Yuan-Sun Chu, Jiun-In Guo |
ISLPED | 5 |
| 2005 | A memory-efficient realization of cyclic convolution and its application to discrete cosine transformabstractThis paper presents a memory-efficient approach to realize the cyclic convolution and its application to the discrete cosine transform (DCT). We adopt the way of distributed arithmetic (DA) computation, exploit the symmetry property of DCT coefficients to merge the elements in the matrix of DCT kernel, separate the kernel to be two perfect cyclic forms, and partition the content of ROM into groups to facilitate an efficient realization of a one-dimensional (1-D) N-point DCT kernel using (N-1)/2 adders or subtractors, one small ROM module, a barrel shifter, and ((N-1)/2)+1 accumulators. The proposed memory-efficient design technique is characterized by rearranging the content of the ROM using the conventional DA approach into several groups such that all the elements in a group can be accessed simultaneously in accumulating all the DCT outputs for increasing the ROM utilization. Considering an example using 16-bit coefficients, the proposed design can save more than 57% of the delay-area product, as compare with the existing DA-based designs in the case of the 1-D seven-point DCT. Finally, a 1-D DCT chip was implemented to illustrate the efficiency associated with the proposed approach. Hun-Chen Chen, Jiun-In Guo, Tian-Sheuan Chang, Chein-Wei Jen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2005 | An Energy-Aware IP Core Design for the Variable-Length DCT/IDCT Targeting at MPEG4 Shape-Adaptive TransformsabstractThis paper proposes a flexible hardware solution and the associated energy-aware IP core design for computing the variable-length discrete cosine transform/inverse discrete cosine transform (DCT/IDCT) required in the MPEG4 shape-adaptive DCT/IDCT (SA-DCT/IDCT). The proposed IP core has been developed based on the design concept of programmable processors to provide the flexibility in dynamically configuring the hardware. To achieve good performance both in area and speed, we optimize the proposed IP core both in the algorithmic computational complexity and hardware complexity. Furthermore, the proposed IP core possesses the feature of energy-aware design flexibility. The simulation shows that the proposed design has 44% energy reduction at the price of 0.3-dB signal quality degradation for the image compression applications. The implementation results show that the proposed IP core costs about 3100 gates along with 16 words (1 word = 16 bits) of memory, which can achieve the real-time processing of the texture coding in MPEG4 SP@L3 and ACE@L2 CODEC system for the CIF format video at 30 frames/s with 4:2:0 color format. Kuan-Hung Chen, Jiun-In Guo, Jinn-Shyan Wang, Chingwei Yeh |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2004 | A high-performance MPEG4 bitstream processing coreabstractEntropy coding is one of the main techniques for lossless data compression in many multimedia applications. In video and image coding standards such as JPEG, MPEG1/2/4, and H.26x, the most often used entropy coding techniques include run-length coding and variable-length coding. Due to the demands of high quality video compression like DTV or HDTV, the complexity of entropy coding increases significantly. This fact motivates the design of the high performance entropy decoder proposed in this paper. The proposed IP core exploits the concurrent operations in the MPEG4 entropy decoding to achieve high performance in terms of limited hardware cost. The implementation results show that the proposed IP core operates at 125 MHz clock frequency (i.e. 2000 Mbit/s) with the cost of 0.22 mm/sup 2/ silicon area based on a 0.25 um CMOS technology. Tai-Lun Chang, Ying-Ming Tsai, Chih-Da Chien, Chien-Chang Lin, Jiun-In Guo |
ICME | 5 |
| 2004 | A power-aware SNR-progressive DCT/IDCT IP core design for multimedia transform codingabstractA power-aware SNR progressive DCT/IDCT IP core design for multimedia transform coding is proposed. The proposed IP core possesses the feature of power-aware design flexibility, allowing the trade-off of lower power consumption with less demand of data precision in developing its instruction library. Relationships of energy reduction and data quality degradation in the examples of both JPEG still images and MPEG4 video sequences have been analyzed. Since the proposed IP core is developed based on the concept of programmable processors, we can select a DCT/IDCT firmware library of different precisions of cosine coefficients according to the accuracy requirement of various applications. This design has been realized based on a 0.35-/spl mu/m CMOS technology and costs about 2175 gates with 8 words of RAM, which can achieve real-time processing of the texture coding in an MPEG4 SP@L3 codec system for CIF video at 30 frames per second (fps). Kuan-Hung Chen, Jiun-In Guo, Jinn-Shyan Wang, Chingwei Yeh |
ICME | 2 |
| 2004 | An efficient 2-D DCT/IDCT core design using cyclic convolution and adder-based realizationabstractThis paper proposes an efficient two-dimensional (2-D) discrete cosine and inverse discrete cosine transform (DCT/IDCT) core design. Adopting the row-column decomposition technique for computing 2-D DCT/IDCT, we formulate the one-dimensional (1-D) DCT/IDCT into cyclic convolution by properly arranging the input sequence, optimize the multiplications based on the concept of common subexpression sharing, and carry out the multiplications through carry-save adders (CSAs). Using cyclic convolution is helpful in exploiting the word-level data sharing in computing different DCT/IDCT outputs. Adopting the common subexpression sharing is beneficial to the bit-level data sharing in computing the outputs. As compared with some existing approaches of realizing DCT/IDCT, the proposed approach can save on average 20%/spl sim/33% in the delay-area product (gate-count * time-unit) based on a 0.35-/spl mu/m CMOS technology under the data word-lengths ranging from 16/spl sim/24 b. Besides, we have also proposed an IP generator for designing the 2-D DCT/IDCT based on the proposed approach. It provides a design-automation environment with parameter configurations in designing a 2-D DCT/IDCT core that is suitable for most image and video compression applications. Jiun-In Guo, Rei-Chin Ju |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2001 | A generalized architecture for the one-dimensional discrete cosine and sine transformsabstractIn this paper, we propose a generalized architecture for the 1-D discrete cosine transform (DCT), discrete sine transform (DST), and their inverses based on a general formulation. This architecture provides the flexibility to adaptively select different transform functions through a simple control signal. As compared with other unified architectures for the DCT/DST, the proposed architecture possesses better performance in the hardware cost and the number of I/O channels. The proposed design can be efficiently realized by using the distributed arithmetic approach in the VLSI implementation. Jiun-In Guo, Chih-Chen Li |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2000 | An efficient design for one dimensional discrete cosine transform using parallel addersabstractThis paper proposes an efficient parallel adder based design for 1-D any-length discrete Cosine transform (DCT). Using the similar idea to the Chirp-Z transform, we develop an algorithm formulating the 1-D any-length DCT as cyclic convolutions. The proposed design using this algorithm not only owns higher flexibility in the transform length, but also possesses low hardware cost by using the parallel adder implementation. Considering an example using 16-bits coefficients, the proposed design can save much gate area as compared with other designs in the longer transform length applications. Jiun-In Guo |
ISCAS | 1 |
| 2000 | A new chaotic key-based design for image encryption and decryptionabstractIn this paper, an image encryption/decryption algorithm and its VLSI architecture are proposed. According to a chaotic binary sequence, the gray level of each pixel is XORed or XNORed bit-by-bit to one of the two predetermined keys. Its features are as follows: (1) low computational complexity, (2) high security, and (3) no distortion. In order to implement the algorithm, its VLSI architecture with low hardware cost, high computing speed, and high hardware utilization efficiency is also designed. Moreover, the architecture of integrating the scheme with MPEG2 is proposed. Finally, simulation results are included to demonstrate its effectiveness. Jui-Cheng Yen, Jiun-In Guo |
ISCAS | 2 |
| 1998 | A new k-winners-take-all neural network and its array architectureabstractIn this paper, a new neural-network model called WINSTRON and its novel array architecture are proposed. Based on a competitive learning algorithm that is originated from the coarse-fine competition, WINSTRON can identify the k larger elements or the k smaller ones in a data set. We will then prove that WINSTRON converges to the correct state in any situation. In addition, the convergence rates of WINSTRON for three special data distributions will be derived. In order to realize WINSTRON, its array architecture with low hardware complexity and high computing speed is also detailed. Finally, simulation results are included to demonstrate its effectiveness and its advantages over three existing networks. Jui-Cheng Yen, Jiun-In Guo, Hun-Chen Chen |
IEEE Trans. Neural Networks | 2 |
| 1994 | A novel VLSI array design for the discrete Hartley transform using cyclic convolutionabstractThis paper presents a novel VLSI array design for the 1-D discrete Hartley transform (DHT). Using the similar idea to the Chirp-Z transform, we develop an algorithm which can formulate the 1-D any-length DHT as cyclic convolutions. This algorithm owns higher flexibility in the transform length as compared with the existing approach of Guo, Liu, and Jen (see Proc. IEEE International Conference on Acoustics; Speech, and Signal Processing, p.v.621-v.624, 1992 and IEEE Transactions on Circuits and Systems-II: Analog and Digital Signal Processing, vol.30, p.723-733, Oct. 1992). Moreover, we use the memory-based approach to realize the cyclic convolutions by systolic arrays and implement the multiplications by small ROMs and adders. The presented array not only outperforms the distributed arithmetic (DA) architectures in the hardware area, but also owns low input/output (I/O) cost, power dissipation, high computing speeds and flexibility in transform length.> Jiun-In Guo, Chi-Min Liu, Chein-Wei Jen |
ICASSP (2) | 1 |
| 1994 | A General Approach to Design VLSI Arrays for the Multi-dimensional Discrete Hartley TransformabstractIn this paper, a general memory-based approach to design VLSI arrays for the multi-dimensional (M-D) discrete Hartley transform (DHT) with any length is proposed. There are four parts of this approach: (1) a new M-D DHT formulation, (2) cyclic convolution representation, (3) systolic array realization, and (4) memory-based implementation. Deriving a new M-D DHT formulation avoids the undesirable overhead required in formal designs. Taking cyclic convolution provides high computing parallelism and low computation complexity. Using systolic array realisation results in high computing speeds and low I/O cost. Adopting the memory-based implementation yields low hardware cost and low power dissipation. In summary the proposed approach will lead to efficient and high-performance VLSI array designs for the M-D DHT.> Jiun-In Guo, Chi-Min Liu, Chein-Wei Jen |
ISCAS | 1 |
| 1993 | A CORDIC-based VLSI Array for Computing 2-D Discrete Hartley Transform
Jiun-In Guo, Chi-Min Liu, Chein-Wei Jen |
ISCAS | 1 |
| 1993 | A Multi-phase Shared Bus Structure for the Fast Fourier Transform
Jiun-In Guo, C. Bernard Shung, Chein-Wei Jen |
ISCAS | 2 |
| 1992 | A memory-based approach to design and implement systolic arrays for DFT and DCTabstractA new design and implementation approach for the discrete Fourier transform (DFT) and discrete cosine transform (DCT) is discussed. The approach, called memory-based approach, is based on a design concept of combining and exploiting both the advantages induced from the systolic array architectures and the efficient implementation techniques of substituting multipliers by read-only-memory (ROM) together. Based on this approach, the proposed arrays feature high computing speeds, low complexity and hardware cost of the processing elements (PEs), and low I/O cost. Moreover, high regularity both among and inside the PEs can be found in the proposed arrays. These merits make the proposed designs much more feasible for VLSI implementation.> Jiun-In Guo, Chi-Min Liu, Chein-Wei Jen |
ICASSP | 1 |