Gwo Giun Lee

dblp:37/1792 · also Gwo Giun Chris Lee · DBLP profile ↗
← Back
33ranked-venue papers
19as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 7 first-authorSystems, architecture and hardware · 14 · 10 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 first-authorArtificial intelligence and machine learning · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 46% Performance modeling and evaluation · 23% Hardware accelerators and domain-specific architectures · 19%
Computer graphics and multimedia
1 paper
Image and video processing · 62% Image and video coding · 38%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › parallel programming models
degree of parallelism
0.112012
Quantifying Intrinsic Parallelism Using Linear Algebra for Algorithm/Architecture Coexploration · IEEE Trans. Parallel Distributed Syst. 2012
Parallel and multicore computing › parallel algorithms
parallel algorithm design
0.112012
Quantifying Intrinsic Parallelism Using Linear Algebra for Algorithm/Architecture Coexploration · IEEE Trans. Parallel Distributed Syst. 2012
Performance modeling and evaluation
workload characterization
0.112012
Quantifying Intrinsic Parallelism Using Linear Algebra for Algorithm/Architecture Coexploration · IEEE Trans. Parallel Distributed Syst. 2012
Image and video processing
motion estimation
0.112007
Algorithm/Architecture Co-Design of 3-D Spatio-Temporal Motion Estimation for Video Coding · IEEE Trans. Multim. 2007
Hardware accelerators and domain-specific architectures
algorithm-hardware co-design
0.112007
Algorithm/Architecture Co-Design of 3-D Spatio-Temporal Motion Estimation for Video Coding · IEEE Trans. Multim. 2007
Integrated circuit design › digital circuit design
VLSI architecture
0.112007
Algorithm/Architecture Co-Design of 3-D Spatio-Temporal Motion Estimation for Video Coding · IEEE Trans. Multim. 2007
Image and video coding › video coding standards
H.264/AVC
0.012007
Algorithm/Architecture Co-Design of 3-D Spatio-Temporal Motion Estimation for Video Coding · IEEE Trans. Multim. 2007
Image and video coding
video coding standards
0.012007
Algorithm/Architecture Co-Design of 3-D Spatio-Temporal Motion Estimation for Video Coding · IEEE Trans. Multim. 2007

Methods — techniques the papers use, named apart from their topics

rank theorem · 0.1optimization theory · 0.1one-at-a-time search · 0.1linear algebra · 0.1
YearPublicationVenuePosition
2019 AIFood: A Large Scale Food Images Dataset for Ingredient Recognition
abstract
In this paper, we introduce a large-scale food images dataset namely AIFood, which is constructed to aim ingredient recognition in food image research. AIFood dataset includes 24 categories and totally 372,095 food images around the world. We collect food images from eight existing food image datasets and a food website. The food images are relabeled using 24 categories. We preliminarily label each image using existing food information, e.g. dish name or ingredient information. Next, we manually check food images to find out undiscovered ingredients and relabel them. Every image can be labeled more than one category. In addition, food images may have color cast or uneven contrast problems, which may disturb performance of image recognition system. So, we applied preprocessing method which contains automatic white balancing and contrast limited adaptive histogram equalization methods to improve visual quality of food images. We set constraints which are defined by luminance and chrominance of image to determine if the image is to be preprocessed.
Gwo Giun Lee, Chin-Wei Huang, Jia-Hong Chen, Shih-Yu Chen, Hsiu-Ling Chen
TENCON1
2019 Classification of Alzheimer's Disease, Mild Cognitive Impairment, and Cognitively Normal Based on Neuropsychological Data via Supervised Learning
abstract
Dementia is neurodegenerative or vascular disorder which is characterized by declining mental function including a combination of symptoms for abnormal activity, behavior and cognitive. Alzheimer's disease (AD) is the most common type of dementia, it accounts for 60 to 80 percent of dementia cases. Their brain neuron are die and loss the connection with each other neurons which stop firing and establishing the brain network, and it is easily considered as normal aging processes when the subject is actually in the early stages of AD. This paper proposed a machine learning algorithm based on neuropsychological data to classify subjects into Alzheimer's disease, mild cognitive impairment (MCI) and cognitively normal (CN). We acquired neuropsychological data from 678 participants with clinical diagnosis information from Alzheimer's Disease Neuroimaging Initiative (ADNI) database, and we focus on Mini-Mental State Examination (MMSE) which is the most extensively used psychometric examination in the clinical practice. Through computer-aided diagnosis (CAD) system which is called as “smart doctor”, this algorithm can help doctors diagnose patients with Alzheimer's disease. The medical AI brings accurate and rapid diagnosis and prediction. As to follow up the patient's care and social resources can be given faster. The main result of this paper is that we found two valuable features from MMSE, orientation and recall, which have the same ability as the entire MMSE, to detect the Alzheimer's disease.
Gwo Giun Lee, Po-Wei Huang, Yi-Ru Xie, Ming-Chyi Pai
TENCON1
2016 Efficient nuclei segmentation based on spectral graph partitioning
abstract
Biomedical image processing that offers computer-aided diagnosis is much more popular due to the availability of high quality and large quantity of medical data. Our well-developed biomedical image computing system, which automatically extracts and segments the nucleus and cytoplasm of cell in medical images, is no doubt following this idea. Nonetheless, even though previous system provide good algorithmic performance, its throughput is limited by high computation load and data dependency. Therefore, we deploy spectral graph partitioning to improve computation speed of the most complex module, maker-controlled watershed transform for nuclei detection. By modeling our problem as a graph and embedding architectural costs as the attributes in vertices and edges, we equally distribute workload among processors and reduce overhead in data transfer rate. We deploy the proposed approach on Intel Core i7-930 CPU with four cores and eight threads and test 153 medical images; as a consequence, we achieve less data transfer and better load balance as compared to conventional workload distribution through clustering and other graph partitioning methods.
Gwo Giun Lee, Shi-Yu Hung, Tai-Ping Wang, Chun-Fu Chen 0001, Chi-Kuang Sun, Yi-Hua Liao
ISCAS1
2015 Implementation of Gabor feature extraction algorithm for electrocardiogram on FPGA
abstract
This paper implements an electrocardiogram (ECG) feature extraction system onto a Field Programmable Gate Array (FPGA). The algorithm extracts ECG features including R peaks, QRS onsets, Q peaks, S peaks, QRS offsets, P onsets, P peaks, P offsets, T onsets, T peaks, and T offsets based on multi-scale analysis in Gabor Wavelet Transform (GWT); and estimates the amplitudes of P, R, T peaks and Q, S depths. Subjects' ECG signals are acquired through the Texas Instrument (TI) ADS1298R-ECGPDK ECG signal acquisition, and then transferred to Xilinx XUPV5-LX110T Evaluation Platform for ECG feature extraction via Serial Peripheral Interface (SPI). Consequently, the results that are obtained by the feature extraction process will be stored in the CompactFlash Card of the evaluation platform and displayed on the monitor via the National Instrument (NI) data acquisition. The experimental results show that the proposed system that achieves low comparative error surpasses the performance in other related works.
Gwo Giun Lee, Zuo-Jheng Huang, Chih-Yuan Chen, Chun-Fu Chen 0001
ISCAS1
2015 Efficient Multi-training Framework of Image Deep Learning on GPU Cluster
abstract
In this paper, we develop a pipelining schema for image deep learning on GPU cluster to leverage heavy workload of training procedure. In addition, it is usually necessary to train multiple models to obtain a good deep learning model due to the limited a priori knowledge on deep neural network structure. Therefore, adopting parallel and distributed computing appears is an obvious path forward, but the mileage varies depending on how amenable a deep network can be parallelized and the availability of rapid prototyping capabilities with low cost of entry. In this work, we propose a framework to organize the training procedures of multiple deep learning models into a pipeline on a GPU cluster, where each stage is handled by a particular GPU with a partition of the training dataset. Instead of frequently migrating data among the disks, CPUs, and GPUs, our framework only moves partially trained models to reduce bandwidth consumption and to leverage the full computation capability of the cluster. In this paper, we deploy the proposed framework on popular image recognition tasks using deep learning, and the experiments show that the proposed method reduces overall training time up to dozens of hours compared to the baseline method.
Chun-Fu Chen 0001, Gwo Giun Lee, Yinglong Xia, Wan-Yi Sabrina Lin, Toyotaro Suzumura, Ching-Yung Lin
ISM2
2014 Content-adaptive depth map enhancement based on motion distribution
abstract
This paper provides a motion-based content-adaptive depth map enhancement algorithm to enhance the quality of the depth map and reduce the artifacts in the synthesized views. The proposed algorithm extracts depth cues from the motion distribution at the specific scenario of camera movement to align the distribution of depth and motion. In real world scenarios, when the camera is panning in horizontal direction, the nearer distance between the object and the camera, the larger motion will be, and vice versa; therefore, we could interpret the depth from motion in this. Moreover, in the scenario of fixed camera, the depth cue from motion could be derived in the same approach, and the depth variation within one moving object shall be small. Hence, the depth values of moving object should not be rapidly changing. In addition, this paper also employs the bi-directional motion-compensated infinite impulse response low-pass filter to stabilize the consistency of depth maps over time. As a consequence, the algorithm so introduced not only aligns the depth map to depth cues from motion but also enhance stability and consistency of depth maps in the spatial-temporal domain. Experiment results via enhanced depth maps show that the synthesized results would be better in both objective and subjective measurement in comparison with the results using original depth maps and the state-of-the-art depth enhancement algorithms.
Gwo Giun Lee, Bo-Syun Li, Chun-Fu Chen 0001
VCIP1
2013 4×4/8×8/12×12 reconfigurable MIMO detector on multi-core DSP based on Eigen decomposition of dataflow graph
abstract
This paper presents a 4×4/8×8/12×12 and QPSK/16-QAM/64-QAM reconfigurable MIMO detector on a multi-core DSP platform for different MIMO detection algorithms, including multiple-candidate-selection QRSIC (MCS-QRSIC), distributed K-best, and sorting-reduced K-best detectors. This study uses an Eigen-value decomposition method to investigate the intrinsic degree of parallelism of MIMO detectors and allocate the processing elements to DSP cores. MCS-QRSIC outperforms other detectors in 4 × 4 detection, and distributed K-best has the best performance in 8×8 and 12×12 MIMO systems. The normalized throughput achieves 188/153/142 Mbps for 64-QAM 4×4/8×8/12×12 MIMO detections at a 1GHz clock rate with 448 cores.
Sheng-Hung Wu, Chien-Yu Kao, Jen-Yuan Hsu, Pangan Ting, Gwo Giun Lee, Yuan-Hao Huang
ICASSP5
2013 Depth map enhancement based on Z-displacement of objects
abstract
A depth map enhancement algorithm based on Z-displacement of objects is presented in this paper. Most depth map enhancement algorithms utilize either the spatial or temporal information only. In this paper, information in another dimension, Z-displacement, is utilized to enhance the depth map of videos by estimating the scale change of objects. Our experimental results show that the proposed scale changing detection method has higher accuracy and lower complexity compared with the conventional scale changing detection approaches such as normalized gradient correlation. Based on the estimated scale change, the proposed depth map enhancement corrects the depth map by estimating the Z-displacement of objects in the 3D space. Experimental results demonstrate that better depth maps and visual quality of view synthesis are achieved by the proposed method as compared to the state-of-arts.
Gwo Giun Lee, Ciao-Siang Siao, Chunhui Cui, Chun-Fu Chen 0001, Huan-Hsiang Lin
ISCAS1
2013 Motion-based depth estimation for 2D-to-3D video conversion
abstract
This paper presents a motion-based depth estimation algorithm for automatic 2D-to-3D video conversion algorithm by employing the co-occurrence matrix of motion vectors (MVCM). Video scenes possess distinct signatures of MVCM, which enables exploiting the corresponding motion-depth relation for depth generation. The subsequent motion-compensated depth updating scheme provides stable and comfort 3D visual quality as synthesized by depth-image-based rendering. The simulation results of several high-definition image sequences indicate that the proposed algorithm produces better and more reasonable depth than two motion-based depth estimation algorithms. With the adaptive depth estimation scheme using MVCM, the proposed 2D-to-3D video conversion algorithm can accommodate a great variety of visual contents. It thus provides an efficient and reliable solution towards the problem of automatic 3D video content creation.
Ming-Jiun Wang, Chun-Fu Chen 0001, Gwo Giun Lee
VCIP3
2012 Cell segmentation and NC ratio analysis of third harmonic generation virtual biopsy images based on marker-controlled gradient watershed algorithm
abstract
Traditional biopsy procedure requires invasive tissue removal from a living subject followed by time-consuming complicatedly processing, so noninvasive in vivo virtual biopsy is a highly desired technique which own ability to obtain exhaustive tissue images without removing tissues from subjects. Some sets of in vivo virtual biopsy images provided by some healthy volunteers are processed by our cell segmentation approach based on marker-controlled gradient watershed algorithm to isolate the nuclei and cytoplasm and also evaluate their Nuclear-to-Cytoplasmic (NC) ratio. From our experimental results, our algorithm has significant potential for in vivo cell segmentation and NC ratio analysis to identify or detect the early symptoms of some skin diseases with abnormal NC ratios, such as skin cancers in clinical diagnosis.
Huan-Hsiang Lin, Ming-Rung Tsai, Chun-Fu Chen 0001, Szu-Yu Chen, Yi-Hua Liao, Gwo Giun Lee, Chi-Kuang Sun
ISCAS6
2012 Quantifying Intrinsic Parallelism Using Linear Algebra for Algorithm/Architecture Coexploration
abstract
Degree of parallelism (DoP) is an essential complexity metric that characterizes the number of independent operation sets (IOSs) that can be concurrently executed within an algorithm. This paper presents a generic framework to identify IOSs and to quantify the DoP based on rank theorem in linear algebra. This framework is applied to extract algorithmic parallelisms at various granularities, namely, multigrain parallelism. Our parallelism is intrinsic and platform independent and can provide insights into architectural information, thus facilitating mapping onto generic platforms and early back annotation for modifying algorithms. It plays a significant role in the concurrent optimization of both algorithms and architectures, referred to as Algorithm/Architecture Coexploration (AAC), by trading off between the DoP and the number of operations (NoO). This paper reports three case studies for AAC. The case study on an IDCT reveals that our framework accurately quantifies the parallelism for mapping the algorithm onto generic platforms, including FPGA and multicore systems. The IDCT parallelized by our technique surpasses a conventional spectral parallelization. By exploiting fine-grain parallelism, this paper presents a better porting of a discrete wavelet transform (DWT) onto single instruction multiple data (SIMD) machines compared with a commercial compiler. A high-quality deinterlacer is implemented on a low-cost multicore platform for real-time high-definition applications by analyzing the multigrain parallelism. These case studies reveal the effectiveness of our parallel analysis framework which is applicable to generic systems. Compared with traditional graph traversal techniques, our linear algebraic approach impressively features low complexity and is practical for complicated algorithms.
Gwo Giun Lee, He-Yuan Lin, Chun-Fu Chen 0001, Tsung-Yuan Huang
IEEE Trans. Parallel Distributed Syst.1
2011 A high throughput parallel AVC/H.264 context-based adaptive binary arithmetic decoder
abstract
In this paper, based on the proposed parallelization scheme of binary arithmetic decoding, a parallel AVC/H.264 con text-based adaptive binary arithmetic coding (CABAC) de coder with high throughput is proposed. Following the top down design methodology, algorithm analyzing and data flow modeling in both high and low granularities are per formed to achieve the proposed architecture. According to the analysis for algorithm, the similarity between CABAC decoder and Viterbi decoder is found to extend the degree of parallelism for binary arithmetic decoding. The application of proposed design is specified to support AVC/H.264 High Profile, 4.2 Level, and 1920 × 1088 resolution at 64 frames per second. By increasing the degree of parallelism of bin decoding, the throughput of the proposed architecture is shown by the experiments to have improved 3.5 times as compared to the sequential bin decoding, and the decoded bin per second can reach 378M at clock speed 108MHz.
Jia-Wei Liang, He-Yuan Lin, Gwo Giun Lee
ICASSP3
2011 Reconfigurable inverse transform architecture for multiple purpose video coding
abstract
In this paper, an area efficient reconfigurable inverse transformation architecture for multiple standards is proposed. We present a top-down design methodology with complexity analysis, commonalities extraction, and dataflow modeling to systematically design reconfigurable architecture. By exporting and sharing the commonalities, the adder usage of the proposed reconfigurable inverse transform processing element can be reduced 44% compared with the total amount of adders in performing target inverse transform types. Then, the reconfigurable architecture is synthesized using TSMC 0.18 urn library. The working frequency is 108Mhz, which is derived from the dataflow scheduling. The area synthesis result is 32k gates, which indicates that the proposed design has more efficient area than other documented design in VLSI implementation. In addition, the proposed architecture also satisfies the accuracy requirement. Therefore, the proposed design have lower cost and enough flexibility for multi-standard purposes with 1920×1088 resolution and 64 frames per second and the color format is 4:2:0 for real time processing.
Tsung-Yuan Huang, He-Yuan Lin, Chun-Fu Chen 0001, Gwo Giun Lee
ISCAS4
2011 Stereoscopic video coding in AVS
abstract
This paper is an overview for AVS stereoscopic video coding technology, including two channels based inter-view prediction coding and stereo packing mode coding. The first one utilizes inter-view prediction to efficiently exploit the redundancy between the two channels of stereoscopic video. The superior coding performance of the inter-view prediction scheme benefits from an enhanced block prediction algorithm, which includes an improvement of direct mode for B-picture and motion vector prediction for P-picture. In addition, stereo packing mode, including side by side and top bottom, is adopted in AVS stereoscopic video coding to support the stereoscopic video service deployments based on the frame-compatible approach. Furthermore, an enhanced stereo packing mode is also developed to allow the prediction between signals coming from different channels within one packed frame. The simulation results demonstrate that the adopted techniques in AVS stereoscopic video are able to improve the compression efficiency of stereoscopic videos compared to simulcast one.
Xiangyang Ji, Yongbing Zhang 0002, Lu Yu 0003, Gwo Giun Lee
VCIP4
2010 Reconfigurable architecture design of motion compensation for multi-standard video coding
abstract
This paper proposes a reconfigurable video decoder architecture of motion compensation for multi-standard video coding including MPEG-2, MPEG-4 and H.264. Through top-down design methodology, we analyze the motion compensation algorithm of the targeted applications and extract the commonality of motion compensation algorithms among the three different standards. To design a reconfigurable processing element to perform the integer-sample and fractional-sample interpolation operations to simultaneously support the main video standards, a regular data flow is arranged during the design space exploration. In addition, the bandwidth reduction strategies are also adopted to reduce the memory access times and power consumption of motion compensation operations for high bandwidth requirement, especially in H.264. The design implementation of the proposed architecture is synthesized using TMSC 0.18um technology library and can operate at 108HMz to achieve the real time motion compensation coding of 1920×1088 at 30 frames per second in the three video standards.
Gwo Giun Lee, Wei-Chiao Yang, Min-Shan Wu, He-Yuan Lin
ISCAS1
2009 An RVC dataflow description of the AVC Constrained Baseline Profile decoder
abstract
Video codec applications become more and more complex to design. To ease the description of such applications, MPEG creates a Framework called Reconfigurable Video Coding (RVC). All existing codecs in MPEG have a similar structure, they are based on a hybrid decoding structure and some of their part can be reused on other design. In RVC, the dataflow is expressed using a network of components also called Functional Units (FUs) interconnected by FIFOs. An FU, programmed in CAL Language, includes the processing and the internal states. This paper puts the focus on a parallel dataflow description of the most complex MPEG RVC decoder available called MPEG4-AVC Constrained Baseline Profile (CBP) decoder.
Jérôme Gorin, Mickaël Raulet, Yuan-Long Cheng, He-Yuan Lin, Nicolas Siret, Kazuo Sugimoto, Gwo Giun Lee
ICIP7
2009 Algorithm and Architecture Design for Wide Range ELA Deinterlacer
abstract
In this paper we introduce a novel algorithm that can detect local features and choose a proper interpolation method for de-interlacing. An edge is a high frequency pattern with certain direction which is a noticeable feature in video sequences. We proposed a wide range ELA (WRELA) algorithm capable of accurately detecting edge directions. The edge direction can be acquired from an optimized procedure. Finally, we can interpolate the missing pixel in the edge with the direction which has the highest correspondence. We also implement the architecture of the proposed de-interlacing algorithm by UMC 0.18 mum technology. This design is capable of real-time de-interlacing for high definition 720i sequences with the clock speed running at 54 MHz, and the gate count is acceptable. Experimental results show that our proposed de-interlacer provides not only high objective performance in terms of PSNR but also impressive visual quality especially for edges.
Rong-Lai Lai, Bo-Han Chen, Gwo Giun Lee, He-Yuan Lin, Ming-Jiun Wang, Yuan-Long Cheng, Jia-Wei Liang
ISCAS3
2009 Low Complexity and High Throughput VLSI Architecture for AVC/H.264 CAVLC Decoding
abstract
This paper introduces a low complexity VLSI hardware architecture for entropy coding with increased throughput, based on the study of the statistical properties of the context-based adaptive variable length coding (CAVLC) in AVC/H.264. These enhanced designs are due to the results of the statistical analyses, in which better symbol length prediction was achieved by breaking the recursive dependency among codewords for multi-symbol decoder implementation. The proposed CAVLC decoder can also easily meet real-time requirements for high definition (HD) (1920times1080) applications, while the clock speed is operated only at 13 MHz under the best case scenario.
Gwo Giun Lee, Chia-Cheng Lo, Yuan-Ching Chen, Sheau-Fang Lei, He-Yuan Lin, Ming-Jiun Wang
ISCAS1
2009 A Motion-compensated Spectrum-adaptive Deinterlacing Algorithm
abstract
This paper presents a motion-compensated deinterlacing algorithm featuring spectrum-adaptive interpolation of interlaced field. Using motion-compensated reference pictures, the proposed spectrum-adaptive filter tactically identifies the baseband via spectrum analysis and removes the replicas of interlaced sampling, making overall algorithm adapt to versatile video scene and different degree of motion compensation scenarios. The experimental results indicate that our proposed algorithm has better objective performance than other motion-compensated and non-motion-compensated algorithms do especially in complex moving textures. The subjective results also support the benefits of our spectrum-adaptive filter.
Gwo Giun Lee, Ming-Jiun Wang, He-Yuan Lin, Ching-Jui Hsiao
ISCAS1
2009 Rate control algorithm based on intra-picture complexity for H.264/AVC
abstract
An efficient rate control algorithm based on the content-adaptive initial quantisation parameter (QP) setting scheme and the peak signal-to-noise ratio (PSNR) variation-limited bit-allocation strategy for low-complexity mobile applications is presented. This algorithm can efficiently measure the residual complexity of intra-pictures without performing the computation-intensive intra-prediction and mode decision in H.264/AVC, based on the structural and statistical features of local textures. This can adaptively set proper initial QP values for versatile video contents. In addition, this bit-allocation strategy can effectively distribute bit-rate budgets based on the monotonic property to enhance overall coding efficiency while maintaining the consistency of visual quality by limiting the variation of quantisation distortion. The experimental results reveal that the proposed algorithm surpasses the conventional rate control approaches in terms of the average PSNR from 0.34 to 0.95 dB. Moreover, this algorithm provides more impressive visual quality and more robust buffer controllability when compared with other algorithms.
Gwo Giun Lee, He-Yuan Lin, Ming-Jiun Wang
IET Image Process.1
2009 Algorithm/Architecture Co-Exploration of Visual Computing on Emergent Platforms: Overview and Future Prospects
abstract
Concurrently exploring both algorithmic and architectural optimizations is a new design paradigm. This survey paper addresses the latest research and future perspectives on the simultaneous development of video coding, processing, and computing algorithms with emerging platforms that have multiple cores and reconfigurable architecture. As the algorithms in forthcoming visual systems become increasingly complex, many applications must have different profiles with different levels of performance. Hence, with expectations that the visual experience in the future will become continuously better, it is critical that advanced platforms provide higher performance, better flexibility, and lower power consumption. To achieve these goals, algorithm and architecture co-design is significant for characterizing the algorithmic complexity used to optimize targeted architecture. This paper shows that seamless weaving of the development of previously autonomous visual computing algorithms and multicore or reconfigurable architectures will unavoidably become the leading trend in the future of video technology.
Gwo Giun Lee, Yen-Kuang Chen, Marco Mattavelli, E. S. Jang
IEEE Trans. Circuits Syst. Video Technol.1
2008 Textural complexity-based rate control algorithm
abstract
This paper presents an efficient rate control algorithm based on our content-adaptive initial quantization parameter setting scheme for H.264/AVC. For versatile video scenes, our algorithm can adaptively set an appropriate initial QP based on the textural complexity estimated from the first picture. In addition, our bit-allocation strategy effectively distributes the bit-rate budget based on the monotonic property to enhance the coding efficiency. Our proposed algorithm surpasses JVT-H014 rate control algorithm and Cauchy-density-based bit-allocation scheme in terms of average PSNR for about 0.47 dB and 0.81 dB respectively. Besides, our algorithm provides more impressive visual quality and more robust buffer controllability.
Gwo Giun Lee, He-Yuan Lin, Ming-Jiun Wang
ICME1
2008 On the efficient algorithm/architecture co-exploration for complex video processing
abstract
Targeted for highly sophisticated visual signal processing, we introduce in this paper complexity metrics or measures of algorithms which featuring architectural information are feedback or back annotated in early design stages to facilitate concurrent exploration of both algorithmic and architectural optimizations. With application to 3D spatio-temporal motion estimation for video coding, we have demonstrated significant reduction in design cost while the algorithmic performance still surpasses recent published works and even full search under many circumstances. Moreover, we have also shown the importance and substantiality of this complexity analysis technique in the extraction of features common to various different de-interlacing algorithms adapted for versatile video content, in designing highly efficient reconfigurable video architectures. As such this novel algorithm/architecture co-exploration methodology forms the basis for dataflow models with more accurate software/hardware partitioning resulting in multi-million gate and/or instruction software simulation platforms and fast prototyping hardware platforms for the next generation electronic system level design of SoC’s.
Gwo Giun Lee, Ming-Jiun Wang, He-Yuan Lin, Ron-Lai Lai
ICME1
2008 A high-quality spatial-temporal content-adaptive deinterlacing algorithm
abstract
This paper introduced a spatial-temporal content-adaptive algorithm, which can precisely select an appropriate interpolation technique for high-quality deinterlacing according to the spectral, edge-oriented and statistical features of local video content. Our algorithm employs a linear-phase statistical-adaptive vertical-temporal filter to deal with generic video scenes and adopts a modified edge-based line-averaging interpolation to efficiently recover moving edges. In addition, annoying flickering artifacts are efficiently suppressed by a flickering detection and a field-averaging filter. As a result, our algorithm outperforms other non-motion compensated methods in terms of objective PSNR and reveals more impressive subjective visual quality.
Gwo Giun Lee, He-Yuan Lin, Ming-Jiun Wang, Rong-Lai Lai, Chih Wen Jhuo
ISCAS1
2007 Motion Adaptive Deinterlacing via Edge Pattern Recognition
abstract
In this paper, a novel edge pattern recognition (EPR) deinterlacing algorithm with successive 4-field enhanced motion detection is introduced. The EPR algorithm surpasses the performance of ELA-based and other conventional methods especially at textural scenes. In addition, the current 4-field enhanced motion detection scheme overcomes conventional motion missing artifacts by gaining good motion detection accuracies and suppression of "motion missing" detection errors efficiently. Furthermore, with the incorporation of our new successive 4-field enhanced motion detection, the interpolation technique of EPR algorithm is capable of flexible adaptation in achieving better performance on textural scenes in generic video sequences.
Gwo Giun Lee, Hsin-Te Li, Ming-Jiun Wang, He-Yuan Lin
ISCAS1
2007 Algorithm/Architecture Co-Design of 3-D Spatio-Temporal Motion Estimation for Video Coding
abstract
This paper presents a new spatio-temporal motion estimation algorithm and its VLSI architecture for video coding based on algorithm and architecture co-design methodology. The algorithm consists of the new strategies of spatio-temporal motion vector prediction, modified one-at-a-time search scheme, and multiple update paths derived from optimization theory. The hardware specification is for high-definition video coding. We applied the ME algorithm to H.264 reference software. Our algorithm surpasses recently published research and achieves close performance to full search. The VLSI implementation proves the low cost feature of our algorithm. The algorithm and architecture co-design concept is highly emphasized in this paper. We provide some quantitative example to show the necessity of algorithm and architecture co-design
Gwo Giun Lee, Ming-Jiun Wang, He-Yuan Lin, Drew Wei-Chi Su, Bo-Yun Lin
IEEE Trans. Multim.1
2006 A 3D Spatio-Temporal Motion Estimation Algorithm for Video Coding
abstract
This paper presents a new spatio-temporal motion estimation algorithm for video coding. The algorithm is based on optimization theory and consists of the strategies including 3D spatio-temporal motion vector prediction, modified one-at-a-time search scheme, and multiple update paths. The simulation results indicate our algorithm is better than other recently proposed ones under the same computational budget and is very close to full search. The low-cost feature and regular demand of computational resource make our algorithm suitable for VLSI implementation. The algorithm also makes single chip solution for high-definition coding feasible
Gwo Giun Lee, Ming-Jiun Wang, He-Yuan Lin, Drew Wei-Chi Su, Bo-Yun Lin
ICME1
2006 Multiresolution-based texture adaptive motion detection for de-interlacing
abstract
Motion-adaptive de-interlacing algorithm selects from inter-field and intra-field interpolations according to motion. Correct determination of motion information is essential for this purpose. Fine textures, having high local pixel variation, tend to cause false detection of motion. This paper proposed a texture detection mechanism utilizing multiresolution technique to improve the correctness of detection. A recursive 2-field algorithm is also proposed to reduce the memory cost in 4-field method. This algorithm provides better perceptual visual quality than other motion adaptive de-interlacing as shown by the experimental results.
Gwo Giun Lee, Drew Wei-Chi Su, He-Yuan Lin, Ming-Jiun Wang
ISCAS1
2006 Model-based optimal rate control algorithm for real-time hybrid video encoder
abstract
This paper presents a frame level optimal rate control scheme based on the proposed rate and distortion functions. A linear rate-quantizer model and a linear distortion-quantizer model are proposed to perform rate-distortion optimization. The coefficients of rate distortion models are estimated by linear regression with the past rate distortion characteristics. We propose the two-stage strategy for rate-distortion optimization. First we use Lagrange multiplier optimization approach to obtain the closed-form solution based on our rate distortion models. Then we take the inter-frame dependency into account to further adjust bit rate allocation. We apply our optimal rate control algorithm to H.264, and the proposed algorithm outperforms JVT-G012 rate control scheme in terms of average PSNR.
He-Yuan Lin, Gwo Giun Lee, Ming-Jiun Wang, Drew Wei-Chi Su, Bo-Yun Lin
ISCAS2
2006 A unified systolic architecture for combined inter and intra predictions in H.264/AVC decoder
abstract
This paper presents a unified systolic architecture for inter and intra predictions in H.264/AVC decoder. To increase hardware utilization and minimize cost, we combine inter and intra predictions by a reprogrammable FIR filter, which is further implemented using systolic array. For intra prediction, the boundary pixels are reshuffled before feeding into the systolic array. For inter prediction, the 2-D interpolation is conducted through separable 1-D filtering. As compared with the state-of-the-art approaches, our architecture provides higher performance while maintaining relatively lower cost and input bandwidth. Specifically, up to 4x throughput improvement has been achieved. Moreover, the input bandwidth is significantly reduced. Further, combining inter and intra predictions saves the cost by 22~88%.
Chih-Hung Li, Chih-Chieh Chen, Drew Wei-Chi Su, Ming-Jiun Wang, Wen-Hsiao Peng, Tihao Chiang, Gwo Giun Lee
IWCMC7
1997 On Digital Mammogram Segmentation and Microcalcification Detection Using Multiresolution Wavelet Analysis
Chi Hau Chen, Gwo Giun Lee
CVGIP Graph. Model. Image Process.2
1996 A multiresolution wavelet analysis of digital mammograms
abstract
This paper discusses the significance of image segmentation via the combination of both statistical and nonstatistical methods based on the hierarchical framework of multiresolution wavelet analysis (MWA) and Gaussian Markov random fields (GMRF). Microcalculations and subtle mass regions are segmented via a fuzzy c-means (FCM) algorithm using localized features. For further enhancement, expected maximization and constrained optimization is applied to a Gibbs distribution defined from the FCM clustered image labels under a Bayesian framework. The effectiveness of this novel algorithm has been clearly illustrated by real mammographic images.
Chi Hau Chen, Gwo Giun Lee
ICPR2
1993 Neural networks for ultrasonic NDE signal classification using time-frequency analysis
Chi Hau Chen, Gwo Giun Lee
ICASSP (1)2