Jhing-Fa Wang

dblp:24/2878 · DBLP profile ↗
← Back
134ranked-venue papers
22as first author
9since 2021 · last 2026
0009-0000-2816-2480ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 51 · 7 first-author · 3 since 2021Artificial intelligence and machine learning · 39 · 8 first-author · 3 since 2021Systems, architecture and hardware · 32 · 3 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 9 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 1 since 2021Computer networks · 4 · 1 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Speech recognition and synthesis · 100%
Computer graphics and multimedia
7 papers
Multimedia analysis and retrieval · 38% Audio and music processing · 35% Image and video processing · 12%
Computer architecture, parallel and distributed computing, and storage systems
5 papers
Electronic design automation · 40% Integrated circuit design · 31% Hardware accelerators and domain-specific architectures · 22%

Topics — the 30 heaviest of 44, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Speech recognition and synthesis
text-to-speech synthesis
0.432017
Personalized Spontaneous Speech Synthesis Using a Small-Sized Unsegmented Semispontaneous Speech · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Variable-Length Unit Selection in TTS Using Structural Syntactic Cost · IEEE Trans. Speech Audio Process. 2007
Voice conversion using duration-embedded bi-HMMs for expressive speech synthesis · IEEE Trans. Speech Audio Process. 2006
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
speaker adaptation
0.312017
Personalized Spontaneous Speech Synthesis Using a Small-Sized Unsegmented Semispontaneous Speech · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Natural language and speech › Speech recognition and synthesis › text-to-speech synthesis
spontaneous speech synthesis
0.312017
Personalized Spontaneous Speech Synthesis Using a Small-Sized Unsegmented Semispontaneous Speech · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Natural language and speech › Speech recognition and synthesis
speech synthesis
0.212016
Candidate Expansion and Prosody Adjustment for Natural Speech Synthesis Using a Small Corpus · IEEE ACM Trans. Audio Speech Lang. Process. 2016
Multimedia analysis and retrieval › multimedia browsing
video browsing
0.112009
A Novel Video Summarization Based on Mining the Story-Structure and Semantic Relations Among Concept Entities · IEEE Trans. Multim. 2009
Multimedia analysis and retrieval
video summarization
0.112009
A Novel Video Summarization Based on Mining the Story-Structure and Semantic Relations Among Concept Entities · IEEE Trans. Multim. 2009
Natural language and speech › Speech recognition and synthesis › text-to-speech synthesis
prosody modeling
0.112007
Variable-Length Unit Selection in TTS Using Structural Syntactic Cost · IEEE Trans. Speech Audio Process. 2007
Natural language and speech › Speech recognition and synthesis › text-to-speech synthesis
unit selection
0.112007
Variable-Length Unit Selection in TTS Using Structural Syntactic Cost · IEEE Trans. Speech Audio Process. 2007
Image and video processing › image restoration
image inpainting
0.112007
A Hybrid Algorithm With Artifact Detection Mechanism for Region Filling After Object Removal From a Digital Photograph · IEEE Trans. Image Process. 2007
Audio and music processing › speaker diarization
speaker change detection
0.112007
Unsupervised Speaker Change Detection Using SVM Training Misclassification Rate · IEEE Trans. Computers 2007
Audio and music processing
speech processing
0.112007
Unsupervised Speaker Change Detection Using SVM Training Misclassification Rate · IEEE Trans. Computers 2007
Natural language and speech › Speech recognition and synthesis › speech synthesis
expressive speech synthesis
0.112006
Voice conversion using duration-embedded bi-HMMs for expressive speech synthesis · IEEE Trans. Speech Audio Process. 2006
Natural language and speech › Speech recognition and synthesis
voice conversion
0.112006
Voice conversion using duration-embedded bi-HMMs for expressive speech synthesis · IEEE Trans. Speech Audio Process. 2006
Audio and music processing › audio analysis
audio segmentation
0.112005
Automatic Scene Change Detection for Composed Speech and Music Sound Under Low SNR Noisy Environment · IEEE Trans. Speech Audio Process. 2005
Audio and music processing
computational auditory scene analysis
0.112005
Automatic Scene Change Detection for Composed Speech and Music Sound Under Low SNR Noisy Environment · IEEE Trans. Speech Audio Process. 2005
Multimedia analysis and retrieval › video content analysis
scene change detection
0.112005
Automatic Scene Change Detection for Composed Speech and Music Sound Under Low SNR Noisy Environment · IEEE Trans. Speech Audio Process. 2005
Integrated circuit design
digital circuit design
0.012002
Chip design of portable speech memopad suitable for persons with visual disabilities · IEEE Trans. Speech Audio Process. 2002
Hardware accelerators and domain-specific architectures › signal processing accelerator
speech recognition accelerator
0.012002
Chip design of portable speech memopad suitable for persons with visual disabilities · IEEE Trans. Speech Audio Process. 2002
Image and video coding › error resilience
packet loss recovery
0.012001
A voicing-driven packet loss recovery algorithm for analysis-by-synthesis predictive speech coders over Internet · IEEE Trans. Multim. 2001
Audio and music processing
speech coding
0.012001
A voicing-driven packet loss recovery algorithm for analysis-by-synthesis predictive speech coders over Internet · IEEE Trans. Multim. 2001
Image and video processing
document image analysis
0.012000
Segmentation of Single- or Multiple-Touching Handwritten Numeral String Using Background and Foreground Analysis · IEEE Trans. Pattern Anal. Mach. Intell. 2000
Electronic design automation
hardware verification and test
0.031992
Enhancing the multiple-fault detection of single-fault test sets · Comput. Aided Des. 1992
A Fault Analysis Method for Synchronous Sequential Circuits · DAC 1990
A New Approach to Derive Robust Sets for Stuck-open Faults in CMOS Combinational Logic Circuits · DAC 1989
Visual content generation and editing
texture synthesis
0.012007
A Hybrid Algorithm With Artifact Detection Mechanism for Region Filling After Object Removal From a Digital Photograph · IEEE Trans. Image Process. 2007
Electronic design automation › hardware verification and test
fault detection
0.011992
Enhancing the multiple-fault detection of single-fault test sets · Comput. Aided Des. 1992
Electronic design automation › hardware verification and test › fault detection
multiple fault detection
0.011992
Enhancing the multiple-fault detection of single-fault test sets · Comput. Aided Des. 1992
Electronic design automation › hardware verification and test
fault analysis
0.011990
A Fault Analysis Method for Synchronous Sequential Circuits · DAC 1990
Integrated circuit design › digital circuit design › VLSI architecture
viterbi decoder
0.011990
A new transform algorithm for Viterbi decoding · IEEE Trans. Commun. 1990
Integrated circuit design › VLSI design
VLSI array
0.011990
A new transform algorithm for Viterbi decoding · IEEE Trans. Commun. 1990
Algorithms and data structures › numerical algorithms
transform algorithm
0.011990
A new transform algorithm for Viterbi decoding · IEEE Trans. Commun. 1990
Coding theory › error-correcting codes › convolutional codes › convolutional code decoding
viterbi decoding
0.011990
A new transform algorithm for Viterbi decoding · IEEE Trans. Commun. 1990

Methods — techniques the papers use, named apart from their topics

prosodic word smoothing · 0.3modulation spectrum postfilter · 0.3average voice model · 0.3unit selection · 0.2prosody adjustment · 0.2candidate expansion · 0.2maximum entropy criterion · 0.1graph entropy model · 0.1concept expansion · 0.1weighted interpolation · 0.1support vector machine · 0.1probabilistic context-free grammar · 0.1latent semantic analysis · 0.1kirsch edge detector · 0.1dynamic programming · 0.1artifact detection · 0.1semantic clustering · 0.1key frame extraction · 0.1
YearPublicationVenuePosition
2026 LQ-FJS: A logical query-digging fake-news judgment system with structured video-summarization engine using LLM
Jhing-Fa Wang, Din-Yuen Chan, Hsin-Chun Tsai, Bo-Xuan Fang
Data Knowl. Eng.1
2026 Disentangled Representation and Contrastive Learning with Adaptive Affinity Squeeze-Excitation for Multimodal Emotion Recognition
abstract
Multimodal Emotion Recognition (MER) aims to leverage information from multiple modalities to enhance the accuracy of emotion classification. However, existing methods struggle with properly disentangling shared and modality-specific features while maintaining discriminative power across modalities. To address this, we propose a novel contrastive learning-based MER framework that enforces inter-modal alignment while preserving intra-modal distinctiveness. Our method employs disentangled representation learning to extract both shared and private representations, ensuring that complementary features are effectively utilized. Additionally, we introduce an Adaptive Affinity Squeeze-Excitation (AASE) mechanism to dynamically refine multimodal representations by capturing inter-channel relationships, further enhancing the model’s ability to identify crucial emotional cues. Experimental results demonstrate that our method improves emotion classification accuracy to 84.6% on benchmark datasets. Accordingly, the proposed MER framework provides a robust and highly adaptive solution for multimodal emotion analysis, effectively integrating modality-specific features while enhancing discriminability across modalities.
Jhing-Fa Wang, Hsin-Chun Tsai, Jie-Ming Liu
Int. J. Pattern Recognit. Artif. Intell.1
2026 Effective personalized cross-lingual TTS with speech synthesis of prosodic naturalness and vivid emotion
Din-Yuen Chan, Jun-Hao Zhong, Jhing-Fa Wang, Hsin-Chun Tsai
Multim. Tools Appl.3
2025 A Fine-tuning free GenAI method with iterative backpropagation algorithm and single-object reference for quantity-faithfulness
Din-Yuen Chan, Yu-Min Hsu, Jhing-Fa Wang
Multim. Tools Appl.3
2025 An innovative ESGH-RAG module with ChatGPT-4o for automatic ESG-report generation
Jhing-Fa Wang, Wen-Yuan Zhang, Shih-Pang Tseng
J. Supercomput.1
2024 A new speaker-diarization technology with denoising spectral-LSTM for online automatic multi-dialogue recording
Din-Yuen Chan, Jhing-Fa Wang, Hsu-Ting Chin
Multim. Tools Appl.2
2023 Cascading AB-YOLOv5 and PB-YOLOv5 for rib fracture detection in frontal and oblique chest X-ray images
abstract
Abstract Convolutional deep learning models have shown comparable performance to radiologists in detecting and classifying thoracic diseases. However, research on rib fractures remains limited compared to other thoracic abnormalities. Moreover, existing deep learning models primarily focus on using frontal chest X‐ray (CXR) images. To address these gaps, the authors utilised the EDARib‐CXR dataset, comprising 369 frontal and 829 oblique CXRs. These X‐rays were annotated by experienced radiologists, specifically identifying the presence of rib fractures using bounding‐box‐level annotations. The authors introduce two detection models, AB‐YOLOv5 and PB‐YOLOv5, and train and evaluate them on the EDARib‐CXR dataset. AB‐YOLOv5 is a modified YOLOv5 network that incorporates an auxiliary branch to enhance the resolution of feature maps in the final convolutional network layer. On the other hand, PB‐YOLOv5 maintains the same structure as the original YOLOv5 but employs image patches during training to preserve features of small objects in downsampled images. Furthermore, the authors propose a novel two‐level cascaded architecture that integrates both AB‐YOLOv5 and PB‐YOLOv5 detection models. This structure demonstrates improved metrics on the test set, achieving an AP30 score of 0.785. Consequently, the study successfully develops deep learning‐based detectors capable of identifying and localising fractured ribs in both frontal and oblique CXR images.
Hsin-Chun Tsai, Nan-Han Lu, Kuo-Ying Liu, Chuan-Han Lin, Jhing-Fa Wang
IET Comput. Vis.5
2023 Guest Editorial Introduction to the Special Issue on AI-Empowered Trajectory Analytics in Intelligent Transportation Systems
abstract
With the rapid growth of location sensing in the Internet of Things (IoT) and Internet of Vehicles (IoV) techniques, trajectory data has been generated that can be used to describe diversity and characteristics of moving objects. The analysis and management of trajectory patterns has become an important issue in recent decades, as it supports efficient strategies and decisions based on discovered patterns and knowledge from the mobility behavior of customers or citizens in many fields and applications (e.g., smart city, intelligent transportation, location-based services, health management, etc.). Since the recent development in AI, it is possible to use AI-based techniques to analyze trajectory data at an unprecedented scale to address applicable issues of effectiveness, efficiency, accuracy, and privacy in Intelligent Transportation Systems (ITS), therefore, it is a highly competitive area to propose innovative methods, principles, procedures, techniques, frameworks, theories, and applications to address the aforementioned challenges of trajectory data in ITS. This Special Issue is intended to provide a forum for all researchers from academia and industry to share their original, creative, innovative, cutting-edge insights, theories, ideas, and developments for the analysis of trajectory data using AI-empowered techniques in ITS. In this special issue, we received 52 submissions, and finally accepted 17 articles to be published in the special issue. Below is a brief introduction to each of them.
Jerry Chun-Wei Lin, Gautam Srivastava 0001, Jhing-Fa Wang
IEEE Trans. Intell. Transp. Syst.3
2022 Guest Editorial Special Issue on Secure Data Analytics for Emerging Internet of Things
abstract
The rapid developments in hardware, software, and communication technologies have facilitated the spread of interconnected sensors, actuators, and heterogeneous devices such as single board computers, which collect and exchange a large amount of data to offer a new class of advanced services characterized by being available anywhere, at any time, and for anyone. This ecosystem is widely referred to as the Internet of Things (IoT). In the past years, the number of deployments both for sensor networks and the IoT grew significantly. This continuous and exponential growth is facilitated by investments and research activities originating from industry, academia, and governments while the penetration of these technologies is also driven by the high technology acceptance rates of both consumers and technologists across disciplines. Such networks collect, store, and exchange a large volume of heterogeneous data. Nevertheless, their rapid and widespread deployment, along with their participation in the provisioning of potentially critical services (e.g., safety applications, healthcare, and manufacturing) raise numerous issues related to the security, data analysis, and energy awareness of the performed operations and provided services.
Sachin Shetty, Jhing-Fa Wang, Uttam Ghosh, Schahram Dustdar
IEEE Internet Things J.2
2020 Automatic drug pills detection based on enhanced feature pyramid network and convolution neural networks
abstract
Drug pill detection is one of the most important tasks in medication safety. The correct identification of drug based on the visual appearance is a key step towards the improvement of medication safety. Previous studies have aimed to recognise a drug based on the front or back view of the drug under a fixed viewing angle. In cases with multiple drugs and randomly placed drugs, the previous methods have difficulties in detecting and recognising different drugs in practical applications. A convolution neural network‐based detector is proposed in this work to overcome the difficulties and to assist patients in drug identification. The proposed system includes a localisation stage and a classification stage. The enhanced feature pyramid network (EFPN), is proposed for drug localisation, and Inception‐ResNet v2 is used in drug classification. The proposed Drug Pills Image Database contains a collection of 612 categories of drug datasets for deep learning research in the pharmaceutical field. The proposed EFPN achieves over 96% accuracy in the localisation experiment. In the complete system evaluation, the proposed system has obtained the Top‐1, Top‐3, and Top‐5 accuracies of 82.1, 92.4, and 94.7%, respectively.
Yang-Yen Ou, An-Chao Tsai, Xuan-Ping Zhou, Jhing-Fa Wang
IET Comput. Vis.4
2020 Fuzzy Obstacle Avoidance for the Mobile System of Service Robots
abstract
This study implements Fuzzy logic-based obstacle avoidance and human tracking on an omnidirectional mobile system for service robots. The mobile system could be separated and combined with the robot which can be controlled remotely and switched to go forward and avoid obstacles in an indoor environment automatically. The system is able to track and go to the user according to the user’s position. The omnidirectional wheel was adapted in the power system to perform translating and spinning movements. The translating movement enables the robot to avoid obstacles faster and flexibly in paths. With the spinning movement, the robot can quickly find the direction of the object. Finally, the experiments show that the proposed system has good performance in service environments.
Shih-Pang Tseng, Che-Wen Chen, Ta-Wen Kuan, Yao-Tsung Hsu, Jhing-Fa Wang
Wirel. Commun. Mob. Comput.5
2017 Introduction to special section on advances of orange technologies
Lei Xie 0001, Jhing-Fa Wang
Frontiers Comput. Sci.2
2017 Spoken dialog summarization system with HAPPINESS/SUFFERING factor recognition
Yang-Yen Ou, Ta-Wen Kuan, Anand Paul 0001, Jhing-Fa Wang, An-Chao Tsai
Frontiers Comput. Sci.4
2017 Personalized Spontaneous Speech Synthesis Using a Small-Sized Unsegmented Semispontaneous Speech
abstract
A systematic approach is proposed to synthesizing personalized spontaneous speech using a small-sized unsegmented speech corpus of the target speaker. First, an automatic segmentation algorithm is employed to segment and label a collected semispontaneous speech corpus of the target speaker. Then, a pretrained average voice model is adapted to the voice model of the target speaker by using the segmented data. A postfilter based on modulation spectrum is adopted to further improve the speaker similarity of the synthesized speech as well as alleviate the over-smoothing problem of the synthesized speech. For generating spontaneous speech, a smoothing method applied at the prosodic word level is proposed to improve speech fluency. For objective evaluation on spontaneous speech segmentation, the segmentation accuracy of the proposed method is superior to that of Viterbi-based forced alignment. The results of subjective listening test also show that the proposed method can improve the spontaneity and speaker similarity of the synthesized speech compared to the maximum likelihood linear regression based speaker adaptation method.
Yi-Chin Huang, Chung-Hsien Wu 0001, Yan-You Chen, Ming-Ge Shie, Jhing-Fa Wang
IEEE ACM Trans. Audio Speech Lang. Process.5
2016 Multivoxel analysis for functional magnetic resonance imaging (fMRI) based on time-series and contextual information: relationship between maternal love and brain regions as a case study
Bo-Wei Chen, Yang-Yen Ou, Chun-Chia Kung, Ding-Ruey Yeh, Seungmin Rho, Jhing-Fa Wang
Multim. Tools Appl.6
2016 Candidate Expansion and Prosody Adjustment for Natural Speech Synthesis Using a Small Corpus
abstract
This study proposes a hybrid approach to natural-sounding speech synthesis based on candidate expansion, unit selection, and prosody adjustment using a small corpus. The proposed method is more specific to tonal language, in particular Mandarin. In conventional speech synthesis studies, the quality of synthesized speech depends heavily on the size of the speech corpus. However, it is highly time-consuming and labor-intensive to prepare a large labeled corpus. In this work, candidate expansion is proposed to retrieve potential candidates that are unlikely to be retrieved using only linguistic features. The optimal unit sequence is then obtained from the expanded candidates by using the proposed unit selection mechanism at the phoneme and prosodic word levels. Finally, a prosodic word-level prosody adjustment is proposed to improve the continuity and smoothness of the prosody of the synthesized speech. To evaluate the proposed method, the Tsing-Hua corpus of speech synthesis was adopted. The results of an objective evaluation demonstrate the effectiveness of candidate expansion and the improvement of the continuity and smoothness of the prosody of the synthesized speech. The results of a subjective evaluation also show the proposed system could synthesize the speech with improved quality and naturalness, in particular for a small-sized or resource-limited corpus.
Yan-You Chen, Chung-Hsien Wu 0001, Yi-Chin Huang, Shih-Lun Lin, Jhing-Fa Wang
IEEE ACM Trans. Audio Speech Lang. Process.5
2016 Fixed-Point Computing Element Design for Transcendental Functions and Primary Operations in Speech Processing
abstract
This brief presents a fixed-point architecture based on a reconfigurable scheme for integrating several commonly used mathematical operations of speech signal processing. The proposed design can perform two transcendental mathematical operations called logarithm and powering, and three commonly used computations with similar operations named polynomial calculation, filtering, and windowing. By analyzing the adopted algorithms of the above five operations, a simplified computing unit is designed. This unit can combine six types of operations by reconfiguring the data paths, and the same multiply-add architecture can be reused for reducing the redundant usage of logic gates. The experimental results reveal that the proposed design can work at a 200-MHz clock rate, and its gate count only has 11.9k. Compared with the results of the floating-point function, the median errors of the proposed design for computing the powering and logarithmic functions are 0.57% and 0.11%, respectively. Such results indicate that this simple architecture can be effectively used in most speech processing applications.
Chung-Hsien Chang, Shi-Huang Chen, Bo-Wei Chen, Wen Ji 0003, K. Bharanitharan, Jhing-Fa Wang
IEEE Trans. Very Large Scale Integr. Syst.6
2016 A New Binary-Halved Clustering Method and ERT Processor for ASSR System
abstract
This paper presents an automatic speech–speaker recognition (ASSR) system implemented in a chip which includes a built-in extraction, recognition, and training (ERT) core. For VLSI design (here, ASSR system), the hardware cost and time complexity are always the important issues which are improved in this proposed design in two levels: 1) algorithmic and 2) architecture. At the algorithm level, a newly binary-halved clustering (BHC) is proposed to achieve low time complexity and low memory requirement. In addition, at the architecture level, a new ERT core is proposed and implemented based on data dependence and reuse mechanism to reduce the time and hardware cost as well. Finally, the chip implementation is synthesized, placed, and routed using TSMC 90-nm technology library. To verify the performance of the proposed BHC method, a case study is performed based on nine speakers. Moreover, the validation of the ASSR system is examined in two parts: 1) speech recognition and 2) speaker recognition. The results show that the proposed system can achieve 93.38% and 87.56% of recognition rates during speech and speaker recognition, respectively. Furthermore, the proposed ASSR chip includes 396k gate counts, and consumes power in 8.74 mW. Such results demonstrate that the performance of the proposed ASSR system is superior to the conventional systems.
Chih-Hung Chou, Ta-Wen Kuan, Shovan Barma, Bo-Wei Chen, Wen Ji 0003, Chih-Hsiang Peng, Jhing-Fa Wang
IEEE Trans. Very Large Scale Integr. Syst.7
2015 Subspace-based DOA with linear phase approximation and frequency bin selection preprocessing for interactive robots in noisy environments
Sheng-Chieh Lee, Bo-Wei Chen, Jhing-Fa Wang, Min-Jian Liao
Comput. Speech Lang.3
2015 User-centric incremental learning model of dynamic personal identification for mobile devices
Hsin-Chun Tsai, Bo-Wei Chen, K. Bharanitharan, Anand Paul 0001, Jhing-Fa Wang, Hung-Chieh Tai
Multim. Syst.5
2015 Quantitative Measurement of Split of the Second Heart Sound (S2)
abstract
This study proposes a quantitative measurement of split of the second heart sound (S2) based on nonstationary signal decomposition to deal with overlaps and energy modeling of the subcomponents of S2. The second heart sound includes aortic (A2) and pulmonic (P2) closure sounds. However, the split detection is obscured due to A2-P2 overlap and low energy of P2. To identify such split, HVD method is used to decompose the S2 into a number of components while preserving the phase information. Further, A2s and P2s are localized using smoothed pseudo Wigner-Ville distribution followed by reassignment method. Finally, the split is calculated by taking the differences between the means of time indices of A2s and P2s. Experiments on total 33 clips of S2 signals are performed for evaluation of the method. The mean ± standard deviation of the split is 34.7 ± 4.6 ms. The method measures the split efficiently, even when A2-P2 overlap is ≤ 20 ms and the normalized peak temporal ratio of P2 to A2 is low (≥ 0.22). This proposed method thus, demonstrates its robustness by defining split detectability (SDT), the split detection aptness through detecting P2s, by measuring up to 96 percent. Such findings reveal the effectiveness of the method as competent against the other baselines, especially for A2-P2 overlaps and low energy P2.
Shovan Barma, Bo-Wei Chen, Ka Lok Man, Jhing-Fa Wang
IEEE ACM Trans. Comput. Biol. Bioinform.4
2015 Low-Complexity Hardware Design for Fast Solving LSPs With Coordinated Polynomial Solution
abstract
This paper presents a low-complexity algorithm and the corresponding hardware based on the coordinated polynomial solutions for solving line spectrum pairs (LSPs). To improve the computation of LSPs, the enhanced Tschirnhaus transform (ETT) is proposed to accelerate the coordinated polynomial solution. The proposed ETT can replace fractional multiplication with addition and shift operations, so unnecessary operations are avoided. To further simplify the hardware of the ETT, three designs are presented: the preprocessing block (PPB), the iterative root-finding block (IRFB), and the closed-form solution block (CFSB). The PPB provides a design with less gate counts that can effectively transform LPCs into general-form polynomials. Such polynomials can be further decomposed into roots using the proposed IRFB based on the Birge-Vieta method. A pipeline-recursive framework is implemented in the IRFB to save calculations. To improve hardware utilization, this paper also analyzes the coefficients relationship of the ETT by introducing the data dependency graph to design the proposed functional blocks in CFSB. The experimental results show that the proposed hardware achieves a 40-fold improvement in throughput and reduces 1.16% of gate counts at the hardware synthesis level; the chip area is 1.29 mm2. The precision analysis indicates the average log spectral distance is 0.310. Moreover, the ETT in the proposed hardware only requires 29.9% of multiplication compared with the original one. Such results reveal that the proposed work is superior to the baseline work, thereby demonstrating the effectiveness of the proposed design.
Chung-Hsien Chang, Bo-Wei Chen, Shi-Huang Chen, Jhing-Fa Wang, Yu-Hao Chiu
IEEE Trans. Very Large Scale Integr. Syst.4
2015 Trainable and Low-Cost SMO Pattern Classifier Implemented via MCMC and SFBS Technologies
abstract
This paper presents a multicore and multichannel (MCMC) technology and a synchronous and forward-backward scheduling (SFBS) for the cost reduction of sequential minimal optimization trainable pattern classifier. The MCMC technology uses multiple processing cores that are self-reconfigurable and preconfigurable. For different functions, five self-configurable modes and four preconfigurable modes can be combined to achieve high flexibility. A multichannel hierarchical architecture enables different transfer rates. To minimize communication cost, the SFBS uses synchronous and forward-backward counting for data scheduling. For implementation in reconfigurable FPGAs, MCMC and SFBS are combined for use in synthesis, placement, and routing. Compared with the baseline design, the emulation results show that the proposed architecture has a low area and low power costs (5755 logic elements and 195 mW), respectively. The experimental results confirm the cost improvement achieved by the proposed architecture and methods.
Chih-Hsiang Peng, Ta-Wen Kuan, Po-Chuan Lin, Jhing-Fa Wang, Guo-Ji Wu
IEEE Trans. Very Large Scale Integr. Syst.4
2014 Speech-driven talking face using embedded confusable system for real time mobile multimedia
Po-Yi Shih, Anand Paul 0001, Jhing-Fa Wang, Yi-Hung Chen
Multim. Tools Appl.3
2014 Enhanced long-range personal identification based on multimodal information of human features
Hsin-Chun Tsai, Bo-Wei Chen, Jhing-Fa Wang, Anand Paul 0001
Multim. Tools Appl.3
2014 REC-STA: Reconfigurable and Efficient Chip Design With SMO-Based Training Accelerator
abstract
Sequential minimal optimization (SMO) and Karush-Kuhn-Tucker condition are often used to solve learning problems in support vector machines. However, during hardware implementation of the SMO algorithm, enhancing chip performance without excessively increasing chip area is often a crucial issue. The solution proposed in this paper is a novel reconfigurable and efficient chip design with SMO-based training accelerator (REC-STA). Two novel methods used in the proposed REC-STA are trimode coarse-grained reconfigurable architecture (TCRA) and triple finite-state-machine with dynamic scheduling The first method modifies the baseline SMO design by developing trimode reconfigurable architectures with parallel and pipeline computing capabilities. The second method provides a schedule for efficient reconfiguration of the TCRA. Use of these methods can remove kernel cache design. For chip manufacturing, the implementation of the REC-STA is synthesized, placed, and routed using the TSMC 0.18-μm technology library. The core size is 2.94 mm × 2.94 mm and the power consumption is 77.3 mW. Compared with the baseline design, the FPGA simulation results show that the proposed architecture requires 50% less memory and 31% fewer gate counts but provides a 16-fold improvement in training performance. The experimental results confirm the efficacy of the proposed architecture and methods.
Chih-Hsiang Peng, Bo-Wei Chen, Ta-Wen Kuan, Po-Chuan Lin, Jhing-Fa Wang, Nai-Sheng Shih
IEEE Trans. Very Large Scale Integr. Syst.5
2013 High-efficient hardware design based on enhanced Tschirnhaus transform for solving the LSPs
abstract
This work presents a novel hardware design based on the enhanced Tschirnhaus transform (ETT) to solve the 8-order line spectral pairs (LSPs). To reduce high-complexity problems caused by fractional multiplication, the ETT is proposed to replace such operations with integer-based shift and addition operations of the original Tschirnhaus transform. Also, the data dependency graph (DDG) of the ETT is analyzed for designing hardware units and reducing computation cycles. The proposed hardware has two key blocks: the mixture computation unit (MCU) and the multiplier-free pipelined square-root unit (PSRU). The first block is designed to fast calculate multiplication and summation operations in the ETT with the use of a two-stage pipeline architecture. The second is developed to speed up square-root operations after 8-order LSPs are decomposed into two 4-order LSPs. It can also timely process the result of the first block within limited cycles. The experimental results show that compared with the Chebyshev-based research, the proposed hardware can reduce the cycle times by 98.1% and also saved about 49.7% of gate counts. In the precision evaluation, the result indicates that 95% of the computation errors are within 0.02 and proves that the proposed hardware is capable of quantizing LSPs almost as accurately as computers do. Such results reveal that the proposed work is superior to the other Chebyshev-based methods, thereby demonstrating the effectiveness of the proposed design.
Chung-Hsien Chang, Shi-Huang Chen, Bo-Wei Chen, Chih-Hsiang Peng, Jhing-Fa Wang
ISCAS5
2013 Video search and indexing with reinforcement agent for interactive multimedia services
abstract
In this study, we present a video search and indexing system based on the state support vector (SVM) network, video graph, and reinforcement agent for recognizing and organizing video events. In order to enhance the recognition performance of the state SVM network, two innovative techniques are presented: state transition correction and transition quality estimation. The classification results are also merged into the video indexing graph, which facilitates the search speed. A reinforcement algorithm with an efficient scheduling scheme significantly reduces both the power consumption and time. The experimental results show the proposed state SVM network was able to achieve a precision rate as high as 83.83% and the query results of the indexing graph reached 80% accuracy. The experiments also demonstrate the performance and feasibility of our system.
Anand Paul 0001, Bo-Wei Chen, K. Bharanitharan, Jhing-Fa Wang
ACM Trans. Embed. Comput. Syst.4
2013 Smart Homecare Surveillance System: Behavior Identification Based on State-Transition Support Vector Machines and Sound Directivity Pattern Analysis
abstract
This study presents a smart homecare surveillance system, which utilizes sound-steered cameras to identify behavior of interest. First of all, to detect multiple source locations, a new direction-of-arrival (DOA) algorithm is proposed by introducing cascaded frequency filters, which can quickly calculate directions without creating much complexity. This method can also locate and separate different signals at the same time. Second, after the camera points in the direction of the estimated angle, the proposed state-transition support vector machine is used to provide favorable discriminability for human behavior identification. A new Markov random field (MRF) function based on the localized contour sequence (LCS) is also presented while the system computes transition probabilities between states. Such LCS-based MRF functions can effectively smooth transitions and enhance recognition. The experimental results show that the average error of DOA decreases to around 7°, which is better than those of the baselines. Also, our proposed behavior identification system can reach an 88.3% accuracy rate. The aforementioned results have therefore demonstrated the feasibility of the proposed method.
Bo-Wei Chen, Chen-Yu Chen, Jhing-Fa Wang
IEEE Trans. Syst. Man Cybern. Syst.3
2012 Novel Binaural Spectro-temporal Algorithm for Speech Enhancement in Low SNR Environments
abstract
A novel BInaural Spectro-Temporal (BIST) algorithm is proposed in this paper to increase the speech intelligibility in low or negative SNR noisy environments. The BIST algorithm consists of two modules. One is the spatial mask for receiving sound from the specific direction, and the other is the spectro-temporal modulation filter for noise reduction. Most speech enhancement algorithms are not applicable in harsh environments because the energy of speech is covered by the noise. To increase the speech intelligibility in low or negative SNR noisy environments, a distinctive approach is proposed to solve this problem. First, the BIST algorithm takes binaural auditory processing as a spatial mask to separate the speech and noise according to their locations. Next, the modulation filter is applied to reduce the noise source in the scale-rate (spectro-temporal modulation) domain according to their different acoustic feature. It works like the spectro-temporal receptive field (STRF) which is the perception response of human auditory cortex. The experimental results demonstrate that the proposed BIST speech enhancement algorithm can improve 20% from the noisy speech at SNR-10dB.
Po-Hsun Sung, Bo-Wei Chen, Ling-Sheng Jang, Jhing-Fa Wang
ICME4
2012 A new stereo packing format based on checkerboard sub-sampling for efficient stereo video coding
abstract
Since the 3D coding and transmission standard is not available nowadays, we need to transfer the 3D contents through the existing coding standards. To achieve this, the most common way is to combine both left and right eye frames into one single frame, which is the so-called stereo packing. So far, there are many packing methods and each of them has different characteristics and advantages. To obtain the better coding efficiency and reconstructed quality, we propose an adaptive checkerboard-based stereo coding system to retain the advantages of most different packing methods. The proposed system first computes the features of the stereo image pairs then finds the most suitable packing mode for efficient transmission. In the decoder side, the proposed de-packing mechanism is first adopted to find the correct packing mode. Afterward, an edge-dependent interpolation method is performed to attain better reconstructed quality. Experimental results reveal that the proposed method can achieve better results when comparing with other existing methods.
An-Ti Chiang, Hung-Ming Wang, Jar-Ferr Yang, Jhing-Fa Wang
ISCAS4
2012 A new hybrid and dynamic fusion of multiple experts for intelligent porch system
Ta-Wen Kuan, Hsin-Chun Tsai, Jhing-Fa Wang, Jia-Ching Wang, Bo-Wei Chen, Zong-You Lin
Expert Syst. Appl.3
2012 Effective Search Point Reduction Algorithm and its VLSI Design for HDTV H.264/AVC Variable Block Size Motion Estimation
abstract
Variable block size motion estimation (VBSME) in H.264/AVC has greatly led to achieve an optimal inter frame encoding. However, the computation burden of the VBSME becomes the bottleneck of the H.264/AVC encoders. The conventional architecture in hardware realization is hard to adopt a fast software algorithm suitable to reduce the VBSME computation burden. Therefore, this paper presents a search point reduction (SPR) algorithm with an efficient hardware design, able to decrease the motion estimation time while maintaining the coding performance of H.264. The effectiveness of the proposed method is compared with those of existing methods with respect to chip area, operation frequency, and throughput rate. The proposed SPR algorithm increases the coding speed by around 90%; with a peak signal-to-noise ratio drop of less than 0.1 dB than that achieved by the JM reference software. The proposed SPR algorithm can operate at 200 MHz with 191 k gate count, which supports high-definition television 1280 720 format.
An-Chao Tsai, K. Bharanitharan, Jhing-Fa Wang, Kuan-I Lee
IEEE Trans. Circuits Syst. Video Technol.3
2012 Parallel Reconfigurable Computing-Based Mapping Algorithm for Motion Estimation in Advanced Video Coding
abstract
Computational load of motion estimation in advanced video coding (AVC) standard is significantly high and even worse for HDTV and super-resolution sequences. In this article, a video processing algorithm is dynamically mapped onto a new parallel reconfigurable computing (PRC) architecture which consists of multiple dynamic reconfigurable computing (DRC) units. First, we construct a directed acyclic graph (DAG) to represent video coding algorithms in which motion estimation is the focus. A novel parallel partition approach is then proposed to map motion estimation DAG onto the multiple DRC units in a PRC system. This partitioning algorithm is capable of design optimization of parallel processing reconfigurable systems for a given number of processing elements in different search ranges. This speeds up the video processing with minimum sacrifice.
Anand Paul 0001, Yung-Chuan Jiang, Jhing-Fa Wang, Jar-Ferr Yang
ACM Trans. Embed. Comput. Syst.3
2012 VLSI Design of an SVM Learning Core on Sequential Minimal Optimization Algorithm
abstract
The sequential minimal optimization (SMO) algorithm has been extensively employed to train the support vector machine (SVM). This work presents an efficient application specific integrated circuit chip design for sequential minimal optimization. This chip is implemented as an intellectual property core, suitable for use in an SVM-based recognition system on a chip. The proposed SMO chip was tested and found to be fully functional, using a prototype system based on the Altera DE2 board with a Cyclone II 2C70 field-programmable gate array.
Ta-Wen Kuan, Jhing-Fa Wang, Jia-Ching Wang, Po-Chuan Lin, Gaung-Hui Gu
IEEE Trans. Very Large Scale Integr. Syst.2
2011 Emotion Detection Based on Concept Inference and Spoken Sentence Analysis for Customer Service
Ren-Ying Fang, Bo-Wei Chen, Jhing-Fa Wang, Chung-Hsien Wu 0001
INTERSPEECH3
2011 An Efficient Pre-Processing Scheme to Improve the Sound Source Localization System in Noisy Environment
Sheng-Chieh Lee, K. Bharanitharan, Bo-Wei Chen, Jhing-Fa Wang, Chung-Hsien Wu 0001, Min-Jian Liao
INTERSPEECH4
2011 Interactional Style Detection for Versatile Dialogue Response Using Prosodic and Semantic Features
Wei-Bin Liang, Chung-Hsien Wu 0001, Chih-Hung Wang, Jhing-Fa Wang
INTERSPEECH4
2011 Candidate Generation for ASR Output Error Correction Using a Context-Dependent Syllable Cluster-Based Confusion Matrix
Chao-Hong Liu, Chung-Hsien Wu 0001, David Sarwono, Jhing-Fa Wang
INTERSPEECH4
2011 Hardware/software co-design for fast-trainable speaker identification system based on SMO
abstract
Embedded speaker identification system is a popular research, but most of current systems can not provide fast training ability. Because of the low computational ability in the embedded environment, a large amount of waiting time usually makes the human-machine interface not friendly. This paper presents a hardware and software (HW/SW) co-design solution for fast-trainable speaker identification system. Fast training ability makes this embedded speaker identification system possess high flexibility and enhances the convenience to a wide range of real-world applications. The proposed system consists of a training phase and a multiclass identification phase. The sequential minimal optimization (SMO) training algorithm occupies the heaviest computational load and is realized as a dedicated VLSI module, i.e., the hardware component. The other processes such as speech preprocess, speech feature extraction, and SVM voting strategy are implemented by software. Moreover, a data-packed mechanism is presented to improve the bandwidth utilization. Compared with the embedded C code based on ARM processor, our system reduces 90% of the training time and achieves 89.9% identification rate with the NIST 2010 speaker recognition database. The proposed system was tested and found to be fully functional working on a Socle CDK prototype system with an AMBA based Xilinx FPGA and an ARM926EJ processor.
Jhing-Fa Wang, Jr-Shiang Peng, Jia-Ching Wang, Po-Chuan Lin, Ta-Wen Kuan
SMC1
2011 Robust sound recognition applied to awareness for health/children/elderly care
abstract
This paper presented a robust sound recognition work applied to awareness for health/children/elderly care. Specific sound awareness services can be activated based on recognized sound classes for detecting human activities as health care. To attain this goal, this study developed key technologies as follows: 1) SNR-aware subspace signal enhancement, 2) pitch and power density-based sound/speech discrimination, 3) HMM-based speech recognition, 4) sound recognition with ICA-transformed MFCCs feature and frame-based multiclass SVMs. Each classified sound event is response to human with predefined processes as sound awareness info. Simulations and an experiment are given to illustrate the performance of the proposed robust sound recognition system in a real-world home environment, Aspire Home, NCKU. The overall average resulting accuracy rate was approximately 90.97%.
Jhing-Fa Wang, Po-Yi Shih, Zhong-Hua Fu, Sheng-Chieh Lee
SMC1
2011 Robust several-speaker speech recognition with highly dependable online speaker adaptation and identification
Po-Yi Shih, Po-Chuan Lin, Jhing-Fa Wang, Yuan-Ning Lin
J. Netw. Comput. Appl.3
2011 Efficient Inter Mode Prediction Based on Model Selection and Rate Feedback for H.264/AVC
abstract
H.264/AVC is a standard developed for various low-complexity video applications and high-definition television. To improve coding performance, H.264/AVC may optionally adopt the rate-distortion optimization (RDO) method to find the best encoding mode among various inter and intra modes. However, the exhaustive RDO search among different modes increases the H.264/AVC encoder complexity and limits its application. In this paper, we propose an inter mode prediction algorithm for P slices based on spatial and temporal consistency analysis to reduce the complexity of the RDO computation. We apply the stochastic method to analyze the spatial consistency and use rate information for temporal consistency. The experimental results show a 0.03 peak signal-to-noise ratio loss, a 0.87% bit rate increase, and a 58.39% encoding time reduction on average.
Kuan-I Lee, An-Chao Tsai, Jhing-Fa Wang, Jar-Ferr Yang
IEEE Trans. Circuits Syst. Video Technol.3
2010 Speech presence probability estimation based on integrated time-frequency minimum tracking for speech enhancement in adverse environments
abstract
Speech enhancement under nonstationary environments is a challenging problem. This paper addresses the problem of speech presence probability (SPP) estimation. According to the fact that speech is approximately sparse in time-frequency domain, we integrate time and frequency minimum tracking results to estimate the noise power spectral density and the a posteriori signal-to-noise ratio. A sparseness measure is proposed to adjust the SPP estimates. By applying Bayes rule, we present the final SPP estimates, which control the time varying smoothing of the noise power spectrum. We show that under slowly and highly nonstationary noise conditions, the integrated minimum tracking (IMT) approach can update the noise estimates faster than the competitive methods. When integrated into a speech enhancement system, it achieves improved speech quality and lower residual noise.
Zhong-Hua Fu, Jhing-Fa Wang
ICASSP2
2010 Kernel-Based Lip Shape Clustering with Phoneme Recognition for Real-Time Voice Driven Talking Face
Po-Yi Shih, Jhing-Fa Wang, Zong-You Chen
ISNN (2)2
2010 Dynamic Fixed-Point Arithmetic Design of Embedded SVM-Based Speaker Identification System
Jhing-Fa Wang, Ta-Wen Kuan, Jia-Ching Wang, Ta-Wei Sun
ISNN (2)1
2010 Noisy Environment-Aware Speech Enhancement for Speech Recognition in Human-Robot Interaction Application
abstract
In this study, we introduce a noisy environment-aware speech enhancement system, which can be used in human-robot interaction (HRI) application for command recognition. In order to effectively filter different noises and improve speech recognition rates, the proposed system adopts automatic noise cancellation that is combined with independent component analysis (ICA) and subspace speech enhancement (SSE). Furthermore, it can automatically decide when to use noise reduction according to SNRs of the detected noisy speeches at any time (using proposed noisy environment-aware determination). The experimental results show that our proposed system is suitable for various types of noisy environments, and it is capable of improving the speech quality for recognition. Our proposed system can enhance SNRs by about 20dB, which is higher than those of original noisy speeches.
Sheng-Chieh Lee, Bo-Wei Chen, Jhing-Fa Wang
SMC3
2010 Classified Multifilter Up-Sampling Algorithm in Spatial Scalability for H.264/SVC Encoder
abstract
H.264/scalable video coding (SVC) has spatial scalability which can provide various resolution sequences for a single encoded bit-stream. The various resolutions can be achieved by up-sampling a lower image resolution sequence. The up-sampling method employed in H.264/SVC is directly proportional to the video quality. In order to improve the video quality with an effective up-sampling method, we propose a classified multifilter up-sampling algorithm for H.264/SVC in spatial scalability which classifies an image region as edges and nonedges. An appropriate filter is then applied to up-sample the image. In the proposed scheme, we applied the Wiener filter for edges and conventional filter for nonedges within the defined window size to preserve the edges in the up-sampled image sequence. The experimental results show that the average peak signal-to-noise ratio improvement and bit-rate reduction are 0.5 dB and 9%, respectively, which confirm that the performance of the proposed method is better than that of the existing H.264/SVC standard.
An-Chao Tsai, K. Bharanitharan, Jhing-Fa Wang, Jar-Ferr Yang
IEEE Trans. Circuits Syst. Video Technol.3
2009 SVM-based state transition framework for dynamical human behavior identification
abstract
This investigation proposes an SVM-based state transition framework (named as STSVM) to provide better performance of discriminability for human behavior identification. The STSVM consists of several state support vector machines (SSVM) and a state transition probability (STPM). The intra-structure information and inter-structure information of a human activity are analyzed and correlated by the SSVM and STPM, respectively. The integration of the SSVM and the STPM effectively provides human behavior understanding. With a database consisting of five kinds of human behaviors: raising hand, standing up, squatting down, falling down, and sitting, the proposed algorithm has been demonstrated with a significant recognition rate of 88.6%.
Chen-Yu Chen, Ta-Cheng Wang, Jhing-Fa Wang, Li Pang Shieh
ICASSP3
2009 Noise robust features for speech/music discrimination in real-time telecommunication
abstract
While many efforts have been made in the audio signal classification field, the noise interruption problem is seldom concerned so far, especially in many telecommunication applications, where a real-time and noise robust approach is needed. This paper addresses this problem by proposing two novel robust features: average pitch density (APD) and relative tonal power density (RTPD). APD refers to the differences in tone characteristics of music and speech signals, and RTPD especially focuses on the distinct properties of the percussion instruments. The comparison experiments are implemented on two databases. The first one is reorganized from the corpus collected,. The second one consists of data collected from various recording situations. The novel features are compared with several state-of-the-art features and are found to achieve significant robustness.
Zhong-Hua Fu, Jhing-Fa Wang, Lei Xie 0001
ICME2
2009 Video Knowledge Augmentation based on Summarized Contents and Online Media
abstract
Exploration techniques of video knowledge have been proposed for years to help people discover the details about videos. However, existing systems still yield limited information for users. In this paper, we present a video knowledge browsing system, which can establish the framework of a video based on its summarized contents and expand them by using online correlated media. Thus, users can not only browse key points of a video efficiently but also focus on what they are interested in. In order to construct the fundamental system, we make use of our previous proposed approaches to transforming a video into a graph. After the relational graph is built up, the social network analysis is then performed to explore online relevant resources. We also apply the Markov clustering algorithm to enhance the results of the network analysis. The experiments demonstrate that our system can achieve better performance than the traditional systems.
Bo-Wei Chen, Jhing-Fa Wang, Jia-Ching Wang
ISCAS2
2009 VLSI Design of Sequential Minimal Optimization Algorithm for SVM Learning
abstract
The sequential minimal optimization (SMO) algorithm has been widely used for training the support vector machine (SVM). In this paper, we present the first chip design for sequential minimal optimization. This chip is implemented as an intellectual property (IP) core, suitable to be utilized in an SVM-based recognition system on a chip. The proposed SMO chip has been tested to be fully functional, using a prototype system based on the Altera DE2 board with Cyclone II 2C70 FPGA (field-programmable gate array).
Ta-Wen Kuan, Jhing-Fa Wang, Jia-Ching Wang, Gaung-Hui Gu
ISCAS2
2009 A Design of Far-Field Speaker Localization System Using Independent Component Analysis with Subspace Speech Enhancement
abstract
In this paper, we propose a design of far-field single speaker localization system including noisy speech separation and enhancement with two separated channel microphones. First, we perform noisy speech separation simulations using a fast-ICA (independent component analysis) with subspace-based speech enhancement. The fast-ICA can be used to separate the two original source signals-noise and speech from their mixtures. Subspace-based speech enhancement is then utilized to reduce the residual background noise of the separated speech. Second, we calculate time-difference-of-arrival (TDOA) estimation for five directions of both right and left sides in the horizontal plane to realize speaker localization. The TDOA method estimates the time delay difference between the incoming speech signals to microphones using the average magnitude difference function (AMDF) algorithm. Finally we evaluate the localization system under various noisy environments. The experimental results reveal that the applicability of a system with noisy speech separation and enhancement, particularly for the front directions.
Shi-Huang Chen, Jhing-Fa Wang, Miao-Hai Chen, Zheng-Wei Sun, Min-Jian Liao, Shun-Chieh Lin, Shyang-Jye Chang
ISM2
2009 A Novel Video Summarization Based on Mining the Story-Structure and Semantic Relations Among Concept Entities
abstract
Video summarization techniques have been proposed for years to offer people comprehensive understanding of the whole story in the video. Roughly speaking, existing approaches can be classified into the two types: one is static storyboard, and the other is dynamic skimming. However, despite that these traditional methods give brief summaries for users, they still do not provide with a concept-organized and systematic view. In this paper, we present a structural video content browsing system and a novel summarization method by utilizing the four kinds of entities: who, what, where, and when to establish the framework of the video contents. With the assistance of the above-mentioned indexed information, the structure of the story can be built up according to the characters, the things, the places, and the time. Therefore, users can not only browse the video efficiently but also focus on what they are interested in via the browsing interface. In order to construct the fundamental system, we employ maximum entropy criterion to integrate visual and text features extracted from video frames and speech transcripts, generating high-level concept entities. A novel concept expansion method is introduced to explore the associations among these entities. After constructing the relational graph, we exploit graph entropy model to detect meaningful shots and relations, which serve as the indices for users. The results demonstrate that our system can achieve better performance and information coverage.
Bo-Wei Chen, Jia-Ching Wang, Jhing-Fa Wang
IEEE Trans. Multim.3
2008 SVM-based one-against-many algorithm for liveness face authentication
abstract
Illegal users are not permitted to operate within a secure environment. To establish the legality of authentication, the authentication system must perceive and refuse a fake biometric over liveness face authentication. In order to achieve reliable liveness face authentication, the intended purpose of the proposed framework should have two major parts: liveness detection and face authentication. The proposed liveness detection describes illuminative variations on the face, which is especially applicable in artificial shadow estimation; face authentication should also employ a one-against-many classification algorithm based on Support Vector Machine (SVM) to obtain individual subsets, then estimates authenticated performances. Based on experiments on liveness XM2VTS database and photographs from the Google Picasa database, we achieved the liveness accuracy rate of 96.5%, the false rejection rate of 1.17% and the false acceptance rate of 1.69%.
Cheng-Ho Huang, Jhing-Fa Wang
SMC2
2008 Ubiquitous and Robust Text-Independent Speaker Recognition for Home Automation Digital Life
Jhing-Fa Wang, Ta-Wen Kuan, Jia-Chang Wang, Gaung-Hui Gu
UIC1
2008 A Long-Distance Time Domain Sound Localization
Jhing-Fa Wang, Jia-Chang Wang, Bo-Wei Chen, Zheng-Wei Sun
UIC1
2008 Robust Environmental Sound Recognition for Home Automation
abstract
This work presents a robust environmental sound recognition system for home automation. Specific home automation services can be activated based on identified sound classes. Additionally, when the sound category is human speech, such speech can be recognized for detecting human intentions as in conventional research on home automation. To attain this ambitious goal, this study uses two key techniques: signal-to-noise ratio-aware subspace-based signal enhancement and sound recognition with independent component analysis mel-frequency cepstral coefficients and a frame-based multiclass support vector machines, respectively. Simulations and an experiment in a real-world environment are given to illustrate the performance of the proposed robust sound recognition system.
Jia-Ching Wang, Hsiao Ping Lee, Jhing-Fa Wang, Cai-Bei Lin
IEEE Trans Autom. Sci. Eng.3
2008 Intensity Gradient Technique for Efficient Intra-Prediction in H.264/AVC
abstract
This study presents an intensity gradient approach for intra-prediction in H.264 encoding system, which enhances the performance and efficiency of previous fast algorithms. We propose a preprocessing stage in which eight orientation features are extracted from a macro block by the intensity gradient filters. The orientation features are utilized to select a subset of prediction modes to be involved in the rate-distortion calculation so that the encoding time can be reduced. The simulation results indicate that the intensity gradient based algorithm for intra-prediction contributes better tradeoff between rate-distorion performance and encoding complexity than the previous algorithms. Compared to H.264 reference software, the proposed algorithm introduces slight PSNR degradation and bit rate increase but saves around 76% of the total encoding time with all intra-frame coding.
An-Chao Tsai, Anand Paul 0001, Jia-Ching Wang, Jhing-Fa Wang
IEEE Trans. Circuits Syst. Video Technol.4
2008 Effective Subblock-Based and Pixel-Based Fast Direction Detections for H.264 Intra Prediction
abstract
In H.264/AVC intra-frame coding, the rate-distortion optimization (RDO) is employed to select the optimal coding mode to achieve the minimum rate-distortion cost. Due to a large number of combinations of coding modes, the computational burden becomes extremely high in intra-prediction. In this paper, we propose two fast, efficient but reliable direction detection algorithms by computing subblock and pixel direction differences. Both proposed methods effectively estimate the edge direction inside the block to narrow down the predictive modes to reduce the RDO computation. Experimental results show that the proposed methods can reduce the encoding time by about 60% with negligible loss of coding performance. For hardware realization, a fast mode decision VLSI circuit for intra-prediction with the silicon core size of 0.12 times 0.12 mm2at 0.18-mum CMOS technology is implemented. The fast mode decision VLSI in three-stage pipelined architecture operated at 173 MHz can encode 30 fps real-time videos up to Level 5.
An-Chao Tsai, Jhing-Fa Wang, Jar-Ferr Yang, Wei-Guang Lin
IEEE Trans. Circuits Syst. Video Technol.2
2007 A Simple and Robust Direction Detection Algorithm for Fast H.264 Intra Prediction
abstract
In H.264/AVC intra frame coding, the rate-distortion optimization is employed to select the optimal coding mode to achieve the minimum rate-distortion cost. Due to a large number of combinations of coding modes, the computational burden becomes extremely high in intra prediction, hi this paper, we propose a fast, efficient but reliable vector based detection algorithm for H.264 intra frame coding system. The proposed algorithm effectively estimates the edge direction inside the block to narrow down the predictive modes to reduce the RDO calculations. Experimental results show that the proposed method can reduce the encoding time by about 60% with negligible coding loss.
An-Chao Tsai, Jhing-Fa Wang, Wei-Guang Lin, Jar-Ferr Yang
ICME2
2007 Efficient Intra Prediction in H.264 Based on Intensity Gradient Approach
abstract
This study presents an intensity gradient approach to intra prediction in H.264 encoding system, which enhances the performance and efficiency by means of edge orientation of the gradient filter. We propose a pre-processing stage in which eight-orientation feature are extracted from a macro block that selects four modes to be applied to the block among a set of predefined modes. It is shown that by choosing number of modes used in rate-distortion calculation lead to significant enhancement in the performance of intra prediction. The validity of this algorithm is confirmed experimentally. The simulation results indicate that the intensity gradient-based algorithm for intra prediction contributes to the bit-rate reduction compared to that of previous algorithms and saves around 53% of the total encoding time in H.264.
An-Chao Tsai, Anand Paul 0001, Jia-Ching Wang, Jhing-Fa Wang
ISCAS4
2007 Event-Based Segmentation of Sports Video Using Motion Entropy
abstract
An event-based segmentation method for sports videos is presented. A motion entropy criterion is employed to characterize the level of intensity of relevant object motion in individual frames of a video sequence. The resulting motion entropy curve then is approximated with a piece-wise linear model using a homoscedastic error model based time series change point detection algorithm. It is observed that interesting sports events are correlated with specific patterns of the piece-wise linear model. A set of empirically derived classification rules then is derived based on these observations. Application of these rules to the motion entropy curve leads to this motion entropy curve, one is able to segment the corresponding video sequence into individual sections, each consisting of a semantically relevant event. The proposed method is tested on six hours of sports videos including basketball, soccer and tennis. Excellent experimental results are observed.
Chen-Yu Chen, Jia-Ching Wang, Jhing-Fa Wang, Yu Hen Hu
ISM3
2007 Variable-Length Unit Selection in TTS Using Structural Syntactic Cost
abstract
This paper presents a variable-length unit selection scheme based on syntactic cost to select text-to-speech (TTS) synthesis units. The syntactic structure of a sentence is derived from a probabilistic context-free grammar (PCFG), and represented as a syntactic vector. The syntactic difference between target and candidate units (words or phrases) is estimated by the cosine measure with the inside probability of PCFG acting as a weight. Latent semantic analysis (LSA) is applied to reduce the dimensionality of the syntactic vectors. The dynamic programming algorithm is adopted to obtain a concatenated unit sequence with minimum cost. A syntactic property-rich speech database is designed and collected as the unit inventory. Several experiments with statistical testing are conducted to assess the quality of the synthetic speech as perceived by human subjects. The proposed method outperforms the synthesizer without considering syntactic property. The structural syntax estimates the substitution cost better than the acoustic features alone
Chung-Hsien Wu 0001, Chi-Chun Hsia, Jiun-Fu Chen, Jhing-Fa Wang
IEEE Trans. Speech Audio Process.4
2007 Unsupervised Speaker Change Detection Using SVM Training Misclassification Rate
abstract
This work presents an unsupervised speaker change detection algorithm based on support vector machines (SVM) to detect speaker change (SC) in a speech stream. The proposed algorithm is called the SVM training misclassification rate (STMR). The STMR can identify SCs with less speech data collection, making it capable of detecting speaker segments with short duration. According to experiments on the NIST Rich Transcription 2005 Spring Evaluation (RT-05S) corpus, the STMR has a missed detection rate of only 19.67 percent.
Po-Chuan Lin, Jia-Ching Wang, Jhing-Fa Wang, Hao-Ching Sung
IEEE Trans. Computers3
2007 A Fast Mode Decision Algorithm and Its VLSI Design for H.264/AVC Intra-Prediction
abstract
In this paper, we present a fast mode decision algorithm and design its VLSI architecture for H.264 intra-prediction. A regular spatial domain filtering technique is proposed to compute the dominant edge strength (DES) to reduce the possible predictive modes. Experimental results revealed that the proposed fast intra-algorithm reduces 40% computation with slight peak signal-to-noise ratio (PSNR) degradation. The designed DES VLSI engine comprises a zigzag converter, a DES finite-state machine (FSM), and a DES core. The former two units handle memory allocation and control flow while the last performs pseudoblock computation, edge filtering, and dominant edge strength extraction. With semicustom design fabricated by 0.18 mum CMOS single-poly-six-metal technology, the realized die size is roughly 0.15 times 0.15 mm2and can be operated at 66 MHz.
Jia-Ching Wang, Jhing-Fa Wang, Jar-Ferr Yang, Jang-Ting Chen
IEEE Trans. Circuits Syst. Video Technol.2
2007 A Hybrid Algorithm With Artifact Detection Mechanism for Region Filling After Object Removal From a Digital Photograph
abstract
This work aims to develop a novel function of the digital camera, i.e., region-filling after object removal from a digital photograph. This problem is defined as how to "guess" the Lacuna region after removal of an object by replicating a part from the remainder of the whole image with visually plausible quality. We propose a hybrid region-filling algorithm composed of a texture synthesis technique and an efficient interpolation method with a refinement approach: 1) the "subpatch texture synthesis technique" can synthesize the Lacuna region with significant accuracy; 2) the "weighted interpolation method" is applied to reduce computaion time; and 3) the "artifact detection mechanism" integrates the Kirsch edge detector and color ratio gradients to detect the artifact blocks in the filled region after the first pass of filling the Lacuna region when the result may not be satisfactory. This can lead to resynthesizing the artifact blocks without user intervention and can be compliant to the unsuccessful results of other algorithms. In the procedure of region-filling, color texture distribution analysis is used to choose whether the subpatch texture synthesis technique or the weighted interpolation method should be applied. In the subpatch texture synthesis technique, the actual pixel values of the Lacuna region are synthesized by adaptively sampling from the source region. The experimental results show that our proposed algorithm can achieve a better performance than previous methods. Particularly, the regular computation of our proposed algorithm is more suitable for implementation of the hardware in a digital camera.
Han-Jen Hsu, Jhing-Fa Wang, Shang-Chia Liao
IEEE Trans. Image Process.2
2007 Temporal Partitioning Data Flow Graphs for Dynamically Reconfigurable Computing
abstract
FPGA-based configurable computing machines are evolving rapidly in large signal processing applications due to flexibility and high performance. In this paper, given a reconfigurable processing unit (RPU) with a logic capacity of ARPUand a computational task represented by a data flow graph G = (V, E, W), we propose a network flow-based multiway task partitioning algorithm to minimize communication costs for temporal partitioning. The proposed algorithm obtains an optimal solution with minimum interconnection under area constraints. The optimal solution is a cut set. In our approach, two techniques are applied. In the initial partition, any feasible min-cut is produced by the proposed network flow-based algorithm, so a set of feasible min-cuts is obtained. From the feasible solutions, the scheduling technique selects an optimal global solution.
Yung-Chuan Jiang, Jhing-Fa Wang
IEEE Trans. Very Large Scale Integr. Syst.2
2006 An ARM-Based Embedded System Design for Speech-to-Speech Translation
Shun-Chieh Lin, Jhing-Fa Wang, Jia-Ching Wang, Hsueh-Wei Yang
EUC2
2006 An Intelligent Guiding Bulletin Board System with Real-Time Vision and Multi-Keyword Spotting Multimedia Human-Computer Interaction
abstract
This paper presents an intelligent guiding bulletin board system (iGBBS), which is based on vision-interactive and multiple key word-spotting technology. The system is aimed to provide different kinds of multimedia human-computer interaction (MMHCI) for users under different requirements. At first, a real-time front-view face detection using Harr-like features is used to decide when iGBBS should wake up and become interactive with the user. After system initialization, some feature points within the detected face area are going to be found. Then the orientation of user's head will be estimated via pyramidal Lucas-Kanade optical flow tracking. In addition, spotting the keyword from user's utterance with some related augmented reality responses would be provided as well. The performance of vision-interaction in iGBBS could be reached to 20 fps under Pentium IV 1G Hz PC. The error rate of multiple key word-spotting interaction in iGBBS is about 36.2% and people can get the right response in 2.76 times search averagely. With the comparison to the traditional guiding system, bulletin board, or other non-vision-based input devices system, such like gloves or markers, our system offers a simple, useful and economical solution for the realtime interaction between the user and computer
Cheng-Yu Chang, Chung-Hsien Yang, You-Sheng Yeh, Pau-Choo Chung, Jhing-Fa Wang, Jar-Ferr Yang
ICME5
2006 Robust Speaker Recognition using SNR-Aware Subspace-Based Enhancement and Probabilistic SVMs
abstract
In this paper, we present a robust text-independent speaker recognition system. The proposed system mainly includes an SNR-aware subspace-based enhancement technique and probabilistic support vector machines (SVMs). First, we construct a perceptual filterbank from psycho-acoustic model and incorporate it with the subspace-based enhancement approach. The prior SNR of each subband within the perceptual filterbank is taken to decide the estimator's gain to effectively suppress environmental background noises. Next, this study uses probabilistic SVMs to identify the speaker from the enhanced speech. The superiority of the proposed system has been demonstrated by twenty speaker recognition from AURORA-2 database with in-car noises
Jia-Ching Wang, Jhing-Fa Wang, Wai-He Kuok, Hsiao Ping Lee, Chung-Hsien Yang
ICME2
2006 Environmental Sound Classification using Hybrid SVM/KNN Classifier and MPEG-7 Audio Low-Level Descriptor
abstract
In this paper, we present a new environmental sound classification architecture. The proposed sound classifier is performed in frame level and fuses the support vector machine (SVM) and the k nearest neighbor rule (KNN). In feature selection, three MPEG-7 audio low-level descriptors, spectrum centroid, spectrum spread, and spectrum flatness are used as the sound features to exploit their ability in sound classification. Experiments carried out on 12-class sound database can achieve an 85.1% accuracy rate. The The The performance comparison between the HMM sound classifier using audio spectrum projection features demonstrates the superiority of the proposed scheme.
Jia-Ching Wang, Jhing-Fa Wang, Wai-He Kuok, Cheng-Shu Hsu
IJCNN2
2006 A novel fast algorithm for intra mode decision in H.264/AVC encoders
abstract
This paper presents a fast mode decision algorithm for H.264 intra prediction based on dominant edge strength (DES). In H.264 intra prediction, the computation-extensive rate distortion optimization (RDO) technique with full intra mode search is used to select the best mode for each macroblock. To reduce the computational load in mode decision, the DES which is corresponding to a decision mode is detected first. In accordance with the detected dominant edge, a subset of the prediction modes is then chosen for RDO calculation. The proposed algorithm only searches 4 modes instead of 9 for the 4/spl times/4 luma blocks. As for the 16/spl times/16 or 8/spl times/8 chroma blocks, instead of 4 modes, only 2 modes are required to be searched. Experimental results revealed that the computation time of the proposed fast intra prediction algorithm is averagely reduced to 40% of the full search method with slight PSNR degradation.
Jhing-Fa Wang, Jia-Ching Wang, Jang-Ting Chen, An-Chao Tsai, Anand Paul 0001
ISCAS1
2006 An Embedded System Design for Ubiquitous Speech Interactive Applications Based on a Cost Effective SPCE061A Micro Controller
Po-Chuan Lin, Jhing-Fa Wang, Shun-Chieh Lin, Ming-Hua Mo
UIC2
2006 Voice conversion using duration-embedded bi-HMMs for expressive speech synthesis
abstract
This paper presents an expressive voice conversion model (DeBi-HMM) as the post processing of a text-to-speech (TTS) system for expressive speech synthesis. DeBi-HMM is named for its duration-embedded characteristic of the two HMMs for modeling the source and target speech signals, respectively. Joint estimation of source and target HMMs is exploited for spectrum conversion from neutral to expressive speech. Gamma distribution is embedded as the duration model for each state in source and target HMMs. The expressive style-dependent decision trees achieve prosodic conversion. The STRAIGHT algorithm is adopted for the analysis and synthesis process. A set of small-sized speech databases for each expressive style is designed and collected to train the DeBi-HMM voice conversion models. Several experiments with statistical hypothesis testing are conducted to evaluate the quality of synthetic speech as perceived by human subjects. Compared with previous voice conversion methods, the proposed method exhibits encouraging potential in expressive speech synthesis.
Chung-Hsien Wu 0001, Chi-Chun Hsia, Te-Hsien Liu, Jhing-Fa Wang
IEEE Trans. Speech Audio Process.4
2006 Efficient news video querying and browsing based on distributed news video servers
abstract
This study presents an efficient news video querying and browsing system based on distributed news video servers. The proposed architecture includes distributed news video preprocessing (NVP) server and visualized querying/browsing (VQB) server. The distributed NVP server receives the news video story from news video web (NVW) server and generates the story abstract, i.e., key frames and key sentences. These story abstract are then combined with the news script of a news story, then sent to the VQB server for news video querying and browsing. The VQB server performs semantic clustering using news script to categorize all the news video stories, it also provides a visualized interface displaying the story abstract so that users can fast grasp the main idea of a news story. The superiority of the proposed system has been demonstrated using news video obtained from NVW servers of EraNews, ETtoday and Formosa TV stations in Taiwan.
Chen-Yu Chen, Jia-Ching Wang, Jhing-Fa Wang
IEEE Trans. Multim.3
2005 Advanced Ubiquitous Media for Creative Cyberspace
abstract
Summary form only given. In this paper we discuss how to construct a modern creative cyberspace and to design advanced living products through the developments of ubiquitous digital contents services, wireless sensor networks with emotion detection and media perception capabilities. This includes inventing, developing, and applying technologies that will be vital to obtaining fundamental results that influence the direction of the above research topics. In the past five years for international achievements, many research centers, such as MIT Media Labs, IBM Research Labs, and Microsoft Research have respectively created the related researches on well-known Oxygen, DreamSpace, and EasyLiving projects. The Oxygen, which enables pervasive, human-centered computing through a combination of specific user and system technologies, highly utilizes the user technologies to directly provide human needs. Speech and vision technologies enable a people to simply communicate with the Oxygen as if we're interacting with another person such we can save unnecessary time and effort for media interfacing. Automation, individualized knowledge access, and collaboration technologies help us perform a wide variety of tasks that we want to do in the ways we like to do them. DreamSpace allows users to collaborate in a shared space. The system "hears" users' voice commands and "sees" their gestures and body positions. Interactions are natural, more like human-to-human interactions. The "computer" understands the user, and - just as important - other users understand. Users are free to focus on virtual objects and information and understanding and thinking, with minimal constraints or distractions by the "computer", which is present only as wall-sized 3D images and sounds (but no keyboard, mouse, wires, wands, etc.). EasyLiving is developing a prototype of architecture and technologies for building intelligent environments.
Jhing-Fa Wang
AINA1
2005 A novel framework for object removal from digital photograph
abstract
This work aims for a novel function for smart camera-redundant object removal from digital photograph. The proposed novel framework can fill the left lacuna region in the digital image. In previous related researches, texture synthesis and image inpainting construct the fundamentals of filling the lost image region. Texture synthesis can be used to fill the large hole of input texture, while image inpainting can be used to repair the small image gaps. In this paper, we propose an object removal framework by the sub-patch texture synthesis algorithm and weighted interpolation method with automatic repainting mechanism. In the filling process, the color distribution analysis is used to choose different methods. The exhaustive computation time is reduced by the weighted interpolation method. In order to repaint the faulty texture region intelligently, we use the color ratio gradients to detect the synthesized artifact region. The automatic artifact detection can lead repainting the faulty region without user intervention. The proposed algorithm can achieve better performance with seamless output images. The regular computation is also suitable for hardware architecture different from previous existing algorithms.
Jhing-Fa Wang, Han-Jen Hsu, Shang-Chia Liao
ICIP (2)1
2005 Automatic Scene Change Detection for Composed Speech and Music Sound Under Low SNR Noisy Environment
abstract
As the amount of available audiovisual data in digital formats is increasing, automatic scene change detection becomes paramount. Many studies have been proposed to treat it recently. Nevertheless, none of these techniques consider the audio signals under a low SNR noisy environment. In this paper, a hierarchical scene change-detection scheme is adopted to detect the scene change automatically under a low SNR noisy environment. The proposed algorithm contains three steps: A statistical model-based audio activity detection scheme that employs the likelihood ratio test is used to segment the audio signal into pure noise segments and noisy audio segments in the first step. In the second step, a noisy robustness feature is adopted to construct a K-Nearest Neighbor (KNN) based classifier and segments the noisy audio segments into speech and music segments. The feature we proposed is called the likelihood ratio crossing rate (LRCR) derived from the likelihood ratio waveform that is obtained in the first step. In last step, a novel speaker change detection based on modified Bayesian information criterion (BIC) is performed to detect the speaker change points for each speech segment. A series of tests were conducted showing the advantage of the proposed scheme. Furthermore, it also shows the robustness under a low SNR noisy environment.
Chaug-Ching Huang, Jhing-Fa Wang, Dian-Jia Wu
IEEE Trans. Speech Audio Process.2
2004 Subspace tracking for speech enhancement in car noise environments
abstract
A signal subspace speech enhancement based on a subspace tracking algorithm is presented. The proposed method incorporates a perceptual filterbank which is derived from a psycho-acoustic model for subband processing. The experiments were performed using the TAICAR in-car noisy speech database. Subjective and objective tests show that our method outperforms other existing signal subspace methods.
Jhing-Fa Wang, Chung-Hsien Yang, Kai-Hsing Chang
ICASSP (2)1
2004 Constrained texture synthesis by scalable sub-patch algorithm
abstract
We propose a novel scalable sub-patch algorithm for texture synthesis. Texture synthesis is an important research topic in computer graphics and image processing. We present an efficient algorithm which can paste a patch each time without boundary processing technique. In particular, the proposed sub-patch algorithm can be used in constrained texture synthesis to fill the large lost region. On the other hand, the searching neighborhood pixels can be changed due to different patch size in image reconstruction. The comparison of previous algorithms is also provided in this paper. We present a first algorithm used in constrained texture synthesis by regular synthesizing order.
Jhing-Fa Wang, Han-Jen Hsu, Hong-Ming Wang
ICME1
2004 Intelligent Sub-patch Texture Synthesis Algorithm for Smart Camera
Jhing-Fa Wang, Han-Jen Hsu, Hong-Ming Wang
KES1
2003 A novel stroke extraction method for Chinese characters using Gabor filters
Yih-Ming Su, Jhing-Fa Wang
Pattern Recognit.2
2003 A learning process to the identification of feature points on Chinese characters
abstract
The paper describes a novel stroke extraction approach to identify the feature points of a character, using line-filtering and learning-based techniques. The line-filtering technique based on convolution operations with a set of one-dimensional (1D) Gabor templates efficiently extracts the stroke segments from noisy and degraded characters. Furthermore, the relationship between endpoints of stroke segments is modeled as junction structure during a learning process. Finally, each endpoint is identified as a feature point to determine the junction structure by the learning-based technique, rather than rule-based techniques with manual rule creation. Experimental results indicate that the learning-based technique can generalize learning knowledge to identify 1200 feature points with an average identification rate of 93.58% for test set, using k-fold cross-validation testing.
Yih-Ming Su, Jhing-Fa Wang
IEEE Trans. Syst. Man Cybern. Part A2
2002 Noise suppression based on approximate KLT with wavelet packet expansion
abstract
In this paper, we perform the noise suppression based on approximate Karhunen-Loeve transform (KL T). The discrete cosine transform(DCT) has been a good candidate for approximate KLT when the signal is modeled as an autoregressive process. However, for nonstationary signals, wavelet transform is more capable than DCT while approximating KLT. To calculate approximate KLT, we first represent the signal by using wavelet packet based on a basis search algorithm, then eigenvectors are evaluated from the basis. A linear estimator based on these eigenvectors can be constructed and used to perform noise reduction. We evaluate the performance of this method by using the Aurora-2 database. The SNR improvement is calculated. Some waveforms and spectrograms of enhanced speech are also shown. Finally. the enhanced speech is tested for speech recognition. These experimental results show that this method achieves satisfactory enhancement of speech.
Chung-Hsien Yang, Jhing-Fa Wang
ICASSP2
2002 A study of multi-speaker dialogue system for mobile information retrieval
Hsien-Chang Wang, Chieh-Yi Huang, Chung-Hsien Yang, Jhing-Fa Wang
INTERSPEECH4
2002 Chip design of MFCC extraction for speech recognition
Jia-Ching Wang, Jhing-Fa Wang, Yu-Sheng Weng
Integr.2
2002 Chip design of portable speech memopad suitable for persons with visual disabilities
abstract
This paper presents the design of a speech recognition and compression chip for portable memopad devices, especially suitable for use by the visually impaired. The proposed chip design is based on several cores of which they can be regarded as intellectual property (IP) cores to be used for a variety of speech-related application systems. A cepstrum extraction core and a dynamic warping core are designed for mapping the speech recognition algorithms. In the cepstrum extraction core, a novel architecture computes the autocorrelation between the overlapping frames using two pairs of shift registers and an intelligent accumulation procedure. The architecture of the dynamic time warping core uses only a single processing element, and is based on our extensive study of the relationship among the nodes in the dynamic time warping lattice. Bit rate is the key factor affecting the memory size for speech compression; therefore, a very low bit-rate speech coder is used. The speech coder exploits a line-spectrum-based interpolation method, which yields fine quality synthesized speech despite the low 1.6 kbps bit rate. The 1.6 kbps vocoder core is cost-effective, and it integrates both encoder and decoder algorithms. The proposed design has been tested via hardware simulations on Xilinx Virtex series FPGAs and a semi-custom chip fabricated by 0.35 /spl mu/m CMOS single-poly-four-metal technology on a die size approximately 4.46/spl times/4.46 mm/sup 2/.
Jhing-Fa Wang, Jia-Ching Wang, Han-Chiang Chen, Tai-Lung Chen, Chin-Chan Chang, Ming-Chi Shih
IEEE Trans. Speech Audio Process.1
2001 Extraction of pitch information in noisy speech using wavelet transform with aliasing compensation
abstract
Although many wavelet-based pitch detection methods have been proposed in the literature, there still remains a need to investigate new wavelet-based methods for more accurate and more robust pitch determination. In this paper, an improved wavelet-based method is developed for extraction of pitch information in noisy speech. At each decomposition in the wavelet transform, an aliasing compensation algorithm is applied to approximate and detail signals, in which the distortion of aliasing due to downsampling and upsampling operations of the wavelet transform is eliminated. In addition, this paper utilizes the concept of spatial correlation function used in signal denoising to improve the performance of pitch detection in a noisy environment. It is shown in various experimental results that this new type of method has a considerable performance improvement compared with other conventional methods and wavelet-based methods.
Shi-Huang Chen, Jhing-Fa Wang
ICASSP2
2001 A voicing-driven packet loss recovery algorithm for analysis-by-synthesis predictive speech coders over Internet
abstract
In this paper, a novel voice-driven adaptive packet loss recovery algorithm is proposed to lessen the possible voice degradation and error propagation for analysis-by-synthesis speech coders in Internet applications. After voicing classification, we adaptively adopt random noise generation, multiresolution excitation generation, or pulse tracking procedure to recover the lost packets, By applying the algorithm to the G.723.1 coder, simulation results show that the proposed algorithm is superior to the recovery algorithm embedded in the G.723.1 standard through the subjective evaluation.
Jhing-Fa Wang, Jia-Ching Wang, Jar-Ferr Yang, Jian-Jia Wang
IEEE Trans. Multim.1
2001 An on-chip march pattern generator for testing embedded memory cores
abstract
In this correspondence, we propose an effective approach to integrate 40 existing march algorithms into an embedded low hardware overhead test pattern generator to test the various kinds of word-oriented memory cores. Each march algorithm is characterized by several sets of up/down address orders, read/write signals, read/write data, and lengths of read/write operations. These characteristics are stored on chip so that any desired march algorithm can be generated with very little external control. An efficient procedure to reduce the memory storage for these characteristics is presented. We use only two programmable cyclic shift registers to generate the various read/write signals and data within the steps of the algorithms. Therefore, the proposed pattern generator is capable of generating any march algorithm with small area overhead.
Wei-Lun Wang, Kuen-Jong Lee, Jhing-Fa Wang
IEEE Trans. Very Large Scale Integr. Syst.3
2000 Chip design of mel frequency cepstral coefficients for speech recognition
abstract
The mel frequency cepstral coefficients (MFCC) is one of the mast important features, which is required among various kinds of speech applications. The chip for speech features extraction based on the MFCC algorithm is first proposed. The chip is designed with area efficient consideration and can achieve the following: (1) the reduction of table size and multiplication complexity by means of the symmetric property of the cosine function, (2) the decrease of the multiplication load and required constant memory in the calculation of the weighted energy spectrum by applying the mapping relationship between the mel scale and the frequency scale, (3) the minimization of the look-up table size for logarithm operations by modifying the partitioned table look-up method. The chip is fabricated with 0.6 /spl mu/m double-metal CMOS technology. It contains approximately 10,000 gates occupying 3.2/spl times/3.3 mm/sup 2/ area and the maximum clock rate is 50 MHz.
Jia-Ching Wang, Jhing-Fa Wang, Yu-Sheng Weng
ICASSP2
2000 Segmentation of Handwritten Connected Numeral String Using Background and Foreground Analysis
abstract
An approach to segmentation of handwritten connected numeral strings is proposed. Most algorithms for segmenting connected digits mainly focus on the analysis of foreground pixels. Some others concentrated on the analysis of background pixels only. We use background and foreground analysis to segment handwritten connected numeral strings. Thinning of both foreground and background pixels are first processed on the image of the numeral stroke and the feature points on foreground and background skeletons are extracted separately. Several possible segmentation paths are constructed and useless strokes are removed. Then, the parameters of geometric properties of each possible segmentation path are determined and these parameters are analyzed by mixture Gaussian probability functions to find the best segmentation path. For preliminary experimentation, we collected 150 images of handwritten connected numeral strings for test and 134 are correctly segmented (correct rate is 96.4%). 5 are in error (error rate is 3.6%) and 11 are rejected (rejected rate is 7.3%).
Yi-Kai Chen, Jhing-Fa Wang
ICPR2
2000 Domain-unconstrained language understanding based on CKIP-auto tag, how-net, and ART
Jhing-Fa Wang, Hsien-Chang Wang, Kin-Nan Lee, Chieh-Yi Huang
INTERSPEECH1
2000 Single chip implementation of the 1.6 kbps speech vocoder
abstract
In this paper, we propose a low bit rate speech vocoder and its corresponding VLSI implementation. The vocoder exploits the interpolation property so that the fine quality in synthesized speech is obtained even though the bit rate is as low as 1.6 kbps. Two novel methods including pitch detection and LSP decoding which are suitable for VLSI implementation are also proposed. The heuristic pitch detection algorithm avoids the heavy computational load introduced by the traditional normalized autocorrelation method. The memory storing triangular function value is no longer needed after adopting the new LSP decoding process. The chip is designed with area effective feature and is suitable for stand alone application.
Jia-Ching Wang, Jhing-Fa Wang, Han-Chiang Chen
ISCAS2
2000 Domain Unconstrained Language Understanding Based on How-net
Jhing-Fa Wang, Hsien-Chang Wang, Chin-Nan Lee
PACLIC1
2000 Segmentation of Single- or Multiple-Touching Handwritten Numeral String Using Background and Foreground Analysis
abstract
An approach of segmenting a single- or multiple-touching handwritten numeral string (two-digits) is proposed. Most algorithms for segmenting connected digits mainly focus on the analysis of foreground pixels. Some concentrated on the analysis of background pixels only and others are based on a recognizer. We combine background and foreground analysis to segment single- or multiple-touching handwritten numeral strings. Thinning of both foreground and background regions are first processed on the image of connected numeral strings and the feature points on foreground and background skeletons are extracted. Several possible segmentation paths are then constructed and useless strokes are removed. Finally, the parameters of geometric properties of each possible segmentation paths are determined and these parameters are analyzed by the mixture Gaussian probability function to decide the best segmentation path or reject it. Experimental results on NIST special database 19 (an update of NIST special database 3) and some other images collected by ourselves show that our algorithm can get a correct rate of 96 percent with rejection rate of 7.8 percent, which compares favorably with those reported in the literature.
Yi-Kai Chen, Jhing-Fa Wang
IEEE Trans. Pattern Anal. Mach. Intell.2
2000 Skew detection and reconstruction based on maximization of variance of transition-counts
Yi-Kai Chen, Jhing-Fa Wang
Pattern Recognit.2
1999 A C/V segmentation algorithm for Mandarin speech signal based on wavelet transforms
abstract
This paper proposes a new consonant/vowel (C/V) segmentation algorithm for Mandarin speech signal. Since the Mandarin phoneme structure is a combination of a consonant (may be null) followed by a vowel, the C/V segmentation is an important part in the Mandarin speech recognition system. Based on the wavelet transform, the proposed method can directly search for the C/V segmentation point by using a product function and energy profile. The product function is generated from the appropriate wavelet and scaling coefficients of the input speech signal, and it can be applied to indicate the C/V segmentation point. With this product function and the additional verification of the energy profile, the C/V segmentation can be accurately pointed out with a low computation complexity. Experiments are provided that demonstrate the superior performance of the proposed algorithm. An overall accuracy rate of 97.2% is achieved. This algorithm is suitable for Mandarin speech recognition task.
Jhing-Fa Wang, Shi-Huang Chen
ICASSP1
1999 A Large-Vocabulary Bilingual Speech Recognition System for Chinese and Japanese Language
Jyh-Shing Shyuu, Jhing-Fa Wang
PACLIC2
1998 A telephone number inquiry system with dialog structure
abstract
A telephone number inquiry system (TNIS) answers caller the phone number he/she wants to know. The traditional system requires the caller to know the full name of the party. In this paper, we propose a novel TNIS with dialog structure that can let caller use a more flexible method while inquiring, i.e., the caller may interact with our system to inquire the phone number by providing just the working, researching area, the surname, or the title, etc. Our system takes the telephone speech as input, after generating the word sequence, it performs a maximum likelihood key-feature matching with the knowledge base. If necessary information is not derived, interactive dialog manager is activated to resolve the caller's requirement. The experimental results show that our novel approach can make the system more natural.
Hsien-Chang Wang, Jhing-Fa Wang
ICASSP2
1998 An algorithm for automatic generation of Mandarin phonetic balanced corpus
Jyh-Shing Shyuu, Jhing-Fa Wang
ICSLP2
1997 VLSI architecture and implementation for FS1016 CELP decoder with reduced power and memory requirements
An-Nan Suen, Jhing-Fa Wang, Jia-Lang Lin
Integr.2
1997 A vowel-driven Mandarin speech autodialer with adaptation ability
abstract
A vowel-driven Mandarin speech autodialer based on the special characteristics of Mandarin digits is introduced. Additionally, an effective speaker adaptation technique based on the generalized probabilistic descent (GPD) algorithm is derived and integrated into the speech autodialer. Experimental results show that an encouraging performance is obtained.
Ruey-Ching Shyu, Jhing-Fa Wang, Jau-Yien Lee
IEEE Signal Process. Lett.2
1996 A programmable application-specific CELP processor with parallel architectures
abstract
The code excited linear predictive (CELP) coder has been widely used as the most effective technique among various linear predictive coding methods for speech compression. However, it is computationally intensive and general-purpose DSP chips are usually not powerful enough to handle such coding algorithms. The CELP processor architecture and a VLSI implementation are presented. A programmable application-specific single chip design for the CELP algorithm will drastically reduce the cost and achieve real-time performance. The CELP processor is programmable and contains a specific modular design for the codebook searches. On the whole, the chip can process 40 MHz sampled speech data. The FS1016 CELP coder was implemented on this processor, that is we can encode the speech data at 4.8 kbps in real-time using this single chip. Fabricated in 0.8 /spl mu/m double-metal CMOS technology, the chip size is 6.3/spl times/6.1 mm/sup 2/ and is the first chip designed for CELP.
An-Nan Suen, Jhing-Fa Wang, Bor-Yueh Liu
ICASSP2
1996 A Mandarin Voice Organizer Based on a Template-Matching Speech Recognizer
Jhing-Fa Wang, Jyh-Shing Shyuu, Chung-Hsien Wu 0001
PACLIC1
1995 Computer-aided analysis and classification of heart sounds based on neural networks and time analysis
abstract
This paper describes a computer-aided heart sound analysis and classification system (CHACS) based on neural networks and time analysis. In this system, two subsystems in both time and frequency domains are proposed. In the first subsystem, a multilayer perceptron neural network is adopted to classify heart sound patterns. In the second subsystem, a set of heuristic rules is used to characterize heart sounds. The individual classification results of these two subsystems are combined to give the final suggestion. Using this system, heart sounds can be selectively stored, retrieved, enhanced, and replayed. Besides, the CHACS provides an online display of the heart beat rate and allows an objective and reliable classification of heart sounds. Experimental results show that a classification rate of 95.6% is obtained.
Chung-Hsien Wu 0001, Ching-Wen Lo, Jhing-Fa Wang
ICASSP3
1995 A multi-layer classifier for recognition of unconstrained handwritten numerals
abstract
A hierarchical architecture for recognition of the unconstrained handwritten numerals is proposed. In the first stage of preclassification, a set of structural features named four-zone codes is adopted to preclassify the numerals. Due to the large degree of data and distortion of characters, it is possible to classify two different numerals with same features into a class. A secondary preclassification that utilizes topological stroke features is presented to solve this ambiguity. In order to promote the recognition rate to be a practical OCR system, a three layer Bayesian neural network with 20 dimensional global feature vectors is designed for fine classification of the confusing classes. Experimental results show that the recognition rate of the proposed hierarchical OCR system for handwritten numerals is over 99.82% based on 15423 samples.
Gwo-En Wang, Jhing-Fa Wang
ICDAR2
1995 A Cepstrum Chip: Architecture and Implementation
abstract
The cepstrum coefficients have been widely used for speech signal representation and play a very important role in recognition accuracies. We present a low cost architecture for VLSI implementation of LPC-based cepstrum algorithm. The circuit performs the cepstrum operation for each frame of the speech data. A pipelining architecture leads to high speed performance up to speech recognition rate. The cepstrum chip is fabricated in 1.2 /spl mu/m double-metal CMOS technology after the physical design and circuit verification. On the whole, the chip can process 18.3 MHz sampled data and it contains about 24000 transistors which occupy 227.5/spl times/213.3 mils/sup 2/ area. It has been shown to be fully functional and is the first working cepstrum chip.
An-Nan Suen, Jhing-Fa Wang, Yuen-Lin Chiang
ISCAS2
1994 A Robust Stroke Extraction Method for Handwritten Chinese Characters
abstract
The stroke analysis method is an effective approach for handwritten Chinese character recognition. But as we know, it is very difficult to accurately extract the strokes. In this paper, a robust stroke extraction method is proposed. First, smoothing and thinning processes are applied to smooth the shape and to obtain the skeleton of the observed character. Then the end point, internal point and fork point are detected by calculating their own crossing numbers while the corner points are determined by a knowledge-based iterative method. Virtual-end-points are introduced for separating a stroke into a certain number of line segments without losing the connection relations among them. By representing each line segment as a vertex and the connection relation of two segments as an edge, the observed character can be represented by an attributed graph. Finally, a stroke extraction procedure is proposed to extract the strokes from the global structures of the character. After each stroke of a character is extracted, the cross points can also be determined. Experimental results have shown that the proposed method is more effective than the other methods.3,5−6
Hong-De Chang, Jhing-Fa Wang
Int. J. Pattern Recognit. Artif. Intell.2
1994 A Bayesian neural network for separating similar complex handwritten Chinese characters
Hong-De Chang, Jhing-Fa Wang, Shye-Chorng Kuo
Pattern Recognit. Lett.2
1993 Motion oriented picture interpolation with the consideration of human perception
Liang-Wei Lee, Jhing-Fa Wang, Jau-Yien Lee, Chuen-Cherng Chen
ICASSP (5)2
1993 High throughput pipelined data path synthesis by conserving the regularity of nested loops
abstract
We present a new technique to synthesize high throughput pipelined data paths for those algorithms containing nested loops. Given an initiation interval constraint, the objective is to synthesize a low cost data path for the problem in register transfer level (RTL). Mapping algorithms for processor array synthesis which do not take initiation interval into account cannot be applied to our case; while traditional pipeline synthesis techniques suffer from complex interconnection and high register cost. Our contributions include proposing (1) an approach which conserves the regularities of nested loops, and (2) an architecture which possesses the advantages of both the highly multiplexed and lowly multiplexed architecture styles. Experiments on several algorithms in the image and DSP applications show this approach is very efficient.
Yuan-Long Jeang, Yu-Chin Hsu, Jhing-Fa Wang, Jau-Yien Lee
ICCAD3
1993 Dynamic handwritten Chinese signature verification
abstract
A dynamic handwritten Chinese signature verification system based upon a Bayesian neural network is presented. Due to a great deal of variability of handwritten Chinese signatures, the proposed Bayesian neural network is trained by an incremental learning vector quantization (ILVQ) algorithm, which endows this system with incremental learning ability, and outputs a posteriori probability to give a more reliable distance estimation. The performance analysis was based upon a set of signature data consisting of 800 true specimens, 200 simple forgeries and 200 skilled forgeries. The experimental results show the type I error is about 2% and the type II error rates are about 0.1% and 2.5% for simple and skilled forgeries, respectively.>
Hong-De Chang, Jhing-Fa Wang, Hong-Ming Suen
ICDAR2
1993 A new method for the segmentation of mixed handprinted Chinese/English characters
abstract
Describes a new method to segment mixed handprinted Chinese/English characters. First, connected component analysis is performed on each text line. A pre-merging step combines those vertical-neighboring bounding rectangles into a new one, but leaves alone the horizontal-neighboring ones which may belong to different characters. Then the feature-complexity analysis of each bounding rectangle is used to classify it into an alphanumeric character, an isolated Chinese radical or a complete Chinese character. Finally, based on the recognition result of an isolated Chinese radical and heuristic formation rules, an isolated Chinese radical may be further combined with adjacent bounded rectangle(s) to make one single character rectangle for a separate character. Implementation and experimental results of the proposed method are described.>
Hsing-Hung Kuo, Jhing-Fa Wang
ICDAR2
1993 A High Throughput-Rate Architecture for 8*8 2-D DCT
Ming-Hwa Sheu, Jau-Yien Lee, Jhing-Fa Wang, An-Nan Suen, Lian-Ying Liu
ISCAS3
1993 An Expandable Chip Desing for Gray-scale Morphological Operations
Ming-Hwa Sheu, Jhing-Fa Wang, Jau-Yien Lee, Lian-Ying Liu
ISCAS2
1993 The determination of the cycle length in high level synthesis
Ming-Hwa Sheu, Yuan-Long Jeang, Jhing-Fa Wang, Jau-Yien Lee
Integr.3
1993 Motion Oriented Picture Interpolation with Human Perceptual Consideration
Liang-Wei Lee, Jhing-Fa Wang, Jau-Yien Lee, Chuen-Cherng Chen
J. Vis. Commun. Image Represent.2
1993 Preclassification for handwritten chinese character recognition by a peripheral shape coding method
Hong-De Chang, Jhing-Fa Wang
Pattern Recognit.2
1993 Dynamic search-window adjustment and interlaced search for block-matching algorithm
abstract
A technique called dynamic search-window adjustment is proposed to improve the performance of three-step searches (TSS) and to prevent the search direction from being easily misdirected by insufficient information. An interlaced-search technique is presented for the purpose of reducing the search positions. A fast search algorithm using both techniques is proposed. It is shown that the average displaced frame difference and search positions of the proposed algorithm are about 1-7% and 24-44% fewer than TSS, respectively.>
Liang-Wei Lee, Jhing-Fa Wang, Jau-Yien Lee, Jung-Dar Shie
IEEE Trans. Circuits Syst. Video Technol.2
1992 Overall consideration of scan design and test generation
abstract
A complete system which takes the test generation algorithm, the scan cell selection strategy and the structure of the scan chain into account is proposed. It is totally different from the traditional approaches which try to enhance the ability of the individual subject. The goal of this research is to reduce the extra costs caused by the scan design, especially the test application time. Experimental results show that the overall consideration of scan design and test generation can speed up test generation and greatly reduce the amount of test application time.>
Pao-Chuan Chen, Bin-Da Liu, Jhing-Fa Wang
ICCAD3
1992 A Fast Testing Method for Sequential Circuits at the State Trasition Level
abstract
In this paper an efficient method called the fast augmented state transition (FAST) test method is proposed to alleviate the testing problem of sequential circuits at the state transition level. By adding some extra logic gates to a sequential circuit under test the FAST method guarantees that each state of the augmented circuit has both the shortest distinguishing and synchronizing sequences, hence the testing complexity can be greatly reduced. The test length of the FAST method is shorter than any other exhaustive testing approaches based on the state transition level. Furthermore the test set for the augmented circuit can be easily identified.
Wei-Lun Wang, Jhing-Fa Wang, Kuen-Jong Lee
ITC2
1992 Enhancing the multiple-fault detection of single-fault test sets
Tah-Yuan Kuo, Jhing-Fa Wang, Jau-Yien Lee
Comput. Aided Des.2
1991 Integrating neural nets and one-stage dynamic programming for speaker independent continuous Mandarin digit recognition
abstract
A Bayesian neural network; a one-stage dynamic programming algorithm, and a Hopfield time-alignment network are integrated to form a speaker-independent continuous Mandarin digit recognizer. In this system, a Bayesian network trained with a splitting LVQ (learning vector quantisation) and the LVQ2 algorithms gives the a posteriori probability. The one-stage algorithm is then employed for coarse recognition. Finally, the Hopfield time-alignment network is used to eliminate unreasonable candidates. Experimental evaluation of this system, using 53 speakers (28 male, 25 female), each speaking 20 digit strings of varying length (1-7 digits/string) and at varying speaking rate (150-240 digits/min), gave an average recognition accuracy of 94.3%, with 1.4% insertion, 1.1% deletion, and 3.2% substitution errors.>
Jhing-Fa Wang, Chung-Hsien Wu 0001, Chaug-Ching Haung, Jau-Yien Lee
ICASSP1
1991 Speaker-Independent Recognition of isolated Words using concatenated Neural Networks
abstract
A speaker-independent isolated word recognizer is proposed. It is obtained by concatenating a Bayesian neural network and a Hopfield time-alignment network. In this system, the Bayesian network outputs the a posteriori probability for each speech frame, and the Hopfield network is then concatenated for time warping. A proposed splitting Learning Vector Quantization (LVQ) algorithm derived from the LBG clustering algorithm and the Kohonen LVQ algorithm is first used to train the Bayesian network. The LVQ2 algorithm is subsequently adopted as a final refinement step. A continuous mixture of Gaussian densities for each frame and multi-templates for each word are employed to characterize each word pattern. Experimental evaluation of this system with four templates/word and five mixtures/frame, using 53 speakers (28 males, 25 females) and isolated words (10 digits and 30 city names) databases, gave average recognition accuracies of 97.3%, for the speaker-trained mode and 95.7% for the speaker-independent mode, respectively. Comparisons with K-means and DTW algorithms show that the integration of the splitting LVQ and LVQ2 algorithms makes this system well suited to speaker-independent isolated word recognition. A cookbook approach for the determination of parameters in the Hopfield time-alignment network is also described.
Chung-Hsien Wu 0001, Jhing-Fa Wang, Chaug-Ching Huang, Jau-Yien Lee
Int. J. Pattern Recognit. Artif. Intell.2
1991 A shunting multilayer perceptron network for confusing/composite pattern recognition
Chung-Hsien Wu 0001, Jhing-Fa Wang, Wen-Horng Wu
Pattern Recognit.2
1990 A Fault Analysis Method for Synchronous Sequential Circuits
abstract
In this paper we extend the use of the fault analysis method dealing with combinational circuits[1] to synchronous sequential circuits. Using the iterative array model, extended forward propagation and backward implication are performed. based on the observed values at primary outputs, to deduce the actual values of each line to determine its fault status. Any stuck fault can be identified, even in a circuit without any initialization sequence. A fault which is covered is tested unconditionally; thus the results obtained would not be invalidated in the presence of untested or untestable lines. Examples will be given to demonstrate the ability of our method.
Tah-Yuan Kuo, Jau-Yien Lee, Jhing-Fa Wang
DAC3
1990 A new transform algorithm for Viterbi decoding
abstract
Implementation of the Viterbi decoding algorithm has attracted a great deal of interest in many applications, but the excessive hardware/time consumption caused by the dynamic and backtracking decoding procedures make it difficult to design efficient VLSI circuits for practical applications. A transform algorithm for maximum-likelihood decoding is derived from trellis coding and Viterbi decoding processes. Dynamic trellis search operations are paralleled and well formulated into a set of simple matrix operations referred to as the Viterbi transform (VT). Based on the VT, the excessive memory accesses and complicated data transfer scheme demanded by the trellis search are eliminated. Efficient VLSI array implementations of the VT have been developed. Long constraint length codes can be decoded by combining the processors as the building blocks.>
Kuei-Ann Wen, Ting-Shiun Wen, Jhing-Fa Wang
IEEE Trans. Commun.3
1989 A New Approach to Derive Robust Sets for Stuck-open Faults in CMOS Combinational Logic Circuits
abstract
In this paper, we address the problem of deriving robust tests for single stuck-open faults in CMOS combinational circuits. We first examine the characteristics of the transition of the two patterns belonging to a two-pattern test. Then, a sufficient and necessary condition for a test to be robust is given. According to the given condition, we propose a new method to derive robust tests. Robustness verification of the derived tests is no longer required when using our approach.
Jhing-Fa Wang, Tah-Yuan Kuo, Jau-Yien Lee
DAC1
1989 An Adaptive Reduction Procedure for the Piecewise Linear Approximation of Digitized Curves
abstract
A new algorithm is presented for the piecewise linear approximation of two-dimensional digitized curves against a square grid. The algorithm utilizes an adaptive reduction procedure in two approximation phases to select the critical points of a digitized curve such that the deviation, from the digitized curve to its final approximated curve, is bounded by a uniform error tolerance. The time complexity of this algorithm is O(m/sup 2/) rather than O(n/sup 2/) on the theoretical plane. In the experiments of fixing the initial and the final processing points, the performance of the algorithm has been compared to those of three prominent other algorithms regarding the required number of critical points and the total execution time of the program. Of the four algorithms compared, the present algorithm consistently has the shortest execution time of the program, and it tends most to require as few critical points as the optimum algorithm that was tested.>
Chin-Shyurng Fahn, Jhing-Fa Wang, Jau-Yien Lee
IEEE Trans. Pattern Anal. Mach. Intell.2
1989 Graph theoretic characterization and reliability of the generalized Boolean n-cube network
Tsung-Chuan Huang, Jhing-Fa Wang, Chu-Sing Yang, Jau-Yien Lee
Parallel Comput.2
1988 A topology-based component extractor for understanding electronic circuit diagrams
Chin-Shyurng Fahn, Jhing-Fa Wang, Jau-Yien Lee
Comput. Vis. Graph. Image Process.2
1986 Fast execution for circuit consistency verification
L. G. Chen, Jau-Yien Lee, Jhing-Fa Wang, K. T. Chen
Integr.3