Frank Y. Shih

dblp:s/FrankYShih · also Frank Yeong-Chyang Shih · DBLP profile ↗
← Back
135ranked-venue papers
68as first author
21since 2021 · last 2026
0000-0002-1507-0285ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 91 · 47 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 16 first-author · 2 since 2021Databases, data management, data science and information retrieval · 14 · 8 first-authorHuman-computer interaction and ubiquitous computing · 5 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1Security and privacy · 1
YearPublicationVenuePosition
2026 Enhancing Copy-Move Forgery Detection via Attention-Based Similarity Modeling
abstract
In copy-move forgery detection (CMFD), distinguishing between source (copied) and target (pasted) regions is essential for understanding the nature and intent of the manipulation. A key component in existing CMFD methods is the similarity modeling branch, which typically relies on self-correlation to identify duplicated regions. However, self-correlation is often limited in detecting small, subtle manipulations due to its rigid spatial comparison. To address this, we introduce a self-attention mechanism in place of self-correlation, enabling the model to capture asymmetric and context-aware dependencies across feature maps. In this paper, we integrate this modification into the BusterNet architecture, which is the first framework to explicitly model source–target disambiguation through its dual-branch design. Experimental results demonstrate that our proposed attention-enhanced model achieves higher recall and F1-scores in pixel-level evaluations, particularly for fine-grained forgeries, while maintaining comparable computational cost. Importantly, this modification is general and can be applied to other CMFD architectures as well.
Mehri R. Aliabadi, Frank Y. Shih
Int. J. Pattern Recognit. Artif. Intell.2
2025 Frequency-Guided Contextual Image Captioning
abstract
Both a deep understanding of visual cues and their contextual importance are demanded by effective image captioning. However, seamlessly integrating balanced contextual information continues to be a substantial challenge. In this paper, we present FreConCap, a novel Frequency-guided Con-textual Image Captioning framework, to overcome the challenge using high-frequency and background features, along with object-level region features. We transform grid features into frequency domain and filter out low-frequency components by a cutoff ratio that enhances fine details critical for detailed visual understanding. Multi-Stream Cross Attention is developed to reduce the modality gap between vision and language, and to capture the interaction of text features with high-frequency local features, objects, context, and their relationships. Our experiments on the MS COCO image captioning benchmark show the superiority of our approach as compared with existing methods for enhanced image captions with more contextual information.
Al Shahriar Rubel, Frank Y. Shih, Fadi P. Deek
ICIP2
2025 Development of Specialized Deep-Learning Models for Crop Freshness Assessment to Mitigate Post-Harvest Loss
abstract
Traditional methods for evaluating crop ripeness are critiqued for their inefficiency and potential harm to produce. The use of image-processing and deep-learning techniques can solve these issues as a trend in non-destructive methods. However, an overfitting problem arises when optimization and generalization are used to estimate the parameters of the next epoch. In this paper, we develop specialized models with a high volume of training images for a single type of crop to achieve the goal of 100% accuracy for both test and validation datasets. This development contributes insights into leveraging deep learning for crop assessment, emphasizing its potential application in diverse agricultural scenarios. Experimental results show that the proposed models are superior to several existing available methods.
Wellington Cunha, Arashdeep Kaur, Frank Y. Shih
Int. J. Pattern Recognit. Artif. Intell.3
2025 BFC-Cap: Background and Frequency-guided Contextual Image Captioning
abstract
Effective image captioning relies on both visual understanding and contextual relevance. In this paper, we present two approaches, BFC-Capb–a novel background-based image captioning and its extension BFC-Capf–frequency-guided, to achieve the above goals. First, we develop an Object-Background Attention (OBA) module to capture the interaction and relationship between objects and background features. Then, we incorporate feature fusion with spatial shift operation, enabling alignment with neighbors and avoiding potential redundancy. This framework is extended to transform grid features into frequency domain and filter out low-frequency components to enhance fine details. Our approaches are evaluated using traditional and recent metrics on MS COCO image captioning benchmark. Experimental results show the effectiveness of our proposed approaches, achieving better quantitative scores as compared to the relevant existing methods. Furthermore, our methods show improved qualitative captions with more background and concise contextual information, including more accurate information regarding the objects and their attributes.
Al Shahriar Rubel, Frank Y. Shih, Fadi P. Deek
Int. J. Pattern Recognit. Artif. Intell.2
2025 A Novel Adaptive Data Transformation for Contrastive Learning
abstract
Although contrastive learning has been playing a critical role in pattern recognition, how to optimize positive pairs through data transformation is still not well developed up to now. In this paper, we propose a novel Adaptive Data Transformation, named ADTrans, to identify an optimal sequence of data transformations, which enables generating high-quality positive pairs adaptively during contrastive training. Extensive experiments on benchmark datasets have shown that ADTrans can improve the performance of representation learning on downstream tasks significantly, including image classification, instance segmentation, and object detection. It can achieve a classification accuracy of 12% and 9% higher than the existing MOCO v2, SimSiam, and BYOL on the STL 10 and TinyImageNet datasets, respectively, with the ResNet-18 backbone. Moreover, it outperforms MOCO v2 on COCO instance segmentation, object detection, and Pascal VOC instance segmentation.
Yucong Shen, NhatHai Phan, Frank Y. Shih
Int. J. Pattern Recognit. Artif. Intell.3
2025 Test Time Prompt Tuning by Optimal Transport for Machine Learning
abstract
Recently, self-supervised learning has drawn lots of attention from researchers. CLIP is a vision-language model that performs cross-modality contrastive pre-training. In this paper, we propose a novel method of prompt tuning by optimal transport to improve zero-shot generalization of the CLIP pre-trained model. Existing entropy-based approaches fail to consider the global structure of output distribution, and cannot align distributions effectively across domains. We develop the Optimal Transport-Test Time Prompt Tuning, named OT-TPT, to resolve this issue. With the help of optimal transport, it can directly align distributions to provide a global regularization effect, and therefore improve robustness against noise and distribution shifts. Moreover, a Sinkhorn regularization term is adopted to provide an efficient and smooth approximation that reduces distribution shifts while improving zero-shot generalization. Experimental results show that the proposed OT-TPT can achieve higher classification accuracies over existing state-of-the-art approaches.
Yucong Shen, Frank Y. Shih
Int. J. Pattern Recognit. Artif. Intell.2
2024 Medical X-Ray Image Enhancement Using Global Contrast-Limited Adaptive Histogram Equalization
abstract
In medical imaging, accurate diagnosis heavily relies on effective image enhancement techniques, particularly for X-ray images. Existing methods often suffer from various challenges such as sacrificing global image characteristics over local image characteristics or vice versa. In this paper, we present a novel approach, called G-CLAHE (Global-Contrast Limited Adaptive Histogram Equalization), which perfectly suits medical imaging with a focus on X-rays. This method adapts from Global Histogram Equalization (GHE) and Contrast Limited Adaptive Histogram Equalization (CLAHE) to take both advantages and avoid weakness to preserve local and global characteristics. Experimental results show that it can significantly improve current state-of-the-art algorithms to effectively address their limitations and enhance the contrast and quality of X-ray images for diagnostic accuracy.
Sohrab Namazi Nia, Frank Y. Shih
Int. J. Pattern Recognit. Artif. Intell.2
2024 A Novel Multi-Data-Augmentation and Multi-Deep-Learning Framework for Counting Small Vehicles and Crowds
abstract
Counting small pixel-sized vehicles and crowds in unmanned aerial vehicles (UAV) images is crucial across diverse fields, including geographic information collection, traffic monitoring, item delivery, communication network relay stations, as well as target segmentation, detection, and tracking. This task poses significant challenges due to factors such as varying view angles, non-fixed drone cameras, small object sizes, changing illumination, object occlusion, and image jitter. In this paper, we introduce a novel multi-data-augmentation and multi-deep-learning framework designed for counting small vehicles and crowds in UAV images. The framework harnesses the strengths of specific deep-learning detection models, coupled with the convolutional block attention module and data augmentation techniques. Additionally, we present a new method for detecting cars, motorcycles, and persons with small pixel sizes. Our proposed method undergoes evaluation on the test dataset v2 of the 2022 AI Cup competition, where we secured the first place on the private leaderboard by achieving the highest harmonic mean. Subsequent experimental results demonstrate that our framework outperforms the existing YOLOv7-E6E model. We also conducted comparative experiments using the publicly available VisDrone datasets, and the results show that our model outperforms the other models with the highest AP50 score of 52%.
Chun-Ming Tsai, Frank Y. Shih
Int. J. Pattern Recognit. Artif. Intell.2
2024 The Deep Hybrid Neural Network and an Application on Polyp Detection
abstract
Mathematical morphology and convolution operators are two different methods to extract the characteristics and structures of images. Over the past decades, Deep Convolutional Neural Networks (DCNN) have been proven to be more powerful than traditional image-processing approaches. In this paper, we propose a novel structure called Deep Hybrid Neural Network (DHNN) by taking advantage of the convolution and morphological neural layers. Its practical application to polyp detection in medical images is illustrated. For experimental completeness, we adopt nine polyp image datasets, including publicly available data and our own collected data. For performance comparisons, we select three backbone models. Experimental results show that our DHNN achieves the best performance in comparisons in terms of computational complexity and accurate performance.
Yi-Ta Wu, Frank Y. Shih, Kuang-Ting Hsiao, You-Cheng Liu, Fu-Chieh Chang 0003, En-Da Yu
Int. J. Pattern Recognit. Artif. Intell.2
2023 FPA-Net: Frequency-Guided Position-Based Attention Network for Land Cover Image Segmentation
abstract
Land cover segmentation has been a significant research area because of its multiple applications including the infrastructure development, forestry, agriculture, urban planning, and climate change research. In this paper, we propose a novel segmentation method, called Frequency-guided Position-based Attention Network (FPA-Net), for land cover image segmentation. Our method is based on encoder–decoder improved U-Net architecture with position-based attention mechanism and frequency-guided component. The position-based attention block is used to capture the spatial dependency among different feature maps and obtain the relationship among relevant patterns across the image. The frequency-guided component provides additional support with high-frequency features. Our model is simple and efficient in terms of time and space complexities. Experimental results on the Deep Globe, GID-15, and Land Cover AI datasets show that the proposed FPA-Net can achieve the best performance in both quantitative and qualitative measures as compared against other existing approaches.
Al Shahriar Rubel, Frank Y. Shih
Int. J. Pattern Recognit. Artif. Intell.2
2023 Drug Toxicity Prediction by Machine Learning Approaches
abstract
Drug property prediction, especially toxicity, helps reduce risks in a range of real-world applications. In this paper, we aim to apply various machine-learning models for solving the drug toxicity prediction problem. Among various machine-learning approaches, we select five suitable representatives: random forest, multi-layer perceptron, logistic regression, graph convolutional neural network, and graph isomorphism network (GIN) for conducting experiments on six datasets for toxicity prediction, including Tox 21, ClinTox, ToxCast, SIDER, HIV, and BACE. We design the GIN with four hidden layers and select the Adam optimizer with the learning rate [Formula: see text] and the batch size [Formula: see text]. Furthermore, we use a batch norm layer inside each of the GIN hidden layers. Experimental results show that the designed GIN model is most efficient in distinguishing between safe and toxic drugs and outperforms the others under the supervision of ROC AUC score and recall.
Yucong Shen, Frank Y. Shih
Int. J. Pattern Recognit. Artif. Intell.2
2022 Deep Morphological Neural Networks
abstract
Mathematical morphology intends to extract object features such as geometric and topological structures in digital images. Given a set of target images and original images, it is cumbersome and time-consuming to determine the suitable morphological operations and structuring elements. In this paper, we propose deep morphological neural networks, which include a nonlinear feature extraction layer to learn the structuring element correctly and an adaptive layer to select appropriate morphological operations automatically. We demonstrate the applications of object recognition, including hand-written digits, geometric shapes, traffic signs, and brain tumor. Experimental results show the higher computational efficiency and higher accuracy of our developed model as compared against existing convolutional neural network models.
Yucong Shen, Frank Y. Shih, Xin Zhong 0001, I-Cheng Chang
Int. J. Pattern Recognit. Artif. Intell.2
2022 An Efficient Detection and Recognition System for Multiple Motorcycle License Plates Based on Decision Tree
abstract
The automatic detection and recognition for motorcycle license plates present a very challenging task since they appear more compact and versatile than vehicle license plates. In this paper, we present an efficient detection and recognition system for motorcycle license plates based on decision tree and deep learning. It can be successfully carried out under various conditions, such as frontal, horizontally or vertically skewed, blurry, poor illumination, large viewing distances or angles, distortions, multiple license plates in an image, at night or interfered with brake lights, and headlights. Experimental results show that our system performs the best when testing with multiple license plates images under different conditions as compared against six state-of-the-art methods. Furthermore, our detection and recognition system have shown more accurate results than three commercial automatic license plate recognition systems in evaluation using accuracy, precision, recall, and F1 rates.
Chun-Ming Tsai, Frank Y. Shih
Int. J. Pattern Recognit. Artif. Intell.2
2022 Adaptive Image Reconstruction for Defense Against Adversarial Attacks
abstract
Adversarial attacks can fool convolutional networks and make the systems vulnerable to fraud and deception. How to defend against malicious attacks is a critical challenge in practice. Adversarial attacks are often conducted by adding tiny perturbations on images to cause network misclassification. Noise reduction can defend the attacks; however, it is not suited for all the cases. Considering that different models have different tolerance abilities on adversarial attacks, we develop a novel detecting module to remove noise by adaptive process and detect adversarial attacks without modifying the models. Experimental results show that by comparing the classification results on adversarial samples of MNIST and two subclasses of ImageNet datasets, our models can successfully remove most of the noise and obtain detection accuracies of 97.71% and 92.96%, respectively. Furthermore, our adaptive module can be assembled into different networks to achieve detection accuracies of 70.83% and 71.96%, respectively, on the white-box adversarial attacks of ResNet18 and SCD01MLP images. The best accuracy of 62.5% is obtained for both networks when dealing with the black-box attacks.
Frank Y. Shih, I-Cheng Chang
Int. J. Pattern Recognit. Artif. Intell.2
2022 Defense Against Adversarial Attacks Based on Stochastic Descent Sign Activation Networks on Medical Images
abstract
Machine learning techniques in medical imaging systems are accurate, but minor perturbations in the data known as adversarial attacks can fool them. These attacks make the systems vulnerable to fraud and deception, and thus a significant challenge has been posed in practice. We present the gradient-free trained sign activation networks to detect and deter adversarial attacks on medical imaging AI systems. Experimental results show that a higher distortion value is required to attack our proposed model than the other existing state-of-the-art models on MRI, Chest X-ray, and Histopathology image datasets, where our model outperforms the best and is even twice superior. The average accuracy of our model in classifying the adversarial examples is 88.89%, whereas those for MLP and LeNet are 81.48%, and that of ResNet18 is 38.89%. It is concluded that the sign network is a solution to defend adversarial attacks due to high distortion and high accuracy on transferability. Our work is a significant step towards safe and secure medical AI systems.
Frank Y. Shih, Usman Roshan
Int. J. Pattern Recognit. Artif. Intell.2
2021 Classification of Chest X-Ray Images Using Novel Adaptive Morphological Neural Networks
abstract
The chest X-ray images are difficult to classify for the radiologists due to the noisy nature. The existing models based on convolutional neural networks contain a giant number of parameters, and thus require multi-advanced GPUs to deploy. In this paper, we are the first to develop the adaptive morphological neural networks to classify chest X-ray images, such as pneumonia and COVID-19. A novel structure, which can self-learn morphological dilation and erosion, is proposed to determine the most suitable depth of the adaptive layer. Experimental results on the chest X-ray and the COVID-19 datasets show that the proposed model can achieve the highest classification rate as compared against the existing models. Moreover, it can significantly reduce the computational parameters of the existing models by 97%. The advantage makes the developed model more attractive than others to deploy in the internet and other device platforms.
Frank Y. Shih, Xin Zhong 0001
Int. J. Pattern Recognit. Artif. Intell.2
2021 Joint Learning for Pneumonia Classification and Segmentation on Medical Images
abstract
Chest X-ray images are notoriously difficult to analyze due to the noisy nature. Automatic identification of pneumonia on medical images has attracted intensive study recently. In this paper, a novel joint-task architecture that can learn pneumonia classification and segmentation simultaneously is presented. Two modules, including an image preprocessing module and an attention module, are developed to improve both the classification and segmentation accuracies. Results from the experiments performed on the massive dataset of the Radiology Society of North America have confirmed its superiority over the other existing methods. The classification test accuracy is improved from 0.89 to 0.95, and the segmentation model achieves an improved mean precision result of 0.58–0.78. Finally, two weakly supervised learning methods, class-saliency map and Grad-CAM, are used to highlight the corresponding pixels or areas which have significant influence on the classification model, such that the refined segmentation can focus on the correct areas with high confidence.
Xin Zhong 0001, Frank Y. Shih
Int. J. Pattern Recognit. Artif. Intell.3
2021 Land Cover Image Segmentation Based on Individual Class Binary Masks
abstract
Remote sensing techniques have been developed over the past decades to acquire data without being in contact of the target object or data source. Their application on land-cover image segmentation has attracted significant attention during recent years. With the help of satellites, scientists and researchers can collect and store high-resolution image data that can be further be processed, segmented, and classified. However, these research results have not yet been synthesized to provide coherent guidance on the effect of variant land-cover segmentation processes. In this paper, we present a novel model that augments segmentation using smaller networks to segment individual classes. The combined network is trained on the same data but with the masks, combined and trained using categorical cross entropy. Experimental results show that the proposed method produces the highest mean IoU (Intersection of Union) as compared against several existing state-of-the-art models on the DeepGlobe dataset.
Sathyanarayanan Somasunder, Frank Y. Shih
Int. J. Pattern Recognit. Artif. Intell.2
2021 35th Anniversary of IJPRAI
Patrick Shen-Pei Wang, Xiaoyi Jiang 0001, Frank Y. Shih, Terence Sim
Int. J. Pattern Recognit. Artif. Intell.3
2021 Automatic Image Pixel Clustering based on Mussels Wandering Optimization
abstract
Image pixel clustering or segmentation intends to identify pixel groups on an image without any preliminary labels. It remains a challenging task in computer vision since the size and shape of object segments are varied. Moreover, determining the segment number in an image without prior knowledge of the image content is an NP-hard problem. In this paper, we present an automatic image pixel clustering scheme based on mussels wandering optimization. An activation variable is applied to determine the number of clusters automatically with the cluster centers optimization. We revise the within- and between-class sum of squares ratio for random natural image content and develop a novel fitness function for the image pixel clustering task. Our proposed scheme is compared against existing state-of-the-art techniques using both synthetic data and real ASD dataset. Experimental results show the superiority performance of the proposed scheme.
Xin Zhong 0001, Frank Y. Shih
Int. J. Pattern Recognit. Artif. Intell.2
2021 An Automated and Robust Image Watermarking Scheme Based on Deep Neural Networks
abstract
Digital image watermarking is the process of embedding and extracting a watermark covertly on a cover-image. To dynamically adapt image watermarking algorithms, deep learning–based image watermarking schemes have attracted increased attention during recent years. However, existing deep learning–based watermarking methods neither fully apply the fitting ability to learn and automate the embedding and extracting algorithms, nor achieve the properties of robustness and blindness simultaneously. In this paper, a robust and blind image watermarking scheme based on deep learning neural networks is proposed. To minimize the requirement of domain knowledge, the fitting ability of deep neural networks is exploited to learn and generalize an automated image watermarking algorithm. A deep learning architecture is specially designed for image watermarking tasks, which will be trained in an unsupervised manner to avoid human intervention and annotation. To facilitate flexible applications, the robustness of the proposed scheme is achieved without requiring any prior knowledge or adversarial examples of possible attacks. A challenging case of watermark extraction from phone camera–captured images demonstrates the robustness and practicality of the proposal. The experiments, evaluation, and application cases confirm the superiority of the proposed scheme.
Xin Zhong 0001, Pei-Chi Huang, Spyridon Mastorakis, Frank Y. Shih
IEEE Trans. Multim.4
2020 Accurate and adversarially robust classification of medical images and ECG time-series with gradient-free trained sign activation neural networks
abstract
Adversarial attacks in medical AI imaging systems can lead to misdiagnosis and insurance fraud as recently highlighted by Finlayson et. al. in Science 2019. They can also be carried out on widely used ECG time-series data as shown in Han et. al. in Nature Medicine 2020. At the heart of adversarial attacks are imperceptible distortions that are visually and statistically undetectable but cause the machine learning model to misclassify data. Recent empirical studies have shown that a gradient-free trained sign activation neural network ensemble model requires a larger distortion than state of the art models. We apply them on medical data in this study as a potential solution to detect and deter adversarial attacks. We show on chest X-ray and histopathology images, and on two ECG datasets that this model requires a greater distortion to be fooled than full-precision, binary, and convolutional neural networks, and random forests. We show that adversaries targeting the gradient-free sign networks are visually distinguishable from the original data and thus likely to be detected by human inspection. Since the sign network distortions are higher we expect an automated method could be developed to detect and deter attacks in advance. Our work here is a significant step towards safe and secure medical machine learning.
Yunzhe Xue, Frank Y. Shih, Justin W. Ady, Usman Roshan
BIBM4
2020 Classification of Ecological Data by Deep Learning
abstract
Ecologists have been studying different computational models in the classification of ecological species. In this paper, we intend to take advantages of variant deep-learning models, including LeNet, AlexNet, VGG models, residual neural network, and inception models, to classify ecological datasets, such as bee wing and butterfly. Since the datasets contain relatively small data samples and unbalanced samples in each class, we apply data augmentation and transfer learning techniques. Furthermore, newly designed inception residual and inception modules are developed to enhance feature extraction and increase classification rates. As comparing against currently available deep-learning models, experimental results show that the proposed inception residual block can avoid the vanishing gradient problem and achieve a high accuracy rate of 92%.
Frank Y. Shih, Gareth Russell, Kimberly Russell, NhatHai Phan
Int. J. Pattern Recognit. Artif. Intell.2
2020 Deep Learning Classification on Optical Coherence Tomography Retina Images
abstract
This paper presents a novel deep learning classification technique applied on optical coherence tomography (OCT) retinal images. We propose the deep neural networks based on Vgg16 pre-trained network model. The OCT retinal image dataset consists of four classes, including three most common retina diseases and one normal retina scan. Because the scale of training data is not sufficiently large, we use the transfer learning technique. Since the convolutional neural networks are sensitive to a little data change, we use data augmentation to analyze the classified results on retinal images. The input grayscale OCT scan images are converted to RGB images using colormaps. We have evaluated different types of classifiers with variant parameters in training the network architecture. Experimental results show that testing accuracy of 99.48% can be obtained as combined on all the classes.
Frank Y. Shih, Himanshu Patel
Int. J. Pattern Recognit. Artif. Intell.1
2019 Development of Deep Learning Framework for Mathematical Morphology
abstract
Mathematical morphology has been applied as a collection of nonlinear operations related to object features in images. In this paper, we present morphological layers in deep learning framework, namely MorphNet, to perform atomic morphological operations, such as dilation and erosion. For propagation of losses through the proposed deep learning framework, we approximate the dilation and erosion operations by differential and smooth multivariable functions of the softmax function, and therefore enable the neural network to be optimized. The proposed operations are analyzed by the derivative of approximation functions in the deep learning framework. Experimental results show that the output structuring element of a morphological neuron and the target structuring element are matched to confirm the efficiency and correctness of the proposed framework.
Frank Y. Shih, Yucong Shen, Xin Zhong 0001
Int. J. Pattern Recognit. Artif. Intell.1
2019 Discrimination of Computer Generated and Photographic Images Based on CQWT Quaternion Markov Features
abstract
In this paper, an effective method based on the color quaternion wavelet transform (CQWT) for image forensics is proposed. Compared to discrete wavelet transform (DWT), the CQWT provides more information, such as the quaternion’s magnitude and phase measures, to discriminate between computer generated (CG) and photographic (PG) images. Meanwhile, we extend the classic Markov features into the quaternion domain to develop the quaternion Markov statistical features for color images. Experimental results show that the proposed scheme can achieve the classification rate of 92.70%, which is 6.89% higher than the classic Markov features.
Frank Y. Shih
Int. J. Pattern Recognit. Artif. Intell.3
2019 A High-Capacity Reversible Watermarking Scheme Based on Shape Decomposition for Medical Images
abstract
We present a high-capacity reversible, fragile, and blind watermarking scheme for medical images in this paper. A bottom-up saliency detection algorithm is applied to automatically locate the multiple arbitrarily-shaped regions of interest (ROIs). The iterative square-production algorithm is developed to generate different sizes of squares for shape decomposition on the regions of noninterest (RONIs). This scheme of combining the frequency-domain watermarking and arbitrarily-shaped ROI methods can significantly increase the watermarking capacity, whereas the embedded image fidelity is preserved. Extensive experiments were carried out on the OASIS medical image dataset, which consists of a cross-sectional collection of 416 subjects, aged from 18 to 96 years old. The results show that the proposed scheme outperforms six existing state-of-the-art schemes in terms of watermarking capacity and embedded image fidelity.
Xin Zhong 0001, Frank Y. Shih
Int. J. Pattern Recognit. Artif. Intell.2
2019 An Efficient Saliency Detection Model Based on Wavelet Generalized Lifting
abstract
Saliency detection refers to the segmentation of all visually conspicuous objects from various backgrounds. The purpose is to produce an object-mask that overlaps the salient regions annotated by human vision. In this paper, we propose an efficient bottom-up saliency detection model based on wavelet generalized lifting. It requires no kernels with implicit assumptions and prior knowledge. Multiscale wavelet analysis is performed on broadly tuned color feature channels to include a wide range of spatial-frequency information. A nonlinear wavelet filter bank is designed to emphasize the wavelet coefficients, and then a saliency map is obtained through linear combination of the enhanced wavelet coefficients. This full-resolution saliency map uniformly highlights multiple salient objects of different sizes and shapes. An object-mask is constructed by the adaptive thresholding scheme on the saliency maps. Experimental results show that the proposed model outperforms the existing state-of-the-art competitors on two benchmark datasets.
Xin Zhong 0001, Frank Y. Shih
Int. J. Pattern Recognit. Artif. Intell.2
2019 Robust Multibit Image Watermarking Based on Contrast Modulation and Affine Rectification
abstract
In this paper, we present a robust multibit image watermarking scheme to undertake the common image-processing attacks as well as affine distortions. This scheme combines contrast modulation and effective synchronization for large payload and high robustness. We analyze the robustness, payload, and the lower bound of fidelity. Regarding watermark resynchronization under affine distortions, we develop a self-referencing rectification method to detect the distortion parameters for reconstruction by the center of mass in affine covariant regions. The effectiveness and advantages of the proposed scheme are confirmed by experimental results, which show the superior performance as comparing against several state-of-the-art watermarking methods.
Xin Zhong 0001, Frank Y. Shih
Int. J. Pattern Recognit. Artif. Intell.2
2018 A Visual Secret Sharing Scheme Based on Improved Local Binary Pattern
abstract
A visual secret sharing (VSS) scheme is intended to share secret information in a group to avoid potential treat of interruption and modification. In this paper, we present a novel VSS scheme based on the improved local binary pattern (LBP) operator. It makes full use of local contrast features of LBP for concealing secret image data into different image shares, which can be used to recover the secret easily and exactly. By varying LBP extensions, we can design various kinds of VSS schemes for sharing secret information. Compared to the currently available VSS algorithms, the proposed scheme demonstrates better randomness in shares with less pixel expansion and exact determination in reconstruction with lower computational cost.
Wenyin Zhang, Frank Y. Shih, Shunbo Hu, Muwei Jian
Int. J. Pattern Recognit. Artif. Intell.2
2018 An adjustable-purpose image watermarking technique by particle swarm optimization
Frank Y. Shih, Xin Zhong 0001, I-Cheng Chang, Shin'ichi Satoh 0001
Multim. Tools Appl.1
2017 Achieving Image Watermarking Robustness by Geometric Rectification
abstract
Image watermarking techniques have been widely used for copyright protection, broadcast monitoring, and data authentication. In this paper, we present a novel watermarking scheme which allows automatic selection of multiple regions-of-interest (ROIs) with robustness against geometric distortion. The fidelity of watermarked images is ensured by preserving salient foreground objects. The proposed scheme achieves watermarking robustness by geometric rectification, which is based on matching feature points between the salient foreground objects of a host image and its distorted stego-image. Experimental results show that the proposed technique can successfully obtain high fidelity and high robustness on an image dataset of multiple salient foreground objects.
Frank Y. Shih, Xin Zhong 0001
Int. J. Pattern Recognit. Artif. Intell.1
2017 Automated Counting and Tracking of Vehicles
abstract
A robust traffic surveillance system is crucial in improving the control and management of traffic systems. Vehicle flow processing primarily involves counting and tracking vehicles; however, due to complex situations such as brightness changes and vehicle partial occlusions, traditional image segmentation methods are unable to segment and count vehicles correctly. This paper presents a novel framework for vision-based vehicle counting and tracking, which consists of four main procedures: foreground detection, feature extraction, feature analysis, and vehicles counting/tracking. Foreground detection intends to generate regions of interest in an image, which are used to produce significant feature points. Vehicles counting and tracking are achieved by analyzing clusters of feature points. As for testing on recorded traffic videos, the proposed framework is verified to be able to separate occluded vehicles and count the number of vehicles accurately and efficiently. By comparing with other methods, we observe that the proposed framework achieves the highest occlusion segment rate and the counting accuracy.
Frank Y. Shih, Xin Zhong 0001
Int. J. Pattern Recognit. Artif. Intell.1
2017 An Efficient Image Stitching Method for Heterogeneous Car Videos Based on Bounding Boxes of Features
abstract
Heterogeneous car video recorders can capture scene information with different modalities including viewing angles, resolutions, and lens sensors. Traditional methods cannot accurately perform image stitching on the images captured by heterogeneous cameras. This paper presents an efficient method to stitch heterogeneous images by allowing a driver to view an ultra-wide angle without blind spots. It extracts bounding boxes of brake lights and license plate numbers as feature points to be matched. A homography matrix is computed to stitch the heterogeneous video images. Experimental results show that our proposed method can stitch images accurately and efficiently, which is superior to the existing methods.
Chun-Ming Tsai, Frank Y. Shih
Int. J. Pattern Recognit. Artif. Intell.2
2016 Intelligent Watermarking for High-Capacity Low-Distortion Data Embedding
abstract
Image watermarking intends to hide secret data for the purposes of copyright protection, image authentication, data privacy, and broadcast monitoring. The ultimate goal is to achieve highest embedding capacity and lowest image distortion. In this paper, we present an intelligent watermarking scheme which can automatically analyze an image content to extract significant regions of interest (ROIs). A ROI is an area involving crucial information, and will be kept intact. The remaining regions of non-interest (RONIs) are collated for embedding watermarks, and will be ranked according to their entropy fuzzy memberships into a degree of importance. They are embedded by different amounts of bits to achieve optimal watermarking. The watermark is compressed and embedded in the frequency domain of the image. Experimental results show that the proposed technique has accomplished high capacity, high robustness, and high PSNR (peak signal-to-noise ratio) watermarking.
Frank Y. Shih, Xin Zhong 0001
Int. J. Pattern Recognit. Artif. Intell.1
2016 High-capacity multiple regions of interest watermarking for medical images
Frank Y. Shih, Xin Zhong 0001
Inf. Sci.1
2016 Active colloids segmentation and tracking
Boyang Gao, Simon Masnou, Liming Chen 0002, Isaac Theurkauff, Cécile Cottin-Bizonne, Frank Y. Shih
Pattern Recognit.8
2015 Copy-Cover Image Forgery Detection in Parallel Processing
abstract
The decrease in camera and smartphone sizes, the increase in picture resolutions, and the convenient access to cameras have created tremendous increases in a large volume of images being captured daily. As computer technology continues to improve, smartphones keep getting smaller, mobile processing power continues to significantly increase, and dual- or quad-core processors are becoming more popular. This has created an increase in the availability of software on small portable devices that cannot only take photos, but also edit them directly on the device, including the use of the copy-paste feature. In this paper, we present the copy-cover image forgery detection that can be used to identify a photo, whether it was altered by copying a portion of the image over another portion of itself. We also extend the detection in parallel processing by utilizing a computer with multiple CPU cores and parallel enhancements to the algorithm. Experimental results show there are significant improvements in performance by taking advantage of parallelism. Changing the detection routines in parallel can scale and take advantage of computing hardware equipped with multiple processors and large amounts of memory.
Frank Y. Shih, Jason K. Jackson
Int. J. Pattern Recognit. Artif. Intell.1
2014 A Self-Directed Method for Image Segmentation using a modified Top-Down Region Dividing Approach
abstract
We present a self-directed method for image segmentation using a modified top-down region dividing (TDRD) approach. The TDRD-based image segmentation method solves some of the issues with histogram and region growing-based segmentation techniques. The process is efficient and achieves proper results without over segmentation or spatial-structure destruction. In this paper, we examine seven user-defined parameters of the method. These parameters are converted from human inputs to values derived from in-class information created by the algorithm allowing for autonomous image segmentation, without the need of human input or feedback. Our new autonomous implementation also reduces the computational complexity of the algorithm. This reduction will produce significant savings for the total number of computations the algorithm needs to perform image segmentation. Experimental results show that the images using these new derived values yield superior results as compared to other methods, including the original TDRD method. We compare our results visually and numerically based on the within-class standard deviation (WCSD) and the number of connected components (NCC).
Walter S. Wehner Jr., Frank Y. Shih
Int. J. Pattern Recognit. Artif. Intell.2
2014 Retinal vessels segmentation based on level set and region growing
Frank Y. Shih
Pattern Recognit.4
2013 An efficient expanding block algorithm for image copy-move forgery detection
Gavin Lynch, Frank Y. Shih, Hong-Yuan Mark Liao
Inf. Sci.2
2012 A Level-Set Method Based on Global and Local Regions for Image Segmentation
abstract
This paper presents a new level-set method based on global and local regions for image segmentation. First, the image fitting term of Chan and Vese (CV) model is adapted to detect the image's local information by convolving a Gaussian kernel function. Then, a global term is proposed to detect large gradient amplitude at the outer region. The new energy function consists of both local and global terms, and is minimized by the gradient descent method. Experimental results on both synthetic and real images show that the proposed method can detect objects in inhomogeneous, low-contrast, and noisy images more accurately than the CV model, the local binary fitting model, and the Lankton and Tannenbaum model.
Frank Y. Shih, Gang Yu 0004
Int. J. Pattern Recognit. Artif. Intell.3
2011 Categorization of Camera Captured Documents Based on Logo Identification
Venkata Gopal Edupuganti, Frank Y. Shih, Suryaprakash Kompalli
CAIP (2)2
2011 An Improved Feature Vocabulary Based Method for Image Categorization
abstract
The bags of feature and feature vocabulary based approaches have been presented for image categorization due to their simplicity and competitive performance. Some modified versions have been subsequently proposed, incorporating the methods such as adapted vocabularies, fast indexing, and Gaussian mixture models. In this paper, we propose an improvement of replacing the Harris-affine detection method by a random sampling procedure together with an increased number of sample points. Experimental results show that this new method improves categorization accuracy on a five-category problem using the Caltech-4 dataset. It is concluded that random sampling produces higher attainable point density and better categorization performance.
Frank Y. Shih, Alexander Sheppard
Int. J. Pattern Recognit. Artif. Intell.1
2011 A kernel-based parametric method for conditional density estimation
Frank Y. Shih, Haimin Wang
Pattern Recognit.2
2010 Passive Detection of Paint-Doctored JPEG Images
Frank Y. Shih, Yun Q. Shi 0001
IWDW2
2009 A differential evolution based algorithm for breaking the visual steganalytic system
Frank Y. Shih, Venkata Gopal Edupuganti
Soft Comput.1
2008 Extracting Faces and Facial Features from Color Images
abstract
In this paper, we present image processing and pattern recognition techniques to extract human faces and facial features from color images. First, we segment a color image into skin and non-skin regions by a Gaussian skin-color model. Then, we apply mathematical morphology and region filling techniques for noise removal and hole filling. We determine whether a skin region is a face candidate by its size and shape. Principle component analysis (PCA) is used to verify face candidates. We create an ellipse model to locate eyes and mouths areas roughly, and apply the support vector machine (SVM) to classify them. Finally, we develop knowledge rules to verify eyes. Experimental results show that our algorithm achieves the accuracy rate of 96.7% in face detection and 90.0% in facial feature extraction.
Frank Y. Shih, Shouxian Cheng, Chao-Fa Chuang, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.1
2008 Performance Comparisons of Facial Expression Recognition in Jaffe Database
abstract
Facial expression provides an important behavioral measure for studies of emotion, cognitive processes, and social interaction. Facial expression recognition has recently become a promising research area. Its applications include human-computer interfaces, human emotion analysis, and medical care and cure. In this paper, we investigate various feature representation and expression classification schemes to recognize seven different facial expressions, such as happy, neutral, angry, disgust, sad, fear and surprise, in the JAFFE database. Experimental results show that the method of combining 2D-LDA (Linear Discriminant Analysis) and SVM (Support Vector Machine) outperforms others. The recognition rate of this method is 95.71% by using leave-one-out strategy and 94.13% by using cross-validation strategy. It takes only 0.0357 second to process one image of size 256 × 256.
Frank Y. Shih, Chao-Fa Chuang, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.1
2008 Decision Combination of Multiple Classifiers
abstract
In order to improve the performance in pattern classification, we utilize multiple classifiers and combine their individual decisions to make a final decision. In this paper, we present the combination using Bayesian method and compare minimum errors. This method requires the posteriori probabilities from all classifiers, which may be difficult to calculate in real world because tremendous amounts of training samples are needed. Alternatively, a confusion matrix is developed for approximation. We also use different combining rules for comparisons and apply them to handwritten digit recognition.
Frank Y. Shih
Int. J. Pattern Recognit. Artif. Intell.1
2008 Approximate Image Quality Measure in Low-Dimensional Domain Based on Random Projection
abstract
Image Quality Measure (IQM) is used to automatically measure the degree of image artifacts such as blocking, ringing and blurring effects. It is calculated traditionally in the image spatial domain. In this paper, we present a new method of transforming an image into a low-dimensional domain based on random projection, so we can efficiently obtain the compatible IQM. From the transformed domain, we can calculate the Peak Signal-to-Noise Ratio (PSNR) and apply fuzzy logic to generate a Low-Dimensional Quality Index (LDQI). Experimental results show that the LDQI can approximate the IQM in the image spatial domain. We observe that the LDQI is suited for measuring the compression blur due to its relatively low distortion. The relative error is about 0.15 as the compression blur increases.
Frank Y. Shih, Yan-Yu Fu
Int. J. Pattern Recognit. Artif. Intell.1
2008 A distance-based separator representation for pattern classification
Frank Y. Shih, Kai Zhang 0047
Image Vis. Comput.1
2008 A top-down region dividing approach for image segmentation
Yi-Ta Wu, Frank Y. Shih, Jiazheng Shi, Yih-Tyng Wu
Pattern Recognit.2
2008 Automatic Detection of Magnetic Flux Emergings in the Solar Atmosphere From Full-Disk Magnetogram Sequences
abstract
In this paper, we present a novel method to detect Emerging Flux Regions (EFRs) in the solar atmosphere from consecutive full-disk Michelson Doppler Imager (MDI) magnetogram sequences. To our knowledge, this is the first developed technique for automatically detecting EFRs. The method includes several steps. First, the projection distortion on the MDI magnetograms is corrected. Second, the bipolar regions are extracted by applying multiscale circular harmonic filters. Third, the extracted bipolar regions are traced in consecutive MDI frames by Kalman filter as candidate EFRs. Fourth, the properties, such as positive and negative magnetic fluxes and distance between two polarities, are measured in each frame. Finally, a feature vector is constructed for each bipolar region using the measured properties, and the Support Vector Machine (SVM) classifier is applied to distinguish EFRs from other regions. Experimental results show that the detection rate of EFRs is 96.4% and of non-EFRs is 98.0%, and the false alarm rate is 25.7%, based on all the available MDI magnetograms in 2001 and 2002.
Frank Y. Shih, Haimin Wang
IEEE Trans. Image Process.2
2007 Locating object contours in complex background using improved snakes
Frank Y. Shih, Kai Zhang 0047
Comput. Vis. Image Underst.1
2007 Machine assessment of neonatal facial expressions of acute pain
Sheryl Brahnam, Chao-Fa Chuang, Randall S. Sexton, Frank Y. Shih
Decis. Support Syst.4
2007 Shape-Based Image Retrieval Using Two-Level Similarity Measures
abstract
In this paper, we present a novel method of using two-level similarity measures for shape-based image retrieval. We first identify the dominant points of a given shape, and then calculate their geometric moments and the distances between two consecutive dominant points. A spectrum representing the normalized geometric moments versus normalized distances is generated, and its area and curve length are computed. We use these two values as similarity features for the indexes in coarse-grained shape retrieval. Furthermore, we use the cross-sectional area and curve length distribution for the indexes in fine-grained shape retrieval. Experimental results show that the proposed method is simple and efficient and can reach the accuracy rate of 95%.
Wai-Tak Wong, Frank Y. Shih, Te-Feng Su
Int. J. Pattern Recognit. Artif. Intell.2
2007 A hierarchical refinement algorithm for fully automatic gridding in spotted DNA microarray image processing
Marc Q. Ma, Kai Zhang 0047, Frank Y. Shih
Inf. Sci.4
2007 Shape-based image retrieval using support vector machines, Fourier descriptors and self-organizing maps
Wai-Tak Wong, Frank Y. Shih, Jung Liu
Inf. Sci.2
2007 An improved incremental training algorithm for support vector machines using active query
Shouxian Cheng, Frank Y. Shih
Pattern Recognit.2
2007 Digital watermarking based on chaotic map and reference register
Yi-Ta Wu, Frank Y. Shih
Pattern Recognit.2
2007 Automatic Detection of Prominence Eruption Using Consecutive Solar Images
abstract
Prominences are clouds of relatively cool and dense gas in the solar atmosphere. In this paper, we present a new method to detect and characterize the prominence eruptions. The input is a sequence of consecutive Halpha solar images, and the output is a list of prominence eruption events detected. We extract the limb events and measure their associated properties by applying image processing techniques. First, we perform image normalization and noise removal. Then, we isolate the limb objects and identify the prominence features. Finally, we apply pattern recognition techniques to classify the eruptive prominences. The characteristics of prominence eruptions, such as brightness, angular width, radial height and velocity are measured. The method presented can lead to automatic monitoring and characterization of solar events
Frank Y. Shih, Haimin Wang
IEEE Trans. Circuits Syst. Video Technol.2
2007 ELB-Q: A New Method for Improving the Robustness in DNA Microarray Image Quantification
abstract
Reliable and robust quantification of signal intensities is a critical step in microarray-based biomedical studies. However, traditional techniques for microarray image processing would face significant challenges if the number of pixels used for the quantification of the local background and the foreground decreases dramatically. We have developed a new method, ELB-Q, which, by design, is well suited for the image quantification of microarrays with very high density of spot layout (large number of spots arranged in unit area). In ELB-Q, a large extended local background (ELB) interspot region excluding those "noise of the background" pixels is used for estimating the local background, and the quantification of spot intensities (mean and median) in the putative target spot regions is performed after further excluding background pixels in these areas based on the cutoff values established during the ELB calculation. ELB-Q takes advantage of the abundant spatial information around each spot of interest, makes no assumption of the shape and size of the spots, and needs no sophisticated adjustment. We show results of image processing using ELB-Q on both the simulated data and real DNA microarrays, which compare favorably in robustness and accuracy against those obtained with GenePix Pro 6.0 (Axon Instruments, 1999) and the Markov random field (MRF) modeling approach. The ELB-Q software is developed in Matlab, and is available upon request.
Marc Q. Ma, Kai Zhang 0047, Hui-Yun Wang, Frank Y. Shih
IEEE Trans. Inf. Technol. Biomed.4
2006 Machine recognition and representation of neonatal facial displays of acute pain
Sheryl Brahnam, Chao-Fa Chuang, Frank Y. Shih, Melinda R. Slack
Artif. Intell. Medicine3
2006 Thinning algorithms based on quadtree and octree representations
Wai-Tak Wong, Frank Y. Shih, Te-Feng Su
Inf. Sci.2
2006 A system for rotational velocity computation from image sequences
Frank Y. Shih, Tobias Klotz, Werner Brockmann
Image Vis. Comput.1
2006 Recognizing facial action units using independent component analysis and support vector machine
Chao-Fa Chuang, Frank Y. Shih
Pattern Recognit.2
2006 Genetic algorithm based methodology for breaking the steganalytic systems
abstract
Steganalytic techniques are used to detect whether an image contains a hidden message. By analyzing various image features between stego-images (the images containing hidden messages) and cover-images (the images containing no hidden messages), a steganalytic system is able to detect stego-images. In this paper, we present a new concept of developing a robust steganographic system by artificially counterfeiting statistic features instead of the traditional strategy by avoiding the change of statistic features. We apply genetic algorithm based methodology by adjusting gray values of a cover-image while creating the desired statistic features to generate the stego-images that can break the inspection of steganalytic systems. Experimental results show that our algorithm can not only pass the detection of current steganalytic systems, but also increase the capacity of the embedded message and enhance the peak signal-to-noise ratio of stego-images.
Yi-Ta Wu, Frank Y. Shih
IEEE Trans. Syst. Man Cybern. Part B2
2006 Special issue: mobile IP
Han-Chieh Chao, Lorna Uden, Frank Y. Shih
Wirel. Commun. Mob. Comput.3
2005 Decomposition of binary morphological structuring elements based on genetic algorithms
Frank Y. Shih, Yi-Ta Wu
Comput. Vis. Image Underst.1
2005 Support vector machine networks for multi-class classification
abstract
The support vector machine (SVM) has recently attracted growing interest in pattern classification due to its competitive performance. It was originally designed for two-class classification, and many researchers have been working on extensions to multiclass. In this paper, we present a new framework that adapts the SVM with neural networks and analyze the source of misclassification in guiding our preprocessing for optimization in multiclass classification. We perform experiments on the ORL database and the results show that our framework can achieve high recognition rates.
Frank Y. Shih, Kai Zhang 0047
Int. J. Pattern Recognit. Artif. Intell.1
2005 Multi-view face identification and pose estimation using B-spline interpolation
Frank Y. Shih, Camel Fu, Kai Zhang 0047
Inf. Sci.1
2005 Geometric modeling and representation based on sweep mathematical morphology
Frank Y. Shih, Vijayalakshmi Gaddipati
Inf. Sci.1
2005 Robust watermarking and compression for medical images based on genetic algorithms
Frank Y. Shih, Yi-Ta Wu
Inf. Sci.1
2005 Automatic seeded region growing for color image segmentation
Frank Y. Shih, Shouxian Cheng
Image Vis. Comput.1
2005 Enhancement of image watermark retrieval based on genetic algorithms
Frank Y. Shih, Yi-Ta Wu
J. Vis. Commun. Image Represent.1
2005 Improved feature reduction in input and feature spaces
Frank Y. Shih, Shouxian Cheng
Pattern Recognit.1
2005 Decomposition of arbitrary gray-scale morphological structuring elements
Frank Y. Shih, Yi-Ta Wu
Pattern Recognit.1
2005 A modified regulated morphological corner detector
Frank Y. Shih, Chao-Fa Chuang, Vijayalakshmi Gaddipati
Pattern Recognit. Lett.1
2004 Solar flare tracking using image processing techniques
abstract
The automatic property measurement of solar flares through their complete cyclic development is valuable in the studies of solar flares. From the analysis of solar H/spl alpha/ images, we are able to use a support vector machine (SVM) to automatically detect flares, and apply image segmentation techniques to compute the properties of solar flares. We also present our solution for automatically tracking the motion of two-ribbon flares and measuring their movement direction and speed.
Frank Y. Shih, Ju Jing, Haimin Wang
ICME2
2004 A novel fragile watermarking technique
abstract
Several fragile watermarking techniques cannot detect vector quantization (VQ) attacks; the block-based approach, for example. Celik et al. (IEEE Trans. Image Processing, vol. 11, no. 6, pp. 585-595, 2002) proposed a hierarchical watermarking approach to reveal the VQ attacks. In order to improve the hierarchical watermarking approach, we present a new fragile watermarking technique based on genetic algorithms (GAs) to embed watermarks into the frequency domain of a host image. Different from existing schemes, the adjustment in the host image by GAs will achieve the embedding of fragile watermarks. The technique can provide a fundamental platform for other fragile watermarking techniques. Furthermore, the embedding of watermarks into frequency domain enhances the security in fragile watermarking.
Frank Y. Shih, Yi-Ta Wu
ICME1
2004 Automatic Solar Flare Tracking
Frank Y. Shih, Ju Jing, Haimin Wang, David Rees
KES2
2004 Fast Euclidean distance transformation in two scans using a 3 × 3 neighborhood
Frank Y. Shih, Yi-Ta Wu
Comput. Vis. Image Underst.1
2004 Forward And Backward Chain-Code Representation For Motion Planning Of Cars
abstract
The path-planning problem is presented to show a car of any desired shape moving from a starting position to a destination in a finite space with arbitrarily shaped obstacles in it. In this paper, a new chain-code representation is developed to record the motion path when forward and backward movements are allowed. By placing the smooth turning-angle constraint, we can obtain more realistic results to the actual motion of cars. Meanwhile, by combining rotational mathematical morphology and distance transformation, we can obtain the shortest collision-free path. As soon as the distance map and the collision-free codes have been established offline, the shortest paths of cars starting from any location toward the destination can be promptly obtained online. Experimental results show that our algorithm works successfully in different conditions. We also extend our algorithm to the automated parallel parking and three-dimensional path planning.
Frank Y. Shih, Yi-Ta Wu, Brian L. C. Chen
Int. J. Pattern Recognit. Artif. Intell.1
2004 Efficient Contour Detection Based On Improved Snake Model
abstract
Active contour model, also called snake, adapts to edges in an image. A snake is defined as an energy minimizing spline – the snake's energy depends on its shape and location within the image. Problems associated with initialization and poor convergence to boundary concavities, however, have limited its utility. In this paper, we present a new external force field, named gravitation force field, for the snake model. We associate this force field with edge preserving smoothing to drive the snake for solving the problems. Our gravitation force field uses gradient values as particles to construct force field in the whole image. This force field will attract the active contour toward the edge boundary. The locations of the initial contour are very flexible, such that they can be very far away from the objects and can be inside, outside, or the mixture. The improved snake can converge toward the object boundary in a fast pace.
Frank Y. Shih, Kai Zhang 0047
Int. J. Pattern Recognit. Artif. Intell.1
2004 Inter-Frame Interpolation By Snake Model And Greedy Algorithm
abstract
In this paper, we present a novel method to solve the inter-frame interpolation problem in image morphing. We use our improved snake model that is associated with the gravitational force field to locate control points in the object contours. Afterwards, we apply the greedy algorithm in free-form deformations to achieve optimal warps among feature point pairs in starting and ending frames. The new method uses an energy-minimization function under the influence of inter-frames. The energy serves to impose frame-wise and curve-wise constraints among the interpolated frames.
Frank Y. Shih, Kai Zhang 0047
Int. J. Pattern Recognit. Artif. Intell.1
2004 A Hybrid Two-Phase Algorithm For Face Recognition
abstract
Scientists have developed numerous classifiers in the pattern recognition field, because applying a single classifier is not very conducive to achieve a high recognition rate on face databases. Problems occur when the images of the same person are classified as one class, while they are in fact different in poses, expressions, or lighting conditions. In this paper, we present a hybrid, two-phase face recognition algorithm to achieve high recognition rates on the FERET data set. The first phase is to compress the large class number database size, whereas the second phase is to perform the decision-making. We investigate a variety of combinations of the feature extraction and pattern classification methods. Principal Component Analysis (PCA), Linear Discriminant Analysis (LDA), and Support Vector Machine (SVM) are examined and tested using 700 facial images of different poses from FERET database. Experimental results show that the two combinations, LDA+LDA and LDA+SVM, outperform the other types of combinations. Meanwhile, when classifiers are considered in the two-phase face recognition, it is better to adopt the L1 distance in the first phase and the class mean in the second phase.
Frank Y. Shih, Kai Zhang 0047, Yan-Yu Fu
Int. J. Pattern Recognit. Artif. Intell.1
2004 Automatic extraction of head and face boundaries and facial features
Frank Y. Shih, Chao-Fa Chuang
Inf. Sci.1
2004 Adaptive mathematical morphology for edge linking
Frank Y. Shih, Shouxian Cheng
Inf. Sci.1
2004 Three-dimensional Euclidean distance transformation and its application to shortest path planning
Frank Y. Shih, Yi-Ta Wu
Pattern Recognit.1
2004 An adjusted-purpose digital watermarking technique
Yi-Ta Wu, Frank Y. Shih
Pattern Recognit.2
2004 The efficient algorithms for achieving Euclidean distance transformation
abstract
Euclidean distance transformation (EDT) is used to convert a digital binary image consisting of object (foreground) and nonobject (background) pixels into another image where each pixel has a value of the minimum Euclidean distance from nonobject pixels. In this paper, the improved iterative erosion algorithm is proposed to avoid the redundant calculations in the iterative erosion algorithm. Furthermore, to avoid the iterative operations, the two-scan-based algorithm by a deriving approach is developed for achieving EDT correctly and efficiently in a constant time. Besides, we discover when obstacles appear in the image, many algorithms cannot achieve the correct EDT except our two-scan-based algorithm. Moreover, the two-scan-based algorithm does not require the additional cost of preprocessing or relative-coordinates recording.
Frank Y. Shih, Yi-Ta Wu
IEEE Trans. Image Process.1
2003 General sweep mathematical morphology
Frank Y. Shih, Vijayalakshmi Gaddipati
Pattern Recognit.1
2003 Combinational image watermarking in the spatial and frequency domains
Frank Y. Shih, Scott Y. T. Wu
Pattern Recognit.1
2001 Computing unique three-dimensional object aspects representation
Frank Y. Shih, Artur J. Kowalski
Inf. Sci.1
2001 An adaptive algorithm for conversion from quadtree to chain codes
Frank Y. Shih, Wai-Tak Wong
Pattern Recognit.1
2000 Wavelet coefficients clustering using morphological operations and pruned quadtrees
Eduardo Morales, Frank Y. Shih
Pattern Recognit.2
1999 A one-pass algorithm for local symmetry of contours from chain codes
Frank Y. Shih, Wai-Tak Wong
Pattern Recognit.1
1998 A morphological approach to shortest path planning for rotating objects
Soo-Chang Pei, Chin-Lun Lai, Frank Y. Shih
Pattern Recognit.3
1998 Size-invariant four-scan Euclidean distance transformation
Frank Y. Shih, Jenny J. Liu
Pattern Recognit.1
1997 An Efficient Class of Alternating Sequential Filters in Morphology
Soo-Chang Pei, Chin-Lun Lai, Frank Y. Shih
CVGIP Graph. Model. Image Process.3
1996 Design of One-Pass Training Algorithms for Variant Morphological Operations
Jenlong Moh, Frank Y. Shih
Inf. Sci.2
1996 Skeletonization for fuzzy degraded character images
abstract
Most skeletonization algorithms are operated on binary images. To avoid information loss and distortion, a topography-based approach is proposed to apply directly on fuzzy or gray scale images. A membership function is used to indicate the degree of membership of each ridge point with respect to the skeleton. Significant ridge points are linked to form strokes of skeleton. Experimental results show that our algorithm can reduce deformation of junction points anti correctly extract the whole skeleton, although a character may be broken into pieces. For merged characters, the breaking positions can be located by searching for the saddle points. A multiple context confirmation is used to increase the reliability of breaking hypotheses.
Shy-Shyan Chen, Frank Y. Shih
IEEE Trans. Image Process.2
1996 Adaptive document block segmentation and classification
abstract
This paper presents an adaptive block segmentation and classification technique for daily-received office documents having complex layout structures such as multiple columns and mixed-mode contents of text, graphics, and pictures. First, an improved two-step block segmentation algorithm is performed based on run-length smoothing for decomposing any document into single-mode blocks. Then, a rule-based block classification is used for classifying each block into the text, horizontal/vertical line, graphics, or-picture type. The document features and rules used are independent of character font and size and the scanning resolution. Experimental results show that our algorithms are capable of correctly segmenting and classifying different types of mixed-mode printed documents.
Frank Y. Shih, Shy-Shyan Chen
IEEE Trans. Syst. Man Cybern. Part B1
1995 Threshold Decomposition of Gray-Scale Soft Morphology into Binary Soft Morphology
Christopher C. Pu, Frank Y. Shih
CVGIP Graph. Model. Image Process.2
1995 A general purpose model for image operations based on multilayer perceptrons
Jenlong Moh, Frank Y. Shih
Pattern Recognit.2
1995 A skeletonization algorithm by maxima tracking on Euclidean distance transform
Frank Y. Shih, Christopher C. Pu
Pattern Recognit.1
1995 Pipeline architectures for recursive morphological operations
abstract
Introduces efficient pipeline architectures for the recursive morphological operations. The standard morphological operation is applied directly on the original input image and produces an output image. The order of image scanning in which the operator is applied to the input pixels is irrelevant. However, the intent of the recursive morphological operations is to feed back the output at the current scanning pixel to overwrite its corresponding input pixel to be considered into computation at the following scanning pixels. The resultant output image by recursive morphology inherently depends on the image scanning sequence. Two pipelined implementations of the recursive morphological operations are presented. The design of an application-specific systolic array is first introduced. The systolic array uses 3 n cells to process an nxn image in 6 n-2 cycles. The cell utilization rate is 100%. Second, a parallel program implementing the recursive morphological operations and running on distributed-memory multicomputers is described. Performance of the program can be finely tuned by choosing appropriate partition parameters.
Frank Y. Shih, Chung Ta King, Christopher C. Pu
IEEE Trans. Image Process.1
1995 Recursive soft morphological filters
abstract
We present properties of recursive soft morphological filters that use previously filtered outputs as their inputs, cascade combinations of these filters, and the idempotent recursive soft morphological filters. The development allows problems in the implementation of cascaded recursive soft morphological filters to be reduced to the equivalent problems of a single recursive standard morphological filter.
Frank Y. Shih, Padmaja Puttagunta
IEEE Trans. Image Process.1
1995 Fuzzy typographical analysis for character preclassification
abstract
This paper presents a fuzzy-logic approach for analyzing typographical structures of textual blocks in order to be used for character preclassification. An efficient baseline detection method embedded with tolerance analysis is developed for locating precisely the baseline. Fuzzy logic is taken into account when the decision ambiguity of typographical categorization is occurred. The constraints on the fuzzy membership functions are formulated. Their boundary conditions are considered to preserve the continuity. An improved character recognition rate can be achieved by means of the typographical categorization.>
Shy-Shyan Chen, Frank Y. Shih, Peter A. Ng
IEEE Trans. Syst. Man Cybern.2
1995 A new safe-point thinning algorithm based on the mid-crack code tracing
abstract
A new thinning algorithm for binary images, based on the safe-point testing and mid-crack code tracing, is presented. Thinning is treated as the deletion of nonsafe border pixels from the contour to the center of the object layer-by-layer. The deletion is determined by masking a 3/spl times/3 weighted template and by the use of lookup tables. The resulting skeleton does not require cleaning or pruning. The obtained skeleton possesses single-pixel thickness and preserves the object's connectivity. The algorithm is very simple and efficient since only boundary pixels are processed at each iteration and lookup tables are used.>
Frank Y. Shih, Wai-Tak Wong
IEEE Trans. Syst. Man Cybern.1
1994 Analysis and modelling of deformed swept volumes
Denis Blackmore, Ming C. Leu, Frank Y. Shih
Comput. Aided Des.3
1994 An Improved Fast Algorithm for the Restoration of Images Based on Chain Codes Description
Frank Y. Shih, Wai-Tak Wong
CVGIP Graph. Model. Image Process.1
1994 An Automatic Text-Free Speaker Recognition System Based on an Enhanced Art 2 Neural Architecture
J. Thomas Eck, Frank Y. Shih
Inf. Sci.2
1994 Fully parallel thinning with tolerance to boundary noise
Frank Y. Shih, Wai-Tak Wong
Pattern Recognit.1
1994 Exact and approximate algorithms for unordered tree matching
abstract
We consider the problem of comparison between unordered trees, i.e., trees for which the order among siblings is unimportant. The criterion for comparison is the distance as measured by a weighted sum of the costs of deletion, insertion and relabel operations on tree nodes. Such comparisons may contribute to pattern recognition efforts in any field (e.g., genetics) where data can naturally be characterized by unordered trees. In companion work, we have shown this problem to be NP-complete. This paper presents an efficient enumerative algorithm and several heuristics leading to approximate solutions. The algorithms are based on probabilistic hill climbing and bipartite matching techniques. The paper evaluates the accuracy and time efficiency of the heuristics by applying them to a set of trees transformed from industrial parts based on a previously proposed morphological model.>
Dennis E. Shasha, Jason Tsong-Li Wang, Kaizhong Zhang, Frank Y. Shih
IEEE Trans. Syst. Man Cybern.4
1993 Threshold decomposition of soft morphological filters
abstract
The properties of soft morphological operations and the new definitions of binary soft morphological operations are presented. It is shown that soft morphological filtering on an arbitrary signal is equivalent to decomposing the signal into binary signals, filtering each binary signal with a binary soft morphological filter, and then reversing the decomposition. This equivalence allows problems in the analysis and the implementation of soft morphological operations in real time by using only logic gates for binary signals instead of sorting numbers.>
Frank Y. Shih, Christopher C. Pu
CVPR1
1993 On solving exact Euclidean distance transformation with invariance to object size
abstract
A distance transformation converts a digital binary image that consists of object (foreground) and non-object (background) pixels into a gray-scale image in which each object pixel has a value corresponding to the minimum distance from the background by a distance function. Due to its nonlinearity, the global operation of Euclidean distance transformation (EDT) is difficult to decompose into small neighborhood operations. Two efficient algorithms on EDT are presented, using integers of squared Euclidean distances in which the global computations can be equivalent to local 3/spl times/3 neighborhood operations. The first algorithm requires only a limited number of iterations on the chain propagation. The second algorithm can avoid iterations, and simply requires two scans of the image. The complexity of both algorithms is only linearly proportional to image size.>
Frank Y. Shih, Chyuan-Huei T. Yang
CVPR1
1993 Model-based partial shape recognition using contour curvature and affine transformation
Frank Y. Shih, Pingchang Yeh
Inf. Sci.1
1993 Reconstruction of Binary and Gray-Scale Images from Mid-crack Code Descriptions
Frank Y. Shih, Wai-Tak Wong
J. Vis. Commun. Image Represent.1
1992 Restoration of binary and gray-scale images using the contour mid-crack codes description
abstract
The chain code is a widely-used description for a contour image. Recently, a mid-crack code algorithm has been proposed as another more precise method for image representation. A simple and fast algorithm for the restoration of binary images based on the mid-crack codes description is presented. The algorithm developed has the advantages of speed, simplicity, and less storage. The algorithm also can be applied to gray-scale images with multiple regions efficiently.>
Frank Y. Shih, Wai-Tak Wong
ICPR (3)1
1992 Pattern Matching in Unordered Trees
abstract
The problem of comparison between unordered trees, i.e. trees for which the order among siblings is unimportant, is considered. The criterion for comparison is the distance as measured by a weighted sum of the costs of deletion, insertion, and relabel operations on tree nodes. Such comparisons may contribute to pattern recognition efforts in any field (e.g. genetics) where data can naturally be characterized by unordered trees. It is observed that the problem is NP-complete. An enumerative algorithm and several heuristics leading to approximate solutions are given. The algorithms are based on probabilistic hill climbing and bipartite matching techniques. The accuracy and time efficiency of the heuristics are evaluated by applying them to a set of trees transformed from industrial parts based on a previously proposed morphological model.>
Dennis E. Shasha, Jason Tsong-Li Wang, Kaizhong Zhang, Frank Y. Shih
ICTAI4
1992 Optimization on euclidean distance transformation using grayscale morphology
Frank Y. Shih
J. Vis. Commun. Image Represent.1
1992 A new single-pass algorithm for extracting the mid-crack codes of multiple regions
Frank Y. Shih, Wai-Tak Wong
J. Vis. Commun. Image Represent.1
1992 Implementing morphological operations using programmable neural networks
Frank Y. Shih, Jenlong Moh
Pattern Recognit.1
1992 A new art-based neural architecture for pattern classification and image enhancement without prior knowledge
Frank Y. Shih, Jenlong Moh, Fu-Chun Chang
Pattern Recognit.1
1992 Morphological shape description using geometric spectrum on multidimensional binary images
Frank Y. Shih, Christopher C. Pu
Pattern Recognit.1
1992 Decomposition of geometric-shaped structuring elements using morphological transformations on binary images
Frank Y. Shih
Pattern Recognit.1
1992 A mathematical morphology approach to Euclidean distance transformation
abstract
A distance transformation technique for a binary digital image using a gray-scale mathematical morphology approach is presented. Applying well-developed decomposition properties of mathematical morphology, one can significantly reduce the tremendous cost of global operations to that of small neighborhood operations suitable for parallel pipelined computers. First, the distance transformation using mathematical morphology is developed. Then several approximations of the Euclidean distance are discussed. The decomposition of the Euclidean distance structuring element is presented. The decomposition technique employs a set of 3 by 3 gray scale morphological erosions with suitable weighted structuring elements and combines the outputs using the minimum operator. Real-valued distance transformations are considered during the processes and the result is approximated to the closest integer in the final output image.
Frank Y. Shih, Owen Robert Mitchell
IEEE Trans. Image Process.1
1991 A maxima-tracking method for skeletonization from Euclidean distance function
abstract
A skeletonization algorithm based on the Euclidean distance function using the sequential maxima-tracking method is described which, when applied to a connected image, generates a connected skeleton composed of simple digital arcs. With a slight modification, the algorithm can preserve the more important features in the skeletal branches which touch the object boundary at corners. Therefore its application to shape recognition can be easily achieved.>
Frank Y. Shih, Christopher C. Pu
ICTAI1
1991 Decomposition of gray-scale morphological structuring elements
Frank Y. Shih, Owen Robert Mitchell
Pattern Recognit.1
1990 Medial axis transformation with single-pixel and connectivity preservation using Euclidean distance computation
abstract
A novel medial axis transformation (MAT) algorithm extracted from the Euclidean distance transform of a binary image is presented. The extracted MAT satisfies the following properties: reconstructivity, rotation-invariance, connectivity, and single-pixel width. The preservation of properties is proved, and some experimental results are shown. The skeleton is trimmed by removing short branches to make it simpler and useful for object recognition.>
Frank Y. Shih, Christopher C. Pu
ICPR (1)1
1989 Threshold Decomposition of Gray-Scale Morphology into Binary Morphology
abstract
Recently, a superposition property called threshold decomposition and another property called stacking were introduced and shown to apply successfully to gray-scale morphological operations. This property allows gray-scale signals to be decomposed into multiple binary signals. The signals are processed in parallel, and the results are combined to produce the desired gray-scale result. The authors present the threshold decomposition architecture and the stacking property that allows the implementation of this architecture. Gray-scale operations are decomposed into binary operations. This decomposition allows gray-scale morphological operations to be implemented using only logic gates in VLSI architectures that can significantly improve speed as well as give theoretical insight into the operations.>
Frank Y. Shih, Owen Robert Mitchell
IEEE Trans. Pattern Anal. Mach. Intell.1
1988 Automated fast recognition and location of arbitrarily shaped objects by image morphology
abstract
Morphological operations are used for segmentation, feature generation and location extraction. A recursive adaptive thresholding algorithm transforms a gray-level image into a set of multiple level regions of objects. A distance transformation algorithm then is used to transform a binary image into the minimum distance from each object point to the object's boundary. This algorithm uses a morphological erosion with a large structuring element which may correspond to Euclidean, city-block, or chessboard distance measures. A shape library database with hierarchical features is automatically generated. The features extracted are the shape number and the skeletal local-maximum points radii and coordinates. Object recognition is achieved by comparing the shape number and the hierarchical radii. Object location is detected by a hierarchical morphological bandpass filter.>
Frank Y. Shih, Owen Robert Mitchell
CVPR1
1988 Industrial parts recognition and inspection by image morphology
abstract
Application algorithms for industrial parts and tool recognition and inspection by image morphology techniques are discussed. A recursive adaptive thresholding algorithm transforms a gray-level image into a set of multiple-level regions of objects. This algorithm uses a morphological erosion with a large symmetrical concave structuring element. A distance-transformation algorithm transforms these binary image regions into the minimum distance from each object point to the boundary of the object. This algorithm also uses a morphological erosion. From the distance transform, it is possible to compute a shape number and extract the skeleton, which is useful for generic pattern recognition and feature extraction. Corner angles and the radii of circular holes can be located, identified, and estimated by using morphological openings and erosions. The algorithms allow robust tool and part recognition and inspection.>
Frank Y. Shih, Owen Robert Mitchell
ICRA1