VLDB 2026 Research / reviewers in the wild / expert
Wei-Yen Hsu
dblp:79/8790
· DBLP profile ↗
27ranked-venue papers
27as first author
17since 2021 · last 2026
0000-0002-4599-0744ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 15 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 first-author · 7 since 2021Computer networks · 5 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving Driver Emotion Recognition Using Structure-Detail-Aware Blind Super-ResolutionabstractIn intelligent driving environments, accurate driver emotion recognition is critical for enhancing human-machine interaction and ensuring road safety. However, due to the constraints of real-world driving conditions, images captured by in-vehicle cameras frequently suffer from blurring, noise, low illumination, and reduced resolution. These visual degradations significantly impair the performance of emotion recognition systems. To address these challenges, we propose a novel structure preservation and detail recovery blind super-resolution (SpDrBSR) framework specifically designed for driver monitoring in human-machine systems. We design the novel cross-domain structure-detail-aware transformer (CSDFormer), which effectively integrates the intradomain and cross-domain information to achieve global and local structure preservation and detail recovery. Specifically, we first apply Canny edge detection to ensure precise edge localization and structural preservation. Building on this, the proposed CSDFormer leverages long-range intradomain and cross-domain dependencies through two key modules: the cross-domain fusion module, which integrates domain features to maintain structure and detail, and the global mixed cross-domain attention module, which retrieves global information, refines similarity scores, and filters irrelevant signals. A cross-attention mechanism further enhances feature interaction and complementarity, enabling accurate reconstruction of complex image structures. We conduct experiments on real-world datasets to verify the effectiveness of the proposed method. The results indicate that the proposed SpDrBSR outperforms the state-of-the-art approaches in terms of quantitative metrics and achieves better structure preservation and detail recovery than those approaches in terms of visual quality. Furthermore, the results also indicate that the proposed SpDrBSR can significantly enhance driver emotion recognition accuracy by nearly 8%. Wei-Yen Hsu, Yi-Chun Chiu, Hsin-Yun Chang |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2026 | SITFFormer: A Blind Super-Resolution Framework Preserving Structural Integrity and Texture FidelityabstractBlind super-resolution (SR) aims to reconstruct high-resolution images from low-quality inputs under unknown degradation conditions. While numerous blind SR methods have been proposed in recent years, they still face critical limitations. Most approaches perform well under specific degradation patterns but struggle with complex scenarios involving multiple degradation factors and varying noise levels. This often leads to loss of structural integrity and fine details, resulting in suboptimal restoration quality. Furthermore, existing methods typically rely on convolutional neural networks (CNNs) with limited receptive fields, which hinders effective cross-domain information integration. Their inability to capture long-range dependencies compromises the reconstruction of global structures. In edge detection, conventional techniques frequently produce inaccurate or false edges, further degrading the quality of cross-domain integration and image restoration. To address these challenges, we propose Structural Integrity and Texture Fidelity Transformer (SITFFormer), a novel transformer-based framework for blind single-image SR. Our approach incorporates the Canny edge detection algorithm to accurately preserve true edges and suppress noise-induced artifacts, enhancing edge localization in complex and noisy environments. We also introduce the Cross-Domain Structure-Texture-Aware Network (CDSTNet), designed to integrate intra-domain and cross-domain features for comprehensive structure preservation and texture recovery. CDSTNet comprises two key modules: Cross-Domain Integration (CDI) that fuses intra- and cross-domain features to retain structural and textural details. Cross-Domain Learnable Attention (CDLA) that explores global dependencies, adaptively refines feature similarity, and filters out redundant non-local information. Both modules are equipped with a Cross-Attention Mechanism (CAM) to facilitate effective interaction and complementarity between domains, enhancing reconstruction fidelity. Extensive experiments on synthetic, noisy, and real-world datasets demonstrate that SITFFormer surpasses state-of-the-art methods in quantitative performance and visual quality, particularly in preserving structural integrity and recovering fine textures. Wei-Yen Hsu, Hsin-Yun Chang, Chung-Yu Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | Wavelet structure-texture-aware super-resolution for pedestrian detection
Wei-Yen Hsu, Chun-Hsiang Wu |
Inf. Sci. | 1 |
| 2025 | Cross-Attribute Feature-Perceptive Driver Emotion Recognition in Real ScenariosabstractDriver emotion recognition has garnered significant attention due to its applications in driver emotion monitoring and regulation systems. By analyzing the driver’s facial expressions, it can effectively detect their emotional state, thereby enhancing driving safety. However, variations in lighting conditions (such as direct sunlight or shadowing) and facial occlusions (such as hair or sunglasses) still pose challenges to the system’s accuracy. To address these challenges, we propose a novel cross-attribute feature-perceptive network (CaFpNet) to enhance the driver emotion recognition performance in real-world environments by effectively utilizing multi-attribute facial features from global, local, and salient subregions and thus fully exploiting the diverse potential information provided by each facial attribute, in line with the human face perception mechanism that extracts both global and regional information. Specifically, the global-attribute feature (GaF) module focuses on the more important facial features in the overall face by expanding the number of channels to preserve features and assigning weights to different channels. Moreover, the local-attribute feature (LaF) and salient-attribute feature (SaF) modules capture regional feature information from local and salient-attribute features, respectively, focusing on fine-grained regional features and reducing the interference from irrelevant regions in feature extraction. The experimental results indicate that CaFpNet exhibits superior performance compared to various state-of-the-art approaches on two driving-scenario multi-domain emotion datasets—MLI-DER and KMU-FED. Its outstanding performance in driver emotion recognition demonstrates the potential and effectiveness in real-world applications. Wei-Yen Hsu, Ting-Hsuan Chiang |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2025 | Progressive Structure Preservation and Detail Refinement for Remote Sensing Single-Image Super-ResolutionabstractRecent advances in deep-learning-based remote sensing image super-resolution (RSISR) have garnered significant attention. Conventional models typically perform upsampling at the end of the architecture, which reduces computational effort but leads to information loss and limits image quality. Moreover, the structural complexity and texture diversity of remote sensing images pose challenges in detail preservation. While transformer-based approaches improve global feature capture, they often introduce redundancy and overlook local details. To address these issues, we propose a novel progressive structure preservation and detail refinement super-resolution (PSPDR-SR) model, designed to enhance both structural integrity and fine details in RSISR. The model comprises two primary subnetworks: the structure-aware super-resolution (SaSR) subnetwork and the detail recovery and refinement (DR&R) subnetwork. To efficiently leverage multilayer and multiscale feature representations, we introduce coarse-to-fine dynamic information transmission (C2FDIT) and fine-to-coarse dynamic information transmission (F2CDIT) modules, which facilitate the extraction of richer details from low-resolution (LR) remote sensing images. These modules integrate transformers and convolutional long short-term memory (ConvLSTM) blocks to form dynamic information transmission modules (DITMs), enabling effective bidirectional feature transmission both horizontally and vertically. This method ensures comprehensive feature fusion, mitigates redundant information, and preserves essential extracted features within the deep network. Experimental results demonstrate that PSPDR-SR outperforms the state-of-the-art approaches on two benchmark datasets in both quantitative and qualitative evaluations, excelling in structure preservation and detail enhancement across various metrics, including SSIM, MS_SSIM, learned perceptual image patch similarity (LPIPS), deep image structure and texture similarity (DISTS), spatial correlation coefficient (SCC), and spectral angle mapper (SAM). Wei-Yen Hsu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Multi-Attribute Feature-Aware Network for Facial Expression RecognitionabstractFacial expression recognition (FER) has gained popularity as a research topic due to its broad applicability. However, real-world environments present significant challenges to FER, including occlusion, illumination variation, and angle. To address these issues, we propose a novel multi-attribute feature-aware network (MAFaNet) to enhance the performance and accuracy of FER in real-world environments. The proposed MAFaNet enhances the FER performance in real-world environments by effectively utilizing multi-attribute facial features from global, local, and salient subregions and thus fully exploiting the diverse potential information provided by each facial attribute, in line with the human face perception mechanism that extracts both global and regional information. Specifically, the global facial feature (GFF) module focuses on the more important facial features in the overall face by expanding the number of channels to preserve features and assigning weights to different channels. Moreover, the local facial feature (LFF) and salient facial feature (SFF) modules capture regional feature information from local and salient facial features, respectively, focusing on fine-grained regional features and reducing the interference from irrelevant regions in feature extraction. The experimental results indicate that the proposed MAFaNet method achieves the promising FER performance in comparison with the state-of-the-art approaches on several real-world datasets. Wei-Yen Hsu, Yu-Chieh Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Wavelet Pyramid Recurrent Structure-Preserving Attention Network for Single Image Super-ResolutionabstractMany single image super-resolution (SISR) methods that use convolutional neural networks (CNNs) learn the relationship between low- and high-resolution images directly, without considering the context structure and detail fidelity. This can limit the potential of CNNs and result in unrealistic, distorted edges and textures in the reconstructed images. A more effective approach is to incorporate prior knowledge about the image into the model to aid in image reconstruction. In this study, we propose a novel recurrent structure-preserving mechanism that innovatively uses the multiscale wavelet transform (WT) as an image prior, namely, wavelet pyramid recurrent structure-preserving attention network (WRSANet), to process both low- and high-frequency subnetworks at each level separately and recursively. We propose a novel structure scale preservation (SSP) architecture that differs from traditional WTs. This architecture allows us to incorporate and learn structure preservation subnetworks at each level. By using our proposed structure scale fusion (SSF) combined with inverse WT, we can recursively restore and preserve rich low-frequency image structure through the combination of SSP at various levels. Furthermore, we also propose novel low-to-high-frequency information transmission (L2HIT) and detail enhancement (DE) mechanisms to address the issue of detail distortion in high-frequency images by transferring information from structure preservation subnetworks. This allows us to preserve the low-frequency structure while reconstructing high-frequency details, improving detail fidelity and avoiding structural distortion. Finally, a joint loss function is also used to balance the fusion of low- and high-frequency information at different degrees, with hyperparameters being adjusted during training. The experimental results demonstrate that the proposed WRSANet achieves better performance and visual presentation than the state-of-the-art (SOTA) on synthetic and real datasets, especially in terms of context structure and texture details. Wei-Yen Hsu, Pei-Wen Jian |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Multi-Scale and Multi-Layer Lattice Transformer for Underwater Image EnhancementabstractUnderwater images are often subject to color deviation and a loss of detail due to the absorption and scattering of light. The challenge of enhancing underwater images is compounded by variations in wavelength and distance attenuation, as well as color deviation that exist across different scales and layers, resulting in different degrees of color deviation, attenuation, and blurring. To address these issues, we propose a novel multi-scale and multi-layer lattice transformer (MMLattFormer) to effectively eliminate artifacts and color deviation, prevent over-enhancement, and preserve details across various scales and layers, thereby achieving more accurate and natural results in underwater image enhancement. The proposed MMLattFormer model integrates the advantage of LattFormer to enhance global perception with the advantage of “multi-scale and multi-layer” configuration to leverages the differences and complementarities between features of various scales and layers to boost local perception. The proposed MMLattFormer model is comprised of multi-scale and multi-layer LattFormers. Each LattFormer primarily encompasses two modules: Multi-head Transposed-attention Residual Network (MTRN) and Gated-attention Residual Network (GRN). The MTRN module enables cross-pixel interaction and pixel-level aggregation in an efficient manner to extract more significant and distinguishable features, whereas the GRN module can effectively suppress under-informed or redundant features and retain only useful information, enabling excellent image restoration exploiting the local and global structures of the images. Moreover, to better capture local details, we introduce depthwise convolution in these two modules before generating global attention maps and decomposing images into different features to better capture the local context in image features. The qualitative and quantitative results indicate that the proposed method outperforms state-of-the-art approaches in delivering more natural results. This is evident in its superior detail preservation, effective prevention of over-enhancement, and successful removal of artifacts and color deviation on several public datasets. Wei-Yen Hsu, Yu-Yu Hsu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Context-detail-aware United Network for Single Image DerainingabstractImages captured outdoors are often affected by rainy days, resulting in a severe deterioration in the visual quality of the captured images and a decrease in the performance of related applications. Therefore, single image deraining has attracted attention as a challenging research topic. Nowadays, there are two common deraining architectures in single image deraining. The first one is to restore the rain-free image by deducting rain streaks learned by the model from the rain image, but the background structure is easily mistaken for rain streaks and subtracted. The other one is to directly learn the clean background structure through the model using rain images, but it is difficult to completely remove the rain streaks due to the complexity of the information in images with rain. Therefore, current methods cannot balance rain streak removal and rain-free image background restoration in a single architecture and achieve good results. To address this issue, we propose a novel framework, namely, Context-Detail-Aware United Network (CDaUNet), which combines the above two architectures in this study. More specifically, we divide the restoration of the background structure of rain-free images and the learning of rain streaks into two independent sub-networks. The proposed Structure-Aware Rain Removal Network (SaRRN) is to learn the background structure in images to reconstruct clean rain-free images, whereas Detail-Aware Rain Streak Learning Network (DaRLN) is proposed to learn the details of rain streaks in images. Finally, we fuse the results generated by the two sub-networks through our designed Dual Architecture Fusion Network (DAFN) to reconstruct original rain images to effectively fuse the results of the two sub-networks. The experimental results show that CDaUNet achieves satisfactory performance in comparison with the state-of-the-art approaches included in rain streak removal and rain-free image structure restoration architectures on both synthetic and real image datasets, confirming the effectiveness of our method. Wei-Yen Hsu, Hsien-Wen Lin |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | Wavelet Approximation-Aware Residual Network for Single Image DerainingabstractIt has been made great progress on single image deraining based on deep convolutional neural networks (CNNs). In most existing deep deraining methods, CNNs aim to learn a direct mapping from rainy images to clean rain-less images, and their architectures are becoming more and more complex. However, due to the limitation of mixing rain with object edges and background, it is difficult to separate rain and object/background, and the edge details of the image cannot be effectively recovered in the reconstruction process. To address this problem, we propose a novel wavelet approximation-aware residual network (WAAR), wherein rain is effectively removed from both low-frequency structures and high-frequency details at each level separately, especially in low-frequency sub-images at each level. After wavelet transform, we propose novel approximation aware (AAM) and approximation level blending (ALB) mechanisms to further aid the low-frequency networks at each level recover the structure and texture of low-frequency sub-images recursively, while the high frequency network can effectively eliminate rain streaks through block connection and achieve different degrees of edge detail enhancement by adjusting hyperparameters. In addition, we also introduce block connection to enrich the high-frequency details in the high-frequency network, which is favorable for obtaining potential interdependencies between high- and low-frequency features. Experimental results indicate that the proposed WAAR exhibits strong performance in reconstructing clean and rain-free images, recovering real and undistorted texture structures, and enhancing image edges in comparison with the state-of-the-art approaches on synthetic and real image datasets. It shows the effectiveness of our method, especially on image edges and texture details. Wei-Yen Hsu, Wei-Chi Chang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Recurrent wavelet structure-preserving residual network for single image deraining
Wei-Yen Hsu, Wei-Chi Chang |
Pattern Recognit. | 1 |
| 2023 | Wavelet detail perception network for single image super-resolution
Wei-Yen Hsu, Pei-Wen Jian |
Pattern Recognit. Lett. | 1 |
| 2023 | Pedestrian Detection Using Multi-Scale Structure-Enhanced Super-ResolutionabstractPedestrian detection remains a crucial technology for applications such as autonomous driving and gait recognition and continues to be a prominent research area. Despite the development of advanced pedestrian detection techniques, the challenge of detecting pedestrians in low-resolution images persists in real-life scenarios where low-quality imaging devices are still in use. The objective of this study is to enhance the detection of pedestrians in low-resolution (LR) images by improving image quality through super-resolution techniques. To achieve this goal, we propose an end-to-end Multi-scale Structure-Enhanced Super-Resolution (MsSE-SR) method to enlarge LR images into high-resolution (SR) images and utilize Yolov4 for detection, effectively addressing the issue of low detection performance in LR images. To generate an SR image that can accurately distinguish between foreground and background elements while emphasizing pedestrian characteristics, we employ the stationary wavelet transform (SWT) to decompose the image into low and high-frequency sub-images. These sub-images are then processed through different network structures, enabling the network to reconstruct high-frequency details and low-frequency structures with greater precision. Moreover, we propose a high-to-low subnetwork information transfer (H2LSnIT) that incorporates high-frequency edge information into the low-frequency image structure during the reconstruction process, leading to a more detailed reconstruction of the low-frequency structure. We also propose a novel loss function that leverages the characteristics of wavelet decomposition to enhance the network’s focus on reconstructing the image structure, further improving detection performance. The experimental results demonstrate the effectiveness of the proposed MsSE-SR method in significantly enhancing pedestrian detection performance. Wei-Yen Hsu, Pei-Yu Yang |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Recurrent Multi-scale Approximation-Guided Network for Single Image Super-ResolutionabstractSingle-image super-resolution (SISR)is an essential topic in computer vision applications. However, most CNN-based SISR approaches directly learn the relationship between low- and high-resolution images while ignoring the contextual texture and detail fidelity to explore super-resolution; thus, they hinder the representational power of CNNs and lead to the unrealistic, distorted reconstruction of edges and textures in the images. In this study, we propose a novel recurrent structure preservation mechanism with the integration and innovative use of multi-scale wavelet transform,Recurrent Multiscale Approximation-guided Network (RMANet), to recursively process the low-frequency and high-frequency sub-networks at each level separately. Unlike traditional wavelet transform, we propose a novelApproximation Level Preservation (ALP)architecture to import and learn the low-frequency sub-networks at each level. Through proposedApproximation level fusion (ALF)and inverse wavelet transform, rich image structures of low frequency at each level can be recursively restored and greatly preserved with the combination of ALP at each level. In addition, a novel low-frequency to high-frequencydetail enhancement (DE)mechanism is also proposed to solve the problem of detail distortion in high-frequency networks by transmitting low-frequency information to the high-frequency network. Finally, a joint loss function is used to balance low-frequency and high-frequency information with different degrees of fusion. In addition to correct restoration, image details are further enhanced by tuning different hyperparameters during training. Compared with the state-of-the-art approaches, the experimental results on synthetic and real datasets demonstrate that the proposed RMANet achieves better performance in visual presentation, especially in image edges and texture details. Wei-Yen Hsu, Pei-Wen Jian |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2021 | A novel eye center localization method for multiview faces
Wei-Yen Hsu, Chi-Jui Chung |
Pattern Recognit. | 1 |
| 2021 | A Novel Eye Center Localization Method for Head Poses With Large RotationsabstractEye localization is undoubtedly crucial to acquiring large amounts of information. It not only helps people improve their understanding of others but is also a technology that enables machines to better understand humans. Although studies have reported satisfactory accuracy for frontal faces or head poses at limited angles, large head rotations generate numerous defects (e.g., disappearance of the eye), and existing methods are not effective enough to accurately localize eye centers. Therefore, this study makes three contributions to address these limitations. First, we propose a novel complete representation (CR) pipeline that can flexibly learn and generate two complete representations, namely the CR-center and CR-region, of the same identity. We also propose two novel eye center localization methods. This first method employs geometric transformation to estimate the rotational difference between two faces and an unknown-localization strategy for accurate transformation of the CR-center. The second method is based on image translation learning and uses the CR-region to train the generative adversarial network, which can then accurately generate and localize eye centers. Five image databases are employed to verify the proposed methods, and tests reveal that compared with existing methods, the proposed method can more accurately and robustly localize eye centers in challenging images, such as those showing considerable head rotation (both yaw rotation of -67.5° to +67.5° and roll rotation of +120° to -120°), complete occlusion of both eyes, poor illumination in addition to head rotation, head pose changes in the dark, and various gaze interaction. Wei-Yen Hsu, Chi-Jui Chung |
IEEE Trans. Image Process. | 1 |
| 2021 | Ratio-and-Scale-Aware YOLO for Pedestrian DetectionabstractCurrent deep learning methods seldom consider the effects of small pedestrian ratios and considerable differences in the aspect ratio of input images, which results in low pedestrian detection performance. This study proposes the ratio-and-scale-aware YOLO (RSA-YOLO) method to solve the aforementioned problems. The following procedure is adopted in this method. First, ratio-aware mechanisms are introduced to dynamically adjust the input layer length and width hyperparameters of YOLOv3, thereby solving the problem of considerable differences in the aspect ratio. Second, intelligent splits are used to automatically and appropriately divide the original images into two local images. Ratio-aware YOLO (RA-YOLO) is iteratively performed on the two local images. Because the original and local images produce low- and high-resolution pedestrian detection information after RA-YOLO, respectively, this study proposes new scale-aware mechanisms in which multiresolution fusion is used to solve the problem of misdetection of remarkably small pedestrians in images. The experimental results indicate that the proposed method produces favorable results for images with extremely small objects and those with considerable differences in the aspect ratio. Compared with the original YOLOs (i.e., YOLOv2 and YOLOv3) and several state-of-the-art approaches, the proposed method demonstrated a superior performance for the VOC 2012 comp4, INRIA, and ETH databases in terms of the average precision, intersection over union, and lowest log-average miss rate. Wei-Yen Hsu, Wen-Yen Lin |
IEEE Trans. Image Process. | 1 |
| 2019 | Cardiac Left Ventricular Ultrasound Image Sequence Recognition and TrackingabstractA novel approach combining faster region-based convolutional neural network (Faster R-CNN) with active shape model is proposed to automatically recognize and track cardiac ultrasound left ventricular image sequences in this work. Because the shape and appearance of the left ventricle vary considerably between adjacent images, conventional methods cannot effectively identify its position and then accurately track them. We propose a novel method, incorporating Faster R-CNN and active shape model, to automatically recognize, segment and track the cardiac ultrasound left ventricular image sequences. Compared with four state-of-the-art approaches, the proposed method has most of better performance in Sørensen-Dice coefficient, mean absolute deviation, and Hausdorff distance. Wei-Yen Hsu |
BIBE | 1 |
| 2018 | A decision-making mechanism for assessing risk factor significance in cardiovascular diseases
Wei-Yen Hsu |
Decis. Support Syst. | 1 |
| 2015 | Assembling A Multi-Feature EEG Classifier for Left-Right Motor Imagery Data Using Wavelet-Based Fuzzy Approximate Entropy for Improved AccuracyabstractAn EEG classifier is proposed for application in the analysis of motor imagery (MI) EEG data from a brain-computer interface (BCI) competition in this study. Applying subject-action-related brainwave data acquired from the sensorimotor cortices, the system primarily consists of artifact and background removal, feature extraction, feature selection and classification. In addition to background noise, the electrooculographic (EOG) artifacts are also automatically removed to further improve the analysis of EEG signals. Several potential features, including amplitude modulation, spectral power and asymmetry ratio, adaptive autoregressive model, and wavelet fuzzy approximate entropy (wfApEn) that can measure and quantify the complexity or irregularity of EEG signals, are then extracted for subsequent classification. Finally, the significant sub-features are selected from feature combination by quantum-behaved particle swarm optimization and then classified by support vector machine (SVM). Compared with feature extraction without wfApEn on MI data from two data sets for nine subjects, the results indicate that the proposed system including wfApEn obtains better performance in average classification accuracy of 88.2% and average number of commands per minute of 12.1, which is promising in the BCI work applications. Wei-Yen Hsu |
Int. J. Neural Syst. | 1 |
| 2013 | Single-Trial Motor Imagery Classification using Asymmetry Ratio, phase Relation, Wavelet-Based Fractal, and their Selected CombinationabstractAn electroencephalogram (EEG) analysis system is proposed for single-trial classification of motor imagery (MI) data in this study. Applying event-related brain potential (ERP) data acquired from the sensorimotor cortices, the system mainly consists of enhanced active segment selection, feature extraction, feature selection and classification. In addition to the original use of continuous wavelet transform (CWT) and Student's two-sample t-statistics, the 2D anisotropic Gaussian filter is proposed to further refine the selection of active segments. We then extract several features, including spectral power and asymmetry ratio, coherence and phase-locking value, and multiresolution fractal feature vector, for subsequent classification. Next, genetic algorithm (GA) is used to select features from the combination of above-mentioned features. Finally, support vector machine (SVM) is used for classification. Compared with "without enhanced active segment selection," several potential features and linear discriminant analysis (LDA) on MI data from two data sets for 10 subjects, the results indicate that the proposed method achieves 86.7% average classification accuracy, which is promising in BCI applications. Wei-Yen Hsu |
Int. J. Neural Syst. | 1 |
| 2013 | Application of Quantum-behaved Particle Swarm Optimization to Motor imagery EEG ClassificationabstractIn this study, we propose a recognition system for single-trial analysis of motor imagery (MI) electroencephalogram (EEG) data. Applying event-related brain potential (ERP) data acquired from the sensorimotor cortices, the system chiefly consists of automatic artifact elimination, feature extraction, feature selection and classification. In addition to the use of independent component analysis, a similarity measure is proposed to further remove the electrooculographic (EOG) artifacts automatically. Several potential features, such as wavelet-fractal features, are then extracted for subsequent classification. Next, quantum-behaved particle swarm optimization (QPSO) is used to select features from the feature combination. Finally, selected sub-features are classified by support vector machine (SVM). Compared with without artifact elimination, feature selection using a genetic algorithm (GA) and feature classification with Fisher's linear discriminant (FLD) on MI data from two data sets for eight subjects, the results indicate that the proposed method is promising in brain-computer interface (BCI) applications. Wei-Yen Hsu |
Int. J. Neural Syst. | 1 |
| 2012 | Improved watershed transform for tumor segmentation: Application to mammogram image compression
Wei-Yen Hsu |
Expert Syst. Appl. | 1 |
| 2012 | Fuzzy Hopfield neural network clustering for single-trial motor imagery EEG classification
Wei-Yen Hsu |
Expert Syst. Appl. | 1 |
| 2012 | Wavelet-based envelope features with automatic EOG artifact removal: Application to single-trial EEG data
Wei-Yen Hsu, Chao-Hung Lin, Hsien-Jen Hsu, Po-Hsun Chen, I-Ru Chen |
Expert Syst. Appl. | 1 |
| 2012 | Application of Competitive Hopfield Neural Network to Brain-Computer Interface SystemsabstractWe propose an unsupervised recognition system for single-trial classification of motor imagery (MI) electroencephalogram (EEG) data in this study. Competitive Hopfield neural network (CHNN) clustering is used for the discrimination of left and right MI EEG data posterior to selecting active segment and extracting fractal features in multi-scale. First, we use continuous wavelet transform (CWT) and Student's two-sample t-statistics to select the active segment in the time-frequency domain. The multiresolution fractal features are then extracted from wavelet data by means of modified fractal dimension. At last, CHNN clustering is adopted to recognize extracted features. Due to the characteristic of non-supervision, it is proper for CHNN to classify non-stationary EEG signals. The results indicate that CHNN achieves 81.9% in average classification accuracy in comparison with self-organizing map (SOM) and several popular supervised classifiers on six subjects from two data sets. Wei-Yen Hsu |
Int. J. Neural Syst. | 1 |
| 2011 | Continuous EEG Signal Analysis for Asynchronous BCI ApplicationabstractIn this study, we propose a two-stage recognition system for continuous analysis of electroencephalogram (EEG) signals. An independent component analysis (ICA) and correlation coefficient are used to automatically eliminate the electrooculography (EOG) artifacts. Based on the continuous wavelet transform (CWT) and Student's two-sample t-statistics, active segment selection then detects the location of active segment in the time-frequency domain. Next, multiresolution fractal feature vectors (MFFVs) are extracted with the proposed modified fractal dimension from wavelet data. Finally, the support vector machine (SVM) is adopted for the robust classification of MFFVs. The EEG signals are continuously analyzed in 1-s segments, and every 0.5 second moves forward to simulate asynchronous BCI works in the two-stage recognition architecture. The segment is first recognized as lifted or not in the first stage, and then is classified as left or right finger lifting at stage two if the segment is recognized as lifting in the first stage. Several statistical analyses are used to evaluate the performance of the proposed system. The results indicate that it is a promising system in the applications of asynchronous BCI work. Wei-Yen Hsu |
Int. J. Neural Syst. | 1 |