VLDB 2026 Research / reviewers in the wild / expert
Bo-Wei Chen
dblp:46/3809
· DBLP profile ↗
66ranked-venue papers
22as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 7 first-authorGraphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 1 since 2021Computer networks · 12 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Nonconvex Leaky Minimax Concave Embedding via Majorization-Minimization for Robust Visual IoT Sensing With Missing ObservationsabstractVisual Internet-of-Things (VIoT) systems provide pervasive sensing capabilities for monitoring ambient dynamics through the capture, processing, and analysis of visual data. VIoT systems enable edge nodes to perform preliminary screening and tagging data before transmission to clouds. This reduces bandwidth and computation requirements. A promising approach, known as perceptual VIoT sensing, focuses on generating discriminative data signatures with embedded class labels directly from captured data to facilitate intelligent recognition. Nonetheless, partially observed data, such as those affected by noise, occlusion, or missing values, can compromise system performance by introducing biases and reducing robustness. Addressing these challenges is crucial for ensuring reliable VIoT operations. Motivated by these challenges, this study proposes robust perceptual sensing based on nonconvex data embedding with leaky minimax concave (LMC) loss, designed to effectively handle partially observed and corrupted data. For brevity, this approach is referred to as nonconvex LMC embedding. Through the use of perceptual signatures, the VIoT can efficiently recognize the collected data by checking their signatures, even if partially observed data are present. To mitigate biases in both contrastive learning and data representation stages for the signature generation model, nonconvex loss functions are devised to suppress the adverse effects of modeling errors. As solving nonconvex loss directly is challenging, surrogate functions are derived. Nevertheless, the surrogate formulations still retain strict discrete constraints owing to the data embedding process. To overcome this limitation, the optimization problem is reformulated using low-rank semidefinite relaxation, which relaxes strict discrete constraints. The relaxed problem is then efficiently solved through majorization–minimization algorithms, which iteratively refine the surrogate solutions toward optimality. Experimental results on open datasets showed that the proposed method outperformed the baselines, with F1 scores improved by 5.28%, 12.59%, and 7.64%, under the conditions of impulse noise, continuous occlusion, and missing values, respectively. Such findings validated effectiveness and robustness of our method. Bo-Wei Chen, Yun-Fen Ma |
IEEE Internet Things J. | 1 |
| 2025 | Robust Partially Observed Data Sensing via ℓ₂,ₚ Norms With Flexible Adaptive Label Marginal Space for Visual IoTabstractVisual Internet of Things (VIoT) empowers intelligent sensing by equipping terminal devices with the capability to preliminarily screen and tag sensing data for further processing. However, sensing environments are often imperfect, and partially observed data may be collected at the terminals. This causes several challenges. First, the model fitting process may become oversensitive owing to the presence of varying corrupted data, e.g., continuous occlusion, therefore compromising robustness because of fitting biases. Second, model fitting relies on label information, which consists of fixed and equally spaced categorical variables. Such information cannot adequately reflect the underlying distribution of collected data, as it assumes that each categorical variable is uniformly distributed. This deepens the difficulty of data fitting. To solve the aforementioned problems, this study proposes robust data sensing based on$\ell _{2,p}$norms, where data collected by VIoT terminals can be converted into resilient perceptual data signatures while label information is embedded inside. To deal with the problem stemming from rigid label marginal space, this study introduces an adaptive slack variable to the proposed model. Such a slack variable can automatically adapt itself to sensing data during model fitting while adjusting label marginal space by providing flexible labels. Moreover, a new adaptive regulating mechanism is developed to control the slack variables, such that they can consider multiple coeffects from different loss terms and penalties during optimization, creating error-tolerant soft margins. This is conducive to model fitting, especially for partially observed data. In addition to label slack variables, this study also derives slack variables for$\ell _{2,p}$-norm loss that is used to capture the nuances of the data, thereby providing flexible Hamming marginal space for resilient signature generation. Experiments on the open datasets show that the proposed method yields better F1 scores than the baselines, improved by$\mathbf {4.54}\boldsymbol {\%}$,$\mathbf {8.08}\boldsymbol {\%}$, and$\mathbf {6.02}\boldsymbol {\%}$, with respect to various data corruption, including impulse noise, continuous occlusion, and missing values. Such results verify the effectiveness of the proposed method. Bo-Wei Chen |
IEEE Internet Things J. | 1 |
| 2025 | Robust Partially-Observed VIoT Data Sensing via Half Quadratic Loss with Flexible Weighted Groupwise Relaxed Label MarginsabstractOne distinctive feature of the Visual Internet of Things (VIoT) is its ability to prescreen and tag sensing data for efficient distribution to dedicated edge nodes. However, when partially observed data are captured, existing fitting models could become oversensitive. Besides, confining corrupted data to rigid and fixed categorical variable space may lead to biased outcomes, as equally spaced categorical variables could inadequately reflect the underlying distribution of corrupted data. Although state-of-the-art methods devised relaxed fixed categorical variables for enhancing performance, their relaxation strategies failed to consider same-class information. This led to inconsistent adjustments among samples within the same class. Consequently, the adjustments of the margin for the same class could potentially conflict. Besides, those methods were all based on convex$\ell _{2}$-norm loss functions, which limited the robustness against outliers. In contrast, existing robust approaches primarily focused on data modeling through nonconvex loss formulations but did not incorporate relaxed label margins, subsequently leaving the issue of flexible label margins unaddressed. Furthermore, to conquer the above problems, this study proposes robust data sensing based on nonconvex half quadratic loss for tackling oversensitivity stemming from partially observed data. To lift the constraint caused by rigid label variables and to increase regularization capabilities for label marginal space, flexible groupwise relaxation based on half quadratic loss is developed. It provides scaling and translation controls over marginal space. Additionally, class information is incorporated to reshape the marginal space during marginal formation by introducing groupwise dragging variables. The optimization procedure for groupwise dragging variables is specifically designed to account for the nonconvex nature of the half quadratic loss. Moreover, this study also derives flexible groupwise relaxation on Hamming space to accommodate the representation problem of partially observed input data. Experiments on VIoT data showed that the proposed method enhanced F1 scores by at least$\mathbf {6.67}\boldsymbol {\%}$,$\mathbf {8.45}\boldsymbol {\%}$, and$\mathbf {5.51}\boldsymbol {\%}$with respect to various forms of data corruption, including impulse noise, continuous occlusion, and missing values. This has verified the effectiveness of the proposed method. Bo-Wei Chen |
IEEE Internet Things J. | 1 |
| 2024 | Partially Observed Visual IoT Data Reconstruction Based on Robust Half-Quadratic Collaborative FilteringabstractRecent advances and applications in the Visual Internet of Things (VIoT) have witnessed significant progress in various areas, such as smart sensing and environmental monitoring, which rely on sensors to capture data to build analytic models. However, when VIoT devices are deployed in harsh environments, such as the deep ocean and underground, the physical integrity and functionality of these devices may deteriorate. As a result, there can be disruption in the data acquisition process, leading to incomplete data generated in the sink node owing to interference, occlusion, connectivity, or unknown abnormal events. Poor data quality affects the validity and generalizability of the analysis. Although numerous data recovery algorithms have been developed,$\ell _{2}$,$\ell _{1}$, and$\ell _{2,1}$norms were unable to effectively handle oversensitivity during data modeling. Despite the fact that recent half quadratic (HQ) functions have shown improvement in robustness, the models still degraded and produced fitting biases owing to nonflexible loss functions and rigid reweighting when sufficient large fitting errors occurred. In view of such, this study proposes collaborative filtering based on scaled lifted augmented HQ functions to accommodate the above-mentioned problems. First, additional auxiliary variables are introduced to flexibly absorb varying levels of data corruption through derived lifting schemes. Moreover, this study devises the scaled reweighting forms for the proposed lifted augmented HQ functions to further increase resilience against oversensitivity. Experiments on open data sets showed that the proposed method yielded better recoverability than the baselines under different data corruption conditions. Such findings verified the effectiveness of the proposed model. Bo-Wei Chen, Ying-Hsuan Wu |
IEEE Internet Things J. | 1 |
| 2024 | Robust Perceptual Data Sensing Based on Majorization-Minimized Low-Rank Semidefinite Relaxation for Visual IoTabstractFor the Visual Internet of Things (VIoT), terminal devices need to preliminarily screen and tag data for enabling efficient indexing and allocation of sensing data to dedicated edge nodes or cloud data centers for further processing. VIoT terminal devices often face a challenge, where collected data are contaminated by noise or contain partial observed information, which can adversely affect the accuracy and reliability of tagged results. In view of such, this study proposes resilient and perceptual data signature generation for robustly tagging captured data based on low-rank semidefinite relaxation. Such signatures can represent data in a compact way while embedding label information inside at the same time. To increase the robustness and effectiveness of data tagging in the proposed method, nonconvex loss called leaky-minimax concave penalty function is studied. This loss function effectively tackles the challenges posed by partially observed data and instance-pairwise label-space mapping, subsequently improving the reliability and accuracy of the tagging process in the terminal devices. To solve the challenges associated with nonconvex loss functions, this study employs the majorization-minimization technique, which helps conquer the nonconvex optimization problem efficiently. Additionally, low-rank semidefinite relaxation is applied to the formation of data signatures to avoid discrete variable optimization problems caused by Laplacian graph embedding, where data locality learning is performed. The relaxation helps simplify the optimization process and improve the applicability of the method. Experimental evaluations conducted on open datasets have demonstrated the superior performance of the proposed method compared with the baselines, particularly under different data corruption conditions. The results showed an improvement of 19.45%, 8.97%, and 5.45% in F1 scores, with respect to impulse noise, continuous occlusion, and missing values. This validated the effectiveness and practicality of the proposed model in handling VIoT data challenges Bo-Wei Chen, Tzu-Hsuan Wang |
IEEE Internet Things J. | 1 |
| 2024 | Robust Data-Driven Automation Based on Relaxed Supervised Hashing With Self-Optimized LabelsabstractRecently, the Visual Internet of Things (VIoT) has been widely used in data-driven automation, where VIoT devices are used to monitor environmental dynamics and to trigger corresponding actuators after examining event signatures (e.g., hashes). Nonetheless, VIoT devices may collect partially observed data, e.g., largely occluded images, which cause biased labeling and oversensitivity during modeling. Moreover, in typical methods, class labels are rigid and fixed categorical variables, where marginal space between different classes is fixed and equal. In fact, margins may be unequally spaced. Such nonflexibility in class labels inevitably causes fitting difficulty. In light of such, this study proposes relaxed robust supervised hashing (Relaxed RSH) for generating reliable signatures that can simultaneously conquer the above problems caused by incomplete data and rigid margins. To accommodate oversensitivity, this study proposes measuring hash learning loss by robust half quadratic (HQ) functions for modeling incomplete data. In the initial label matrix, slack variables are added to relax binary constraints. Such slack variables can be self-optimized during the learning process and can be used to automatically adjust margins between different classes. Decorrelation, balancing, and normalization constraints based on Relaxed RSH are also devised to provide discriminant and compact codes. Experimental results based on open datasets showed that the proposed method yielded higher mAP and F1 than the baselines. Note to Practitioners—This work is motivated by the problems caused by incomplete data and rigid class margins during event signature (e.g., hash) learning, where hashes are used to trigger automation systems. In existing hashing methods, incomplete data (e.g., continuous occlusion, missing values, and sample-specific outliers) along with rigid class margins cause fitting biases, thereby challenging data-driven automation. This study proposes Relaxed RSH by designing robust HQ loss, self-optimized label learning, and corresponding constraints for generating discriminant hashes. This prevents actuators from being mistriggered by corrupted data. Experiments were conducted on various data corruption. Future research will address the design for incremental/decentralized hash learning. Bo-Wei Chen, Jhao-Yang Huang, Ying-Hsuan Wu |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2023 | Image Inpainting by Mscswin Transformer Adversarial AutoencoderabstractImage inpainting has been researched for years. From deeper and larger models to models that focus on global information, all of them aim to obtain results closer to reality. In this paper, we combine the stripe window and line-by-line feature shift to modify the Vision Transformer (ViT) to reduce the computation cost and obtain global information from the oblique attention. In addition, we design a new loss function to enhance the texture and colors for inpainting. At last, to validate the efficacy of our proposed model, we conduct extensive experiments on commonly seen datasets (Places2 and CelebA) compared with other state-of-the-art methods. The source code and pretrained models are available at https: //github.com/bobo0303/MSCS-Net. Bo-Wei Chen, Tsung-Jung Liu, Kuan-Hsien Liu |
ICIP | 1 |
| 2023 | Enhancing Prediction of Forelimb Movement Trajectory through a Calibrating-Feedback Paradigm Incorporating RAT Primary Motor and Agranular Cortical Ensemble Activity in the Goal-Directed Reaching TaskabstractComplete reaching movements involve target sensing, motor planning, and arm movement execution, and this process requires the integration and communication of various brain regions. Previously, reaching movements have been decoded successfully from the motor cortex (M1) and applied to prosthetic control. However, most studies attempted to decode neural activities from a single brain region, resulting in reduced decoding accuracy during visually guided reaching motions. To enhance the decoding accuracy of visually guided forelimb reaching movements, we propose a parallel computing neural network using both M1 and medial agranular cortex (AGm) neural activities of rats to predict forelimb-reaching movements. The proposed network decodes M1 neural activities into the primary components of the forelimb movement and decodes AGm neural activities into internal feedforward information to calibrate the forelimb movement in a goal-reaching movement. We demonstrate that using AGm neural activity to calibrate M1 predicted forelimb movement can improve decoding performance significantly compared to neural decoders without calibration. We also show that the M1 and AGm neural activities contribute to controlling forelimb movement during goal-reaching movements, and we report an increase in the power of the local field potential (LFP) in beta and gamma bands over AGm in response to a change in the target distance, which may involve sensorimotor transformation and communication between the visual cortex and AGm when preparing for an upcoming reaching movement. The proposed parallel computing neural network with the internal feedback model improves prediction accuracy for goal-reaching movements. Han-Lin Wang, Yun-Ting Kuo, Yu-Chun Lo, Chao-Hung Kuo, Bo-Wei Chen, Ching-Fu Wang, Zu-Yu Wu, Chi-En Lee, Shih-Hung Yang, Sheng-Huang Lin, Po-Chuan Chen, You-Yin Chen |
Int. J. Neural Syst. | 5 |
| 2023 | Robust Data-Aware Sensing Based on Half Quadratic Hashing for the Visual Internet of ThingsabstractIn the Visual Internet of Things (VIoT) such as drones over flying ad-hoc networks (FANETs), terminal nodes need to preliminarily identify sensing data. Thus, data can be rapidly dispatched to dedicated edge nodes for further processing. Nonetheless, terminals may collect partially observed visual data, which cause biased labeling. In view of such, this study proposes half quadratic supervised discrete hashing (HQSDH) that can conquer the biased similarity estimation for partially observed images, especially when large continuous occlusion, missing values, or sample-specific outliers exist. The proposed method can automatically and adaptively adjust itself via HQ auxiliary variables to avoid oversensitivity to large errors caused by partially observed data. In this study, to solve HQ discrete optimization problems in HQSDH, HQ discrete cyclic coordinate descent (HQDCCD) is developed. Second, partially observed images may cause inconsistent binary codes and increase the Hamming distances. Thus, this study devises HQ balancing decorrelation constraints along with normalization regularization to accommodate the problems. Third, a systematic and generic solution that allows various HQ functions to jointly work in the same framework is modeled to provide versatility. Experiments on open image data sets were carried out for evaluation. The results showed that the proposed method provided more resilience against continuous occlusion, missing values, and sample-specific outliers than the baselines when dealing with partially observed images. Bo-Wei Chen, Jhao-Yang Huang |
IEEE Internet Things J. | 1 |
| 2023 | Conquering insufficient/imbalanced data learning for the Internet of Medical Things
Zi-Ching Lan, Guan-Yu Huang, Yun-Pei Li, Seungmin Rho, S. Vimal 0001, Bo-Wei Chen |
Neural Comput. Appl. | 6 |
| 2022 | Neuro-Inspired Reinforcement Learning to Improve Trajectory Prediction in Reward-Guided BehaviorabstractHippocampal pyramidal cells and interneurons play a key role in spatial navigation. In goal-directed behavior associated with rewards, the spatial firing pattern of pyramidal cells is modulated by the animal’s moving direction toward a reward, with a dependence on auditory, olfactory, and somatosensory stimuli for head orientation. Additionally, interneurons in the CA1 region of the hippocampus monosynaptically connected to CA1 pyramidal cells are modulated by a complex set of interacting brain regions related to reward and recall. The computational method of reinforcement learning (RL) has been widely used to investigate spatial navigation, which in turn has been increasingly used to study rodent learning associated with the reward. The rewards in RL are used for discovering a desired behavior through the integration of two streams of neural activity: trial-and-error interactions with the external environment to achieve a goal, and the intrinsic motivation primarily driven by brain reward system to accelerate learning. Recognizing the potential benefit of the neural representation of this reward design for novel RL architectures, we propose a RL algorithm based on [Formula: see text]-learning with a perspective on biomimetics (neuro-inspired RL) to decode rodent movement trajectories. The reward function, inspired by the neuronal information processing uncovered in the hippocampus, combines the preferred direction of pyramidal cell firing as the extrinsic reward signal with the coupling between pyramidal cell–interneuron pairs as the intrinsic reward signal. Our experimental results demonstrate that the neuro-inspired RL, with a combined use of extrinsic and intrinsic rewards, outperforms other spatial decoding algorithms, including RL methods that use a single reward function. The new RL algorithm could help accelerate learning convergence rates and improve the prediction accuracy for moving trajectories. Bo-Wei Chen, Shih-Hung Yang, Chao-Hung Kuo, Yu-Chun Lo, Yun-Ting Kuo, Yi-Chen Lin 0002, Hao-Cheng Chang, Sheng-Huang Lin, Boyi Qu, Shuan-Chu Vina Ro, Hsin-Yi Lai, You-Yin Chen |
Int. J. Neural Syst. | 1 |
| 2021 | Low-Error Data Recovery Based on Collaborative Filtering With Nonlinear Inequality Constraints for Manufacturing ProcessesabstractThis study proposes a data recovery model where substituted values can be further limited by nonlinear and inequality constraints to approximate the ground truth. The objective is to generate substituted values for multifactors while considering their lower/upper bounds, data means, and nonlinearity at the same time. This is critical when data need to fall inside a nonlinear range, e.g., a partial hypersphere centered at a given mean. In view of such, this study proposes collaborative filtering with nonlinear inequality constraints to tackle the problem. The proposed method consists of three steps. First, the system finds class-dependent and box-bounded imputation basis factors for an incomplete data set. Class-dependent bases can reflect data domains well. Second, class-dependent imputation coefficients are located by the proposed nonnegative coefficient discovery with nonlinear inequality constraints. This step limits searching space and avoids generating substituted values out of range. Finally, constrained iterative projection pursuit is proposed for measuring the quality of recovered data by examining reconstruction residuals. By using both the nonlinear inequality constraints and the constrained iterative projection pursuit, the system can recover data while satisfying multifactor nonlinear coeffects required by manufacturers. Experimental results showed that the proposed method was capable of generating substituted values with lower root-mean-squared errors. In addition, errors were reduced by at least 10.06% on average, better than those of the baselines. Furthermore, the classification accuracy of the proposed method after data imputation was higher than that of the baselines. Such findings indicated that the proposed method could approximate the characteristics of data when missing values appeared.Note to Practitioners—This work was motivated by the problem of missing values in industrial heterogeneous sensor readings. When data recovery is performed, the process should consider multifactor nonlinearity, lower/upper bounds, historical references, and the divergence between substituted values and references at the same time in order to reconstruct original sensor readings as many as possible. Existing approaches generally have solutions to linear or nonlinear equality constraints but not the aforementioned nonlinear inequality ones. This work designs a self-dictionary method—class-dependent and box-bounded imputation basis factors along with constrained iterative projection pursuit—for finding substituted values. Real industrial experiments were conducted based on$\mathcal {L}_{2}$-norms. Future research will address the design for other norms. Bo-Wei Chen, Wei-Cheng Ye |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2020 | Cooperative comodule discovery for swarm-intelligent drone arrays
Hsin Chuang, Kuan-Lin Hou, Seungmin Rho, Bo-Wei Chen |
Comput. Commun. | 4 |
| 2020 | Incomplete data classification - Fisher Discriminant Ratios versus Welch Discriminant Ratios
Bo-Wei Chen |
Future Gener. Comput. Syst. | 1 |
| 2020 | Enhancement of Hippocampal Spatial Decoding Using a Dynamic Q-Learning Method With a Relative Reward Using Theta Phase PrecessionabstractHippocampal place cells and interneurons in mammals have stable place fields and theta phase precession profiles that encode spatial environmental information. Hippocampal CA1 neurons can represent the animal’s location and prospective information about the goal location. Reinforcement learning (RL) algorithms such as Q-learning have been used to build the navigation models. However, the traditional Q-learning ([Formula: see text]Q-learning) limits the reward function once the animals arrive at the goal location, leading to unsatisfactory location accuracy and convergence rates. Therefore, we proposed a revised version of the Q-learning algorithm, dynamical Q-learning ([Formula: see text]Q-learning), which assigns the reward function adaptively to improve the decoding performance. Firing rate was the input of the neural network of [Formula: see text]Q-learning and was used to predict the movement direction. On the other hand, phase precession was the input of the reward function to update the weights of [Formula: see text]Q-learning. Trajectory predictions using [Formula: see text]Q- and [Formula: see text]Q-learning were compared by the root mean squared error (RMSE) between the actual and predicted rat trajectories. Using [Formula: see text]Q-learning, significantly higher prediction accuracy and faster convergence rate were obtained compared with [Formula: see text]Q-learning in all cell types. Moreover, combining place cells and interneurons with theta phase precession improved the convergence rate and prediction accuracy. The proposed [Formula: see text]Q-learning algorithm is a quick and more accurate method to perform trajectory reconstruction and prediction. Bo-Wei Chen, Shih-Hung Yang, Yu-Chun Lo, Ching-Fu Wang, Han-Lin Wang, Chen-Yang Hsu, Yun-Ting Kuo, Jung-Chen Chen, Sheng-Huang Lin, Han-Chi Pan, Sheng-Wei Lee, Boyi Qu, Chao-Hung Kuo, You-Yin Chen, Hsin-Yi Lai |
Int. J. Neural Syst. | 1 |
| 2018 | Efficient multiple incremental computation for Kernel Ridge Regression with Bayesian uncertainty modeling
Bo-Wei Chen, Nik Nailah Binti Abdullah, Sang Oh Park, Y. Gu |
Future Gener. Comput. Syst. | 1 |
| 2018 | Privacy-preserved big data analysis based on asymmetric imputation kernels and multiside similarities
Bo-Wei Chen, Seungmin Rho, Laurence T. Yang |
Future Gener. Comput. Syst. | 1 |
| 2018 | A 12-bit 40-MS/s SAR ADC With a Fast-Binary-Window DAC Switching Scheme
Yung-Hui Chung, Chia-Wei Yen, Pei-Kang Tsai, Bo-Wei Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2017 | Kernel weighted Fisher sparse analysis on multiple maps for audio event recognitionabstractThis work presents a novel approach for audio event recognition. The approach develops a weighted kernel fisher sparse analysis method based on multiple maps. The proposed method consists of maps extraction and kernel weighted Fisher sparse analysis. Two maps are firstly extracted from each audio file, i.e. scale-frequency map and damping-frequency map. The scale and frequency of the Gabor atoms are extracted to construct a scale-frequency map. On the other hand, the damping-frequency map is generated according to the frequency and damping factor of damped atoms. Gabor atoms can be utilized to model human auditory perception, and the damped atoms can be used to model commonly observed damped oscillations in natural signals. This work fuses the advantages of these two dictionaries to improve the performance of the system. During the recognition stage, this work constructs a kernel sparse representation-based classifier via the proposed kernel weighted Fisher sparse analysis to enhance separability. The proposed kernel weighted Fisher sparse analysis combines sparse representation with heteroscedastic kernel weighted discriminant analysis (HKWDA), which is useful for providing a discriminative recognition of audio events because a weighted pairwise Chernoff criterion is utilized in the kernel space. Experiments on a 20-class audio event database indicate that the proposed approach can achieve an accuracy rate of 82.70%. Also, integrating the scale-frequency map with MFCCs increases the accuracy rate to 87.70%. Yu-Hao Chin, Bo-Wei Chen, Jia-Ching Wang |
ICASSP | 2 |
| 2017 | Modeling of large-scale social network services based on mechanisms of information diffusion: Sina Weibo as a case study
Seungmin Rho, Bo-Wei Chen, Wandong Cai |
Future Gener. Comput. Syst. | 3 |
| 2017 | Guest Editorial for ACM TECS Special Issue on Effective Divide-and-Conquer, Incremental, or Distributed Mechanisms of Embedded Designs for Extremely Big Data in Large-Scale DevicesabstractNo abstract available. Bo-Wei Chen, Wen Ji 0003, Zhu Li 0001 |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2016 | Divide-and-conquer signal processing, feature extraction, and machine learning for big data
Bo-Wei Chen, Wen Ji 0003, Seungmin Rho |
Neurocomputing | 1 |
| 2016 | Evaluate mobile video quality in hybrid spatial and temporal domain
Wen Ji 0003, Seungmin Rho, Bo-Wei Chen, Yiqiang Chen 0001 |
Multim. Tools Appl. | 4 |
| 2016 | Multivoxel analysis for functional magnetic resonance imaging (fMRI) based on time-series and contextual information: relationship between maternal love and brain regions as a case study
Bo-Wei Chen, Yang-Yen Ou, Chun-Chia Kung, Ding-Ruey Yeh, Seungmin Rho, Jhing-Fa Wang |
Multim. Tools Appl. | 1 |
| 2016 | Online distribution and interaction of video data in social multimedia network
Xiangyang Ji, Qifei Wang, Bo-Wei Chen, Seungmin Rho, C.-C. Jay Kuo, Qionghai Dai |
Multim. Tools Appl. | 3 |
| 2016 | Big data driven decision making and multi-prior models collaboration for media restoration
Feng Jiang 0001, Seungmin Rho, Bo-Wei Chen, Debin Zhao |
Multim. Tools Appl. | 3 |
| 2016 | Updating high-utility pattern trees with transaction modification
Jerry Chun-Wei Lin, Wensheng Gan, Bo-Wei Chen, Seungmin Rho, Tzung-Pei Hong |
Multim. Tools Appl. | 4 |
| 2016 | A semi-supervised privacy-preserving clustering algorithm for healthcare
Meiyu Huang, Yiqiang Chen 0001, Bo-Wei Chen, Junfa Liu, Seungmin Rho, Wen Ji 0003 |
Peer-to-Peer Netw. Appl. | 3 |
| 2016 | Cross-Layer Opportunistic Scheduling for Device-to-Device Video Multicast ServicesabstractIn this article, we address the problem of how to make the wireless device-to-device (D2D) video multicast systems have better quality provision with consideration of internet-of-things (IoT) applications. We propose an opportunistic transmission and fair resource allocation framework, including joint application-layer and physical-layer transmission and optimization. First, we use a parallel subchannels structure by concatenating the Fountain codes and diversity-embedded space-time block codes to provide reliable and flexible transmission in heterogeneous circumstances. Second, we exploit the quality of heterogeneous user experience (quality of experience) metric under D2D video multicast systems, with consideration of various channel states, device capability, video content urgency, and the number of demanding users. Third, we formulate reliable multiple video streams broadcasting to heterogeneous devices as an aggregate maximum utility achieving problem, and we use opportunistic scheduling to select suitable users in each transmission interval to improve the broadcasting utility. Fourth, we use the utility fair scheme to guide rate allocation among multicontent video multicast. Extensive performance comparison and analysis are presented to demonstrate efficiency of the proposed solution. Wen Ji 0003, Bo-Wei Chen, Haiyong Luo, Mucheol Kim, Yiqiang Chen 0001 |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2016 | Large-scale image colorization based on divide-and-conquer support vector machines
Bo-Wei Chen, Wen Ji 0003, Seungmin Rho, Sun-Yuan Kung |
J. Supercomput. | 1 |
| 2016 | Support vector analysis of large-scale data based on kernels with iteratively increasing order
Bo-Wei Chen, Wen Ji 0003, Seungmin Rho, Sun-Yuan Kung |
J. Supercomput. | 1 |
| 2016 | Erratum to: Large-scale image colorization based on divide-and-conquer support vector machines
Bo-Wei Chen, Wen Ji 0003, Seungmin Rho, Sun-Yuan Kung |
J. Supercomput. | 2 |
| 2016 | Optimal filter based on scale-invariance generation of natural images
Feng Jiang 0001, Bo-Wei Chen, Seungmin Rho, Wen Ji 0003, Liqiang Pan, Debin Zhao |
J. Supercomput. | 2 |
| 2016 | Profit Maximization through Online Advertising Scheduling for a Wireless Video Broadcast NetworkabstractIn this paper, we address the problem of how to make the wireless service provider (WSP) earn profits in a wireless video broadcast network with consideration of advertisement insertion. At the beginning, this study examines the profit components by analyzing traffic provision and advertisement insertion. This study considers using two components for profit maximization-one is the function for allocating video rates, and the other is the function for inserting advertisement duration. The maximum achievable profit depends on joint optimization of optimal video-rate vectors and advertisement-duration vectors, which are usually computationally intensive. To resolve such a complexity problem, this work also proposes an effective algorithm for joint optimization. First, the overall profit is formulated as the solution of four local optimization problems through horizontal and vertical decomposition. Second, a theoretic polymatroidal framework is introduced in our work for optimization as this framework is proved effective in profit maximization of multiuser systems. Third, this study shows that the overall profit can be maximized by finding the optimal profit points on the boundary of the rate and duration regions. As a result, the optimum points and the total profit can be obtained through a hierarchical greedy algorithm. Experimental results demonstrate that the proposed method is capable of making maximum profits for WSPs in a wide range of broadcasting rates. Wen Ji 0003, Yingying Chen 0001, Min Chen 0003, Bo-Wei Chen, Yiqiang Chen 0001, Sun-Yuan Kung |
IEEE Trans. Mob. Comput. | 4 |
| 2016 | Fixed-Point Computing Element Design for Transcendental Functions and Primary Operations in Speech ProcessingabstractThis brief presents a fixed-point architecture based on a reconfigurable scheme for integrating several commonly used mathematical operations of speech signal processing. The proposed design can perform two transcendental mathematical operations called logarithm and powering, and three commonly used computations with similar operations named polynomial calculation, filtering, and windowing. By analyzing the adopted algorithms of the above five operations, a simplified computing unit is designed. This unit can combine six types of operations by reconfiguring the data paths, and the same multiply-add architecture can be reused for reducing the redundant usage of logic gates. The experimental results reveal that the proposed design can work at a 200-MHz clock rate, and its gate count only has 11.9k. Compared with the results of the floating-point function, the median errors of the proposed design for computing the powering and logarithmic functions are 0.57% and 0.11%, respectively. Such results indicate that this simple architecture can be effectively used in most speech processing applications. Chung-Hsien Chang, Shi-Huang Chen, Bo-Wei Chen, Wen Ji 0003, K. Bharanitharan, Jhing-Fa Wang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | A New Binary-Halved Clustering Method and ERT Processor for ASSR SystemabstractThis paper presents an automatic speech–speaker recognition (ASSR) system implemented in a chip which includes a built-in extraction, recognition, and training (ERT) core. For VLSI design (here, ASSR system), the hardware cost and time complexity are always the important issues which are improved in this proposed design in two levels: 1) algorithmic and 2) architecture. At the algorithm level, a newly binary-halved clustering (BHC) is proposed to achieve low time complexity and low memory requirement. In addition, at the architecture level, a new ERT core is proposed and implemented based on data dependence and reuse mechanism to reduce the time and hardware cost as well. Finally, the chip implementation is synthesized, placed, and routed using TSMC 90-nm technology library. To verify the performance of the proposed BHC method, a case study is performed based on nine speakers. Moreover, the validation of the ASSR system is examined in two parts: 1) speech recognition and 2) speaker recognition. The results show that the proposed system can achieve 93.38% and 87.56% of recognition rates during speech and speaker recognition, respectively. Furthermore, the proposed ASSR chip includes 396k gate counts, and consumes power in 8.74 mW. Such results demonstrate that the performance of the proposed ASSR system is superior to the conventional systems. Chih-Hung Chou, Ta-Wen Kuan, Shovan Barma, Bo-Wei Chen, Wen Ji 0003, Chih-Hsiang Peng, Jhing-Fa Wang |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2015 | Subspace-based DOA with linear phase approximation and frequency bin selection preprocessing for interactive robots in noisy environments
Sheng-Chieh Lee, Bo-Wei Chen, Jhing-Fa Wang, Min-Jian Liao |
Comput. Speech Lang. | 2 |
| 2015 | Robust skin detection in real-world images
Lei Huang 0010, Zhiqiang Wei 0002, Bo-Wei Chen, Chenggang Yan 0001, Jie Nie, Jian Yin 0003, Baochen Jiang |
J. Vis. Commun. Image Represent. | 4 |
| 2015 | LSH-based semantic dictionary learning for large scale image understanding
Liang Li 0003, Chenggang Yan 0001, Bo-Wei Chen, Shuqiang Jiang, Qingming Huang |
J. Vis. Commun. Image Represent. | 4 |
| 2015 | User-centric incremental learning model of dynamic personal identification for mobile devices
Hsin-Chun Tsai, Bo-Wei Chen, K. Bharanitharan, Anand Paul 0001, Jhing-Fa Wang, Hung-Chieh Tai |
Multim. Syst. | 2 |
| 2015 | Peer-to-peer usage analysis in dynamic databases
Jerry Chun-Wei Lin, Wensheng Gan, Tzung-Pei Hong, Sang Oh Park, Bo-Wei Chen |
Peer-to-Peer Netw. Appl. | 5 |
| 2015 | Speech Emotion Verification Using Emotion Variance Modeling and Discriminant Scale-Frequency MapsabstractThis paper develops an approach to speech-based emotion verification based on emotion variance modeling and discriminant scale-frequency maps. The proposed system consists of two parts-feature extraction and emotion verification. In the first part, for each sound frame, important atoms from the Gabor dictionary are selected by using the matching pursuit algorithm. The scale, frequency, and magnitude of the atoms are extracted to construct a nonuniform scale-frequency map, which supports auditory discriminability by the analysis of critical bands. Next, sparse representation is used to transform scale-frequency maps into sparse coefficients to enhance the robustness against emotion variance and achieve error-tolerance improvement. In the second part, emotion verification, two scores are calculated. A novel sparse representation verification approach based on Gaussian-modeled residual errors is proposed to generate the first score from the sparse coefficients. Such a classifier can minimize emotion variance and improve recognition accuracy. The second score is calculated by using the emotional agreement index (EAI) from the same coefficients. These two scores are combined to obtain the final detection result. Experiments on an emotional database of spoken speech were conducted and indicate that the proposed approach can achieve an average equal error rate (EER) of as low as 6.61%. A comparison among different approaches reveals that the proposed method is superior to the others and confirms its feasibility. Jia-Ching Wang, Yu-Hao Chin, Bo-Wei Chen, Chang-Hong Lin, Chung-Hsien Wu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | Quantitative Measurement of Split of the Second Heart Sound (S2)abstractThis study proposes a quantitative measurement of split of the second heart sound (S2) based on nonstationary signal decomposition to deal with overlaps and energy modeling of the subcomponents of S2. The second heart sound includes aortic (A2) and pulmonic (P2) closure sounds. However, the split detection is obscured due to A2-P2 overlap and low energy of P2. To identify such split, HVD method is used to decompose the S2 into a number of components while preserving the phase information. Further, A2s and P2s are localized using smoothed pseudo Wigner-Ville distribution followed by reassignment method. Finally, the split is calculated by taking the differences between the means of time indices of A2s and P2s. Experiments on total 33 clips of S2 signals are performed for evaluation of the method. The mean ± standard deviation of the split is 34.7 ± 4.6 ms. The method measures the split efficiently, even when A2-P2 overlap is ≤ 20 ms and the normalized peak temporal ratio of P2 to A2 is low (≥ 0.22). This proposed method thus, demonstrates its robustness by defining split detectability (SDT), the split detection aptness through detecting P2s, by measuring up to 96 percent. Such findings reveal the effectiveness of the method as competent against the other baselines, especially for A2-P2 overlaps and low energy P2. Shovan Barma, Bo-Wei Chen, Ka Lok Man, Jhing-Fa Wang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2015 | High-efficient video compression for social multimedia distribution
Xiangyang Ji, Sam Kwong, Bo-Wei Chen, Seungmin Rho |
J. Supercomput. | 3 |
| 2015 | Face hallucination and recognition in social network services
Feng Jiang 0001, Seungmin Rho, Bo-Wei Chen, Xiaodan Du 0003, Debin Zhao |
J. Supercomput. | 3 |
| 2015 | Profit Improvement in Wireless Video Broadcasting System: A Marginal Principle ApproachabstractIn this paper, we address the problem of how to make the wireless service provider have better profits with consideration of user experience provision in wireless video broadcasting systems. We propose a marginal-based pricing and a resource-allocation framework to achieve better resource utilization and profit improvement. The marginal principle includes 1) marginal user principle, in which a pricing mechanism is established on the basis of marginal users, such that the WSP can seek its own maximum profit of each content with a QoE guarantee; 2) marginal profit principle, in which a WSP can earn the maximum profit through multicontent-service provision by regulating rate allocation in limited available bandwidth. Furthermore, we present a two-tier framework consisting of the inner and outer loops. The inner loop focuses on pricing-based service provision based on the notion of marginal user principle. The outer loop concentrates on allocating bandwidth among multiple video contents according to marginal profit principle. For the solution, we model the profit regions of WSPs and end-users as the polymatroid structures and model the corresponding allocated rate regions as the contra-polymatroid structures. Through exploiting the properties of polymatroid and contra-polymatroid structures, the broadcasting profit problem is solved by finding the optimal rate vector on the sum-rate facet which satisfies the maximal achievable profit. Extensive performance comparison and analysis are presented to demonstrate efficiency of the proposed solution. Wen Ji 0003, Bo-Wei Chen, Yiqiang Chen 0001, Sun-Yuan Kung |
IEEE Trans. Mob. Comput. | 2 |
| 2015 | Profit Optimization for Wireless Video Broadcasting Systems Based on Polymatroidal AnalysisabstractThis study addresses the problem of profit maximization between wireless service providers (WSPs) and content providers (CPs) in wireless broadcasting systems , while simultaneously providing high quality of experience for end-users (EUs). We first study the profit model in wireless broadcasting networks with a particular attention to the heterogeneous requirements of EUs, e.g., different display sizes and variable channel conditions. Then, we propose a profit formulation that describes the requirements of wireless service providers and content providers, as well as the satisfaction of EUs that essentially depends on video quality and service charges. We propose a new polymatroidal theoretic framework for maximizing the resulting three-side achievable profit through proper bandwidth allocation. Our framework exploits two particular structures, namely the underlying polymatroidal structure of the profit region and the contra- polymatroidal structure of the rate region. We then propose a profit maximization solution by finding a rate allocation vector on the sum-rate facet that satisfies the maximal achievable profit among the WSP, CPs, and EUs. Experiments on different broadcasting scenarios demonstrate the effectiveness of the proposed method. The WSP is capable of generating more revenues by applying the proposed approach to their marketing strategies while satisfying the demands from CPs and EUs. Wen Ji 0003, Pascal Frossard, Bo-Wei Chen, Yiqiang Chen 0001 |
IEEE Trans. Multim. | 3 |
| 2015 | Low-Complexity Hardware Design for Fast Solving LSPs With Coordinated Polynomial SolutionabstractThis paper presents a low-complexity algorithm and the corresponding hardware based on the coordinated polynomial solutions for solving line spectrum pairs (LSPs). To improve the computation of LSPs, the enhanced Tschirnhaus transform (ETT) is proposed to accelerate the coordinated polynomial solution. The proposed ETT can replace fractional multiplication with addition and shift operations, so unnecessary operations are avoided. To further simplify the hardware of the ETT, three designs are presented: the preprocessing block (PPB), the iterative root-finding block (IRFB), and the closed-form solution block (CFSB). The PPB provides a design with less gate counts that can effectively transform LPCs into general-form polynomials. Such polynomials can be further decomposed into roots using the proposed IRFB based on the Birge-Vieta method. A pipeline-recursive framework is implemented in the IRFB to save calculations. To improve hardware utilization, this paper also analyzes the coefficients relationship of the ETT by introducing the data dependency graph to design the proposed functional blocks in CFSB. The experimental results show that the proposed hardware achieves a 40-fold improvement in throughput and reduces 1.16% of gate counts at the hardware synthesis level; the chip area is 1.29 mm2. The precision analysis indicates the average log spectral distance is 0.310. Moreover, the ETT in the proposed hardware only requires 29.9% of multiplication compared with the original one. Such results reveal that the proposed work is superior to the baseline work, thereby demonstrating the effectiveness of the proposed design. Chung-Hsien Chang, Bo-Wei Chen, Shi-Huang Chen, Jhing-Fa Wang, Yu-Hao Chiu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | Quality of Service Enhancement by Using an Integer Bloom Filter Based Data Deduplication Mechanism in the Cloud Storage Environment
Kuo-Qin Yan, Yung-Hsiang Su, Hsin-Met Chuan, Shu-Chin Wang, Bo-Wei Chen |
NPC | 5 |
| 2014 | Enhanced long-range personal identification based on multimodal information of human features
Hsin-Chun Tsai, Bo-Wei Chen, Jhing-Fa Wang, Anand Paul 0001 |
Multim. Tools Appl. | 2 |
| 2014 | Gabor-Based Nonuniform Scale-Frequency Map for Environmental Sound Classification in Home AutomationabstractThis work presents a novel feature extraction approach called nonuniform scale-frequency map for environmental sound classification in home automation. For each audio frame, important atoms from the Gabor dictionary are selected by using the Matching Pursuit algorithm. After the system disregards phase and position information, the scale and frequency of the atoms are extracted to construct a scale-frequency map. Principal Component Analysis (PCA) and Linear Discriminate Analysis (LDA) are then applied to the scale-frequency map, subsequently generating the proposed feature. During the classification phase, a segment-level multiclass Support Vector Machine (SVM) is operated. Experiments on a 17-class sound database indicate that the proposed approach can achieve an accuracy rate of 86.21%. Furthermore, a comparison reveals that the proposed approach is superior to the other time-frequency methods. Jia-Ching Wang, Chang-Hong Lin, Bo-Wei Chen, Min-Kang Tsai |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2014 | Mixed Sound Event Verification on Wireless Sensor Network for Home AutomationabstractIn this paper, we present the problem of mixed sound event verification in a wireless sensor network for home automation systems. In home automation systems, the sound recognized by the system becomes the basis for performing certain tasks. However, if a target source is mixed with another sound due to simultaneous occurrence, the system would generate poor recognition results, subsequently leading to inappropriate responses. To handle such problems, this study proposes a framework, which consists of sound separation and sound verification techniques based on a wireless sensor network (WSN), to realize sound-triggered automation. In the sound separation phase, we present a convolutive blind source separation system with source number estimation using time-frequency clustering. An accurate mixing matrix can be estimated by the proposed phase compensation technique and used for reconstructing the separated sound sources. In the verification phase, Mel frequency cepstral coefficients and Fisher scores that are derived from the wavelet packet decomposition of signals are used as features for support vector machines. Finally, a sound of interest can be selected for triggering automated services according to the verification result. The experimental results demonstrate the robustness and feasibility of the proposed system for mixed sound verification in WSN-based home environments. Jia-Ching Wang, Chang-Hong Lin, Ernestasia Siahaan, Bo-Wei Chen, Hsiang-Lung Chuang |
IEEE Trans. Ind. Informatics | 4 |
| 2014 | REC-STA: Reconfigurable and Efficient Chip Design With SMO-Based Training AcceleratorabstractSequential minimal optimization (SMO) and Karush-Kuhn-Tucker condition are often used to solve learning problems in support vector machines. However, during hardware implementation of the SMO algorithm, enhancing chip performance without excessively increasing chip area is often a crucial issue. The solution proposed in this paper is a novel reconfigurable and efficient chip design with SMO-based training accelerator (REC-STA). Two novel methods used in the proposed REC-STA are trimode coarse-grained reconfigurable architecture (TCRA) and triple finite-state-machine with dynamic scheduling The first method modifies the baseline SMO design by developing trimode reconfigurable architectures with parallel and pipeline computing capabilities. The second method provides a schedule for efficient reconfiguration of the TCRA. Use of these methods can remove kernel cache design. For chip manufacturing, the implementation of the REC-STA is synthesized, placed, and routed using the TSMC 0.18-μm technology library. The core size is 2.94 mm × 2.94 mm and the power consumption is 77.3 mW. Compared with the baseline design, the FPGA simulation results show that the proposed architecture requires 50% less memory and 31% fewer gate counts but provides a 16-fold improvement in training performance. The experimental results confirm the efficacy of the proposed architecture and methods. Chih-Hsiang Peng, Bo-Wei Chen, Ta-Wen Kuan, Po-Chuan Lin, Jhing-Fa Wang, Nai-Sheng Shih |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | High-efficient hardware design based on enhanced Tschirnhaus transform for solving the LSPsabstractThis work presents a novel hardware design based on the enhanced Tschirnhaus transform (ETT) to solve the 8-order line spectral pairs (LSPs). To reduce high-complexity problems caused by fractional multiplication, the ETT is proposed to replace such operations with integer-based shift and addition operations of the original Tschirnhaus transform. Also, the data dependency graph (DDG) of the ETT is analyzed for designing hardware units and reducing computation cycles. The proposed hardware has two key blocks: the mixture computation unit (MCU) and the multiplier-free pipelined square-root unit (PSRU). The first block is designed to fast calculate multiplication and summation operations in the ETT with the use of a two-stage pipeline architecture. The second is developed to speed up square-root operations after 8-order LSPs are decomposed into two 4-order LSPs. It can also timely process the result of the first block within limited cycles. The experimental results show that compared with the Chebyshev-based research, the proposed hardware can reduce the cycle times by 98.1% and also saved about 49.7% of gate counts. In the precision evaluation, the result indicates that 95% of the computation errors are within 0.02 and proves that the proposed hardware is capable of quantizing LSPs almost as accurately as computers do. Such results reveal that the proposed work is superior to the other Chebyshev-based methods, thereby demonstrating the effectiveness of the proposed design. Chung-Hsien Chang, Shi-Huang Chen, Bo-Wei Chen, Chih-Hsiang Peng, Jhing-Fa Wang |
ISCAS | 3 |
| 2013 | Video search and indexing with reinforcement agent for interactive multimedia servicesabstractIn this study, we present a video search and indexing system based on the state support vector (SVM) network, video graph, and reinforcement agent for recognizing and organizing video events. In order to enhance the recognition performance of the state SVM network, two innovative techniques are presented: state transition correction and transition quality estimation. The classification results are also merged into the video indexing graph, which facilitates the search speed. A reinforcement algorithm with an efficient scheduling scheme significantly reduces both the power consumption and time. The experimental results show the proposed state SVM network was able to achieve a precision rate as high as 83.83% and the query results of the indexing graph reached 80% accuracy. The experiments also demonstrate the performance and feasibility of our system. Anand Paul 0001, Bo-Wei Chen, K. Bharanitharan, Jhing-Fa Wang |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2013 | Smart Homecare Surveillance System: Behavior Identification Based on State-Transition Support Vector Machines and Sound Directivity Pattern AnalysisabstractThis study presents a smart homecare surveillance system, which utilizes sound-steered cameras to identify behavior of interest. First of all, to detect multiple source locations, a new direction-of-arrival (DOA) algorithm is proposed by introducing cascaded frequency filters, which can quickly calculate directions without creating much complexity. This method can also locate and separate different signals at the same time. Second, after the camera points in the direction of the estimated angle, the proposed state-transition support vector machine is used to provide favorable discriminability for human behavior identification. A new Markov random field (MRF) function based on the localized contour sequence (LCS) is also presented while the system computes transition probabilities between states. Such LCS-based MRF functions can effectively smooth transitions and enhance recognition. The experimental results show that the average error of DOA decreases to around 7°, which is better than those of the baselines. Also, our proposed behavior identification system can reach an 88.3% accuracy rate. The aforementioned results have therefore demonstrated the feasibility of the proposed method. Bo-Wei Chen, Chen-Yu Chen, Jhing-Fa Wang |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2012 | Novel Binaural Spectro-temporal Algorithm for Speech Enhancement in Low SNR EnvironmentsabstractA novel BInaural Spectro-Temporal (BIST) algorithm is proposed in this paper to increase the speech intelligibility in low or negative SNR noisy environments. The BIST algorithm consists of two modules. One is the spatial mask for receiving sound from the specific direction, and the other is the spectro-temporal modulation filter for noise reduction. Most speech enhancement algorithms are not applicable in harsh environments because the energy of speech is covered by the noise. To increase the speech intelligibility in low or negative SNR noisy environments, a distinctive approach is proposed to solve this problem. First, the BIST algorithm takes binaural auditory processing as a spatial mask to separate the speech and noise according to their locations. Next, the modulation filter is applied to reduce the noise source in the scale-rate (spectro-temporal modulation) domain according to their different acoustic feature. It works like the spectro-temporal receptive field (STRF) which is the perception response of human auditory cortex. The experimental results demonstrate that the proposed BIST speech enhancement algorithm can improve 20% from the noisy speech at SNR-10dB. Po-Hsun Sung, Bo-Wei Chen, Ling-Sheng Jang, Jhing-Fa Wang |
ICME | 2 |
| 2012 | A new hybrid and dynamic fusion of multiple experts for intelligent porch system
Ta-Wen Kuan, Hsin-Chun Tsai, Jhing-Fa Wang, Jia-Ching Wang, Bo-Wei Chen, Zong-You Lin |
Expert Syst. Appl. | 5 |
| 2011 | Emotion Detection Based on Concept Inference and Spoken Sentence Analysis for Customer Service
Ren-Ying Fang, Bo-Wei Chen, Jhing-Fa Wang, Chung-Hsien Wu 0001 |
INTERSPEECH | 2 |
| 2011 | An Efficient Pre-Processing Scheme to Improve the Sound Source Localization System in Noisy Environment
Sheng-Chieh Lee, K. Bharanitharan, Bo-Wei Chen, Jhing-Fa Wang, Chung-Hsien Wu 0001, Min-Jian Liao |
INTERSPEECH | 3 |
| 2010 | Noisy Environment-Aware Speech Enhancement for Speech Recognition in Human-Robot Interaction ApplicationabstractIn this study, we introduce a noisy environment-aware speech enhancement system, which can be used in human-robot interaction (HRI) application for command recognition. In order to effectively filter different noises and improve speech recognition rates, the proposed system adopts automatic noise cancellation that is combined with independent component analysis (ICA) and subspace speech enhancement (SSE). Furthermore, it can automatically decide when to use noise reduction according to SNRs of the detected noisy speeches at any time (using proposed noisy environment-aware determination). The experimental results show that our proposed system is suitable for various types of noisy environments, and it is capable of improving the speech quality for recognition. Our proposed system can enhance SNRs by about 20dB, which is higher than those of original noisy speeches. Sheng-Chieh Lee, Bo-Wei Chen, Jhing-Fa Wang |
SMC | 2 |
| 2009 | Video Knowledge Augmentation based on Summarized Contents and Online MediaabstractExploration techniques of video knowledge have been proposed for years to help people discover the details about videos. However, existing systems still yield limited information for users. In this paper, we present a video knowledge browsing system, which can establish the framework of a video based on its summarized contents and expand them by using online correlated media. Thus, users can not only browse key points of a video efficiently but also focus on what they are interested in. In order to construct the fundamental system, we make use of our previous proposed approaches to transforming a video into a graph. After the relational graph is built up, the social network analysis is then performed to explore online relevant resources. We also apply the Markov clustering algorithm to enhance the results of the network analysis. The experiments demonstrate that our system can achieve better performance than the traditional systems. Bo-Wei Chen, Jhing-Fa Wang, Jia-Ching Wang |
ISCAS | 1 |
| 2009 | A Novel Video Summarization Based on Mining the Story-Structure and Semantic Relations Among Concept EntitiesabstractVideo summarization techniques have been proposed for years to offer people comprehensive understanding of the whole story in the video. Roughly speaking, existing approaches can be classified into the two types: one is static storyboard, and the other is dynamic skimming. However, despite that these traditional methods give brief summaries for users, they still do not provide with a concept-organized and systematic view. In this paper, we present a structural video content browsing system and a novel summarization method by utilizing the four kinds of entities: who, what, where, and when to establish the framework of the video contents. With the assistance of the above-mentioned indexed information, the structure of the story can be built up according to the characters, the things, the places, and the time. Therefore, users can not only browse the video efficiently but also focus on what they are interested in via the browsing interface. In order to construct the fundamental system, we employ maximum entropy criterion to integrate visual and text features extracted from video frames and speech transcripts, generating high-level concept entities. A novel concept expansion method is introduced to explore the associations among these entities. After constructing the relational graph, we exploit graph entropy model to detect meaningful shots and relations, which serve as the indices for users. The results demonstrate that our system can achieve better performance and information coverage. Bo-Wei Chen, Jia-Ching Wang, Jhing-Fa Wang |
IEEE Trans. Multim. | 1 |
| 2008 | A Long-Distance Time Domain Sound Localization
Jhing-Fa Wang, Jia-Chang Wang, Bo-Wei Chen, Zheng-Wei Sun |
UIC | 3 |
| 2006 | Fatigue-Induced Reversed Hemispheric Plasticity During Motor Repetitions: A Brain Electrophysiological Study
Ling Fu Meng, Chiu-Ping Lu, Bo-Wei Chen, Ching-Horng Chen |
ICONIP (1) | 3 |
| 2000 | Fuzzy/neural congestion control for integrated voice and data DS-CDMA/FRMA cellular networksabstractThe paper proposes congestion control using fuzzy/neural techniques for integrated voice and data direct-sequence code division multiple access/frame reservation multiple access (DS-CDMA/FRMA) cellular networks. The fuzzy/neural congestion controller is constituted by a pipeline recurrent neural network (PRNN) interference predictor, a fuzzy performance indicator, and a fuzzy/neural access probability controller. It regulates the traffic input to the integrated voice and data DS-CDMA/FRMA cellular system by determining proper access probabilities for users so that congestion can be avoided and the throughput can be maximized. Simulation results show that the DS-CDMA/FRMA fuzzy/neural congestion controllers perform better than conventional DS-CDMA/PRMA with channel access function in terms of voice packet dropping ratio, corruption ratio, and utilization. In addition, the neural congestion controller outperforms the fuzzy congestion controller. Chung-Ju Chang, Bo-Wei Chen, Terng-Yuan Liu, Fang-Ching Ren |
IEEE J. Sel. Areas Commun. | 2 |