Sun-Yuan Kung

dblp:25/4640 · DBLP profile ↗
← Back
223ranked-venue papers
40as first author
37since 2021 · last 2027
0000-0002-7314-0720ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 120 · 24 first-author · 13 since 2021Artificial intelligence and machine learning · 61 · 4 first-author · 20 since 2021Systems, architecture and hardware · 24 · 8 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 3 first-author · 5 since 2021Computer networks · 12 · 1 first-author · 1 since 2021Theory of computation · 3Databases, data management, data science and information retrieval · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2027 Auto-metric network for open-set recognition
Shuoshi Li, Yuan Zhou 0006, Ke Zhang 0005, Sun-Yuan Kung
Expert Syst. Appl.4
2026 Large-model-based smart agent for time series anomaly detection in power systems
Bingrui Wang, Yuan Zhou 0006, Leijiao Ge, Sun-Yuan Kung
Expert Syst. Appl.4
2026 MSSFN: Multi-stimulus stereo spatiotemporal fusion network with pattern disentanglement for Alzheimer's disease diagnosis
Peiguang Jing, Yu Liu 0004, Sun-Yuan Kung
Inf. Process. Manag.6
2025 Masked Image Pretraining on Language Assisted Representation
abstract
Self-attention based transformer models have been dominating many computer vision tasks in the past few years. Their superb model qualities heavily depend on the excessively large labeled image datasets. In order to reduce the reliance on large labeled datasets, reconstruction based masked autoencoders are gaining popularity, which learn high quality transferable representations from unlabeled images. For the same purpose, recent weakly supervised image pretraining methods explore language supervision from text captions accompanying the images. In this work, we propose Masked Image pretraining on Language Assisted representatioN, dubbed as MILAN. Instead of predicting raw pixels or low level features, our pretraining objective is to reconstruct the image features with substantial semantic signals that are obtained using caption supervision. Moreover, to accommodate our reconstruction target, we propose a more efficient prompting decoder architecture and a semantic aware mask sampling mechanism, which further advance the transfer performance of the pretrained model. Experimental results demonstrate that MILAN delivers higher accuracy than the previous works. When the masked autoencoder is pretrained and finetuned on ImageNet-1K dataset with an input resolution of 224×224, MILAN achieves a top-1 accuracy of 85.4% on ViT-Base, surpassing previous state-of-the-arts by 1%. In the downstream semantic segmentation task, MILAN achieves 52.7 mIoU using ViT-Base on ADE20K dataset, outperforming previous masked pre-training results by 4 points.
Zejiang Hou, Sun-Yuan Kung
ICASSP2
2025 Adding Conditional Control to Diffusion Models with Reinforcement Learning
abstract
Diffusion models are powerful generative models that allow for precise control over the characteristics of the generated samples. While these diffusion models trained on large datasets have achieved success, there is often a need to introduce additional controls in downstream fine-tuning processes, treating these powerful models as pre-trained diffusion models. This work presents a novel method based on reinforcement learning (RL) to add such controls using an offline dataset comprising inputs and labels. We formulate this task as an RL problem, with the classifier learned from the offline dataset and the KL divergence against pre-trained models serving as the reward functions. Our method, **CTRL** (**C**onditioning pre-**T**rained diffusion models with **R**einforcement **L**earning), produces soft-optimal policies that maximize the abovementioned reward functions. We formally demonstrate that our method enables sampling from the conditional distribution with additional controls during inference. Our RL-based approach offers several advantages over existing methods. Compared to classifier-free guidance, it improves sample efficiency and can greatly simplify dataset construction by leveraging conditional independence between the inputs and additional controls. Additionally, unlike classifier guidance, it eliminates the need to train classifiers from intermediate states to additional controls. The code is available at https://github.com/zhaoyl18/CTRL.
Yulai Zhao 0002, Masatoshi Uehara, Gabriele Scalia, Sun-Yuan Kung, Tommaso Biancalani, Sergey Levine, Ehsan Hajiramezanali
ICLR4
2025 Adaptive motion enhancement for passive non-line-of-sight action recognition
Zhongqi Sun, Yuan Zhou 0006, Shuwei Huo, Sun-Yuan Kung
Neurocomputing4
2025 Illumination guided domain adaptation object detection in thermal imagery
Yuan Zhou 0006, Yu Liu 0004, Sun-Yuan Kung
Neurocomputing5
2025 Module-Pruning-Based Neural Architectural Search for Remote Sensing Image Captioning
abstract
Remote sensing image captioning (RSIC) has garnered significant attention for enhancing the interpretability of aerial imagery through textual descriptions. Conventional approaches employ convolutional neural networks (CNNs) for visual feature extraction paired with recurrent neural networks (RNNs) or transformers for caption generation. However, these architectures suffer from high complexity and computational costs. While neural architecture search (NAS) via network pruning has been extensively studied, module-based pruning for RSIC systems remains largely unexplored. We propose a novel dedicated decoder pruning methodology for sequential caption generators—a module-based pruning method for end-to-end encoder–decoder architectural adaptation. It features two key innovations: 1) structured pruning of a pre-trained ResNet encoder and transformer encoder–decoder components and 2) a cross-entropy-based caption matching strategy replacing conventional prediction training in the decoder’s final layer. The proposed method enables simultaneously enhancing inference efficiency and reducing storage requirements without compromising performance. As evaluated on the RSICD dataset using CIDEr, ROUGE, METEOR, bilingual evaluation understudy (BLEU), and Sm metrics, our method achieves 42.8% model size reduction while improving accuracy, establishing new benchmarks in efficient RSIC.
Yogendra Rao Musunuri, Changwon Kim, Oh-Seol Kwon, Sun-Yuan Kung
IEEE Geosci. Remote. Sens. Lett.4
2024 Message from the General Chairs; IEEE CSCloud/EdgeCom 2024
abstract
Representatives of collaborating organizations, delegates from all supporters, it is our great pleasure and honor as General Chairs of IEEE CSCloud/EdgeCom 2024 to welcome you to the 11th IEEE International Conference on Cyber Security and Cloud Computing (IEEE CSCloud 2024), the 10th IEEE International Conference on Edge Computing and Scalable Cloud (IEEE EdgeCom 2024), and the co-located events. IEEE CSCloud/EdgeCom 2024 Committees organize these two co-located conferences.
Zhihui Lv, Sun-Yuan Kung
CSCloud3
2024 Contrastive learning based open-set recognition with unknown score
Yuan Zhou 0006, Songyu Fang, Shuoshi Li, Boyu Wang 0004, Sun-Yuan Kung
Knowl. Based Syst.5
2024 Cross-Transfer Learning for Enhancing Object Detection in Remote Sensing Images
abstract
Deep learning models have gained widespread acclaim for their effectiveness and precision in object detection systems. However, the development of novel models is often hindered by various challenges, such as limited data availability, scarcity of GPU resources, and the need for specialized knowledge. Despite the adoption of existing transfer learning techniques, these methods have frequently fallen short of achieving the desired results. Consequently, this study introduces the concept of cross-transfer learning (CTL) as an innovative joint learning strategy. CTL integrates curriculum and transfer learning methodologies, along with a difficulty measurer, to leverage knowledge obtained from simpler to more complex datasets across a sequence of subsets for learning new objectives. Data are categorized using a difficulty measurer, which ranks it based on visual recognizability and complexity. Throughout the learning phase, the updated weights from each subset are transferred to the next one, thereby enhancing the model’s capability to extract more robust features. This enhancement significantly improves the model’s generalization and convergence rate. The effectiveness of the proposed method was evaluated by utilizing manually annotated data from three publicly accessible remote sensing datasets. The experimental results demonstrate a significant increase in accuracy when CTL is implemented with YOLOv7, with improvements of 21%, 45%, and 26% on the VEDAI-VISIBLE, VEDAI-IR, and DOTA datasets, respectively.
Yogendra Rao Musunuri, Oh-Seol Kwon, Sun-Yuan Kung
IEEE Geosci. Remote. Sens. Lett.3
2024 Weakly Supervised Video Re-Localization Through Multi-Agent-Reinforced Switchable Network
abstract
The objective of video re-localization (VRL) is to localize a successive sequence of frames, namely, the target moment, from untrimmed reference videos that semantically correspond to a given query video. During training, the weakly supervised setting of VRL provides only coarse-grained video-level rather than frame-level annotations. For the weakly supervised VRL (WS-VRL) task, obtaining effective video feature representations that can be used to evaluate the relevance between videos and localizing the accurate temporal boundaries of the target moment remain challenging. In this paper, a novel multi-agent-reinforced switchable network (MARS) is proposed to address these challenges. MARS can adaptively guide video feature encoding and moment localization using multiple learned agents. Specifically, an agent-controlled switchable encoder is used to obtain effective video feature representations, and an agent-reinforced boundary localizer is used to determine accurate localized moments through progressive refinement. Furthermore, a relevance-oriented reward generator was designed to estimate the relevance of the localized moment to the query video and assign a reward to multiple agents. The effectiveness of the proposed MARS model was verified through extensive experiments on the ActivityNet-VRL dataset.
Yuan Zhou 0006, Axin Guo, Shuwei Huo, Yu Liu 0004, Sun-Yuan Kung
IEEE Trans. Circuits Syst. Video Technol.5
2024 Geometric Variation Adaptive Network for Remote Sensing Image Change Detection
abstract
Change detection identifies surface changes on the earth by comparing two images from the same area at different times. To generate smooth change maps, a common method is fusing information from neighboring areas around each pixel. While the conventional fusion methods primarily rely on fixed regular-shaped neighboring areas, which may be inadequate in capturing the diverse and irregular geometric structures of changed ground objects. To address this limitation, we propose a novel Geometric Variation Adaptive Change Detector (GVA-CD), which adaptively adjusts the shape and size of neighboring areas based on the geometrical structure of ground objects. More specifically, we design a new geometric variation adaptive module (GVAM) as a component of GVA-CD. GVAM captures the structure of the ground objects to constructs geometrically flexible neighboring areas for each pixel, enabling the model to adapt to different ground object structures and generate discriminative difference features. We further propose a new difference measurement module to compute the difference between the features of pre-and post-change images by leveraging the adaptive neighboring areas. In addition, the GVA-CD introduces a multi-stage cross-scale fusion mechanism in both feature extraction and change map generation, to enhance the scale adaption ability of the feature extraction and change map generation. Extensive experiments on three large datasets demonstrate that our GVA-CD can outperform existing methods in change detection.
Shuwei Huo, Yuan Zhou 0006, Lei Zhang 0202, Yanjie Feng, Wei Xiang 0001, Sun-Yuan Kung
IEEE Trans. Geosci. Remote. Sens.6
2024 PSRNet: A Progressive Self-Refine Network for Lightweight Optical Remote Sensing Image Dehazing
abstract
In this article, we proposed a lightweight method for remote sensing image (RSI) dehazing, termed progressive self-refine network (PSR-Net). Image dehazing is an effective means to enhance the quality of images obscured by haze. Given that RSIs typically encompass extensive area with high resolution, RSI dehazing requires a lightweight model to cope with the complex RSIs. Existing methods focused mainly on improving dehazing performance without considering the computational overhead or model size. To remedy the deficiency, the proposed PSR-Net designs a lightweight framework incorporating a predehaze module (PDM) and a progressive self-refine module (PSRM). It first generates the initial dehazing results with low computation requirement, then iteratively refines the initial dehazing results by using the proposed restoration attention block (RAB). Meanwhile, considering the broad scope and varied object sizes inherent in RSIs, we designed a spatial adaptive feature extraction block which dynamically adjust its receptive field according to the input image. Extensive experiments are conducted to demonstrate that our proposed PSR-Net performs favorably against state-of-the-art methods on three popular RS image dehazing benchmark datasets.
Shuoshi Li, Yuan Zhou 0006, Sun-Yuan Kung
IEEE Trans. Geosci. Remote. Sens.3
2024 Dynamic View Aggregation for Multi-View 3D Shape Recognition
abstract
In the field of 3D shape recognition, the view-based approach has achieved state-of-the-art performance. A major challenge that needs to be addressed by the view-based approach is how to effectively aggregate multi-view features to obtain a better 3D shape representation. Existing methods which rely on networks with static parameters for feature aggregation adversely coerce the network to learn a general feature aggregation strategy for all inputs, ignoring the diversity of input 3D shapes in real-world scenarios. In this work, we propose a novelDynamic View Aggregation NetworkcalledDVA-Netto address this challenge. DVA-Net can dynamically adjust the network parameter depending on the input 3D shapes to flexibly fuse multi-view information. The shape-specific parameter adaptation is achieved by our designedDynamic Relation-aware Aggregationmodule, dubbedDRAmodule. It is responsible for learning relations among views and adaptively integrating multi-view features. Comprehensive experiments on benchmark datasets demonstrate that our proposed method achieves state-of-the-art performance for 3D shape classification and retrieval.
Yuan Zhou 0006, Zhongqi Sun, Shuwei Huo, Sun-Yuan Kung
IEEE Trans. Multim.4
2024 Automatic Metric Search for Few-Shot Learning
abstract
Few-shot learning (FSL) aims to learn a model that can identify unseen classes using only a few training samples from each class. Most of the existing FSL methods adopt a manually predefined metric function to measure the relationship between a sample and a class, which usually require tremendous efforts and domain knowledge. In contrast, we propose a novel model called automatic metric search (Auto-MS), in which an Auto-MS space is designed for automatically searching task-specific metric functions. This allows us to further develop a new searching strategy to facilitate automated FSL. More specifically, by incorporating the episode-training mechanism into the bilevel search strategy, the proposed search strategy can effectively optimize the network weights and structural parameters of the few-shot model. Extensive experiments on the miniImageNet and tieredImageNet datasets demonstrate that the proposed Auto-MS achieves superior performance in FSL problems.
Yuan Zhou 0006, Jieke Hao, Shuwei Huo, Boyu Wang 0004, Leijiao Ge, Sun-Yuan Kung
IEEE Trans. Neural Networks Learn. Syst.6
2023 Weakly-supervised content-based video moment retrieval using low-rank video representation
abstract
Content-based video moment retrieval (CVMR) aims to localize a successive sequence of frames in an untrimmed reference video, called target moment, that is semantically corresponding to a given query video. Current state-of-the-art CVMR methods are mainly developed using frame-level annotation, which is often quite expensive to collect. In this paper, we aim to develop a weakly-supervised CVMR method, which uses coarse-grained video-level annotations during training. Under weak supervision, video localizers require more discriminative frame-level video features. To achieve this goal, we proposed a novel prior, termed low-rank prior, based on an observation that the frame-level feature of a video should have low-rank properties. We demonstrated that the low-rank features are more discriminative and are beneficial to accurately localize the action boundaries. To produce a low-rank feature, we designed a low-rank feature reconstruction (LFR) operator. A new differentiable matrix decomposition approach is proposed to generate the low-rank reconstruction of the input matrix, meanwhile ensuring that the matrix decomposition process is differentiable. Based on the LFR, we developed a new weakly-supervised CVMR model which produces low-rank video representation and performs semantic consistency measures to discover the semantically matched segment in the reference video to the query video. Extensive experiments demonstrate that our method outperforms state-of-the-art weakly-supervised methods consistently and even achieves competing performance to fully-supervised baselines.
Shuwei Huo, Yuan Zhou 0006, Wei Xiang 0001, Sun-Yuan Kung
Knowl. Based Syst.4
2023 Progressive Learning for Unsupervised Change Detection on Aerial Images
abstract
This article focuses on unsupervised methods for optical aerial image change detection. Existing unsupervised change detection techniques are mainly categorized as patch-based methods and transfer-learning-based methods. However, the first type ignores the spatial information in the images, and the second type may introduce new errors due to knowledge extracted from additional datasets. To effectively tackle these problems, we propose an unsupervised progressive learning framework (UPLF). We first use original estimated change maps as the labeled samples and choose the reliable regions from samples to train the network. We then propose a progressive learning method to expand the reliable labeled region. Briefly, we apply a label selection filter to filter out incorrect change information from the regions to help rectify incorrect labeling in the regions. This leads to a more reliable labeled region and thus, in turn, more accurate detection results. Compared with the patch-based and transfer-learning-based unsupervised techniques, our method takes the entire map as the training sample to avoid the problem associated with using small patches; moreover, our iterative and progressive methods further enhance the change detection performance without involving external knowledge. Indeed, based on our experimental results on the real datasets, the proposed method demonstrates highly competitive performance compared with the state-of-the-art.
Yuan Zhou 0006, Xiangrui Li, Keran Chen, Sun-Yuan Kung
IEEE Trans. Geosci. Remote. Sens.4
2023 Semantic Relevance Learning for Video-Query Based Video Moment Retrieval
abstract
The task of video-query based video moment retrieval (VQ-VMR) aims to localize the segment in the reference video, which matches semantically with a short query video. This is a challenging task due to the rapid expansion and massive growth of online video services. With accurate retrieval of the target moment, we propose a new metric to effectively assess the semantic relevance between the query video and segments in the reference video. We also develop a new VQ-VMR framework to discover the intrinsic semantic relevance between a pair of input videos. It comprises two key components: a Fine-grained Feature Interaction (FFI) module and a Semantic Relevance Measurement (SRM) module. Together they can effectively deal with both the spatial and temporal dimensions of videos. First, the FFI module computes the semantic similarity between videos at a local frame level, mainly considering the spatial information in the videos. Subsequently, the SRM module learns the similarity between videos from a global perspective, taking into account the temporal information. We have conducted extensive experiments on two key datasets which demonstrate noticeable improvements of the proposed approach over the state-of-the-art methods.
Shuwei Huo, Yuan Zhou 0006, Ruolin Wang, Wei Xiang 0001, Sun-Yuan Kung
IEEE Trans. Multim.5
2023 Privacy Enhancing Machine Learning via Removal of Unwanted Dependencies
abstract
The rapid rise of IoT and Big Data has facilitated copious data-driven applications to enhance our quality of life. However, the omnipresent and all-encompassing nature of the data collection can generate privacy concerns. Hence, there is a strong need to develop techniques that ensure the data serve only the intended purposes, giving users control over the information they share. To this end, this article studies new variants of supervised and adversarial learning methods, which remove the sensitive information in the data before they are sent out for a particular application. The explored methods optimize privacy-preserving feature mappings and predictive models simultaneously in an end-to-end fashion. Additionally, the models are built with an emphasis on placing little computational burden on the user side so that the data can be desensitized on device in a cheap manner. Experimental results on mobile sensing and face datasets demonstrate that our models can successfully maintain the utility performances of predictive models while causing sensitive predictions to perform poorly.
Mert Al, Semih Yagli, Sun-Yuan Kung
IEEE Trans. Neural Networks Learn. Syst.3
2022 CHEX: CHannel EXploration for CNN Model Compression
abstract
Channel pruning has been broadly recognized as an effective technique to reduce the computation and memory cost of deep convolutional neural networks. However, conventional pruning methods have limitations in that: they are restricted to pruning process only, and they require a fully pre-trained large model. Such limitations may lead to sub-optimal model quality as well as excessive memory and training cost. In this paper, we propose a novel Channel Exploration methodology, dubbed as CHEX, to rectify these problems. As opposed to pruning-only strategy, we propose to repeatedly prune and regrow the channels throughout the training process, which reduces the risk of pruning important channels prematurely. More exactly: From intra-Layer's aspect, we tackle the channel pruning problem via a well-known column subset selection (CSS) formulation. From inter-Layer's aspect, our regrowing stages open a path for dynamically re-allocating the number of channels across all the layers under a global channel sparsity constraint. In addition, all the exploration process is done in a single training from scratch without the need of a pre-trained large model. Experimental results demonstrate that CHEX can effectively reduce the FLOPs of diverse CNN architectures on a variety of computer vision tasks, including image classification, object detection, instance segmentation, and 3D vision. For example, our compressed ResNet-50 model on ImageNet dataset achieves 76% top-l accuracy with only 25% FLOPs of the original ResNet-50 model, outperforming previous state-of-the-art channel pruning methods. The checkpoints and code are available at here.
Zejiang Hou, Minghai Qin, Fei Sun 0002, Kun Yuan 0001, Yi Xu 0008, Yen-Kuang Chen, Rong Jin 0001, Yuan Xie 0001, Sun-Yuan Kung
CVPR10
2022 3D-FM GAN: Towards 3D-Controllable Face Manipulation
Yuchen Liu 0002, Zhixin Shu, Yijun Li 0001, Zhe Lin 0001, Richard Zhang 0001, Sun-Yuan Kung
ECCV (15)6
2022 Evolving transferable neural pruning functions
abstract
Structural design of neural networks is crucial for the success of deep learning. While most prior works in evolutionary learning aim at directly searching the structure of a network, few attempts have been made on another promising track, channel pruning, which recently has made major headway in designing efficient deep learning models. In fact, prior pruning methods adopt human-made pruning functions to score a channel's importance for channel pruning, which requires domain knowledge and could be sub-optimal. To this end, we pioneer the use of genetic programming (GP) to discover strong pruning metrics automatically. Specifically, we craft a novel design space to express high-quality and transferable pruning functions, which ensures an end-to-end evolution process where no manual modification is needed on the evolved functions for their transferability after evolution. Unlike prior methods, our approach can provide both compact pruned networks for efficient inference and novel closed-form pruning metrics which are mathematically explainable and thus generalizable to different pruning tasks. While the evolution is conducted on small datasets, our functions shows promising results when applied to more challenging datasets, different from those used in the evolution process. For example, on ILSVRC-2012, an evolved function achieves state-of-the-art pruning results.
Yuchen Liu 0002, Sun-Yuan Kung, David Wentzlaff
GECCO2
2022 Multi-Dimensional Model Compression of Vision Transformer
abstract
Vision transformers (ViT) have recently attracted considerable attentions, but the huge computational cost remains an issue for practical deployment. Previous ViT pruning methods tend to prune the model along one dimension solely, which may suffer from excessive reduction and lead to sub-optimal model quality. In contrast, we advocate a multi-dimensional ViT compression paradigm, and propose to harness the redundancy reduction from attention head, neuron and sequence dimensions jointly. We firstly propose a statistical dependence based pruning criterion that is generalizable to different dimensions for identifying deleterious components. Moreover, we cast the multi-dimensional compression as an optimization, learning the optimal pruning policy across the three dimensions that maximizes the compressed model's accuracy under a computational budget. The problem is solved by our adapted Gaussian process search with expected improvement. Experimental results show that our method effectively reduces the computational cost of various ViT models. For example, our method reduces 40% FLOPs without top-1 accuracy loss for DeiT and T2T-ViT models, outperforming previous state-of-the-arts.
Zejiang Hou, Sun-Yuan Kung
ICME2
2022 Class-Discriminative CNN Compression
abstract
Compressing convolutional neural networks (CNNs) by pruning and distillation has received ever-increasing focus. In particular, designing a class-discrimination based approach would be desired as it fits seamlessly with the CNNs training objective. In this paper, we propose class-discriminative compression (CDC), which injects class discrimination in both pruning and distillation to facilitate the CNNs training goal. We first study the effectiveness of a group of discriminant functions for channel pruning, where we include well-known single-variate binary-class statistics like Student’s T-Test in our study via an intuitive generalization. We then propose a novel layer-adaptive hierarchical pruning approach, where we use a coarse class discrimination scheme for early layers and a fine one for later layers. This method naturally accords with the fact that CNNs process coarse semantics in the early layers and extract fine concepts at the later. Moreover, we leverage discriminant component analysis (DCA) to distill knowledge of intermediate representations in a subspace with rich discriminative information, which enhances hidden layers’ linear separability and classification accuracy of the student. Combining pruning and distillation, CDC is evaluated on CIFAR and ILSVRC-2012, where we consistently outperform the state-of-the-art results.
Yuchen Liu 0002, David Wentzlaff, Sun-Yuan Kung
ICPR3
2022 Multi-Dimensional Dynamic Model Compression for Efficient Image Super-Resolution
abstract
Modern single image super-resolution (SR) system based on convolutional neural networks achieves substantial progress. However, most SR deep networks are computationally expensive and require excessively large activation memory footprints, impeding their effective deployment to resource-limited devices. Based on the observation that the activation patterns in SR networks exhibit high input-dependency, we propose Multi-Dimensional Dynamic Model Compression method that can reduce both spatial and channel wise redundancy in an SR deep network for different input images. To reduce the spatial-wise redundancy, we propose to perform convolution on scaled-down feature-maps where the down-scaling factor is made adaptive to different input images. To reduce the channel-wise redundancy, we introduce a low-cost channel saliency predictor for each convolution to dynamically skip the computation of unimportant channels based on the Gumbel-Softmax. To better capture the feature-maps information and facilitate input-adaptive decision, we employ classic image processing metrics, e.g., Spatial Information, to guide the saliency predictors. The proposed method can be readily applied to a variety of SR deep networks and trained end-to-end with standard super-resolution loss, in combination with a sparsity criterion. Experiments on several benchmarks demonstrate that our method can effectively reduce the FLOPs of both lightweight and non-compact SR models with negligible PSNR loss. Moreover, our compressed models achieve competitive PSNR-FLOPs Pareto frontier compared with SOTA NAS-based SR methods.
Zejiang Hou, Sun-Yuan Kung
WACV2
2022 Cross-Scale Residual Network: A General Framework for Image Super-Resolution, Denoising, and Deblocking
abstract
In general, image restoration involves mapping from low-quality images to their high-quality counterparts. Such optimal mapping is usually nonlinear and learnable by machine learning. Recently, deep convolutional neural networks have proven promising for such learning processing. It is desirable for an image processing network to support well with three vital tasks, namely: 1) super-resolution; 2) denoising; and 3) deblocking. It is commonly recognized that these tasks have strong correlations, which enable us to design a general framework to support all tasks. In particular, the selection of feature scales is known to significantly impact the performance on these tasks. To this end, we propose the cross-scale residual network to exploit scale-related features among the three tasks. The proposed network can extract spatial features across different scales and establish cross-temporal feature reusage, so as to handle different tasks in a general framework. Our experiments show that the proposed approach outperforms state-of-the-art methods in both quantitative and qualitative evaluations for multiple image restoration tasks.
Yuan Zhou 0006, Xiaoting Du, Mingfei Wang, Shuwei Huo, Yeda Zhang, Sun-Yuan Kung
IEEE Trans. Cybern.6
2022 Exploiting Operation Importance for Differentiable Neural Architecture Search
abstract
Recently, differentiable neural architecture search (NAS) methods have made significant progress in reducing the computational costs of NASs. Existing methods search for the best architecture by choosing candidate operations with higher architecture weights. However, architecture weights cannot accurately reflect the importance of each operation, that is, the operation with the highest weight might not be related to the best performance. To circumvent this deficiency, we propose a novel indicator that can fully represent the operation importance and, thus, serve as an effective metric to guide the model search. Based on this indicator, we further develop a NAS scheme for "exploiting operation importance for effective NAS" (EoiNAS). More precisely, we propose a high-order Markov chain-based strategy to slim the search space to further improve search efficiency and accuracy. To evaluate the effectiveness of the proposed EoiNAS, we applied our method to two tasks: image classification and semantic segmentation. Extensive experiments on both tasks provided strong evidence that our method is capable of discovering high-performance architectures while guaranteeing the requisite efficiency during searching.
Yuan Zhou 0006, Xukai Xie, Sun-Yuan Kung
IEEE Trans. Neural Networks Learn. Syst.3
2022 XNAS: A Regressive/Progressive NAS for Deep Learning
abstract
Deep learning has achieved great and broad breakthroughs in many real-world applications. In particular, the task of training the network parameters has been masterly handled by back-propagation learning. However, the pursuit on optimal network structures remains largely an art of trial and error. This prompts some urgency to explore an architecture engineering process, collectively known as Neural Architecture Search (NAS). In general, NAS is a design software system for automating the search of effective neural architecture. This article proposes an X-learning NAS (XNAS) to automatically train a network’s structure and parameters. Our theoretical footing is built upon the subspace and correlation analyses between the input layer, hidden layer, and output layer. The design strategy hinges upon the underlying principle that the network should be coerced to learn how to structurally improvethe input/output correlation successively (i.e., layer by layer). It embraces both Progressive NAS (PNAS) and Regressive NAS (RNAS). For unsupervised RNAS, Principal Component Analysis (PCA) is a classic tool for subspace analyses. By further incorporating teacher’s guidance, PCA can be extended to Regression Component Analysis (RCA) to facilitate supervised NAS design. This allows the machine to extract components most critical to the targeted learning objective. We shall further extend the subspace analysis from multi-layer perceptrons to convolutional neural networks, via introduction of Convolutional-PCA (CPCA) or, more simply, Deep-PCA (DPCA). The supervised variant of DPCA will be named Deep-RCA (DRCA). The subspace analyses allow us to compute optimal eigenvectors (respectively, eigen-filters) and principal components (respectively, eigen-channels) for optimal NAS design of multi-layer perceptrons (respectively, convolutional neural networks). Based on the theoretical analysis, an X-learning paradigm is developed to jointly learn the structure and parameters of learning models. The objective is to reduce the network complexity while retaining (and sometimes improving) the performance. With carefully pre-selected baseline models, X-learning has shown great successes in numerous classification-type and/or regression-type applications. We have applied X-learning to the ImageNet datasets for classification and DIV2K for image enhancements. By applying X-learning to two types of baseline models, MobileNet and ResNet, both the low-power and high-performance application categories can be supported. Our simulations confirm that X-learning is by and large very competitive relative to the state-of-the-art approaches.
Sun-Yuan Kung
ACM Trans. Sens. Networks1
2021 Parameter Efficient Dynamic Convolution via Tensor Decomposition
Zejiang Hou, Sun-Yuan Kung
BMVC2
2021 Content-Aware GAN Compression
abstract
Generative adversarial networks (GANs), e.g., StyleGAN2, play a vital role in various image generation and synthesis tasks, yet their notoriously high computational cost hinders their efficient deployment on edge devices. Directly applying generic compression approaches yields poor results on GANs, which motivates a number of recent GAN compression works. While prior works mainly accelerate conditional GANs, e.g., pix2pix and Cycle-GAN, compressing state-of-the-art unconditional GANs has rarely been explored and is more challenging. In this paper, we propose novel approaches for unconditional GAN compression. We first introduce effective channel pruning and knowledge distillation schemes specialized for unconditional GANs. We then propose a novel content-aware method to guide the processes of both pruning and distillation. With content-awareness, we can effectively prune channels that are unimportant to the contents of interest, e.g., human faces, and focus our distillation on these regions, which significantly enhances the distillation quality. On StyleGAN2 and SN-GAN, we achieve a substantial improvement over the state-of-the-art compression method. Notably, we reduce the FLOPs of StyleGAN2 by 11× with visually negligible image quality loss compared to the full-size model. More interestingly, when applied to various image manipulation tasks, our compressed model forms a smoother and better disentangled latent manifold, making it more effective for image editing.
Yuchen Liu 0002, Zhixin Shu, Yijun Li 0001, Zhe Lin 0001, Federico Perazzi, Sun-Yuan Kung
CVPR6
2021 Meta-Learning with Attention for Improved Few-Shot Learning
abstract
We consider few-shot learning (FSL), where a model learns from very few labeled examples such that it can generalize to unseen examples. Model-agnostic meta-learning (MAML) has been proposed to solve FSL. However, the low performance of MAML suggests its difficulty in tackle diverse tasks, due to the restriction of sharing a single model initialization for fast adaptation. In this paper, we propose meta-learning with attention mechanisms. Our method meta-learns attention modules to instantiate task-specific model initialization for fast adaptation, which can obtain high-quality solution to a new task using few gradient descent steps. To further improve generalization during inference, we propose to incorporate an entropy regularizer into the adaptation objective to penalize the Shannon entropy of prediction probability. Extensive experiments under various FSL scenarios show that our method achieves state-of-the-art performance on the mini-ImageNet and tiered-ImageNet.
Zejiang Hou, Anwar Elwalid, Sun-Yuan Kung
ICASSP3
2021 Shape autotuning activation function
Yuan Zhou 0006, Shuwei Huo, Sun-Yuan Kung
Expert Syst. Appl.4
2021 Image super-resolution based on dense convolutional auto-encoder blocks
Yuan Zhou 0006, Yeda Zhang, Xukai Xie, Sun-Yuan Kung
Neurocomputing4
2021 Guest editorial: Special issue on next generation AI approaches for intelligent internet of things
Meikang Qiu, Sun-Yuan Kung
J. Syst. Archit.2
2021 Adversarial Learning for Multiscale Crowd Counting Under Complex Scenes
abstract
In this article, a multiscale generative adversarial network (MS-GAN) is proposed for generating high-quality crowd density maps of arbitrary crowd density scenes. The task of crowd counting has many challenges, such as severe occlusions in extremely dense crowd scenes, perspective distortion, and high visual similarity between the pedestrians and background elements. To address these problems, the proposed MS-GAN combines a multiscale convolutional neural network (generator) and an adversarial network (discriminator) to generate a high-quality density map and accurately estimate the crowd count in complex crowd scenes. The multiscale generator utilizes the fusion features from multiple hierarchical layers to detect people with large-scale variation. The resulting density map produced by the multiscale generator is processed by a discriminator network trained to solve a binary classification task between a poor quality density map and real ground-truth ones. The additional adversarial loss can improve the quality of the density map, which is critical to accurately estimate the crowd counts. The experiments were conducted on multiple datasets with different crowd scenes and densities. The results showed that the proposed method provided better performance compared to current state-of-the-art methods.
Yuan Zhou 0006, Jianxing Yang, Sun-Yuan Kung
IEEE Trans. Cybern.5
2021 Temporal Action Localization Using Long Short-Term Dependency
abstract
Temporal action localization in untrimmed videos is an important but difficult task. Difficulties are encountered in the application of existing methods when modeling the temporal structures of videos. In the present study, we develop a novel method, referred to as the Gemini Network, for effective modeling of temporal structures and achieving high-performance temporal action localization. The significant improvements afforded by the proposed method are due to three major factors. First, temporal dependencies are explicitly distinguished as long-term temporal dependencies and short-term temporal dependencies and are separately captured by two dedicated subnets. Second, a long-range temporal dependency capture module combined with a self-adaptive pooling module is proposed to capture long-term temporal dependency. Third, the proposed method uses auxiliary supervision, with the auxiliary classifier losses affording additional constraints for improving the modeling capability of the network. As a demonstration of its effectiveness, the Gemini Network is used to achieve a state-of-the-art temporal action localization performance on two challenging datasets, namely, THUMOS14 and ActivityNet.
Yuan Zhou 0006, Ruolin Wang, Sun-Yuan Kung
IEEE Trans. Multim.4
2020 Scalable Kernel Learning Via the Discriminant Information
abstract
Kernel approximation methods create explicit, low-dimensional kernel feature maps to deal with the high computational and memory complexity of standard techniques. This work studies a supervised kernel learning methodology to optimize such mappings. We utilize the Discriminant Information criterion, a measure of class separability with a strong connection to Discriminant Analysis. By generalizing this measure to cover a wider range of kernel maps and learning settings, we develop scalable methods to learn kernel features with high discriminant power. Experimental results on several datasets showcase that our techniques can improve optimization and generalization performances over state of the art kernel learning methods.
Mert Al, Zejiang Hou, Sun-Yuan Kung
ICASSP3
2020 Efficient Image Super Resolution Via Channel Discriminative Deep Neural Network Pruning
abstract
Deep convolutional neural networks (CNN) have demonstrated superior performance in image super-resolution (SR) problem. However, CNNs are known to be heavily over-parameterized, and suffer from abundant redundancy. The growing size of CNNs may be incompatible with their deployment on mobile or embedded devices. Network pruning has benefited classification tasks by removing redundant parameters and associated computation. However, it has rarely been studied for SR, because existing methods assume the channel-wise features are of equal importance to the final reconstruction. On the contrary, we show the existence of uninformative feature-maps with no contribution to the task. In order to identify and remove such uninformative channels, we propose a new pruning criterion, Discriminant Information, by characterizing the dependency of the output w.r.t to the hidden-layer feature-maps. Empirically, our DI-based channel pruning algorithm is able to trim the state-of-the-art SR networks significantly (e.g. 8.7x model size compression and 3.6x CPU acceleration on SRResNet), with no quantitative or visual performance loss.
Zejiang Hou, Sun-Yuan Kung
ICASSP2
2020 Exploring Highly Efficient Compact Neural Networks For Image Classification
abstract
Group convolution works well with many lightweight convolutional neural networks (CNNs) that can effectively reduce the number of parameters and computational cost. However, feature maps of different groups cannot communicate, which restricts their representation capability. To address this issue, in this work, we propose a novel convolution operation named Hierarchical Group Convolution (HGC) for creating computationally efficient neural networks. Different from standard group convolution which blocks the inter-group information exchange and induces the severe performance degradation, HGC hierarchically fuses the feature maps from each group and leverages the inter-group information effectively. Taking advantage of the proposed operation, we introduce an efficient compact network named HGCNet. Extensive experimental results on image classification task demonstrate that HGCNet obtain significant reduction of computational cost and the number of parameters, while achieving comparable performance over the prior CNN architectures designed for mobile devices.
Xukai Xie, Yuan Zhou 0006, Sun-Yuan Kung
ICIP3
2020 Hierarchically Aggregated Residual Transformation for Single Image Super Resolution
abstract
Visual patterns usually appear at different scales/sizes in natural images. Multi-scale feature representation is of great importance for the single-image super-resolution (SISR) task to reconstruct image objects at different scales. However, such characteristic has been rarely considered by CNN-based SISR methods. In this work, we propose a novel building block, i.e. hierarchically aggregated residual transformation (HART), to achieve multi-scale feature representation in each layer of the network. Within each HART block, we connect multiple convolutions in a hierarchical residual-like manner, which efficiently provides a wide range of effective receptive fields at a more granular level to detect both local and global image features. To theoretically understand the proposed HART block, we recast SISR as an optimal control problem and show that HART effectively approximates the classical 4th-order Runge-Kutta method, which has the merit of small local truncation error for solving numerical ordinary differential equation. By cascading the proposed HART blocks, we establish our high-performing HARTnet. Through extensive experiments on various benchmark datasets under different degradation models, we demonstrate that HARTnet compares favourably against existing state-of-the-art methods (including those in the NTIRE 2019 SR Challenge leaderboard) in terms of both quantitative metrics and visual quality. Moreover, the same HARTnet architecture achieves promising performance on such other image restoration tasks as image denoising and low-light image enhancement.
Zejiang Hou, Sun-Yuan Kung
ICPR2
2020 A Discriminant Information Approach to Deep Neural Network Pruning
abstract
Network pruning has become the de facto tool to accelerate deep neural networks for mobile and edge applications. Recently, feature-map discriminant based channel pruning has shown promising results, as it aligns well with the CNN's objective of differentiating multiple classes and offers better interpretability of the pruning decision. However, existing discriminant-based methods are challenged by computation inefficiency, as there is a lack of theoretical guidance on quantifying the feature-map discriminant power. In this paper, we develop a mathematical formulation to accurately and efficiently quantify the feature-map discriminativeness, which gives rise to a novel criterion, Discriminant Information (DI). We analyze the theoretical property of DI, specifically the non-decreasing property, that makes DI a valid channel selection criterion. By measuring the differential discriminant, we can identify and remove those channels with minimum influence to the discriminant power. The versatility of DI criterion also enables an intra-layer mixed precision quantization to further compress the network. Moreover, we propose a DI-based greedy pruning algorithm and structure distillation technique to automatically decide the pruned structure that satisfies certain resource budget, which is a common requirement in reality. Extensive experiments demonstrate the effectiveness of our method: our pruned ResNet50 on ImageNet achieves 44% FLOPs reduction without any Top-1 accuracy loss compared to unpruned model.
Zejiang Hou, Sun-Yuan Kung
ICPR2
2020 Intelligent security and optimization in Edge/Fog Computing
Meikang Qiu, Sun-Yuan Kung, Keke Gai
Future Gener. Comput. Syst.2
2020 Adaptive Irregular Graph Construction-Based Salient Object Detection
abstract
Saliency detection represents a vital pre-processing stage of computer vision. Most existing propagation-based salient object detection methods construct a k-regular graph for saliency propagation. Applying a regular graph to a vast smooth region is potentially prone to unnecessary or prolonged propagation errors, leading to the excessive highlighting of the background regions. To mitigate such problems, we substitute the conventional k-regular graph with an adaptive irregular graph for saliency value propagation, thereby avoiding unnecessary iterations over a vast smooth region. We first perform a clustering analysis based on the smoothness, color, and other features of regions. The new graph boosts an adaptive link density by considering the clustering result. In addition, we propose a seeding strategy for the propagation. Based on our experimental studies of six major benchmark datasets, our method performed favorably against the other state-of-the-art methods, both quantitatively and qualitatively.
Yuan Zhou 0006, Shuwei Huo, Chunping Hou, Sun-Yuan Kung
IEEE Trans. Circuits Syst. Video Technol.5
2019 Methodical Design and Trimming of Deep Learning Networks: Enhancing External BP Learning with Internal Omnipresent-supervision Training Paradigm
abstract
Back-propagation (BP) is now a classic learning paradigm whose source of supervision is exclusively from the external (input/output) nodes. Consequently, BP is easily vulnerable to curse-of-depth in (very) Deep Learning Networks (DLNs). This prompts us to advocate Internal Neuron’s Learnablility (INL) with (1)internal teacher labels (ITL); and (2)internal optimization metrics (IOM) for evaluating hidden layers/nodes. Conceptually, INL is a step beyond the notion of Internal Neuron’s Explainablility (INE), championed by DARPA’s XAI (or AI3.0). Practically, INL facilitates a structure/parameter NP-iterative learning for (supervised) deep compression/quantization: simultaneously trimming hidden nodes and raising accuracy. Pursuant to our simulations, the NP-iteration appears to outperform several prominent pruning methods in the literature.
Sun-Yuan Kung, Zejiang Hou, Yuchen Liu 0002
ICASSP1
2019 A Kernel Discriminant Information Approach to Non-linear Feature Selection
abstract
Feature selection has become a de facto tool for analyzing high-dimensional data, especially in bioinformatics. It is effective in improving learning algorithms' scalability and facilitating feature generalization or interpretability by removing noise and redundancy. Our focus is placed on the paradigm of supervised feature selection, which aims to find an optimal feature subset to best predict the target. We propose a nonlinear approach for finding a feature subset that achieves the highest inter-class separability in terms of the kernel Discriminant Information (KDI) measure. Theoretically, we prove the existence of good prediction hypotheses for feature subsets with high KDI value. We also establish the equivalency between maximizing the KDI statistic and minimizing a functional dependency measure of label variable on data. Moreover, we asymptotically prove the concentration property of the optimal feature subset found by maximizing the KDI measure. Practically, we provide an efficient gradient optimization algorithm for solving the KDI feature selection problem. We evaluate the proposed method based on 19 benchmark datasets in various domains, and demonstrates a noticeable improvement against state-of-the-art baselines on the majority of classification and regression tasks. Notably, our method is robust to the choice of hyper-parameters, works well with various downstream classifiers, has competitive computational complexity among the kernel based methods considered, and scales well the large-scale object recognition dataset, with generalization enhancement on CIFAR.
Zejiang Hou, Sun-Yuan Kung
IJCNN2
2019 Semi-Supervised Salient Object Detection Using a Linear Feedback Control System Model
abstract
To overcome the challenging problems in saliency detection, we propose a novel semi-supervised classifier which makes good use of a linear feedback control system (LFCS) model by establishing a relationship between control states and salient object detection. First, we develop a boundary homogeneity model to estimate the initial saliency and background likelihoods, which are regarded as the labeled samples in our semi-supervised learning procedure. Then in order to allocate an optimized saliency value to each superpixel, we present an iterative semi-supervised learning framework which integrates multiple saliency cues and image features using an LFCS model. Via an innovative iteration method, the system gradually converges an optimized stable state, which is associating with an accurate saliency map. This paper also covers comprehensive simulation study based on public datasets, which demonstrates the superiority of the proposed approach.
Yuan Zhou 0006, Shuwei Huo, Wei Xiang 0001, Chunping Hou, Sun-Yuan Kung
IEEE Trans. Cybern.5
2019 Salient Object Detection via Fuzzy Theory and Object-Level Enhancement
abstract
This paper proposes a bottom-up saliency detection method via effective integration of regional saliency measure and object-level information using fuzzy theory. First, we generate an initial saliency map by fusing multiple prior maps. Second, to emphasize the object-level concept of saliency, we further generate many object proposals of the input image. A fuzzy set theory is then applied to measure the objectness score of the object proposals and integrate them into an objectness map. Third, an optimization framework is proposed to effectively fuse various prior saliency cues and object-level information to produce a clean and uniform saliency map as well as to maintain the salient object completeness. Experimental studies in several benchmark datasets confirmed the superiority of the proposed method over state-of-the-art saliency detection methods.
Yuan Zhou 0006, Ailing Mao, Shuwei Huo, Jianjun Lei 0001, Sun-Yuan Kung
IEEE Trans. Multim.5
2019 Semisupervised Learning Based on a Novel Iterative Optimization Model for Saliency Detection
abstract
In this paper, we propose a novel iterative optimization model for bottom-up saliency detection. By exploring bottom-up saliency principles and semisupervised learning approaches, we design a high-performance saliency analysis method for wide ranging scenes. The proposed algorithm consists of two stages: 1) we develop a boundary homogeneity model to characterize the general position and the contour of the salient objects and 2) we propose a novel iterative optimization model, termed gradual saliency optimization, for further performance improvement. Our main contribution falls on the second stage, where we propose an iterative framework with self-repairing mechanisms for refining saliency maps. In this framework, we further develop a more comprehensive optimization function applying a novel semisupervised learning scheme to enhance the traditional saliency measure. More elaborately, the iterative method can gradually improve the output in each iteration and finally converge to high-quality saliency maps. Based on our experiments on four different public data sets, it can be demonstrated that our approach significantly outperforms the state-of-the-art methods.
Shuwei Huo, Yuan Zhou 0006, Wei Xiang 0001, Sun-Yuan Kung
IEEE Trans. Neural Networks Learn. Syst.4
2019 Editorial: IEEE Transactions on Sustainable Computing, Special Issue on Smart Data and Deep Learning in Sustainable Computing
abstract
The twelve papers in this special section focus on smart data and deep learning in sustainable computing. We are living in a data-driven era in which numerous infrastructure can be connected and the interconnected systems can perform “smart” when the large pool of the data are well utilized. Finding the way of well utilizing the large volume of data has an urgent demand in multiple realms, including academics, industries, and education. The force behind the data can be pushed out from a variety of data-driven techniques, such as machine learning and deep learning, which is a great potential for generating successful model, framework, and method for achieving sustainable computing. Therefore, gathering recent achievements in smart data and deep learning in sustainable computing is meaningful and valuable for powering the capability of datadriven domain and the various applications, implementations, and innovations in different disciplines and fields. This special issue focuses on two aspects considering the perspective of sustainable computing, which include smart data and deep learning. The smart data covers all dimensions of data usage lifecycles, such as data selections and collections, data preprocessing, data mining, and data analytics, in various application scenarios. The other aspect, deep learning, emphasizes the intelligent performance of applying data-driven techniques in practices and research explorations. Thus, this special issue aims at collecting updated outstanding papers that illustrate the latest achievements and development updates concerning the smart data and deep learning solutions, issues, applications, trends, and implementations in sustainable computing.
Meikang Qiu, Sun-Yuan Kung, Qing Yang 0001
IEEE Trans. Sustain. Comput.2
2019 Editorial: IEEE Transactions on Sustainable Computing, Special Issue on Secure Sustainable Green Smart Computing
abstract
The booming development of cloud computing has resulted in a remarkable growth of multiple industries in which dramatic demands of green computing and sustainability are addressed. The platform of smart computing has provided an efficient approach for connecting various infrastructure such that many new technologies are eventually formed, such as Internet-of-Thing and ubiquitous computing. The concept of sustainable green smart computing has become a significant issue for those enterprises or practitioners who are engaging the implementations of smart computing aiming at a longer-term strategy. Considering the achievement of the real sustainability, one of the crucial values is to ensure all operations across different computing sources are under a secure executive environment. For reaching a high performance of securing sustainable green smart computing, many problems need to be solved. For example, one of the main challenges is to balance the costs among security, energy, performance, and sustainable requirements. The distribution of the computing resources in this issue is a great challenge because the real-time executions are usually constrained by multiple elements. An efficient approach of providing an adaptive and scalable service as well as addressing sustainability is an urgent research direction for current advanced cloud computing applications. Thus, this special issue aims at collecting updated outstanding papers that illustrate the latest achievements and development updates concerning the security solutions, issues, applications, trends, and implementations in sustainable green smart computing.
Meikang Qiu, Sun-Yuan Kung, Qing Yang 0001
IEEE Trans. Sustain. Comput.2
2018 Multi-Kernel, Deep Neural Network and Hybrid Models for Privacy Preserving Machine Learning
abstract
The rapid rise of IoT and Big Data can facilitate the use of data to enhance our quality of life. However, the omnipresent and sensitive nature of data can simultaneously generate privacy concerns. Hence, there is a strong need to develop techniques that ensure the data serve the intended purposes, but not for prying into one's sensitive information. We address this challenge via utility maximizing lossy compression of data. Our techniques combine the mathematical rigor of Kernel Learning models with the structural richness of Deep Neural Networks, and lead to the novel Multi-Kernel Learning and Hybrid Learning models. We systematically construct the proposed models in progressive stages, as motivated by the cumulative improvement in the experimental results from the two previously non-intersecting regimes, namely, Kernel Learning and Deep Neural Networks. The final experimental results of the three proposed models on three mobile sensing datasets show that, not only are our methods able to improve the utility prediction accuracies, but they can also cause sensitive predictions to perform nearly as bad as random guessing, resulting in a win-win situation in terms of utility and privacy.
Mert Al, Thee Chanyaswad, Sun-Yuan Kung
ICASSP3
2018 Outlier Removal for Enhancing Kernel-Based Classifier Via the Discriminant Information
abstract
Pattern recognition on big data can be challenging for kernel machines as the complexity grows with the squared number of training samples. In this work, we overcome this hurdle via the outlying data sample removal pre-processing step. This approach removes less-informative data samples and trains the kernel machines only with the remaining data, and hence, directly reduces the complexity by reducing the number of training samples. To enhance the classification performance, the outlier removal process is done such that the discriminant information of the data is mostly intact. This is achieved via the novel Outlier-Removal Discriminant Information (ORDI) metric, which measures the contribution of each sample toward the discriminant information of the dataset. Hence, the ORDI metric can be used together with the simple filter method to effectively remove insignificant outliers to both reduce the computational cost and enhance the classification performance. We experimentally show on two real-world datasets at the sample removal ratio of 0.2 that, with outlier removal via ORDI, we can simultaneously (1) improve the accuracy of the classifier by 1 %, and (2) provide significant saving on the total running time by 1.5x and 2x on the two datasets. Hence, ORDI can provide a win-win situation in this performance-complexity tradeoff of the kernel machines for big data analysis.
Thee Chanyaswad, Mert Al, Sun-Yuan Kung
ICASSP3
2018 Cross-Layer Design for Network Lifetime Maximization in Underwater Wireless Sensor Networks
abstract
This paper investigates the cross-layer design problem with the goal of maximizing the network lifetime for energy-constrained underwater wireless sensor networks (UWSNs). We first jointly consider link scheduling, transmission power and transmission rate in a proposed optimization problem with the adoption of time division multiple access (TDMA) schedules. Then, we propose an iterative algorithm to solve the optimization problem. It alternates between (1) link scheduling and (2) computation of transmission powers and transmission rates. In fact, the convergence of such iterative algorithm can be mathematically and empirically supported. We evaluate our algorithm for several network topologies. Extensive simulation results demonstrate the superiority of the proposed approach.
Yuan Zhou 0006, Yu Hen Hu, Boyu Wang 0004, Sun-Yuan Kung
ICC5
2018 Super-resolution Imaging Based on Global Interpolation and Structural Similarities
abstract
In this paper, we propose a double dictionary learning method for image super-resolution (SR) reconstruction. Different from existing dictionary learning based super-resolution, we combine both self-similarity and external images to construct a double dictionary learning method. A new optimization model is established using self-similarities and external-similarities as regularization terms. Furthermore, we propose a global interpolation method to reconstruct an accurate initial estimation at the edges. Experimental results show that the proposed algorithm can produce high-quality reconstruction results both perceptually and quantitatively in terms of peak signal-to-noise ratio (PSNR) and structural similarity (SSIM), as compared to existing algorithms.
Yuan Zhou 0006, Shuwei Huo, Sun-Yuan Kung
ICPR4
2018 Multi-scale Generative Adversarial Networks for Crowd Counting
abstract
We investigate generative adversarial networks as an effective solution to the crowd counting problem. These networks not only learn the mapping from crowd image to corresponding density map, but also learn a loss function to train this mapping. There are many challenges to the task of crowd counting, such as severe occlusions in extremely dense crowd scenes, perspective distortion, and high visual similarity between pedestrians and background elements. To address these problems, we proposed multi-scale generative adversarial network to generate high-quality crowd density maps of arbitrary crowd density scenes. We utilized the adversarial loss from discriminator to improve the quality of the estimated density map, which is critical to accurately predict crowd counts. The proposed multi-scale generator can extract multiple hierarchy features from the crowd image. The results showed that the proposed method provided better performance compared to current state-of-the-art methods.
Jianxing Yang, Yuan Zhou 0006, Sun-Yuan Kung
ICPR3
2018 Blurred Image Region Detection based on Stacked Auto-Encoder
abstract
In this study, we address a fundamental yet challenging problem on detection and classification of blurred regions in partially blurred images. we propose to learn a latent feature representation with stacked auto-encoder (SAE) network to perform blur region detection. Most previous approaches focus on extracting a few blur features in image gradient, Fourier domain, and data-driven local filters. We extract a latent high-level feature representation from such low-level features using the stacked auto-encoder network, thereby improve the accuracy of blur region classification. This high accuracy enables us to successfully separate the clear and blurred regions. Experimental results demonstrate that the proposed method significantly outperforms the state-of-the-arts methods in detecting and classifying blur regions in partially blurred images.
Yuan Zhou 0006, Jianxing Yang, Sun-Yuan Kung
ICPR4
2018 Guest Editor's Introduction to the Special Issue on Security and Privacy on Clouds
abstract
The fifteen papers in this special section focus on cloud computer security and privacy. The emerging paradigm of cloud computing provides a new way to address the constraints of limited energy, capabilities, and resources. Researchers and practitioners have embraced cloud computing as a new approach that has the potential for a profound impact in our daily life and world economy. However, security and privacy protection is a critical concern in the development and adoption of cloud computing. To avoid system fragility and defend against vulnerabilities exploration from cyber attacker, various cyber security techniques and tools have been developed for cloud systems.
Meikang Qiu, Sun-Yuan Kung
IEEE Trans. Cloud Comput.2
2018 CLAss-Specific Subspace Kernel Representations and Adaptive Margin Slack Minimization for Large Scale Classification
abstract
In kernel-based classification models, given limited computational power and storage capacity, operations over the full kernel matrix becomes prohibitive. In this paper, we propose a new supervised learning framework using kernel models for sequential data processing. The framework is based on two components that both aim at enhancing the classification capability with a subset selection scheme. The first part is a subspace projection technique in the reproducing kernel Hilbert space using a CLAss-specific Subspace Kernel representation for kernel approximation. In the second part, we propose a novel structural risk minimization algorithm called the adaptive margin slack minimization to iteratively improve the classification accuracy by an adaptive data selection. We motivate each part separately, and then integrate them into learning frameworks for large scale data. We propose two such frameworks: the memory efficient sequential processing for sequential data processing and the parallelized sequential processing for distributed computing with sequential data acquisition. We test our methods on several benchmark data sets and compared with the state-of-the-art techniques to verify the validity of the proposed techniques.
Yinan Yu, Konstantinos I. Diamantaras, Tomas McKelvey, Sun-Yuan Kung
IEEE Trans. Neural Networks Learn. Syst.4
2017 Salient object detection via a linear feedback control system
abstract
Linear feedback control systems (LFCS) have been widely applied in signal analysis, filtering, and error correction. Many functional properties of LFCS are amenable to numerous object recognition and detection tasks. In fact, there exists an intimate relationship between control states and salient values. This prompts us to adopt the linear feedback control system to detect salient object in static images. Via an innovative iteration method, the system gradually converges an optimized stable state, which is associating with an accurate saliency map. In addition, to initialize the system, we propose the so called boundary homogeneity based on a priori knowledge on the boundary to estimate the background likelihood and indirectly depict a foreground (saliency) map. By our experimental results, we demonstrates that such feedback control model can bring about noticable improvement in salient object detection.
Shuwei Huo, Yuan Zhou 0006, Sun-Yuan Kung
ICIP3
2017 A compressive multi-kernel method for privacy-preserving machine learning
abstract
As the analytic tools become more powerful, and more data are generated on a daily basis, the issue of data privacy arises. This leads to the study of the design of privacy-preserving machine learning algorithms. Given two objectives, namely, utility maximization and privacy-loss minimization, this work is based on two previously non-intersecting regimes - Compressive Privacy and multi-kernel method. Compressive Privacy is a privacy framework that employs utility-preserving lossy-encoding scheme to protect the privacy of the data, while multi-kernel method is a kernel-based machine learning regime that explores the idea of using multiple kernels for building better predictors. In relation to the neural-network architecture, multi-kernel method can be described as a two-hidden-layered network with its width proportional to the number of kernels. The compressive multi-kernel method proposed consists of two stages - the compression stage and the multi-kernel stage. The compression stage follows the Compressive Privacy paradigm to provide the desired privacy protection. Each kernel matrix is compressed with a lossy projection matrix derived from the Discriminant Component Analysis (DCA). The multikernel stage uses the signal-to-noise ratio (SNR) score of each kernel to non-uniformly combine multiple compressive kernels. The proposed method is evaluated on two mobile-sensing datasets - MHEALTH and HAR - where activity recognition is defined as utility and person identification is defined as privacy. The results show that the compression regime is successful in privacy preservation as the privacy classification accuracies are almost at the random-guess level in all experiments. On the other hand, the novel SNR-based multi-kernel shows utility classification accuracy improvement upon the state-of-the-art in both datasets. These results indicate a promising direction for research in privacy-preserving machine learning.
Thee Chanyaswad, J. Morris Chang, Sun-Yuan Kung
IJCNN3
2017 Semi-supervised saliency classifier based on a linear feedback control system model
abstract
Linear feedback control systems (LFCS) are amenable to numerous object recognition and detection tasks on account of its functional properties in signal filtering and error correction. In fact, there exists an intimate relationship between control states and salient values. Therefore, we propose a novel semi-supervised classifier which makes use of linear feedback control theory to improve saliency detection performance. First, we develop a boundary homogeneity model to estimate the initial saliency and background likelihoods, which may lead to the labeled samples in our semi-supervised learning procedure. Then in order to allocate an optimized saliency value to each superpixel, we present an iterative semi-supervised learning framework which integrates multiple saliency cues and image features using a LCSF model. Via an innovative iteration method, the system gradually converges an optimized stable state, which is associating with an accurate saliency map. Based on our experiments on public datasets, it can be demonstrated that our approach significantly outperforms the state-of-the-art methods.
Shuwei Huo, Yuan Zhou 0006, Sun-Yuan Kung
IJCNN3
2017 FUEL-mLoc: feature-unified prediction and explanation of multi-localization of cellular proteins in multiple organisms
abstract
Although many web-servers for predicting protein subcellular localization have been developed, they often have the following drawbacks: (i) lack of interpretability or interpreting results with heterogenous information which may confuse users; (ii) ignoring multi-location proteins and (iii) only focusing on specific organism. To tackle these problems, we present an interpretable and efficient web-server, namely FUEL-mLoc, using eature- nified prediction and xplanation of m ulti- oc alization of cellular proteins in multiple organisms. Compared to conventional localization predictors, FUEL-mLoc has the following advantages: (i) using unified features (i.e. essential GO terms) to interpret why a prediction is made; (ii) being capable of predicting both single- and multi-location proteins and (iii) being able to handle proteins of multiple organisms, including Eukaryota, Homo sapiens, Viridiplantae, Gram-positive Bacteria, Gram-negative Bacteria and Virus . Experimental results demonstrate that FUEL-mLoc outperforms state-of-the-art subcellular-localization predictors. Availability and Implementation: http://bioinfo.eie.polyu.edu.hk/FUEL-mLoc/. Contacts: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Shibiao Wan, Man-Wai Mak, Sun-Yuan Kung
Bioinform.3
2017 Discriminant component analysis for privacy protection and visualization of big data
Sun-Yuan Kung
Multim. Tools Appl.1
2017 Transductive Learning for Multi-Label Protein Subchloroplast Localization Prediction
abstract
Predicting the localization of chloroplast proteins at the sub-subcellular level is an essential yet challenging step to elucidate their functions. Most of the existing subchloroplast localization predictors are limited to predicting single-location proteins and ignore the multi-location chloroplast proteins. While recent studies have led to some multi-location chloroplast predictors, they usually perform poorly. This paper proposes an ensemble transductive learning method to tackle this multi-label classification problem. Specifically, given a protein in a dataset, its composition-based sequence information and profile-based evolutionary information are respectively extracted. These two kinds of features are respectively compared with those of other proteins in the dataset. The comparisons lead to two similarity vectors which are weighted-combined to constitute an ensemble feature vector. A transductive learning model based on the least squares and nearest neighbor algorithms is proposed to process the ensemble features. We refer to the resulting predictor to as EnTrans-Chlo. Experimental results on a stringent benchmark dataset and a novel dataset demonstrate that EnTrans-Chlo significantly outperforms state-of-the-art predictors and particularly gains more than 4% (absolute) improvement on the overall actual accuracy. For readers' convenience, EnTrans-Chlo is freely available online at http://bioinfo.eie.polyu.edu.hk/EnTransChloServer/.
Shibiao Wan, Man-Wai Mak, Sun-Yuan Kung
IEEE ACM Trans. Comput. Biol. Bioinform.3
2017 Cost-Effective Kernel Ridge Regression Implementation for Keystroke-Based Active Authentication System
abstract
In this paper, a fast kernel ridge regression (KRR) learning algorithm is adopted with ( ) training cost for large-scale active authentication system. A truncated Gaussian radial basis function (TRBF) kernel is also implemented to provide better cost-performance tradeoff. The fast-KRR algorithm along with the TRBF kernel offers computational advantages over the traditional support vector machine (SVM) with Gaussian-RBF kernel while preserving the error rate performance. Experimental results validate the cost-effectiveness of the developed authentication system. In numbers, the fast-KRR learning model achieves an equal error rate (EER) of 1.39% with ( ) training time, while SVM with the RBF kernel shows an EER of 1.41% with ( ) training time.
Pei Yuan Wu, Chi-Chen Fang, J. Morris Chang, Sun-Yuan Kung
IEEE Trans. Cybern.4
2017 Collaborative PCA/DCA Learning Methods for Compressive Privacy
abstract
In the Internet era, the data being collected on consumers like us are growing exponentially, and attacks on our privacy are becoming a real threat. To better ensure our privacy, it is safer to let the data owner control the data to be uploaded to the network as opposed to taking chance with data servers or third parties. To this end, we propose compressive privacy , a privacy-preserving technique to enable the data creator to compress data via collaborative learning so that the compressed data uploaded onto the Internet will be useful only for the intended utility and not be easily diverted to malicious applications. For data in a high-dimensional feature vector space, a common approach to data compression is dimension reduction or, equivalently, subspace projection. The most prominent tool is principal component analysis (PCA). For unsupervised learning, PCA can best recover the original data given a specific reduced dimensionality. However, for the supervised learning environment, it is more effective to adopt a supervised PCA, known as discriminant component analysis (DCA), to maximize the discriminant capability. The DCA subspace analysis embraces two different subspaces. The signal-subspace components of DCA are associated with the discriminant distance/power (related to the classification effectiveness), whereas the noise subspace components of DCA are tightly coupled with recoverability and/or privacy protection. This article presents three DCA-related data compression methods useful for privacy-preserving applications: — Utility-driven DCA : Because the rank of the signal subspace is limited by the number of classes, DCA can effectively support classification using a relatively small dimensionality (i.e., high compression). — Desensitized PCA : By incorporating a signal-subspace ridge into DCA, it leads to a variant especially effective for extracting privacy-preserving components. In this case, the eigenvalues of the noise-space are made to become insensitive to the privacy labels and are ordered according to their corresponding component powers. — Desensitized K-means/SOM : Since the revelation of the K-means or SOM cluster structure could leak sensitive information, it is safer to perform K-means or SOM clustering on a desensitized PCA subspace.
Sun-Yuan Kung, Thee Chanyaswad, J. Morris Chang, Pei Yuan Wu
ACM Trans. Embed. Comput. Syst.1
2017 Hyperspectral and Multispectral Image Fusion Based on Local Low Rank and Coupled Spectral Unmixing
abstract
Hyperspectral images (HSIs) usually have high spectral and low spatial resolution. Conversely, multispectral images (MSIs) usually have low spectral and high spatial resolution. The fusion of HSI and MSI aims to create spectral images with high spectral and spatial resolution. In this paper, we propose a fusion algorithm by combining linear spectral unmixing with the local low-rank property. By taking advantage of the local low-rank property, we first partition the corresponding spectral image into patches. For each patch pair, we cast the fusion problem as a coupled spectral unmixing problem that extracts the abundance and the endmembers of MSI and HSI, respectively. It then updates the abundance and the endmember through an alternating update algorithm. In fact, the convergence of the alternative update algorithm can be mathematically and empirically supported. We also propose a multiscale postprocessing procedure to combine fusion results obtained under different patch sizes. In experiments on three data sets, the proposed fusion algorithms outperformed state-of-the-art fusion algorithms in both spatial and spectral domains.
Yuan Zhou 0006, Liyang Feng, Chunping Hou, Sun-Yuan Kung
IEEE Trans. Geosci. Remote. Sens.4
2016 Adaptive margin slack minimization in RKHS for classification
abstract
In this paper, we design a novel regularized empirical risk minimization technique for classification called Adaptive Margin Slack Minimization (AMSM). The proposed method is based on minimizing a regularized upper bound of the misclassification error. Compared to the cost function of the classical L2-SVM, AMSM can be interpreted as minimizing a tighter bound with some additional flexibilities regarding the choice of marginal hyperplane. A hyperparameter-free adaptive algorithm is presented for finding a solution to the proposed risk function. Numerical results shows that AMSM outperforms L2-SVM on the tested standard datasets.
Yinan Yu, Konstantinos I. Diamantaras, Tomas McKelvey, Sun-Yuan Kung
ICASSP4
2016 Sparse regressions for predicting and interpreting subcellular localization of multi-label proteins
abstract
BACKGROUND: Predicting protein subcellular localization is indispensable for inferring protein functions. Recent studies have been focusing on predicting not only single-location proteins, but also multi-location proteins. Almost all of the high performing predictors proposed recently use gene ontology (GO) terms to construct feature vectors for classification. Despite their high performance, their prediction decisions are difficult to interpret because of the large number of GO terms involved. RESULTS: This paper proposes using sparse regressions to exploit GO information for both predicting and interpreting subcellular localization of single- and multi-location proteins. Specifically, we compared two multi-label sparse regression algorithms, namely multi-label LASSO (mLASSO) and multi-label elastic net (mEN), for large-scale predictions of protein subcellular localization. Both algorithms can yield sparse and interpretable solutions. By using the one-vs-rest strategy, mLASSO and mEN identified 87 and 429 out of more than 8,000 GO terms, respectively, which play essential roles in determining subcellular localization. More interestingly, many of the GO terms selected by mEN are from the biological process and molecular function categories, suggesting that the GO terms of these categories also play vital roles in the prediction. With these essential GO terms, not only where a protein locates can be decided, but also why it resides there can be revealed. CONCLUSIONS: Experimental results show that the output of both mEN and mLASSO are interpretable and they perform significantly better than existing state-of-the-art predictors. Moreover, mEN selects more features and performs better than mLASSO on a stringent human benchmark dataset. For readers' convenience, an online server called SpaPredictor for both mLASSO and mEN is available at http://bioinfo.eie.polyu.edu.hk/SpaPredictorServer/.
Shibiao Wan, Man-Wai Mak, Sun-Yuan Kung
BMC Bioinform.3
2016 Mem-mEN: Predicting Multi-Functional Types of Membrane Proteins by Interpretable Elastic Nets
abstract
Membrane proteins play important roles in various biological processes within organisms. Predicting the functional types of membrane proteins is indispensable to the characterization of membrane proteins. Recent studies have extended to predicting single- and multi-type membrane proteins. However, existing predictors perform poorly and more importantly, they are often lack of interpretability. To address these problems, this paper proposes an efficient predictor, namely Mem-mEN, which can produce sparse and interpretable solutions for predicting membrane proteins with single- and multi-label functional types. Given a query membrane protein, its associated gene ontology (GO) information is retrieved by searching a compact GO-term database with its homologous accession number, which is subsequently classified by a multi-label elastic net (EN) classifier. Experimental results show that Mem-mEN significantly outperforms existing state-of-the-art membrane-protein predictors. Moreover, by using Mem-mEN, 338 out of more than 7,900 GO terms are found to play more essential roles in determining the functional types. Based on these 338 essential GO terms, Mem-mEN can not only predict the functional type of a membrane protein, but also explain why it belongs to that type. For the reader's convenience, the Mem-mEN server is available online at http://bioinfo.eie.polyu.edu.hk/MemmENServer/.
Shibiao Wan, Man-Wai Mak, Sun-Yuan Kung
IEEE ACM Trans. Comput. Biol. Bioinform.3
2016 Large-scale image colorization based on divide-and-conquer support vector machines
Bo-Wei Chen, Wen Ji 0003, Seungmin Rho, Sun-Yuan Kung
J. Supercomput.5
2016 Support vector analysis of large-scale data based on kernels with iteratively increasing order
Bo-Wei Chen, Wen Ji 0003, Seungmin Rho, Sun-Yuan Kung
J. Supercomput.5
2016 Erratum to: Large-scale image colorization based on divide-and-conquer support vector machines
Bo-Wei Chen, Wen Ji 0003, Seungmin Rho, Sun-Yuan Kung
J. Supercomput.5
2016 Profit Maximization through Online Advertising Scheduling for a Wireless Video Broadcast Network
abstract
In this paper, we address the problem of how to make the wireless service provider (WSP) earn profits in a wireless video broadcast network with consideration of advertisement insertion. At the beginning, this study examines the profit components by analyzing traffic provision and advertisement insertion. This study considers using two components for profit maximization-one is the function for allocating video rates, and the other is the function for inserting advertisement duration. The maximum achievable profit depends on joint optimization of optimal video-rate vectors and advertisement-duration vectors, which are usually computationally intensive. To resolve such a complexity problem, this work also proposes an effective algorithm for joint optimization. First, the overall profit is formulated as the solution of four local optimization problems through horizontal and vertical decomposition. Second, a theoretic polymatroidal framework is introduced in our work for optimization as this framework is proved effective in profit maximization of multiuser systems. Third, this study shows that the overall profit can be maximized by finding the optimal profit points on the boundary of the rate and duration regions. As a result, the optimum points and the total profit can be obtained through a hierarchical greedy algorithm. Experimental results demonstrate that the proposed method is capable of making maximum profits for WSPs in a wide range of broadcasting rates.
Wen Ji 0003, Yingying Chen 0001, Min Chen 0003, Bo-Wei Chen, Yiqiang Chen 0001, Sun-Yuan Kung
IEEE Trans. Mob. Comput.6
2015 Profit Improvement in Wireless Video Broadcasting System: A Marginal Principle Approach
abstract
In this paper, we address the problem of how to make the wireless service provider have better profits with consideration of user experience provision in wireless video broadcasting systems. We propose a marginal-based pricing and a resource-allocation framework to achieve better resource utilization and profit improvement. The marginal principle includes 1) marginal user principle, in which a pricing mechanism is established on the basis of marginal users, such that the WSP can seek its own maximum profit of each content with a QoE guarantee; 2) marginal profit principle, in which a WSP can earn the maximum profit through multicontent-service provision by regulating rate allocation in limited available bandwidth. Furthermore, we present a two-tier framework consisting of the inner and outer loops. The inner loop focuses on pricing-based service provision based on the notion of marginal user principle. The outer loop concentrates on allocating bandwidth among multiple video contents according to marginal profit principle. For the solution, we model the profit regions of WSPs and end-users as the polymatroid structures and model the corresponding allocated rate regions as the contra-polymatroid structures. Through exploiting the properties of polymatroid and contra-polymatroid structures, the broadcasting profit problem is solved by finding the optimal rate vector on the sum-rate facet which satisfies the maximal achievable profit. Extensive performance comparison and analysis are presented to demonstrate efficiency of the proposed solution.
Wen Ji 0003, Bo-Wei Chen, Yiqiang Chen 0001, Sun-Yuan Kung
IEEE Trans. Mob. Comput.4
2014 Image Restoration via Multi-prior Collaboration
Feng Jiang 0001, Shengping Zhang, Debin Zhao, Sun-Yuan Kung
ACCV (3)4
2014 Ensemble random projection for multi-label classification with application to protein subcellular localization
abstract
The curse of dimensionality severely restricts the predictive power of multi-label classification systems. High-dimensional feature vectors may contain redundant or irrelevant information, causing the classification systems suffer from overfitting. To address this problem, this paper proposes a dimensionality-reduction method that applies random projection (RP) to construct an ensemble of multilabel classifiers. The merits of the proposed method are demonstrated through a multi-label protein classification task. Specifically, high-dimensional feature vectors are extracted from protein sequences using the gene ontology (GO) and Swiss-Prot databases. The feature vectors are then projected onto lower-dimensional spaces by random projection matrices whose elements conform to a distribution with zero mean and unit variance. The transformed low-dimensional vectors are classified by an ensemble of one-vs-rest multi-label support vector machine (SVM) classifiers, each corresponding to one of the RP matrices. The scores obtained from the ensemble are then fused for predicting the subcellular localization of proteins. Experimental results suggest that the proposed method can reduce the dimensions by seven folds and impressively improve the classification performance.
Shibiao Wan, Man-Wai Mak, Bai Zhang, Yue Joseph Wang, Sun-Yuan Kung
ICASSP5
2014 Cost-effective kernel ridge regression implementation for keystroke-based active authentication system
abstract
In this study a keystroke-based authentication system is implemented on a large-scale free-text keystroke data set, where cost effective kernel-based learning algorithms are designed to enable trade-off between computational cost and accuracy performance. The authentication process evaluates the user's typing behavior on a vocabulary of words, where the judgments based on each word are concatenated by weighted votes, whose weights are also trained to provide optimal fusion of independent judgments. A novel truncated-RBF kernel is also implemented to provide better cost-performance trade-off. Experimental results validate the cost-effectiveness of the developed authentication system.
Pei Yuan Wu, Chi-Chen Fang, J. Morris Chang, Stephen B. Gilbert, Sun-Yuan Kung
ICASSP5
2014 Feature reduction based on Sum-Of-SNR (SOSNR) optimization
abstract
Dimensionality reduction plays an important role in machine learning techniques. In classification, data transformation aims to reduce the number of feature dimensions, whereas attempts to enhance the class separability. To this end, we propose a new classifier-independent criterion called “Sum-of-Signal-to-Noise-Ratio” (SoSNR). A framework designed for maximization with respect to this criterion is presented and three types of algorithms, respectively based on (1) gradient, (2) deflation and (3) sparsity, are proposed. The techniques are conducted on standard UCI databases and compared to other related methods. Results show trade-offs between computational complexity and classification accuracy among different approaches.
Yinan Yu, Tomas McKelvey, Sun-Yuan Kung
ICASSP3
2014 Multikernel Least Mean Square Algorithm
abstract
The multikernel least-mean-square algorithm is introduced for adaptive estimation of vector-valued nonlinear and nonstationary signals. This is achieved by mapping the multivariate input data to a Hilbert space of time-varying vector-valued functions, whose inner products (kernels) are combined in an online fashion. The proposed algorithm is equipped with novel adaptive sparsification criteria ensuring a finite dictionary, and is computationally efficient and suitable for nonstationary environments. We also show the ability of the proposed vector-valued reproducing kernel Hilbert space to serve as a feature space for the class of multikernel least-squares algorithms. The benefits of adaptive multikernel (MK) estimation algorithms are illuminated in the nonlinear multivariate adaptive prediction setting. Simulations on nonlinear inertial body sensor signals and nonstationary real-world wind signals of low, medium, and high dynamic regimes support the approach.
Felipe A. Tobar, Sun-Yuan Kung, Danilo P. Mandic
IEEE Trans. Neural Networks Learn. Syst.2
2013 From green computing to big-data learning: A kernel learning perspective
abstract
Summary form only given. The SVM learning model has been successfully applied to an enormously broad spectrum of application domains and has become a main stream of the modern machine learning technologies. Unfortunately, along with its success and popularity, there also raises a grave concern on it suitability for big data learning applications. For example, in some biomedical applications, the sizes may be hundreds of thousands. In social media application, the sizes could be easily in the order of millions. This curse of dimensionality represents a new challenge calling for new learning paradigm as well as application-specific parallel and distributed hardware and software. This talk will explore cost-effective design on kernel-based machine learning and classification for big data learning applications. It will present a recursive tensor based classification algorithm, especially amenable to systolic/wavefront array processors, which may potentially expedite realtime prediction speed by orders of magnitude. For time-series analysis, with nonstationary environment, it is vital to develop time-adaptive learning algorithms so as to allow incremental and active learning. The talk will tackle the active learning problems from two kernel-induced perspectives, one in intrinsic space and another in empirical space. The talk will show, if time permits, an algorithmic example highlighting the application of Map-Reduce technologies to supervised kernel (Slackmin) learning under a parallel and distributed processing framework.
Sun-Yuan Kung
ASAP1
2013 An ensemble classifier with random projection for predicting multi-label protein subcellular localization
abstract
In protein subcellular localization prediction, a predominant scenario is that the number of available features is much larger than the number of data samples. Among the large number of features, many of them may contain redundant or irrelevant information, causing the prediction systems suffer from overfitting. To address this problem, this paper proposes a dimensionality-reduction method that applies random projection (RP) to construct an ensemble multi-label classifier for predicting protein subcellular localization. Specifically, the frequencies of occurrences of gene-ontology terms are used as feature vectors, which are projected onto lower-dimensional spaces by random projection matrices whose elements conform to a distribution with zero mean and unit variance. The transformed low-dimensional vectors are classified by an ensemble of one-vs-rest multi-label support vector machine (SVM) classifiers, each corresponding to one of the RP matrices. The scores obtained from the ensemble are then fused for making the final decision. Experimental results on two recent datasets suggest that the proposed method can reduce the dimensions by six folds and remarkably improve the classification performance.
Shibiao Wan, Man-Wai Mak, Bai Zhang, Yue Joseph Wang, Sun-Yuan Kung
BIBM5
2013 Adaptive thresholding for multi-label SVM classification with application to protein subcellular localization prediction
abstract
Multi-label classification has received increasing attention in computational proteomics, especially in protein subcellular localization. Many existing multi-label protein predictors suffer from over-prediction because they use a fixed decision threshold to determine the number of labels to which a query protein should be assigned. To address this problem, this paper proposes an adaptive thresholding scheme for multi-label support vector machine (SVM) classifiers. Specifically, each one-vs-rest SVM has an adaptive threshold that is a fraction of the maximum score of the one-vs-rest SVMs in the classifier. Therefore, the number of class labels of the query protein depends on the confidence of the SVMs in the classification. This scheme is integrated into our recently proposed subcellular localization predictor that uses the frequency of occurrences of gene-ontology terms as feature vectors and one-vs-rest SVMs as classifiers. Experimental results on two recent datasets suggest that the scheme can effectively avoid both over-prediction and under-prediction, resulting in performance significantly better than other gene-ontology based subcellular localization predictors.
Shibiao Wan, Man-Wai Mak, Sun-Yuan Kung
ICASSP3
2013 A classification scheme for 'high-dimensional-small-sample-size' data using soda and ridge-SVM with microwave measurement applications
abstract
The generalization performance of SVM-type classifiers severely suffers from the `curse of dimensionality'. For some real world applications, the dimensionality of the measurement is sometimes significantly larger compared to the amount of training data samples available. In this paper, a classification scheme is proposed and compared with existing techniques for such scenarios. The proposed scheme includes two parts: (i) feature selection and transformation based on Fisher discriminant criteria and (ii) a hybrid classifier combining Kernel Ridge Regression with Support Vector Machine to predict the label of the data. The first part is named Successively Orthogonal Discriminant Analysis (SODA), which is applied after Fisher score based feature selection as a preliminary processing for dimensionality reduction. At this step, SODA maximizes the ratio of between-class-scatter and within-class-scatter to obtain an orthogonal transformation matrix which maps the features to a new low dimensional feature space where the class separability is maximized. The techniques are tested on high dimensional data from a microwave measurements system and are compared with existing techniques.
Yinan Yu, Tomas McKelvey, Sun-Yuan Kung
ICASSP3
2013 Kernel SODA: A Feature Reduction Technique Using Kernel Based Analysis
abstract
A feature extraction technique called Successively Orthogonal Discriminant Analysis (SODA) has been recently proposed to overcome the limitation of Linear Discriminant Analysis (LDA), whose objective is to find a projection vector such that the projected values of data from both classes have maximum class separability. However, in LDA, only one such vector can be found due to the rank deficiency for binary classification problems. On the other hand, as a feature extraction technique, the proposed algorithm SODA attempts to obtain a transformation matrix instead of a vector. In this paper, the kernel version of SODA is presented in both intrinsic space and empirical space. To obtain the solution without sacrificing numerical efficiency, we propose a relaxed formulation and data selection for large scale computations. Simulations are conducted on 5 data sets from UCI database to verify and evaluated the new approach.
Yinan Yu, Tomas McKelvey, Sun-Yuan Kung
ICMLA (1)3
2012 On efficient learning and classification kernel methods
abstract
Improving learning and classification efficiency has become increasingly important for machine learning. If the traditional RBF kernel is adopted, the learned kernel-based classifier usually delivers better performance by engaging a large training dataset. However, such a high performance comes at the expense of costly learning and classification complexities, which grow drastically with the training size N. To overcome this curse of dimensionality, we propose a so-called TRBF kernel(with finite intrinsic degree J) which approximates the RBF kernel. The contributions of this paper are as follows. First, the optimal classification efficiency attainable is shown to be J' ≈ J. To improve learning efficiency, we propose a fast PDA algorithm with learning complexity linearly growing with N. We adopt pruned-PDA (PPDA) to improve the accuracy by removing harmful “anti-support” vectors from the training set. Experiments on ECG dataset showed that TRBF-PPDA delivers nearly optimal performance with very low power.
Sun-Yuan Kung, Pei Yuan Wu
ICASSP1
2012 Low-power SVM classifiers for sound event classification on mobile devices
abstract
With the high processing power of today's smartphones, it becomes possible to turn a smartphone into a personal audio surveillance and monitoring system. Ideally, such a system should be able to detect and classify a variety of sound events 24 hours a day and trigger an emergence phone call or message once a specified sound event (e.g., screaming) occurs. To prolong battery life, it is important to trade off the detection accuracy against power consumption. This paper investigates the power consumption of different stages of a sound-event classification system, including segmentation, feature extraction, and SVM scoring. The performance and power consumption of various acoustic features and SVM kernels are compared. This paper advocates the notion of intrinsic complexity through which the scoring function of polynomial SVMs can be written in a matrix-vector-multiplication form so that the resulting complexity becomes independent of the number of support vectors. Results show that this intrinsic complexity can reduce the CPU utilization of polynomial SVMs by 28 times without reducing classification accuracy.
Man-Wai Mak, Sun-Yuan Kung
ICASSP2
2012 GOASVM: Protein subcellular localization prediction based on Gene ontology annotation and SVM
abstract
Protein subcellular localization is an essential step to annotate proteins and to design drugs. This paper proposes a functional-domain based method-GOASVM-by making full use of Gene Ontology Annotation (GOA) database to predict the subcellular locations of proteins. GOASVM uses the accession number (AC) of a query protein and the accession numbers (ACs) of homologous proteins returned from PSI-BLAST as the query strings to search against the GOA database. The occurrences of a set of predefined GO terms are used to construct the GO vectors for classification by support vector machines (SVMs). The paper investigated two different approaches to constructing the GO vectors. Experimental results suggest that using the ACs of homologous proteins as the query strings can achieve an accuracy of 94.68%, which is significantly higher than all published results based on the same dataset. As a user-friendly web-server, GOASVM is freely accessible to the public at http://bioinfo.eie.polyu.edu.hk/mGoaSvmServer/GOASVM.html.
Shibiao Wan, Man-Wai Mak, Sun-Yuan Kung
ICASSP3
2012 Color-frequency-orientation histogram based image retrieval
abstract
This paper proposes a multiclass image retrieval method using combined color-frequency-orientation histogram. Shape information, obtained via edge detector and Hough Transform, is also incorporated into the new feature. The feature has shown advantage in both unsupervised and supervised learning on Corel image dataset containing 10 categories of 1000 complex scenes. In unsupervised learning, comparing with histogram-based method [1], SIMPLIcity [2], FIRM [3], edge-based method [4], multi-resolution-based method [5], our approach respectively shows 25%, 14%, 10%, 7% and 2% improvement in accuracy. In supervised learning, we implement both one-against-one SVM and one-against-all SVM for multiclass classification. One-against-all SVM beats one-against-one SVM, achieving 95% accuracy with sufficient training.
Sun-Yuan Kung
ICASSP3
2012 mGOASVM: Multi-label protein subcellular localization based on gene ontology and support vector machines
abstract
BACKGROUND: Although many computational methods have been developed to predict protein subcellular localization, most of the methods are limited to the prediction of single-location proteins. Multi-location proteins are either not considered or assumed not existing. However, proteins with multiple locations are particularly interesting because they may have special biological functions, which are essential to both basic research and drug discovery. RESULTS: This paper proposes an efficient multi-label predictor, namely mGOASVM, for predicting the subcellular localization of multi-location proteins. Given a protein, the accession numbers of its homologs are obtained via BLAST search. Then, the original accession number and the homologous accession numbers of the protein are used as keys to search against the Gene Ontology (GO) annotation database to obtain a set of GO terms. Given a set of training proteins, a set of T relevant GO terms is obtained by finding all of the GO terms in the GO annotation database that are relevant to the training proteins. These relevant GO terms then form the basis of a T-dimensional Euclidean space on which the GO vectors lie. A support vector machine (SVM) classifier with a new decision scheme is proposed to classify the multi-label GO vectors. The mGOASVM predictor has the following advantages: (1) it uses the frequency of occurrences of GO terms for feature representation; (2) it selects the relevant GO subspace which can substantially speed up the prediction without compromising performance; and (3) it adopts an efficient multi-label SVM classifier which significantly outperforms other predictors. Briefly, on two recently published virus and plant datasets, mGOASVM achieves an actual accuracy of 88.9% and 87.4%, respectively, which are significantly higher than those achieved by the state-of-the-art predictors such as iLoc-Virus (74.8%) and iLoc-Plant (68.1%). CONCLUSIONS: mGOASVM can efficiently predict the subcellular locations of multi-label proteins. The mGOASVM predictor is available online at http://bioinfo.eie.polyu.edu.hk/mGoaSvmServer/mGOASVM.html.
Shibiao Wan, Man-Wai Mak, Sun-Yuan Kung
BMC Bioinform.3
2011 A mutual information based approach for evaluating the quality of clustering
abstract
In this paper, a new method for evaluating the quality of clustering of genes is proposed based on mutual information criterion. Instead of using the conventional histogram-based modeling method to assess clustering performance, we derive a normalized mutual information criterion utilizing the Gaussian kernel density estimator. In the computation of the mutual information, we propose to use only cluster-centroids instead of involving all the members, which offers a huge computational savings. The proposed algorithm not only considers the cluster size but also takes into consideration the homogeneity within a cluster. One major advantage of the proposed algorithm is that, it is capable of estimating an appropriate number of clusters. Extensive experimentation has been carried out on some synthetic data as well as the most widely used Yeast cell cycle gene expression data. Under various clustering conditions it is found that the proposed method provides an excellent performance in terms of measuring the quality of cluster and identifying the true number of cluster.
Shaikh Anowarul Fattah, Chia-Chun Lin, Sun-Yuan Kung
ICASSP3
2011 Improving kernel-energy trade-offs for machine learning in implantable and wearable biomedical applications
abstract
Emerging biomedical sensors and stimulators offer unprecedented modalities for delivering therapy and acquiring physiological signals (e.g., deep brain stimulators). Exploiting these in intelligent, closed loop systems requires detecting specific physiological states using very low power (i.e., 1-10 mW for wearable devices, 10-100 μW for implantable devices). Machine learning is a powerful tool for modeling correlations in physiological signals, but model complexity in typical biomedical applications makes detection too computationally intensive. We analyze the computational energy trade-offs and propose a method of restructuring the computations to yield more favorable trade-offs, especially for typical biomedical applications. We thus develop a methodology for implementing low-energy classification kernels and demonstrate energy reduction in practical biomedical systems. Two applications, arrhythmia detection using electrocardiographs (ECG) from the MIT-BIH database and seizure detection using electroencephalographs (EEG) from the CHB-MIT database, are used. The proposed computational restructuring can be used with very little performance degradation, and it reduces energy by 2627x and 7.0-36.3x (depending on the patient), respectively.
Kyong-Ho Lee, Sun-Yuan Kung, Naveen Verma
ICASSP2
2010 Truncation of protein sequences for fast profile alignment with application to subcellular localization
abstract
We have recently found that the computation time of homology-based subcellular localization can be substantially reduced by aligning profiles up to the cleavage site positions of signal peptides, mitochondrial targeting peptides, and chloro-plast transit peptides [1]. While the method can reduce the profile alignment time by as much as 20 folds, it cannot reduce the computation time spent on creating the profiles. In this paper, we propose a new approach that can reduce both the profile creation time and profile alignment time. In the new approach, instead of cutting the profiles, we shorten the sequences by cutting them at the cleavage site locations. The shortened sequences are then presented to PSI-BLAST to compute the profiles. Experimental results and analysis of profile-alignment score matrices suggest that both profile creation time and profile alignment time can be reduced without sacrificing subcellular localization accuracy. Once a pairwise profile-alignment score matrix has been obtained, a one-vs-rest SVM classifier can be trained. To further reduce the training and recognition time of the classifier, we propose a perturbation discriminant analysis (PDA) technique. It was found that PDA enjoys a short training time as compared to the conventional SVM.
Man-Wai Mak, Sun-Yuan Kung
BIBM3
2010 Speeding up subcellular localization by extracting informative regions of protein sequences for profile alignment
abstract
The functions of proteins are closely related to their subcellular locations. In the post-proteomics era, the amount of gene and protein data grows exponentially, which necessitates the prediction of subcellular localization by computational means. This paper proposes mitigating the computation burden of alignment-based approaches to subcellular localization prediction by using the information provided by the N-terminal sorting signals. To this end, a cascaded fusion of cleavage site prediction and profile alignment is proposed. Specifically, the informative segments of protein sequences are identified by a cleavage site predictor. Then, only the informative segments are applied to a homology-based classifier for predicting the subcellular locations. Experimental results on a newly constructed dataset show that the method can make use of the best property of both approaches and can attain an accuracy higher than using the full-length sequences. Moreover, the method can reduce the computation time by 20 folds. We advocate that the method will be important for biologists to conduct large-scale protein annotation or for bioinformaticians to perform preliminary investigations on new algorithms that involve pairwise alignments.
Man-Wai Mak, Sun-Yuan Kung
CIBCB3
2009 Conditional random fields for the prediction of signal peptide cleavage sites
abstract
Correct prediction of signal peptide cleavage sites has a significant impact on drug design. State-of-the-art approaches to cleavage site prediction typically use generative models (such as HMMs) to represent the statistics of amino acid sequences or use neural networks to detect the changes in short amino-acid segments along a query sequence. By formulating cleavage site prediction as a sequence labeling problem, this paper demonstrates how conditional random fields (CRFs) can be applied to cleavage site prediction. The paper also demonstrates how amino acid properties can be exploited and incorporated into the CRFs to boost prediction performance. Results show that the performance of CRFs is comparable to that of a state-of-the-art predictor (SignalP V3.0). Further performance improvement was observed when the decisions of SignalP and the CRF-based predictor are fused.
Man-Wai Mak, Sun-Yuan Kung
ICASSP2
2009 Feature-based registration of confocal fluorescence endomicroscopy images
abstract
In this paper, we propose a feature-based registration algorithm for the confocal fluorescence images captured by an endomicroscopy system. We first extract a number of feature points from the endomicroscopy images by applying the binarization-thinning process and using the Rutovitz crossing number. These feature points are then post-processed to eliminate the spurious ones. After that we use the purified feature points to complete the image registration between every two consecutive slice images. The aligned image stack will be finally utilized to reconstruct and visualize the 3D structure of the living cell and tissue in real time, which provides the opportunity for the clinicians to diagnose various diseases including the early-stage cancers.
Feng Zhao 0004, Feng Lin 0002, Kemao Qian, Seah Hock Soon, Sun-Yuan Kung
ICIP5
2009 A kernel-based core growing clustering method
abstract
In this paper, a novel clustering method in the kernel space is proposed. It effectively integrates several existing algorithms to become an iterative clustering scheme, which can handle clusters with arbitrary shapes. In our proposed approach, a reasonable initial core for each of the cluster is estimated. This allows us to adopt a cluster growing technique, and the growing cores offer partial hints on the cluster association. Consequently, the methods used for classification, such as support vector machines (SVMs), can be useful in our approach. To obtain initial clusters effectively, the notion of the incomplete Cholesky decomposition is adopted so that the fuzzy c-means (FCM) can be used to partition the data in a kernel defined-like space. Then a one-class and a multiclass soft margin SVMs are adopted to detect the data within the main distributions (the cores) of the clusters and to repartition the data into new clusters iteratively. The structure of the data set is explored by pruning the data in the low-density region of the clusters. Then data are gradually added back to the main distributions to assure exact cluster boundaries. Unlike the ordinary SVM algorithm, whose performance relies heavily on the kernel parameters given by the user, the parameters are estimated from the data set naturally in our approach. The experimental evaluations on two synthetic data sets and four University of California Irvine real data benchmarks indicate that the proposed algorithms outperform several popular clustering algorithms, such as FCM, support vector clustering (SVC), hierarchical clustering (HC), self-organizing maps (SOM), and non-Euclidean norm fuzzy c-means (NEFCM). © 2009 Wiley Periodicals, Inc.4
Tzuen Wuu Hsieh, Jin-Shiuh Taur, Chin-Wang Tao, Sun-Yuan Kung
Int. J. Intell. Syst.4
2009 On Practical Design for Joint Distributed Source and Network Coding
abstract
This paper considers the problem of communicating correlated information from multiple source nodes over a network of noiseless channels to multiple destination nodes, where each destination node wants to recover all sources. The problem involves a joint consideration of distributed compression and network information relaying. Although the optimal rate region has been theoretically characterized, it was not clear how to design practical communication schemes with low complexity. This work provides a partial solution to this problem by proposing a low-complexity scheme for the special case with two sources whose correlation is characterized by a binary symmetric channel. Our scheme is based on a careful combination of linear syndrome-based Slepian-Wolf coding and random linear mixing (network coding). It is in general suboptimal; however, its low complexity and robustness to network dynamics make it suitable for practical implementation.
Yunnan Wu, Vladimir Stankovic 0001, Zixiang Xiong, Sun-Yuan Kung
IEEE Trans. Inf. Theory4
2008 Fusion of cleavage site detection and pairwise alignment for fast subcellular localization
abstract
In recent years, homology-based and signal-based methods have been proposed for predicting the subcellular localization of proteins. While it has been known that homology-based methods can detect more subcellular locations than signal-based methods, the former generally requires a lot more computational resources during both training and prediction. The problem will become intractable for annotating large databases. One possible solution is to reduce the sequence length. This paper proposes to use the cleavage sites detected by signal-based methods (e.g., TargetP) to extract the sequence or profile segments that contain the most localization information for alignment. It was found that the method can reduce computation time of full-length alignment by 27-fold at a cost of only 8% reduction in prediction accuracy. Moreover, the method can increase the accuracy by 0.8% and at the same time reduce the computation time by 41%. Results also show that cutting the sequences at the cleavage sites detected by TargetP is better than cutting them at a fixed position.
Man-Wai Mak, Sun-Yuan Kung
ICASSP2
2008 Fusion of feature selection methods for pairwise scoring SVM
Man-Wai Mak, Sun-Yuan Kung
Neurocomputing2
2008 PairProSVM: Protein Subcellular Localization Based on Local Pairwise Profile Alignment and SVM
abstract
The subcellular locations of proteins are important functional annotations. An effective and reliable subcellular localization method is necessary for proteomics research. This paper introduces a new method---PairProSVM---to automatically predict the subcellular locations of proteins. The profiles of all protein sequences in the training set are constructed by PSI-BLAST and the pairwise profile-alignment scores are used to form feature vectors for training a support vector machine (SVM) classifier. It was found that PairProSVM outperforms the methods that are based on sequence alignment and amino-acid compositions even if most of the homologous sequences have been removed. This paper also demonstrates that the performance of PairProSVM is sensitive (and somewhat proportional) to the degree of its kernel matrix meeting the Mercer's condition. PairProSVM was evaluated on Reinhardt and Hubbard's, Huang and Li's, and Gardy et al.'s protein datasets. The overall accuracies on these three datasets reach 99.3\\%, 76.5\\%, and 91.9\\%, respectively, which are higher than or comparable to those obtained by sequence alignment and by the methods compared in this paper.
Man-Wai Mak, Jian Guo 0002, Sun-Yuan Kung
IEEE ACM Trans. Comput. Biol. Bioinform.3
2007 Feature Selection for Pairwise Scoring Kernels with Applications to Protein Subcellular Localization
abstract
In biological sequence classification, it is common to convert variable-length sequences into fixed-length vectors via pairwise sequence comparison. This pairwise approach, however, can lead to feature vectors with dimension equal to the training set size, causing the curse of dimensionality. This calls for feature selection methods that can weed out irrelevant features to reduce training and recognition time. In this paper, we propose to train an SVM using the full-feature column vectors of a pairwise scoring matrix and select the relevant features based on the support vectors of the SVM. The idea stems from the fact that pairwise scoring matrices are symmetric and support vectors are important for classification. We refer to this approach as vector-index-adaptive SVM (VIA-SVM). We compare VIA-SVM with other feature selection schemes-including SVM-RFE, R-SVM, and a filter method based on symmetric divergence (SD)-in protein subcellular localization. Results show that VIA-SVM is able to automatically bound the number of selected features within a small range. We also found that fusion of VIA-SVM and SD can produce more compact feature subsets without decreasing prediction accuracy, and that while VIA-SVM is superior for large feature-set size, the combination of SD and VIA-SVM performs better at small feature-set size.
Sun-Yuan Kung, Man-Wai Mak
ICASSP (2)1
2007 Environment adaptation for robust speaker verification by cascading maximum likelihood linear regression and reinforced learning
Kwok-Kwong Yiu, Man-Wai Mak, Sun-Yuan Kung
Comput. Speech Lang.3
2007 Probabilistic feature-based transformation for speaker verification over telephone networks
Man-Wai Mak, Kwok-Kwong Yiu, Sun-Yuan Kung
Neurocomputing3
2006 On Consistent Fusion of Multimodal Biometrics
abstract
Audio-visual (AV) biometrics offer complementary information sources, and the use of both voice and facial images for biometric authentication has recently become economically feasible. Therefore, multi-modality adaptive fusion, combining audio and visual information, offers an efficient tool for substantially improving the classification performance. In terms of implementation, we propose to integrate an audio classifier (based on Gaussian mixture models) and a visual classifier (based on FaceIT, a commercially available software) into a well-established mixture-of-expert fusion architecture. In addition, a consistent fusion strategy is introduced as a baseline fusion scheme, which establishes the lower bound of the "consistent region” in the FAR-FRR ROC. Our simulation results indicate that the prediction performance of the proposed adaptive fusion schemes fall in the consistent region. More importantly, the notion of consistent fusion can also facilitate the selection of the best modalities to fuse.
Sun-Yuan Kung, Man-Wai Mak
ICASSP (5)1
2006 A Solution to the Curse of Dimensionality Problem in Pairwise Scoring Techniques
Man-Wai Mak, Sun-Yuan Kung
ICONIP (1)2
2006 Distributed Utility Maximization for Network Coding Based Multicasting: A Shortest Path Approach
abstract
One central issue in practically deploying network coding is the adaptive and economic allocation of network resource. We cast this as an optimization, where the net-utility-the difference between a utility derived from the attainable multicast throughput and the total cost of resource provisioning-is maximized. By employing the MAX of flows characterization of the admissible rate region for multicasting, this paper gives a novel reformulation of the optimization problem, which has a separable structure. The Lagrangian relaxation method is applied to decompose the problem into subproblems involving one destination each. Our specific formulation of the primal problem results in two key properties. First, the resulting subproblem after decomposition amounts to the problem of finding a shortest path from the source to each destination. Second, assuming the net-utility function is strictly concave, our proposed method enables a near-optimal primal variable to be uniquely recovered from a near-optimal dual variable. A numerical robustness analysis of the primal recovery method is also conducted. For ill-conditioned problems that arise, for instance, when the cost functions are linear, we propose to use the proximal method, which solves a sequence of well-conditioned problems obtained from the original problem by adding quadratic regularization terms. Furthermore, the simulation results confirm the numerical robustness of the proposed algorithms. Finally, the proximal method and the dual subgradient method can be naturally extended to provide an effective solution for applications with multiple multicast sessions.
Yunnan Wu, Sun-Yuan Kung
IEEE J. Sel. Areas Commun.2
2006 Adaptive articulatory feature-based conditional pronunciation modeling for speaker verification
Ka-Yee Leung, Man-Wai Mak, Man-Hung Siu, Sun-Yuan Kung
Speech Commun.4
2006 A unification of network coding and tree-packing (routing) theorems
abstract
Given a network of lossless links with rate constraints, a source node, and a set of destination nodes, the multicast capacity is the maximum rate at which the source can transfer common information to the destinations. The multicast capacity cannot exceed the capacity of any cut separating the source from a destination; the minimum of the cut capacities is called the cut bound. A fundamental theorem in graph theory by Edmonds established that if all nodes other than the source are destinations, the cut bound can be achieved by routing. In general, however, the cut bound cannot be achieved by routing. Ahlswede et al. established that the cut bound can be achieved by performing network coding, which generalizes routing by allowing information to be mixed. This paper presents a unifying theorem that includes Edmonds' theorem and Ahlswede et al.'s theorem as special cases. Specifically, it shows that the multicast capacity can still be achieved even if information mixing is only allowed on edges entering relay nodes. This unifying theorem is established via a graph theoretic hardwiring theorem, together with the network coding theorems for multicasting. The proof of the hardwiring theorem implies a new proof of Edmonds' theorem.
Yunnan Wu, Kamal Jain, Sun-Yuan Kung
IEEE Trans. Inf. Theory3
2005 Multi-Class Biclustering and Classification Based on Modeling of Gene Regulatory Networks
abstract
The attempt to elucidate biological pathways and classify genes has led to the development of numerous clustering approaches to gene expression. All these approaches use a single metric to identify genes with similar expression levels. Until now, the correlation between the expression levels of such genes has been based on phenomenological and heuristic correlation functions, rather than on biological models. In this paper, we derive six distinct correlation functions based on explicit thermodynamic modeling of gene regulatory networks. We then combine these correlation functions with novel biclustering algorithms to identify functionally enriched groups. The statistical significance of the identified groups is demonstrated by precision-recall curves and calculated p-values. Furthermore, comparison with chromatin immunoprecipitation data indicates that the performance of the derived correlation functions depends on the specific regulatory mechanisms. Finally, we introduce the idea of multi-class biclustering and with the help of support vector machines we demonstrate its improved classification performance in a microarray dataset.
Ilias Tagkopoulos, Nikolai Slavov, Sun-Yuan Kung
BIBE3
2005 A two-level fusion approach to multimodal biometric verification
abstract
This paper proposes a two-level fusion strategy for audio-visual biometric authentication. Specifically, fusion is performed at two levels: intramodal and intermodal. In intramodal fusion, the scores of multiple samples (e.g. utterances or video shots) obtained from the same modality are linearly combined, where the combination weights depend on the difference between the score values and a client-dependent reference score obtained during enrollment. This is followed by intermodal fusion in which the means of intramodal fused scores obtained from different modalities are either linearly combined or fused by a support vector machine (SVM). Experimental results based on the XM2VTSDB corpus show that intramodal and intermodal fusion are complementary to each other and that SVM-based intermodal fusion is superior to linear combination. 1.
Ming-Cheung Cheung, Man-Wai Mak, Sun-Yuan Kung
ICASSP (5)3
2005 Speaker Verification Using Adapted Articulatory Feature-based Conditional Pronunciation Modeling
abstract
The paper proposes an articulatory feature-based conditional pronunciation modeling (AFCPM) technique for speaker verification. The technique captures the pronunciation characteristics of speakers by modeling the linkage between the actual phones produced by the speakers and the state of articulations during speech production. The speaker models, which consist of conditional probabilities of two articulatory classes, are adapted from a set of universal background models (UBMs) via MAP adaptation. This creates a direct coupling between the speaker and background models, which prevents over-fitting the speaker models when the amount of speaker data is limited. Experimental results demonstrate that MAP adaptation not only enhances the discriminative power of the speaker models but also improves their robustness against handset mismatches. Results also show that fusing the scores derived from an AFCPM-based system and a conventional spectral-based system achieves an error rate that is significantly lower than that which can be achieved by the individual systems. This suggests that AFCPM and spectral features are complementary to each other.
Ka-Yee Leung, Man-Wai Mak, Man-Hung Siu, Sun-Yuan Kung
ICASSP (1)4
2005 Reduced-complexity network coding for multicasting over ad hoc networks
abstract
Network coding generalizes the conventional routing paradigm by allowing nodes to mix information received on its incoming links to generate information to be transmitted to other nodes. As a result, network coding improves throughput, resource efficiency, robustness, manageability, etc., in wired and wireless ad hoc networks. In particular, it was established that network coding can achieve the maximum rate for multicasting information from a source node to multiple destination nodes. The objective of this work is to show how to achieve the aforementioned multicast capacity with lower processing/implementation complexity than that proposed in the literature. We classify the links in a network into two categories: 1) links entering relay nodes; and 2) links entering destinations. We show the same multicast capacity can be achieved by applying (non-trivial) network coding only on the links entering relay nodes. In other words, links entering destinations only require routing, which leads to a saving in the complexity. The novelty of this work lies in a new algorithm, its proof of correctness, and a complexity analysis.
Yunnan Wu, Sun-Yuan Kung
ICASSP (3)2
2005 Bounding the power rate function of wireless ad hoc networks
abstract
Given a wireless ad hoc network and an end-to-end traffic pattern, the power rate function refers to the minimum total power required to support different throughput under a layered model of wireless networks. A critical notion of the layered model is the realizable graphs, which describe possible hit-rate supplies on the links by the physical and medium access layers. Under the layered model, the problem of finding the power rate function can he transformed into finding the minimum-power realizable graph that can provide a given throughput. We introduce a usage conflict graph to represent the conflicts among different uses of the wireless medium. Testing the realizability of a given graph can be transformed into finding the (vertex) chromatic number, i.e., the minimum number of colors required in a proper vertex-coloring, of the associated usage conflict graph. Based on an upper bound of the chromatic number, we propose a linear program that outputs an upper bound of the power rate function. A lower bound of the chromatic number is the clique number. We propose a systematic way of identifying cliques based on a geometric analysis of the space sharing among active links. This leads to another linear program, which yields a lower bound of the power rate function. We further apply greedy vertex-coloring to fine tune the bounds. Simulations results demonstrate that the obtained bounds are tight in the low power and low rate regime.
Yunnan Wu, Qian Zhang 0001, Wenwu Zhu 0001, Sun-Yuan Kung
INFOCOM4
2005 Speaker verification via articulatory feature-based conditional pronunciation modeling with vowel and consonant mixture models
abstract
Articulatory feature-based conditional pronunciation modeling (AFCPM) aims to capture the pronunciation characteristics of speakers by modeling the linkage between the states of articulation during speech production and the actual phones produced by a speaker. Previous AFCPM systems use one discrete density function for each phoneme to model the pronunciation characteristics of speakers. This paper proposes using a mixture of discrete density functions for AFCPM. In particular, the pronunciation characteristics of each phoneme is modeled by two density functions: one responsible for describing the articulatory features that are more relevant to vowels and the other for consonants. Verification scores are the weighted sum of the outputs of the two models. To enhance the resolution of the pronunciation models, four articulatory properties (front-back, liprounding, place of articulation, and manner of articulation) are used for pronunciation modeling. The proposed AFCPM is applied to a speaker verification task. Results show that using four articulatory features achieves a lower error rate as compared to using two features (manner and place of articulation) only. It was also found that dividing the articulatory properties into two groups is an effective means of solving the data-sparseness problem encountered in the training phase of AFCPM systems. 1.
Ka-Yee Leung, Man-Wai Mak, Man-Hung Siu, Sun-Yuan Kung
INTERSPEECH4
2005 Channel robust speaker verification via Bayesian blind stochastic feature transformation
abstract
In telephone-based speaker verification, the channel conditions can be varied significantly from sessions to sessions. Therefore, it is desirable to estimate the channel conditions online and compensate the acoustic distortion without prior knowledge of the channel characteristics. Because no a priori knowledge is used, the estimation accuracy depends greatly on the length of the verification utterances. This paper extends the Blind Stochastic Feature Transformation (BSFT) algorithm that we recently proposed to handle the short-utterance scenario. The idea is to estimate a set of prior transformation parameters from a development set in which a wide variety of channel conditions exists in the verification utterances. The prior transformations are then incorporated into the online estimation of the BSFT parameters in a Bayesian (maximum a posteriori) fashion. The resulting transformation parameters are therefore dependent on both the prior transformations and the verification utterances. For short (long) utterances, the prior transformations play a more (less) important role. We referred the extended algorithm to as Bayesian BSFT (BBSFT) and applied it to the 2001 NIST SRE task. Results show that Bayesian BSFT outperforms BSFT for utterances shorter than or equal to 4 seconds. 1.
Kwok-Kwong Yiu, Man-Wai Mak, Sun-Yuan Kung
INTERSPEECH3
2005 Network planning in wireless ad hoc networks: a cross-Layer approach
abstract
In this paper, the network planning problem in wireless ad hoc networks is formulated as the problem of allocating physical and medium access layer resources or supplies to minimize a cost function, while fulfilling certain end-to-end communication demands, which are given as a collection of multicast sessions with desired transmission rates. We propose an iterative cross-layer optimization, which alternates between: 1) jointly optimizing the timesharing in the medium access layer and the sum of max of flows assignment in the network layer and 2) updating the operational states in the physical layer. We consider two objectives, minimizing aggregate congestion and minimizing power consumption, respectively, corresponding to operating in a bandwidth-limited regime and in an energy-limited regime. The end result is a set of achievable tradeoffs between throughput and energy efficiency, in a given wireless network with a given traffic pattern. We evaluate our approach quantitatively by simulations of community wireless networks and compare with designs that decouple the layers. We demonstrate that significant performance advantages can be achieved by adopting a full-fledged cross-layer optimization. Furthermore, we observe that optimized solutions generally profit from network coding, physical-layer broadcasting, and traffic-dependent physical states.
Yunnan Wu, Philip A. Chou, Qian Zhang 0001, Kamal Jain, Wenwu Zhu 0001, Sun-Yuan Kung
IEEE J. Sel. Areas Commun.6
2005 Minimum-energy multicast in mobile ad hoc networks using network coding
abstract
The minimum energy required to transmit one bit of information through a network characterizes the most economical way to communicate in a network. In this paper, we show that, under a layered model of wireless networks, the minimum energy-per-bit for multicasting in a mobile ad hoc network can be found by a linear program; the minimum energy-per-bit can be attained by performing network coding. Compared with conventional routing solutions, network coding not only allows a potentially lower energy-per-bit to be achieved, but also enables the optimal solution to be found in polynomial time, in sharp contrast with the NP-hardness of constructing the minimum-energy multicast tree as the optimal routing solution. We further show that the minimum energy multicast formulation is equivalent to a cost minimization with linear edge-based pricing, where the edge prices are the energy-per-bits of the corresponding physical broadcast links. This paper also investigates minimum energy multicasting with routing. Due to the linearity of the pricing scheme, the minimum energy-per-bit for routing is achievable by using a single distribution tree. A characterization of the admissible rate region for routing with a single tree is presented. The minimum energy-per-bit for multicasting with routing is found by an integer linear program. We show that the relaxation of this integer linear program, studied earlier in the Steiner tree literature, can now be interpreted as the optimization for minimum energy multicasting with network coding. In short, this paper presents a unifying study of minimum energy multicasting with network coding and routing.
Yunnan Wu, Philip A. Chou, Sun-Yuan Kung
IEEE Trans. Commun.3
2004 Multi-sample data-dependent fusion of sorted score sequences for biometric verification
abstract
In many biometric systems, the scores of multiple samples (e.g. utterances) are averaged and the average score is compared against a decision threshold for decision making. The average score, however, may not be optimal because the distribution of the scores is ignored. To address this limitation, we have recently proposed a fusion model that incorporates the score distribution by making the fusion weights dependent on the dispersion between the frame-based scores and the prior score statistics obtained from training data. As the fusion weights are data-dependent, the positions of scores in the score sequences become detrimental to the final fused scores. We propose to enhance the fusion model by sorting the score sequences before fusion takes place. The fusion model was evaluated on a speaker verification task where each claimant utters two utterances in a verification session. Results demonstrate that fusion of sorted scores has the effect of maximizing the dispersion between the client scores and the impostor scores, making the verification process more reliable. Compared with our previous work, where no sorting was applied, the new approach reduces the equal error rate by 11 %.
Ming-Cheung Cheung, Man-Wai Mak, Sun-Yuan Kung
ICASSP (5)3
2004 A recursive QR approach to semi-blind equalization of time-varying MIMO channels
abstract
This paper presents a novel adaptive equalization algorithm for time-varying, frequency-selective MIMO systems that harnesses the finite alphabet property inherent in digital communication. The algorithm leads to a direct, cost-efficient QR-based recursive updating procedure for the equalizer coefficients that forces adaptation to changing channel characteristics. The proposed method does not require precise channel estimation and uses significantly less pilot symbols than other traditional equalizers, implying a drastic reduction in bandwidth overhead. Simulation results confirm that this approach outperforms the traditional recursive least squares (RLS) adaptive equalizer for this application and rivals the MMSE equalizers with perfect channel knowledge.
Sun-Yuan Kung, Chad L. Myers, Xinying Zhang
ICASSP (4)1
2004 Applying articulatory features to telephone-based speaker verification
abstract
This paper presents an approach that uses articulatory features (AF) derived from spectral features for telephone-based speaker verification. To minimize the acoustic mismatch caused by different handsets, handset-specific normalization is applied to the spectral features before the AF are extracted. Experimental results based on 150 speakers using 10 different handsets show that AF contain useful speaker-specific information for speaker verification and the use of handset-specific normalization significantly lowers the error rates under the handset mismatched conditions. Results also demonstrate that fusing the scores obtained from an AF-based system with those obtained from a spectral feature-based (MFCC) system helps lower the error rates of the individual systems.
Ka-Yee Leung, Man-Wai Mak, Sun-Yuan Kung
ICASSP (1)3
2004 Cross-weighted Fisher discriminant analysis for visualization of DNA microarray data
abstract
Fisher discriminant analysis (DA) has recently shown promise in dimensionality reduction of high dimensional DNA data. However, the 1D projection provided by this method is an optimal Bayesian classifier only when the intraclass data patterns are purely Gaussian distributed. Unfortunately, it has been well recognized that most DNA expression data are much more realistically represented by a Gaussian mixture model (GMM), which allows for multiple cluster centroids per class. When a data set from such a GMM is projected onto a 1D subspace, its inherent multi-modal nature may be partially or completely obscured. Consequently, traditional Fisher DA is quite inadequate when higher dimensional visualization (e.g. 2D or 3D) is necessary. The proposed technique addresses this problem and makes use of combined supervised and unsupervised learning techniques for several DNA microarray signal processing functions, including intraclass cluster discovery, optimal projection, and identification/selection of responsible gene groups. In particular, a cross-weighted Fisher DA is proposed and its abilities to reduce dimensionality and to visualize data sets are evaluated.
Xinying Zhang, Chad L. Myers, Sun-Yuan Kung
ICASSP (5)3
2004 Multi-sample fusion with constrained feature transformation for robust speaker verification
abstract
This paper proposes a single-source multi-sample fusion approach to text-independent speaker verification. In conventional speaker verification systems, the scores obtained from claimant's utterances are averaged and the resulting mean score is used for decision making. Instead of using an equal weight for all scores, this paper proposes assigning a different weight to each score, where the weights are made dependent on the difference between the score values and a speaker-dependent reference score obtained during enrollment. Because the fusion weights depend on the verification scores, a technique called constrained stochastic feature transformation is applied to minimize the mismatch between enrollment and verification data in order to enhance the scores' reliability. Experimental results based on the 2001 NIST evaluation set show that the proposed fusion approach outperforms the equal-weight approach by 22% in terms of equal error rate and 16% in terms of minimum detection cost.
Ming-Cheung Cheung, Kwok-Kwong Yiu, Man-Wai Mak, Sun-Yuan Kung
INTERSPEECH4
2004 Articulatory feature-based conditional pronunciation modeling for speaker verification
Ka-Yee Leung, Man-Wai Mak, Sun-Yuan Kung
INTERSPEECH3
2004 A new approach to channel robust speaker verification via constrained stochastic feature transformation
abstract
This paper proposes a constrained stochastic feature transformation algorithm for robust speaker verification. The algorithm computes the feature transformation parameters based on the statistical difference between a test utterance and a composite GMM formed by combining the speaker and background models. The transformation is then used to transform the test utterance to fit the clean speaker model and background model before verification. By implicitly constraining the transformation, the transformed features can fit both models simultaneously. Experimental results based on the 2001 NIST evaluation set show that the proposed algorithms achieves significant improvement in both equal error rate and minimum detection cost when compared to cepstral mean subtraction and Z-norm. The performance of the proposed transformation approach is also slightly better than the short-time Gaussianization method proposed in [1].
Man-Wai Mak, Kwok-Kwong Yiu, Ming-Cheung Cheung, Sun-Yuan Kung
INTERSPEECH4
2004 Minimum-energy multicast in mobile ad hoc networks using network coding
abstract
The minimum energy required to transmit a bit of information through a network characterizes the most economical way to communicate in a network. In this paper, we show that under a simplified layered model of wireless networks, the minimum-energy multicast problem in mobile ad hoc networks is solvable as a linear program, assuming network coding. Compared with conventional routing solutions, network coding not only promises a potentially lower energy-per-bit, but also enables finding the optimal solution in polynomial time, in sharp contrast with the NP-hardness of constructing the minimum-energy multicast tree as the optimal routing solution.
Yunnan Wu, Philip A. Chou, Sun-Yuan Kung
ITW3
2004 Mobility assisted optimal routing in noninterfering mobile ad hoc networks
abstract
A mobile wireless network experiences random variations due to node mobility, which may potentially be exploited for more cost-effective communications. The pioneering work by Grossglauser and Tse (2001) first demonstrated that a network under sufficient amount of (random) mobility could provide a larger scaling rate of throughput capacity than a static network, at the cost of significant and potentially unbounded end-to-end delay. Subsequent works have addressed the issue of the capacity gain under bounded delay. In this paper, we take a rather different approach in that we explore node mobility in the search for the optimal packet delivery routes subject to the QoS criteria such as delay and energy consumption. This is obtained by adopting a deterministic model in a noninterfering mobile ad hoc network (MANET). We present polynomial time algorithms for finding these optimal routes. Specifically, in a system without power control capability, where the transmission range of each node is fixed, we seek optimal routes with minimal end-to-end delivery time or lowest total power consumption along the path, respectively. For both formulations, we propose hop-expansion based algorithms that carry out the computations inductively over the number of hops. In a system with power control, we seek the minimum energy route, subject to certain end-to-end delay constraint. For this optimization, we propose a layered algorithm that performs the computations inductively over the discrete time periods.
Jihui Zhang 0002, Yunnan Wu, Qian Zhang 0001, Bo Li 0001, Wenwu Zhu 0001, Sun-Yuan Kung
IWQoS6
2004 Accurate detection of aneuploidies in array CGH and gene expression microarray data
abstract
MOTIVATION: Chromosomal copy number changes (aneuploidies) are common in cell populations that undergo multiple cell divisions including yeast strains, cell lines and tumor cells. Identification of aneuploidies is critical in evolutionary studies, where changes in copy number serve an adaptive purpose, as well as in cancer studies, where amplifications and deletions of chromosomal regions have been identified as a major pathogenetic mechanism. Aneuploidies can be studied on whole-genome level using array CGH (a microarray-based method that measures the DNA content), but their presence also affects gene expression. In gene expression microarray analysis, identification of copy number changes is especially important in preventing aberrant biological conclusions based on spurious gene expression correlation or masked phenotypes that arise due to aneuploidies. Previously suggested approaches for aneuploidy detection from microarray data mostly focus on array CGH, address only whole-chromosome or whole-arm copy number changes, and rely on thresholds or other heuristics, making them unsuitable for fully automated general application to gene expression datasets. There is a need for a general and robust method for identification of aneuploidies of any size from both array CGH and gene expression microarray data. RESULTS: We present ChARM (Chromosomal Aberration Region Miner), a robust and accurate expectation-maximization based method for identification of segmental aneuploidies (partial chromosome changes) from gene expression and array CGH microarray data. Systematic evaluation of the algorithm on synthetic and biological data shows that the method is robust to noise, aneuploidal segment size and P-value cutoff. Using our approach, we identify known chromosomal changes and predict novel potential segmental aneuploidies in commonly used yeast deletion strains and in breast cancer. ChARM can be routinely used to identify aneuploidies in array CGH datasets and to screen gene expression data for aneuploidies or array biases. Our methodology is sensitive enough to detect statistically significant and biologically relevant aneuploidies even when expression or DNA content changes are subtle as in mixed populations of cells. AVAILABILITY: Code available by request from the authors and on Web supplement at http://function.cs.princeton.edu/ChARM/
Chad L. Myers, Maitreya J. Dunham, Sun-Yuan Kung, Olga G. Troyanskaya
Bioinform.3
2003 Achievable capacity with sequential equalizers in Rayleigh fading MIMO channels
abstract
Employing multiple antennas at both the transmitter and receiver end offers a promising channel capacity. Unfortunately equalizers, which can deliver better theoretical capacity performance usually, incur higher implementation cost. To facilitate the performance-complexity tradeoff design in practice, this paper explores the asymptotic capacity performance of sequential zero-forcing equalizers under i.i.d. Rayleigh fading channels, in terms of two measurements "capacity loss" and "capacity efficiency". Based on linear algebra and matrix operations, the capacity results are given in analytical form as functions of the coupling terms in the channel transfer function and the signal to noise ratio. The closed-form solutions enable our theoretical work serves as a reference for practical system designers.
Xinying Zhang, Sun-Yuan Kung
GLOBECOM2
2003 Phase-shift-based antenna selection for MIMO channels
abstract
This paper addresses the antenna subset selection in multiple antenna systems with full diversity transmission, for both correlated and uncorrelated channels. To reduce the severe performance degradation of traditional selection/combining schemes, we propose to embed phase-shift-only operations in the RF chains before selection. The resulting system shows a significant advantage in utilizing the multiple antenna diversity under almost any channel condition while incurring only a small hardware overhead. With the optimum phase shifter design given in analytical form, our analysis shows that with more than two branches allowed for selection, the new scheme can achieve the same SNR gain as the full-complexity MRC (maximum-ratio-combining). Even when only one branch is allowed, the performance is still well above the conventional selection scheme and near optimum.
Xinying Zhang, Andreas F. Molisch, Sun-Yuan Kung
GLOBECOM3
2003 A nonlinear recursive least-squares algorithm for the blind separation of finite-alphabet sources
abstract
We present an adaptive algorithm that blindly separates mixtures of finite-alphabet sources given knowledge of the source alphabet and distribution. The algorithm is a nonlinear recursive least-squares procedure that employs a simple and numerically-robust square root Householder update. Simulations verify that the algorithm can separate large-scale noisy mixtures of finite-alphabet sources without any knowledge of the number of sources in the mixture.
Scott C. Douglas, Sun-Yuan Kung
ICASSP (2)2
2003 Robust speaker verification from GSM-transcoded speech based on decision fusion and feature transformation
abstract
In speaker verification, a claimant may produce two or more utterances. Typically, the scores of the speech patterns extracted from these utterances are averaged and the resulting mean score is compared with a decision threshold. Rather than simply computing the mean score, we propose to compute the optimal weights for fusing the scores based on the score distribution of the independent utterances and our prior knowledge about the score statistics. More specifically, we use enrollment data to compute the mean scores of client speakers and impostors and consider them to be the prior scores. During verification, we set the fusion weights for individual speech patterns to be a function of the dispersion between the scores of these speech patterns and the prior scores. Experimental results based on the GSM-transcoded speech of 150 speakers from the HTIMIT corpus demonstrate that the proposed fusion algorithm can increase the dispersion between the mean speaker scores and the mean impostor scores. Compared with a baseline approach where equal weights are assigned to all scores, the proposed approach provides a relative error reduction of 19%.
Man-Wai Mak, Ming-Cheung Cheung, Sun-Yuan Kung
ICASSP (2)3
2003 Detection for MIMO systems with imprecise channel knowledge
abstract
We investigate signal detection for MIMO systems with imprecise channel knowledge. The optimal detector is one which best matches the "total" observation matrix and a "total" signal matrix which has a finite alphabet constraint and a Sylvester structure constraint. An iterative local optimization with interference cancellation (LOIC) algorithm is proposed to achieve low complexity and exploit the finite alphabet constraint. Simulation results show that our proposed algorithms can detect signals with BER close to the case of perfect channel knowledge, if a rough channel estimate is available initially.
Yunnan Wu, Sun-Yuan Kung
ICASSP (4)2
2003 Capacity analysis for parallel and sequential MIMO equalizers
abstract
It is well known that linear MMSE can outperform its zero-forcing counterpart. In combination with a successive interference canceller, MMSE can fully exploit the capacity of MIMO (multiple-input-multiple-output) channels. In practice, however, such an advantage is compromised due to its implementation complexity and the requirement of accurate SNR estimate. Thus other equalizers such as zero-forcing may present an attractive alternative as long as the performance gap is tolerable. This motivates a need to quantify the tradeoff between MMSE and zero-forcing in both parallel and sequential structures. In this paper, the capacity performance of different equalization schemes is investigated, with closed-form formulas provided in terms of two key measures: capacity gaps and ratios. We also conclude that the capacity gain via structural choice (between parallel and sequential) far outweighs that via filter choice (between zero-forcing and MMSE). Indeed, the latter is found to be almost negligible for most practical SNR regions. It is also shown that the sequential zero-forcing equalizers can asymptotically reach the channel capacity when SNR approaches infinity, irrespective of the detection order. Although this paper is focused on the flat-fading channels, the result is directly extendable to the ISI case by slicing the frequency band into infinitesimal stripes, each of which can be treated as flat.
Xinying Zhang, Sun-Yuan Kung
ICASSP (4)2
2003 Computational intelligence approach for gene expression data mining and classification
abstract
The exploration of high dimensional gene expression microarray data demands powerful analytical tools. Our data mining software, visual data analyzer (VISDA) for cluster discovery, reveals many distinguishing patterns among gene expression profiles. The model-supported hierarchical data exploration tool has two complementary schemes: discriminatory dimensionality reduction for structure-focused data visualization, and cluster decomposition by probabilistic clustering. Reducing dimensionality generates the visualization of the complete data set at the top level. This data set is then partitioned into subclusters that can consequently be visualized at lower levels and if necessary partitioned again. These approaches produce different visualizations that are compared against known phenotypes from the microarray experiments. For class prediction on cancers using miroarray data, multilayer perceptrons (MLPs) are trained and optimized, whose architecture and parameters are regularized and initialized by weighted Fisher criterion (wFC)-based discriminatory component analysis (DCA). The prediction performance is compared and evaluated via multifold cross-validation.
Zuyi Wang, Sun-Yuan Kung, Javed I. Khan, Jianhua Xuan, Yue Joseph Wang
ICME2
2003 Detection for MIMO systems with imprecise channel knowledge
abstract
In this work, we investigate the signal detection for MIMO systems with imprecise channel knowledge. The optimal detector is one which best matches the "total" observation matrix and a "total" signal matrix which has a finite alphabet constraint and a Sylvester structure constraint. An iterative local optimization with interference cancellation (LOIC) algorithm is proposed to achieve low complexity and exploit the finite alphabet constraint. Simulation results show that our proposed algorithms can detect the signals with BER close to the case of perfect channel knowledge, if a rough channel estimate is available initially.
Yunnan Wu, Sun-Yuan Kung
ICME2
2003 Capacity analysis for parallel and sequential MIMO equalizers
abstract
It is well known that linear MMSE can outperform its zero-forcing counterpart. In combination with a successive interference canceller, MMSE can fully exploit the capacity of MIMO (multiple-input-multiple-output) channels [A.J. Viterbi, 1986, M.K. Varanasi, T. Guess, 1997]. In practice, however, such an advantage is compromised due to its implementation complexity and the requirement of accurate SNR estimate. Thus other equalizers such as zero-forcing may present an attractive alternative as long as the performance gap is tolerable. This motivates a need to quantify the tradeoff between MMSE and zero-forcing in both parallel and sequential structures. In this paper, the capacity performance of different equalization schemes is investigated, with closed-form formulas provided in terms of two key measures: capacity gaps and ratios. We also conclude that the capacity gain via structural choice (between parallel and sequential) far out-weights that via filter choice (between zero-forcing and MMSE). Indeed, the latter is found to be almost negligible for most practical SNR regions. It is also shown that the sequential zero-forcing equalizers can asymptotically reach the channel capacity when SNR approaches infinity, irrelevant of the detection order. Although this paper is focused on the flat-fading channels, the result is directly extendable to the ISI case by slicing the frequency band into infinitesimal stripes, each of which can be treated as flat.
Xinying Zhang, Sun-Yuan Kung
ICME2
2003 Adaptive decision fusion for multi-sample speaker verification over GSM networks
abstract
In speaker verification, a claimant may produce two or more utterances. In our previous study [1], we proposed to compute the optimal weights for fusing the scores of these utterances based on their score distribution and our prior knowledge about the score statistics estimated from the mean scores of the corresponding client speaker and some pseudo-impostors during enrollment. As the fusion weights depend on the prior scores, in this paper, we propose to adapt the prior scores during verification based on the likelihood of the claimant being an impostor. To this end, a pseudo-imposter GMM score model is created for each speaker. During verification, the claimant's scores are fed to the score model to obtain a likelihood for adapting the prior score. Experimental results based on the GSM-transcoded speech of 150 speakers from the HTIMIT corpus demonstrate that the proposed prior score adaptation approach provides a relative error reduction of 15% when compared with our previous approach where the prior scores are non-adaptive.
Ming-Cheung Cheung, Man-Wai Mak, Sun-Yuan Kung
INTERSPEECH3
2003 Environment adaptation for robust speaker verification
abstract
In speaker verification over public telephone networks, utterances can be obtained from different types of handsets. Different handsets may introduce different degrees of distortion to the speech signals. This paper attempts to combine a handset selector with (1) handset-specific transformations and (2) handset-dependent speaker models to reduce the effect caused by the acoustic distortion. Specifically, a number of Gaussian mixture models are independently trained to identify the most likely handset given a test utterance; then during recognition, the speaker model and background model are either transformed by MLLR-based handset-specific transformation or respectively replaced by a handset-dependent speaker model and a handset-dependent background model whose parameters were adapted by reinforced learning to fit the new environment. Experimental results based on 150 speakers of the HTIMIT corpus show that environment adaptation based on both MLLR and reinforced learning outperforms the classical CMS, Hnorm and Tnorm approaches, with MLLR adaptation achieves the best performance.
Kwok-Kwong Yiu, Man-Wai Mak, Sun-Yuan Kung
INTERSPEECH3
2003 Speaker verification based on g.729 and g.723.1 coder parameters and handset mismatch compensation
abstract
A novel technique for speaker verification over a communication network is proposed. The technique employs cepstral coefficients (LPCCs) derived from G.729 and G.723.1 coder parameters as feature vectors. Based on the LP coefficients derived from the coder parameters, LP residuals are reconstructed, and the verification performance is improved by taking account of the additional speaker-dependent information contained in the reconstructed residuals. This is achieved by adding the LPCCs of the LP residuals to the LPCCs derived from the coder parameters. To reduce the acoustic mismatch between different handsets, a technique combining a handset selector with stochastic feature transformation is employed. Experimental results based on 150 speakers show that the proposed technique outperforms the approaches that only utilize the coder-derived LPCCs.
Eric W. M. Yu, Man-Wai Mak, Chin-Hung Sit, Sun-Yuan Kung
INTERSPEECH4
2003 DFT-based hybrid antenna selection schemes for spatially correlated MIMO channels
abstract
We address the antenna subset selection problem in spatially correlated MIMO channels. To reduce the severe performance degradation of the traditional antenna selection scheme in correlated channels, we propose to embed DFT operations in the RF chains. The resulting system shows a significant advantage both for diversity schemes and for the capacity of spatial multiplexing, while requiring only a minor hardware overhead.
Andreas F. Molisch, Xinying Zhang, Sun-Yuan Kung, Jinyun Zhang
PIMRC3
2003 Impulse radio pulse shaping for ultra-wide bandwidth (UWB) systems
abstract
In this paper, we investigate the design of pulse shaping UK filters for impulse radio ultrawideband (UWB) communications systems. The goal of the shaping is to meet an arbitrary spectrum mask, e.g., the mask mandated by the FCC for UWB emissions. Compared with classical FIR filter designs, the current problem introduces three new challenges: (1) it is minimax with quadratic constraints, (2) a single-sided distortion function is used, (3) delay positions are treated as tuning parameters. We first approach this problem by constructing a least-squares approximation to the minimax problem where the optimization over delays can be easily solved. With the LS solution serving as an initialization, nonlinear optimization techniques are employed to fine tune the solutions.
Yunnan Wu, Andreas F. Molisch, Sun-Yuan Kung, Jinyun Zhang
PIMRC3
2002 Bezout equalization for STBC-MIMO systems
abstract
Space-Time Blocking Coding (STBC) has become very popular as an efficient diversity creation technique. With extra transmission redundancies introduced, the receiver performance can be significantly improved via STBC. In this paper, we study the signal recovery problem for a physical Multiple-Input-Multiple-Output (MIMO) channel via an inverse FIR filterbank together with STBC technique. It can be formulated within a multirate polyphase framework to reach an equivalent virtual MIMO system. It is shown that such a system enables equalization of ill conditioned MIMO channels and can offer performance superior to that with direct equalizations. Based on Generalized Bezout Identity theorems, the recoverability conditions are established. We also address the design problem for optimal noise resilience, which is critical from practical perspective. Furthermore, STBC can flexibly reduce the transmission rate to enhance the equalization SNR. The roles of different STBC and equalizer parameters are analyzed for the optimal tradeoff among transmission rate, diversity gain, and implementation complexity.
Sun-Yuan Kung, Yunnan Wu, Xinying Zhang
ICASSP1
2002 Combining stochastic feature transformation and handset identification for telephone-based speaker verification
abstract
The performance of telephone-based speaker verification systems can be severely degraded by the acoustic mismatch caused by telephone handsets. This paper proposes to combine a handset selector with stochastic feature transformation to reduce the mismatch. Specifically, a GMM-based handset selector is trained to identify the most likely handset used by the claimants, and then handset-specific stochastic feature transformations are applied to the distorted feature vectors. To overcome the non-linear distortion introduced by telephone handsets, a 2nd-order stochastic feature transformation is proposed. Estimation algorithms based on the stochastic matching technique and the EM algorithm are derived. Experimental results based on 150 speakers of the HTIMIT corpus show that the handset selector is able to identify the handsets accurately (98.3%), and that both linear and non-linear transformation reduce the error rate significantly (from 12.37% to 5.49%).
Man-Wai Mak, Sun-Yuan Kung
ICASSP2
2002 Spreading code assignment in an ad hoc DS-CDMA wireless network
abstract
In this paper, we examine the problem of optimal spreading code assignment in an ad hoc DS-CDMA wireless network. The problem is formulated in two ways, with the optimization criteria being minimizing the congestion measure, and minimizing the total consumed power, respectively. Either formulation is converted to an optimization of the weighted total squared cross-correlation (WTSC), after some mathematical simplification. We then devise algorithms to reach some local minima of the optimization. Simulation results show significant reduction in the network congestion measure and power consumption, compared with random code allocation.
Yunnan Wu, Qian Zhang 0001, Wenwu Zhu 0001, Sun-Yuan Kung
ICC4
2002 Divergence-based out-of-class rejection for telephone handset identification
abstract
Research has shown that handset selectors can be used to assist telephone-based speech/speaker recognition. Most handset selectors, however, simply select the most likely handset from a set of known handsets even for speech coming from an ‘unseen’ handset. This paper proposes a divergence-based handset selector with out-of-handset (OOH) rejection capability to identify the ‘unseen’ handsets. This is achieved by measuring the Jensen difference between the selector’s output and a constant vector with identical elements. The resulting handset selector is combined with a feature-based channel compensation algorithm for telephonebased speaker verification. Utterances whose handsets were identified as ‘unseen’ are either transformed by a global bias vector or normalized by cepstral mean subtraction (CMS). On the other hand, if the handset can be identified (considered as ‘seen’), its corresponding transformation parameters will be used to transform the utterances. Experiments based on ten handsets of the HTIMIT corpus show that using the transformation parameters of the ‘seen’ handsets to transform the utterances with correctly identified handsets and processing those utterances with ‘unseen’ handsets by CMS achieve the best result.
Chi-Leung Tsang, Man-Wai Mak, Sun-Yuan Kung
INTERSPEECH3
2002 A Comparative Study on Kernel-Based Probabilistic Neural Networks for Speaker Verification
abstract
This paper compares kernel-based probabilistic neural networks for speaker verification based on 138 speakers of the YOHO corpus. Experimental evaluations using probabilistic decision-based neural networks (PDBNNs), Gaussian mixture models (GMMs) and elliptical basis function networks (EBFNs) as speaker models were conducted. The original training algorithm of PDBNNs was also modified to make PDBNNs appropriate for speaker verification. Results show that the equal error rate obtained by PDBNNs and GMMs is less than that of EBFNs (0.33% vs. 0.48%), suggesting that GMM- and PDBNN-based speaker models outperform the EBFN ones. This work also finds that the globally supervised learning of PDBNNs is able to find decision thresholds that not only maintain the false acceptance rates to a low level but also reduce their variation, whereas the ad-hoc threshold-determination approach used by the EBFNs and GMMs causes a large variation in the error rates. This property makes the performance of PDBNN-based systems more predictable.
Kwok-Kwong Yiu, Man-Wai Mak, Sun-Yuan Kung
Int. J. Neural Syst.3
2001 Hierarchical adaptive regularisation method for depth extraction from planar recording of 3D-integral images
abstract
The paper presents a novel algorithm for object space reconstruction from the planar (2D) recorded data set of a 3D-integral image. The integral imaging system is described and the associated point spread function is given. The space data extraction is formulated as an inverse problem, which proves ill-conditioned, and tackled by using a hierarchical multiresolution strategy and imposing additional conditions to the sought solution. The hierarchical strategy and the two-phase adaptive constrained 3D-reconstruction algorithm based on the use of two sigmoid functions are presented. Finally, illustrative simulation results are given.
Silvia Manolache Cirstea, Malcolm McCormick, Sun-Yuan Kung
ICASSP3
2001 COD: blind path separation with limited antenna size
abstract
This paper proposes a generalized multipath separability condition for subspace processing and derives a novel COD (combined oversampling and displacement) algorithm to utilize both spatial and temporal diversities for path separation and DOA estimation. A unique advantage lies in its ability to cope with the situation where the number of multipaths is much larger than that of antenna elements, which arises in many practical situations. The traditional data matrix or any of its horizontally expanded versions cannot yield a sufficient matrix rank to satisfy the condition, when there is antenna deficiency. Neither can a vertical expansion via oversampling, except when there is no overlapping among intra-user paths (a much stronger condition than the asynchrony condition). The COD strategy solves the antenna deficiency problem by combining vertical expansion with temporal oversampling and horizontal expansion with spatial displacement. Another unique advantage of COD is its multiplicity of eigenvalues which greatly facilitates the later signal recovery processing. The paper first analyzes the theoretical footings for COD and follows with some illustrative simulation results in noisy channels.
Xinying Zhang, Sun-Yuan Kung
ICASSP2
2001 Dynamic resource allocation via video content and short-term traffic statistics
abstract
The reliable and efficient transmission of high-quality variable bit rate (VBR) video through the Internet generally requires network resources be allocated in a dynamic fashion. This includes the determination of when to renegotiate for network resources, as well as how much to request at a given time. The accuracy of any resource request method depends critically on its prediction of future traffic patterns. Such a prediction can be performed using the content and traffic information of short video segments. This paper presents a systematic approach to select the best features for prediction, indicating that while content is important in predicting the bandwidth of a video hit stream, the use of both content and available short-term bandwidth statistics can yield significant improvements. A new framework for traffic prediction is proposed in this paper; experimental results show a smaller mean-square resource prediction error and higher overall link utilization.
Min Wu 0001, Robert A. Joyce, Hau-San Wong, Ling Guan, Sun-Yuan Kung
IEEE Trans. Multim.5
2000 Dynamic Resource Allocation via Video Content and Short-Term Traffic Statistics
abstract
Dynamic resource allocation is critical in the transmission of VBR video. Our study shows that content is one of the major factors that controls the bandwidth of the video bit-stream, yet content alone may not be sufficient in predicting future traffic and in determining how much resource to request. A new framework of traffic prediction is proposed, taking into account both content features and available short-term bandwidth statistics.
Min Wu 0001, Robert A. Joyce, Sun-Yuan Kung
ICIP3
2000 Estimation of elliptical basis function parameters by the EM algorithm with application to speaker verification
abstract
This paper proposes to incorporate full covariance matrices into the radial basis function (RBF) networks and to use the expectation-maximization (EM) algorithm to estimate the basis function parameters. The resulting networks, referred to as elliptical basis function (EBF) networks, are evaluated through a series of text-independent speaker verification experiments involving 258 speakers from a phonetically balanced, continuous speech corpus (TIMIT).We propose a verification procedure using RBF and EBF networks as speaker models and show that the networks are readily applicable to verifying speakers using LP-derived cepstral coefficients as features. Experimental results show that small EBF networks with basis function parameters estimated by the EM algorithm outperform the large RBF networks trained in the conventional approach. The results also show that the equal error rate achieved by the EBF networks is about two-third of that achieved by the vetor quantization (VQ)-based speaker models.
Man-Wai Mak, Sun-Yuan Kung
IEEE Trans. Neural Networks Learn. Syst.2
2000 Probabilistic principal component subspaces: a hierarchical finite mixture model for data visualization
abstract
Visual exploration has proven to be a powerful tool for multivariate data mining and knowledge discovery. Most visualization algorithms aim to find a projection from the data space down to a visually perceivable rendering space. To reveal all of the interesting aspects of multimodal data sets living in a high-dimensional space, a hierarchical visualization algorithm is introduced which allows the complete data set to be visualized at the top level, with clusters and subclusters of data points visualized at deeper levels. The methods involve hierarchical use of standard finite normal mixtures and probabilistic principal component projections, whose parameters are estimated using the expectation-maximization and principal component neural networks under the information theoretic criteria.We demonstrate the principle of the approach on several multimodal numerical data sets, and we then apply the method to the visual explanation in computer-aided diagnosis for breast cancer detection from digital mammograms.
Yue Joseph Wang, Matthew T. Freedman, Sun-Yuan Kung
IEEE Trans. Neural Networks Learn. Syst.4
1999 Adaptive paraunitary filter banks for principal and minor subspace analysis
abstract
Paraunitary filter banks are important for several signal processing tasks. We consider the task of adapting the coefficients of a multichannel FIR paraunitary filter bank via gradient ascent or descent on a chosen cost function. The proposed generalized algorithms inherently adapt the system's parameters in the space of paraunitary filters. Modifications and simplifications of the techniques for spatio-temporal principal and minor subspace analysis are described. Simulations verify one algorithm's useful behavior in this task.
Scott C. Douglas, Shun-ichi Amari, Sun-Yuan Kung
ICASSP3
1999 Hierarchical probabilistic principal component subspaces for data visualization
abstract
Visual exploration has proven to be a powerful tool for multivariate data mining. Most visualization algorithms aim to find a projection from the data space down to a visually perceivable rendering space. To reveal all of the interesting aspects of complex data sets existing in a high-dimensional space, a hierarchical visualization algorithm is introduced, which allows the complete data set to be visualized at the top level, with clusters and subclusters of data points visualized at deeper levels. The methods involve multiple use of standard finite normal mixture models and probabilistic principal component projections, whose parameters are estimated using the expectation-maximization and principal component neural networks under the information theoretic criteria. We demonstrate the principle of the approach on two 3D synthetic data sets.
Yue Joseph Wang, Matthew T. Freedman, Sun-Yuan Kung
IJCNN4
1999 Automatic music score recognition/play system based on decision based neural network
abstract
This paper proposes an automatic music score recognition system based on a hierarchically structured decision based neural network (DBNN), which can classify patterns with nonlinear decision boundaries. Currently, this system yields around a 97% recognition rate for printed music scores.
Toyokazu Hori, Shinichiro Wada, Howzan Tai, Sun-Yuan Kung
MMSP4
1999 Intelligent multimedia signal processing: technology, application, and challenge
abstract
Multimedia technologies will profoundly change the way we access information, conduct business, communicate, educate, learn, and entertain. Multimedia technologies also represent a new opportunity for research interactions among a variety of media such as speech, audio, image, video, text, and graphics. As digitization of image/video have become more affordable, supported with a much richer environment of connectivity and accessibility, computer and web database systems are starting to store voluminous image/video data. This paper addresses emerging research fronts precipitated by such an ultra-scale information processing application and technology: Multimedia Signal Processing Technologies; Neural Networks for Intelligent Multimedia Processing; Video Object Plane (VOP) Extraction and Representation; Implementation of Multimedia Processors.
Sun-Yuan Kung, I-Jong Lin
MMSP1
1999 Automatic video object segmentation via Voronoi ordering and surface optimization
abstract
We combine two theoretical concepts, Voronoi ordering and surface optimization formulation, to form a system for automatic video object segmentation in support of the MPEG-4 standard. Voronoi ordering is a means of projecting the object shape onto the image space; the surface optimization formulation provides a fitness measure for object boundaries in the video sequence. By simultaneously maintaining an invariant Voronoi ordering and maximizing the surface optimization metric, we can automatically extract video object planes. This paper outlines the theoretical aspects of Voronoi ordering and the formulation of video object segmentation as a surface optimization problem. Results from a direct implementation of this theory are shown for three MPEG-4 test sequences (container, coastguard and hallmonitor).
I-Jong Lin, Sun-Yuan Kung
MMSP2
1999 Synergistic modeling and applications of hierarchical fuzzy neural networks
abstract
Many common foundations exist between neural networks and fuzzy inference systems in terms of their mathematical models and system structures. This paper explores such a rich synergy and uses it to form the basis for a unifying framework under which fuzzy logic processing and neural networks may be integrated to achieve more robust information processing. It in turn leads to a family of hierarchical fuzzy neural networks (FNNs) which incorporate an adaptive and modular design of neural networks into the basic fuzzy logic systems. Several important models which are critical to the development of the the hierarchical FNN family are studied. We demonstrate how existing unsupervised and supervised learning strategies can be an integral part of a fuzzy processing framework. In addition, hierarchical structures involving both expert modules and class modules are incorporated into the FNNs. Also presented are some promising application examples.
Sun-Yuan Kung, Jin-Shiuh Taur, Shang-Hung Lin
Proc. IEEE1
1999 A fast rate-optimized motion estimation algorithm for low-bit-rate video coding
abstract
Motion estimation is known to be the main bottleneck in real-time encoding applications, and the search for an effective motion estimation algorithm (in terms of computational complexity and compression efficiency) has been a challenging problem for years. This paper describes a new block-matching algorithm that is much faster than the full search algorithm and occasionally even produces better rate-distortion curves than the full search algorithms. We observe that a piecewise continuous motion field reduces the bit rate for differentially encoded motion vectors. Our motion estimation algorithm exploits the spatial correlations of motion vectors effectively in the sense of producing better rate-distortion curves. Furthermore, we incorporate such correlations in a multiresolution framework to reduce the computational complexity. Simulation shows that this method is successful because of the homogeneous and reliable estimation of the displacement vectors. In nine out of our ten benchmark simulations, the performance of the full search algorithm and that of our subblock multiresolution method is about the same. In one out of our ten benchmark simulations, our method has improvement.
John C.-H. Ju, Yen-Kuang Chen, Sun-Yuan Kung
IEEE Trans. Circuits Syst. Video Technol.3
1998 Extraction of independent components from hybrid mixture: KuicNet learning algorithm and applications
abstract
A hybrid mixture is a mixture of supergaussian, gaussian, and subgaussian independent components (ICs). This paper addresses extraction of ICs from a hybrid mixture. There are two kinds of (single-output vs. all-outputs) kurtosis function to be considered as a contrast function. We advocate the former approach due to its (1) simple and closed-form analysis, and (2) numerical convergence and computational saving. Via this approach, all (and only) the positive local maxima (resp. negative local minima) can yield supergaussian (resp, subgaussian) ICs from any mixture (Kung 1997). We also propose a network algorithm, kurtosis-based independent component network (KuicNet), for recursively extracting ICs. Numerical and convergence properties are analyzed and several application examples demonstrated.
Sun-Yuan Kung, Cristina Mejuto
ICASSP1
1998 A novel learning method by structural reduction of DAGs for on-line OCR applications
abstract
This paper introduces a learning algorithm for a neural structure, directed acyclic graphs (DAGs) that is structurally based, i.e. reduction and manipulation of internal structure are directly linked to learning. This paper extends the concepts of I-Jong Lin and Kung (see IEEE Transactions in Signal Processing Special Issue Neural Networks, 1996) for template matching to a neural structure with capabilities for generalization. DAG-learning is derived from concepts in finite state transducers, hidden Markov models, and dynamic time warping to form an algorithmic framework within which many adaptive signal techniques such as vector quantization, K-means, approximation networks, etc., may be extended to temporal recognition. The paper provides a concept of path-based learning to allow comparison among hidden Markov models (HMMs), finite state transducers (FSTs) and DAG-learning. The paper also outlines the DAG-learning process and provides results from the DAG-learning algorithm over a test set of isolated cursive handwriting characters.
I-Jong Lin, Sun-Yuan Kung
ICASSP2
1998 Frame-rate up-conversion using transmitted true motion vectors
abstract
In this paper, we present a video frame-rate up-conversion scheme that uses transmitted true motion vectors for motion-compensated interpolation. In a past work, we demonstrated that a neighborhood-relaxation motion tracker can provide more accurate true motion information than a conventional minimal-residue block-matching algorithm. Although the technique to estimate the true motion vectors is a novelty in its own right, the strength of this technique can be further demonstrated through various spatio-temporal interpolation applications. In this work, we focus on the particular problem of frame-rate up-conversion. In the proposed scheme, the true motion field is derived by the encoder and transmitted by normal means (e.g., MPEG or H.263 encoding). Then, it is recovered by the decoder and is used not only for motion compensated predictions but also used to reconstruct missing data. It is shown that the use of our neighborhood-relaxation motion estimation provides a method of constructing high quality image sequences in a practical manner.
Yen-Kuang Chen, Anthony Vetro, Huifang Sun, Sun-Yuan Kung
MMSP4
1998 Circular Viterbi based adaptive system for automatic video object segmentation
abstract
Many future video standards such as MPEG-4 are shifting focus from compression to content; video object segmentation is the key technology required by these standards. Unlike still images, motion in video can be used to discriminate between objects. We present an automatic video single object segmentation system whose core is the adaptive version of the circular Viterbi algorithm. The circular Viterbi algorithm fuses the information from motion analysis, edge analysis, active contours (dynamic snake), and temporal correlation into a unified system. Results of pixel-resolution boundaries from our system are shown. Analysis of results, the integration of human-guided data and region-based analysis and a future iterative scheme for multiple object sequences are also discussed.
I-Jong Lin, Sun-Yuan Kung
MMSP2
1998 Neural networks for intelligent multimedia processing
abstract
This paper reviews key attributes of neural processing essential to intelligent multimedia processing (IMP). The objective is to show why neural networks (NNs) are a core technology for the following multimedia functionalities: (1) efficient representations for audio/visual information, (2) detection and classification techniques, (3) fusion of multimodal signals, and (4) multimodal conversion and synchronization. It also demonstrates how the adaptive NN technology presents a unified solution to a broad spectrum of multimedia applications. As substantiating evidence, representative examples where NNs are successfully applied to IMP applications are highlighted. The examples cover a broad range, including image visualization, tracking of moving objects, image/video segmentation, texture classification, face-object detection/recognition, audio classification, multimodal recognition, and multimodal lip reading.
Sun-Yuan Kung, Jenq-Neng Hwang
Proc. IEEE1
1998 A self-stabilized minor subspace rule
abstract
In this letter, we present a minor subspace rule that extracts the subspace that spans the m minor components of a n-dimensional vector stationary random process, m
Scott C. Douglas, Sun-Yuan Kung, Shun-ichi Amari
IEEE Signal Process. Lett.2
1998 Guest Editorial Applications Of Artificial Neural Networks To Image Processing
Rama Chellappa, Kunihiko Fukushima, Aggelos K. Katsaggelos, Sun-Yuan Kung, Yann LeCun, Nasser M. Nasrabadi, Tomaso A. Poggio
IEEE Trans. Image Process.4
1998 Quantification and segmentation of brain tissues from MR images: a probabilistic neural network approach
abstract
This paper presents a probabilistic neural network based technique for unsupervised quantification and segmentation of brain tissues from magnetic resonance images. It is shown that this problem can be solved by distribution learning and relaxation labeling, resulting in an efficient method that may be particularly useful in quantifying and segmenting abnormal brain tissues where the number of tissue types is unknown and the distributions of tissue types heavily overlap. The new technique uses suitable statistical models for both the pixel and context images and formulates the problem in terms of model-histogram fitting and global consistency labeling. The quantification is achieved by probabilistic self-organizing mixtures and the segmentation by a probabilistic constraint relaxation network. The experimental results show the efficient and robust performance of the new algorithm and that it outperforms the conventional classification based approaches.
Yue Joseph Wang, Tülay Adali, Sun-Yuan Kung, Zsolt Szabo
IEEE Trans. Image Process.3
1997 A hierarchical algorithm for image retrieval by sketch
abstract
In this paper, we introduce a hierarchical algorithm for image retrieval by sketch, The application scenario is that the user inputs a rough sketch depicting the prominent edges or contours of objects and wishes to retrieve database images that have similar shapes. We can only expect to get a rough query sketch from the user, which is likely a distorted version of the intended database image, hence it is imperative that tolerance be provided towards sketch distortion. Compared with a previous method that has been adopted by various well-known content-based image indexing and retrieval systems such as the IBM QBIC project, this hierarchical algorithm offers 7 times computation speed-up while demonstrates more tolerance towards distortion in user sketches.
Yin Chan, Sun-Yuan Kung
MMSP2
1997 Rate optimization by true motion estimation
abstract
We propose a rate-optimized motion estimation based on a "true" motion tracker. We observe that the piecewise continuous motion field reduces the bit rate for differentially encoded motion vectors. Hence, a neighborhood relaxation method is proposed. In addition, in current MPEG-4 video VM, each video-object-plane (VOP) is individually coded by a block-based approach. The bit rate can be further improved by the removal of redundancy among the block motion vectors within the same VOP. Therefore, we also propose an object-and-block hybrid coding.
Yen-Kuang Chen, Sun-Yuan Kung
MMSP2
1997 On architectural styles for multimedia signal processors
abstract
After presenting several possible multimedia signal processor (MSP) architecture styles, we propose an architecture style which could provide high performance and high flexibilities, and require less external memory accesses and I/O operations. It is a hierarchical and scalable architecture style which facilitates the hardware-software co-design of MSP circuits and systems.
Sun-Yuan Kung, Yen-Kuang Chen
MMSP1
1997 A recursively structured solution for handwriting and speech recognition
abstract
This paper extends the basic theory of DAGs (Directed Acyclic Graphs) and their DAG-Compare operation to produce a recursive architecture for language recognition systems. Building upon theory and practical implementation, we treat the cases of multiple interacting levels of language recognition. We propose that DAG data structure and its complementary comparison operation are a structural inductive step for a recursive system architecture. We further propose that a recursive system architecture: 1) divide-and-conquers system design at each level of language recognition and seamlessly integrates different levels of contextual information, 2) allows improvements on core algorithms to have a global and compound system improvement, 3) is amenable for high parallelism and 4) can trade off speed for accuracy. We have implemented a simple prototype cursive word recognition system with two levels of recursive structure and a simple DFT basis. The system integrates a curve matching algorithm, letter recognizer and a full dictionary search and interface to grammar checker. Our cursive word recognizer has good, robust performance (93.6%/98.4%, top 1 and top 2 choices word recognition rate, respectively) which can be further improved with training.
I-Jong Lin, Sun-Yuan Kung
MMSP2
1997 Face recognition/detection by probabilistic decision-based neural network
abstract
This paper proposes a face recognition system, based on probabilistic decision-based neural networks (PDBNN). With technological advance on microelectronic and vision system, high performance automatic techniques on biometric recognition are now becoming economically feasible. Among all the biometric identification methods, face recognition has attracted much attention in recent years because it has potential to be most nonintrusive and user-friendly. The PDBNN face recognition system consists of three modules: First, a face detector finds the location of a human face in an image. Then an eye localizer determines the positions of both eyes in order to generate meaningful feature vectors. The facial region proposed contains eyebrows, eyes, and nose, but excluding mouth (eye-glasses will be allowed). Lastly, the third module is a face recognizer. The PDBNN can be effectively applied to all the three modules. It adopts a hierarchical network structures with nonlinear basis functions and a competitive credit-assignment scheme. The paper demonstrates a successful application of PDBNN to face recognition applications on two public (FERET and ORL) and one in-house (SCR) databases. Regarding the performance, experimental results on three different databases such as recognition accuracies as well as false rejection and false acceptance rates are elaborated. As to the processing speed, the whole recognition process (including PDBNN processing for eye localization, feature extraction, and classification) consumes approximately one second on Sparc10, without using hardware accelerator or co-processor.
Shang-Hung Lin, Sun-Yuan Kung, Long-Ji Lin
IEEE Trans. Neural Networks2
1996 Motion-based segmentation by principal singular vector (PSV) clustering method
abstract
Motion-based segmentation has attracted a lot of attention. The task of identifying independent objects is called segmentation. Motion-based segmentation has a broad video application domain. An approach based on principal singular vectors (PSVs) of the image measurement matrix was proposed for separating independent moving objects in Kung and Yun-Ting Lin (1995). After applying SVD (singular value decomposition), feature blocks with different object-based motions tend to form separate clusters on the PSV space. Therefore, a frame can be divided into regions each with consistent motion. Our approach offers several additional features: (1) a multi-candidate feature tracker is adopted. (2) Multiple frames are utilized to facilitate motion-based separation. (3) We would like to achieve not only accurate motion estimation, but also the object regions should retain some neighborhood property (to save the bits for the coding boundary). For this, a neighborhood sensitivity parameter /spl delta/ is introduced. One application of motion-based segmentation is low-bit-rate video compression. In very low bit-rate video coding, only motion vectors of finite regions and the region boundary (coded in prediction error) need to be transmitted. Yet simulations yield quite respectable compensated frames.
Sun-Yuan Kung, Yun-Ting Lin, Yen-Kuang Chen
ICASSP1
1996 A probabilistic decision-based neural network for locating deformable objects and its applications to surveillance system and video browsing
abstract
Detection of a (deformable) pattern or object is an important machine learning and computer vision problem. The task involves finding specific (but locally deformable) patterns in images, such as human faces and eyes/mouths. There are many important commercial applications. This paper presents a decision-based neural network for finding such patterns with specific applications to detecting human faces and locating eyes in the faces. The system built upon the proposal has been demonstrated to be applicable under reasonable variations of orientation and/or lighting, and with the possibility of eye glasses. This method has been shown to be very robust against a large variation of face features and eye shapes. The algorithm takes only 200 ms on a SUN Sparc20 workstation to find human faces in an image with 320/spl times/240 pixels. For a facial image with 320/spl times/240 pixels, the algorithm takes 500 ms to locate two eyes on a SUN Sparc20 workstation. Furthermore, the algorithm can be easily implemented via specialised hardware for real time performance. We have applied this technique to two applications (surveillance system, video browsing) and this paper provides experimental results. Although we have only shown its successful implementation on face detection and eye localization, the proposed technique is meant for more general applications of detection of any (locally deformable) object.
Shang-Hung Lin, Yin Chan, Sun-Yuan Kung
ICASSP3
1996 Video shot classification using human faces
abstract
People usually make up a lot of the information content in videos. The abilities to answer queries and facilitate browsing related to people in videos are crucial. In a single video sequence, a particular person may appear multiple number of times. We propose a scheme to automatically detect the repeated occurrences of the same people to enable fast people related searching. In particular, we propose a video shot classification scheme using human faces, regardless of scale and background. Video shots are classified by clustering facial features extracted from these shots. Potential applications include video indexing and browsing. Employing unsupervised clustering algorithms, this scheme requires no human intervention. Experimental results on a 4-minute news sequence show that it achieves encouraging results.
Yin Chan, Shang-Hung Lin, Yap-Peng Tan, Sun-Yuan Kung
ICIP (3)4
1996 A feature tracking algorithm using neighborhood relaxation with multi-candidate pre-screening
abstract
Tracking of features in video sequences has many applications. Conventionally, the minimum displaced frame difference (referred to as DFD or residue) of a block of pixels is used as the criterion for tracking in block-matching algorithms (BMA). However, such a criterion often misses the true motion vectors, due to many practical factors, e.g. affine warping, image noise, object occlusion, lighting variation, and existence of multiple minimal DFD. Our goal is to find motion vectors of the features for object-based motion tracking, in which (1) any region of an object contains a good number of blocks, whose motion vectors exhibit certain consistency; and (2) only true motion vectors for a few blocks per region are needed. Hence, we propose a new tracking method. (1) At the outset, we disqualify some of the reference blocks which are considered to be unreliable to track. (2) We adopt a multi-candidate pre-screening to provide some robustness in selecting motion candidates. (3) Assuming the true motion field is piecewise continuous, we determine the motion of a feature block by consulting all its neighboring blocks' directions. This allows for the chance that a singular and erroneous motion vector may be corrected by its surrounding motion vectors (just like median filtering). Our method is also designed for tracking more flexible affine-type motions, such as rotation, zooming, sheering, etc. Finally, the performance improvement over other existing methods is demonstrated.
Yen-Kuang Chen, Yun-Ting Lin, Sun-Yuan Kung
ICIP (2)3
1996 Object-based scene segmentation combining motion and image cues
abstract
This paper presents an object-based scene segmentation algorithm which combines the temporal information (e.g. motion) from video and image cues from individual frame. First a motion-based segmentation is decided based on the hierarchical principal component split (HPCS) algorithm for multi-moving-object motion classification. HPCS is a binary-tree-structured recursive procedure which clusters the feature blocks according to their principal component of the feature track matrix. Tracking of feature blocks from multiple frames (/spl ges/2) can be effectively processed and this results in a more accurate rigid motion classification. Experimental result shows that by using motion alone, some mostly homogeneous blocks may fit well to more than one motion classes so that ambiguity occurs. Such blocks are categorized into the so-called "undetermined" region (or U-region) for further processing. An image segmentation scheme using local pixel statistics of blocks in the U-region (U-blocks) is applied to find "valid voting regions" (VVRs). A VVR has a mostly homogeneous interior and is surrounded by a closed contour consisting of relatively high gradient points, which can offer the needed discriminating power for classifying each VVR to its belonging object class by motion voting. By combining the motion-based segmentation with the classification result of VVRs, the final object-based scene segmentation is determined. Simulation results are presented.
Yun-Ting Lin, Yen-Kuang Chen, Sun-Yuan Kung
ICIP (1)3
1995 Bit Level Block Matching Systolic Arrays
abstract
We present two bit-level systolic arrays for block matching which are designed by using a well-known methodology. Hardware complexities and speeds of both bit-level designs and conventional word-level arrays are compared by using synthesis tools. We pay special attention to a class of issues which were somewhat overlooked by previous publications, including power consumption due to high frequency, area due to routing and control, and optimal level of pipelining. Our design offers the following features: (1) The bit-level arrays are estimated to offer 200+% speed-up over word-level arrays. (2) When compared with word-level system with same throughput, the bit-level designs reduce control complexity, bus/routing area, and data buffering. (3) When dynamic power control is desired, these bit-level designs offer the flexibility of disabling some processing elements (for lower significant bits) at slight cost of picture quality. Finally, the potential promises and limitations of bit-level systolic block matching arrays, especially those concerning their integration into codec application system are investigated and discussed.
Yin Chan, Sun-Yuan Kung
ASAP2
1995 Multi-level pixel difference classification methods
abstract
Block matching motion estimation algorithms are useful in many video applications such as the block-based video coding scheme employed in MPEG1/2. A single-chip implementation of a motion estimator (ME) for high quality video compression domains has been the goal of many ongoing research projects. There are several complementary directions along which we can reduce hardware complexity, for example, (1) reduction of search points, and (2) simplification of criterion functions. The last category is what this paper focuses on. We study the algorithmic and architectural potentials of the pixel difference classification (PDC) method and propose a generalisation called multi-level PDC (MPDC). The goal is to examine different hardware-complexity vs performance trade-offs. Moreover, we identify a subset of MPDC, the bit-truncation (BT) method which has the most potential for hardware saving. Experimental results show that it offers attractive trade-offs. Under fixed bit rate constraints, it gives picture quality degradation of less than 0.5 dB, which is non-perceivable, for up to 6-bit truncation. BT results in no complicated data or control flows. Hence the consequent hardware reduction is straightforward. The estimated overall encoder hardware saving ranges from 12% to 35% for 6-bit truncation.
Yin Chan, Sun-Yuan Kung
ICIP (3)2
1995 Decision-based neural network for face recognition system
abstract
This paper proposes a face recognition system based on decision-based neural networks (DBNN). The DBNN adopts a hierarchical network structure with nonlinear basis functions and a competitive credit-assignment scheme. The face recognition system consists of three modules. First, a face detector finds the location of a human face in an image. Then an eye localizer determines the positions of both eyes to help generate size-normalized, reoriented, and reduced-resolution feature vectors. (The facial region proposed contains eyebrows, eyes, and nose, but excluding mouth. Eye-glasses will be permissible.) The last module is a face recognizer. The DBNN can be effectively applied to all the three modules. The DBNN based face recognizer has yielded very high recognition accuracies based on experiments on the ARPA-FERET and SCR-IM databases. In terms of processing speeds and recognition accuracies, the performance of DBNN is superior to that of multilayer perceptron (MLP). The training phase for 100 persons would take around one hour, while the recognition phase (including eye localization, feature extraction, and classification using DBNN) consumes only a fraction of a second (on Sparc10).
Sun-Yuan Kung, M. Fang, S. P. Liou, M. Y. Chiu, Jin-Shiuh Taur
ICIP1
1995 Probabilistic DBNN via expectation-maximization with multi-sensor classification applications
abstract
The original learning rule of the decision based neural network (DBNN) is very much decision-boundary driven. When pattern classes are clearly separated, such learning usually provides very fast and yet satisfactory learning performance. Application examples including OCR and (finite) face/object recognition. Different tactics are needed when dealing with overlapping distribution and/or issues on false acceptance/rejection, which arises in applications such as face recognition and verification. For this, a probabilistic DBNN would be more appealing. This paper investigates several training rules augmenting probabilistic DBNN learning, based largely on the expectation maximization (EM) algorithm. The objective is to establish evidence that the probabilistic DBNN offers an effective tool for multi-sensor classification. Two approaches to multi-sensor classification are proposed and the (enhanced) performance studied. The first involves a hierarchical classification, where sensor information are cascaded in sequential processing stages. The second is multi-sensor fusion, where sensor information are laterally combined to yield improved classification. For the experimental studies, a hierarchical DBNN-based face recognition system is described. For a 38-person face database, the hierarchical classification significantly reduces the false acceptance (from 9.35% to 0%) and false rejection (from 7.29% to 2.25%), as compared to non-hierarchical face recognition. Another promising multiple-sensor classifier fusing face and palm biometric features is also proposed.
Shang-Hung Lin, Sun-Yuan Kung
ICIP (3)2
1995 Decision-based neural networks with signal/image classification applications
abstract
Supervised learning networks based on a decision-based formulation are explored. More specifically, a decision-based neural network (DBNN) is proposed, which combines the perceptron-like learning rule and hierarchical nonlinear network structure. The decision-based mutual training can be applied to both static and temporal pattern recognition problems. For static pattern recognition, two hierarchical structures are proposed: hidden-node and subcluster structures. The relationships between DBNN's and other models (linear perceptron, piecewise-linear perceptron, LVQ, and PNN) are discussed. As to temporal DBNN's, model-based discriminant functions may be chosen to compensate possible temporal variations, such as waveform warping and alignments. Typical examples include DTW distance, prediction error, or likelihood functions. For classification applications, DBNN's are very effective in computation time and performance. This is confirmed by simulations conducted for several applications, including texture classification, OCR, and ECG analysis.
Sun-Yuan Kung, Jin-Shiuh Taur
IEEE Trans. Neural Networks1
1994 Register transfer modeling and simulation for array processors
abstract
This paper presents a register transfer modeling scheme for array processor simulation. Its main goals are to verify the application specific design by real data computation, and to help fine tune the array architecture by precise timing analysis. The data flow graph of the design is translated into a register transfer language which is further combined with a hardware description module. An interactive simulator SISim v2.0 has been implemented to simulate the behavior of such a system. The results are compared with the expected valves to verify the array processor design. The recorded timing information can help the designer to analyze the system and improve the performance and resource utilization.>
W. H. Chou, Sun-Yuan Kung
ASAP2
1994 An SVD Approach to Multi-Camera-Multi-Target 3-D Motion-Shape Analysis
abstract
An SVD approach to the so-called structure-from-motion problem was proposed by Tomasi and Kanade [1992]. The present paper extends the original motion-shape-estimation (MSE) to the multi-camera-multi-target case. The multi-target MSE problem is: given a sequence of 2D video images of multiple moving targets, the problem is to track the 3D motion of the targets and reconstruct their 3D shapes. This is further extended to multi-camera-multi-target MSE, with potential application to the 3D occlusion problem. After collection of feature points (FPs), which are sequentially tracked by a video system, the SVD may be applied to a measurement matrix formed by the FPs. The distribution of singular values would first reveal the information about the number of objects at hand. Then, using an algebraic-based subspace clustering method, the FPs may be mapped onto their corresponding objects. Thereafter, the motion and shape may be estimated from a matrix factorization.>
Sun-Yuan Kung, Jin-Shiuh Taur, M. Y. Chiu
ICIP (1)1
1994 Multilayer neural networks for reduced-rank approximation
abstract
This paper is developed in two parts. First, the authors formulate the solution to the general reduced-rank linear approximation problem relaxing the invertibility assumption of the input autocorrelation matrix used by previous authors. The authors' treatment unifies linear regression, Wiener filtering, full rank approximation, auto-association networks, SVD and principal component analysis (PCA) as special cases. The authors' analysis also shows that two-layer linear neural networks with reduced number of hidden units, trained with the least-squares error criterion, produce weights that correspond to the generalized singular value decomposition of the input-teacher cross-correlation matrix and the input data matrix. As a corollary the linear two-layer backpropagation model with reduced hidden layer extracts an arbitrary linear combination of the generalized singular vector components. Second, the authors investigate artificial neural network models for the solution of the related generalized eigenvalue problem. By introducing and utilizing the extended concept of deflation (originally proposed for the standard eigenvalue problem) the authors are able to find that a sequential version of linear BP can extract the exact generalized eigenvector components. The advantage of this approach is that it's easier to update the model structure by adding one more unit or pruning one or more units when the application requires it. An alternative approach for extracting the exact components is to use a set of lateral connections among the hidden units trained in such a way as to enforce orthogonality among the upper- and lower-layer weights. The authors call this the lateral orthogonalization network (LON) and show via theoretical analysis-and verify via simulation-that the network extracts the desired components. The advantage of the LON-based model is that it can be applied in a parallel fashion so that the components are extracted concurrently. Finally, the authors show the application of their results to the solution of the identification problem of systems whose excitation has a non-invertible autocorrelation matrix. Previous identification methods usually rely on the invertibility assumption of the input autocorrelation, therefore they can not be applied to this case.
Konstantinos I. Diamantaras, Sun-Yuan Kung
IEEE Trans. Neural Networks2
1993 Scheduling partitioned algorithms on processor arrays with limited communication supports
abstract
It is important that array designs, especially the scheduling of partitioned arrays, must cope with various kinds of communication constraints such as interconnection topology, channel bandwidth, and inhomogeneous communication delay. The interprocessor communication requirements can be dictated by the dependence vectors and size of the partitioned tiles. A folded constraint graph is created to describe timing constraints between computation and communication events. An integer-programming based method can then be adopted to find an optimal execution schedule, including offset between the start times of processors, which satisfies these constraints.>
W. H. Chou, Sun-Yuan Kung
ASAP2
1993 Volume rendering by wavefront architecture
abstract
To achieve real time processing for volume rendering, it is necessary to employ parallel processing techniques. Therefore, a ray casting wavefront architecture is proposed. The architecture offers the following advantages. By a special memory organization, the processor array can perform fast rotation along x, y, z axes without incurring costly memory transformation. The projection phase and post-projection phase are assigned to different processors and a significant time-saving can be achieved by the pipelining technique. In the proposed architecture, the projection processors can "skip" the transparent region and step computation when further computation does not affect pixel values any more, e.g., the rays reach opaque surface. Finally, according to the VHDL simulation result, the speedup of the wavefront architecture over several SIMD designs (including CM2 and the Princeton Engine) is about 200% under normal condition.>
Shang-Hung Lin, Sun-Yuan Kung
ASAP2
1993 Fuzzy-decision neural networks
Jin-Shiuh Taur, Sun-Yuan Kung
ICASSP (1)2
1992 Hidden Markov models for character recognition
abstract
A hierarchical system for character recognition with hidden Markov model knowledge sources which solve both the context sensitivity problem and the character instantiation problem is presented. The system achieves 97-99% accuracy using a two-level architecture and has been implemented using a systolic array, thus permitting real-time (1 ms per character) multifont and multisize printed character recognition as well as handwriting recognition.
John A. Vlontzos, Sun-Yuan Kung
IEEE Trans. Image Process.2
1991 An unsupervised neural model for oriented principal component extraction
abstract
The concept of oriented principal component (OPC) analysis is introduced. It is the extension of the GSVD (generalized singular value decomposition) concept to the case of random processes (much like principal component analysis extends SVD for stochastic signals). In the random signal case, OPC analysis is equivalent to matched filtering and can be found useful in many classification and detection applications. The authors propose a corresponding neural model equipped with an efficient training algorithm for estimating the oriented principal component of two stochastic processes without assuming explicit knowledge of their statistics. The algorithm is based on the (normalized) learning rule proposed by Hebb for training the synaptic weights of a network of neurons. Both the theoretical justification and the numerical performance are shown, giving an explicit estimate of the learning rate parameter for best convergence speed.>
Konstantinos I. Diamantaras, Sun-Yuan Kung
ICASSP2
1991 Competition-based supervised learning algorithm for nonlinear discriminant functions
abstract
A basic competition-based model is the now-classic perceptron net using linear discriminant functions. The competition-based learning is extended to the general cases of nonlinear discriminant functions. Generalized perceptron learning rules for the binary-classification and multiple-classification cases are proposed. The convergency properties of the general perceptrons are established. Simulation results on texture classification applications are provided.>
Sun-Yuan Kung, W. D. Mao
ICASSP1
1991 Comparison of several learning subspace methods for classification
abstract
Several competition-based methods for classification are compared. Special attention is paid to subspace methods which are based on computing the projections of the patterns on the principal component vectors of the correlation matrices that span the pattern subspaces. A decision learning rule which updates the correlation matrices can be used to adjust the class boundary and improve the performance of the classification. A learning subspace method is proposed, and some other classification methods are reviewed. In this comparison, all of the methods are applied to a texture classification problem and the performance results are presented.>
Jin-Shiuh Taur, Sun-Yuan Kung
ICASSP2
1991 An Optimal Systolic Array for the Algebraic Path Problem
abstract
A systolic array design for the algebraic path problem (APP) is presented that is both simpler and more efficient than previously proposed configurations. This array uses N/sup 2/ orthogonally connected processing elements and requires 2N I/O connections. Total computation time is 5N-2, which is the minimum time possible in a systolic implementation. The data pipelining rate is one, so no pipeline interleave is required. For multiple problem instances a block pipeline rate of N can be achieved, which is optimal for an array of N/sup 2/ processing elements.>
Paul S. Lewis, Sun-Yuan Kung
IEEE Trans. Computers2
1990 A neural network learning algorithm for adaptive principal component extraction (APEX)
abstract
The problem of the recursive computation of the principal components of a vector stochastic process is discussed. The applications of this problem arise in modeling of control systems, high-resolution spectrum analysis, image data compression, motion estimation, etc. An algorithm called APEX which can recursively compute the principal components using a linear neural network is proposed. The algorithm is recursive and adaptive: given the first m-1 principal components, it can produce the mth component iteratively. The numerical theoretical basis of the fast convergence of the APEX algorithm is given, and its computational advantages over previously proposed methods are demonstrated. Extension to extracting constrained principal components using APEX is also discussed.>
Sun-Yuan Kung, Konstantinos I. Diamantaras
ICASSP1
1990 An object recognition system using stochastic knowledge source and VLSI parallel architecture
abstract
The authors present a system for 2D shape recognition using hidden Markov model (HMM) knowledge sources. The shape is represented by a sequence of curvature values. A ring hidden Markov model (RHMM), which incorporates a ring structure and local connectivity, is proposed. The approach solves both the context sensitivity problem and the pattern instantiation problem. Simulation results on aircraft indicate that the proposed system can achieve almost 100% recognition accuracy at a very fast learning speed. It is shown that the RHMM system can be efficiently implemented in a systolic array, permitting real-time processing.>
W. D. Mao, Sun-Yuan Kung
ICPR (1)2
1990 Orthogonal learning network for constrained principal component problem
abstract
The regular principal components (PC) analysis of stochastic processes is extended to the constrained principal components (CPC) problem. As in the PC analysis, the CPC analysis involves extracting representative components which contain the most information about the original processes. In contrast to the PC problem, the CPC solution has to be extracted from a given constraint subspace. Therefore, the CPC solution may be adopted to best recover the original signal and simultaneously avoid the undesirable noisy or redundant components. This is very appealing in many practical applications. A technique is proposed for finding optimal CPC solutions with an orthogonal learning network (OLN). The underlying numerical analysis for the theoretical proof of the convergency of OLN is discussed. As a byproduct, the same numerical analysis also provides a useful estimate of optimal learning rates, leading to very fast convergence speed. Simulation and application examples are provided
Sun-Yuan Kung
IJCNN1
1989 A unifying algorithm/architecture for artificial neural networks
abstract
A generic iterative model is presented for a wide variety of artificial neural networks (ANNs): single-layer feedback networks, multilayer feed-forward networks, hierarchical competitive networks, and hidden Markov models. Unifying mathematical formulations are provided for both the retrieving and learning phases of ANNs. Based on the unifying mathematical formulation, a programmable universal ring systolic array is derived for both phases. It maximizes the strength of VLSI in terms of intensive and pipelined computing and yet circumvents the limitation on communication. Hardware implementation for the processing units based on CORDIC techniques is discussed.>
Sun-Yuan Kung, Jenq-Neng Hwang
ICASSP1
1989 Hidden Markov models for character recognition
abstract
The authors present a hierarchical system for character recognition with hidden Markov model knowledge sources that solve both the context sensitivity problem and the character instantiation problem. The system achieves 97 to 99% accuracy using a two-level architecture and has been implemented using a systolic array, thus permitting real-time (1 ms per character) multifont and multisize printed character recognition as well as handwriting recognition.>
John A. Vlontzos, Sun-Yuan Kung
ICASSP2
1989 A Unified Systolic Architecture for Artificial Neural Networks
Sun-Yuan Kung, Jenq-Neng Hwang
J. Parallel Distributed Comput.1
1989 Fault-Tolerant Array Processors Using Single-Track Switches
abstract
An array grid model based on single-track switches is proposed. A reconfigurability theorem is developed to provide the theoretical footing for novel reconfiguration algorithms for the fabrication-time and run-time processing. For fabrication-time yield enhancement, the problem of finding a feasible reconfiguration using global control can be reformulated as a maximum independent set problem. An existing algorithm in graph theory is adopted to solve this problem. The simulations conducted indicate that the algorithm is computationally very efficient; therefore, it may also be applicable to certain run-time fault tolerance. In real-time fault tolerance, the propagation time of data/control signals between the host computer incurred in the global control is often prohibitively long; therefore, only distributed processing is feasible. Based on the same reconfigurability theorem, a distributive reconfiguration algorithm is developed for (asynchronous) array processors.>
Sun-Yuan Kung, Shiann-Ning Jean, Chih-Wei Jim Chang
IEEE Trans. Computers1
1988 CORP-a new recovery procedure for VLSI processor arrays
abstract
CORP, a procedure for recovery from transient faults in real-time sensitive medical applications, is introduced. CORP is more effective than the traditional approach because it relies on concurrent retries executed by two neighbor processors in the array (faulty and assistant), instead of successive retries executed only by the faulty processor. Techniques to analyze how the occurrence of transient/intermittent faults disturbs the execution of a parallel algorithm in linear arrays are discussed. An optimal assistant assignment policy is constructed that maximizes the array performance in the presence of faults. The adaptive implementation of the optimal policy in linear-wavefront arrays using local distributed control and near-neighbor communications is presented.>
Elias S. Manolakos, Sun-Yuan Kung
CBMS2
1988 An efficient triarray systolic design for real-time Kalman filtering
abstract
Systolic Kalman (SK) filter designs are presented which are based on a triangular array (triarray) configuration. In order to facilitate the systolic design, the original algorithm for the Kalman filter estimation is reformulated in a new least-squares formulation. The design has advantages in both numerical accuracy and computational efficiency. For the case of white additive noise, the SK-W filter design uses approximately n/sup 2//2 processors and provides a speed-up of n/sup 2//2, with a nearly 100% utilization rate. For the case of colored additive noise, the SK-C filter design also offers comparable speed-up performance.>
Sun-Yuan Kung, Jenq-Neng Hwang
ICASSP1
1988 Efficient modeling for multilayer feed-forward neural nets
abstract
The authors discuss two important aspects in multilayer feed-forward neural nets: the optimal number of hidden units per layer, and the optimal number of synaptic weights between two adjacent layers. On the basis of simulations, they conjecture that the optimal number of hidden units shall be equal to or a little bit more than M-1 for efficient learning, where M is the number of pairs of training patterns used. Locally interconnected nets may be useful for some real applications where geometrical properties are significant. By introducing highway links into the locally interconnected nets, the convergence speed can be improved significantly.>
Sun-Yuan Kung, Jenq-Neng Hwang, S. W. Sun
ICASSP1
1988 Graceful Degradation Schemes for Static/Dynamic Wavefront Arrays
Shiann-Ning Jean, Chih-Wei Jim Chang, Sun-Yuan Kung
ICPP (1)3
1988 Ring systolic designs for artificial neural nets
Sun-Yuan Kung, Jenq-Neng Hwang
Neural Networks1
1988 An algebraic projection analysis for back-propagation learning
Sun-Yuan Kung, Jenq-Neng Hwang
Neural Networks1
1987 A Wavefront Array Processor Using Dataflow Processing Elements
John A. Vlontzos, Sun-Yuan Kung
ICS2
1987 Performance Analysis and Optimization of VLSI Dataflow Arrays
Sun-Yuan Kung, Paul S. Lewis, Sheng-Chun Lo
J. Parallel Distributed Comput.1
1987 Optimal Systolic Design for the Transitive Closure and the Shortest Path Problems
abstract
Due to VLSI technological progress, algorithm- oriented array architectures, such as systolic arrays, appear to be very effective, feasible, and economic. This paper discusses how to design systolic arrays for the transitive closure and the shortest path problems. We shall focus on the Warshall algorithm for the transitive closure problem and the Floyd algorithm for the shortest path problem. These two algorithms share exactly the same structural formulation; therefore, they lead to the same systolic array design. In this paper, we first present a general method for mapping algorithms to systolic arrays. Using this methodology, two new systolic designs for the Warshall-Floyd algorithm will be derived. The first one is a spiral array, which is easy to derive and can be further simplified to a hexagonal array. The other is an orthogonal systolic array which is optimal in terms of pipelining rate, block pipelining rate, and the number of input/output connections.
Sun-Yuan Kung, Sheng-Chun Lo, Paul S. Lewis
IEEE Trans. Computers1
1986 On VLSI array architectures for digital image processing
abstract
VLSI device technology offers a tremendous computing capacity, , programming flexibility, high speed, desirable precision and low power, and VLSI circuit design technique is well supported by of many existing modern CAD tools and software packages. The emergence of such integrated VLSI technology and design technique represents a new opportunity for a digital hardware revolution in image/vision processing systems. For a cohesive design of VLSI image array processors, a cross-disciplinary study on application, algorithm, and architecture is necessary, for this purpose, this paper discusses the algorithm analysis, architecture design and technology evaluation pertinent to VLSI design for image processing systems.
Sun-Yuan Kung
ICASSP1
1986 A Toeplitz approximation approach to coherent source direction finding
abstract
In this paper, the Toeplitz Approximation Method (TAM) of stochastic system identification is applied to the linear equal spaced array narrowband source direction finding problem. The proposed algorithm provides high resolution direction finding capability and is designed for an arbitrary noise, multipath signal environment. As such, it extend existing capability in fields such as passive sonar, radar and communications. A comparitive simulation between TAM and the MUSIC method, using spatial smoothing, is presented which are based on low signal-noise-ratio (SNR) data and a multipath environment.
Sun-Yuan Kung, C. K. Lo, R. Foka
ICASSP1
1986 Timing Analysis and Design Optimization of VLSI Data Flow Arrays
Sun-Yuan Kung, Sheng-Chun Lo, Paul S. Lewis
ICPP1
1986 Real-Time Configuration for Fault-Tolerant VLSI Array Processors
Sun-Yuan Kung, Chih-Wei Jim Chang, Chein-Wei Jen
RTSS1
1985 Hierarchical flowgraph integration for VLSI array processors
abstract
The structural properties of parallel recursive algorithms point to the feasibility of a Hierarchical Flow-graph integration (HIFI) design method for VLSI array processor design. The Hierarchical approach allows the designer to focus attention at the appropriate level of detail. Flow-graphs are used because they offer a powerful and convenient tool for describing many signal processing algorithms. Integration is used here to mean top-down integration - from algorithm analysis to VLSI array - as opposed to merely an integration of electronic components. The HIFI method is proposed as a design and description tool aiming specially at VLSI arrays for signal processing algorithms. The major issues involved are: the recursive algorithm decomposition, abstract notations, flow-graph (structural) and functional (behavior) description, temporal and structural decomposition, bi-directional mapping between graphic and textual codes, simulation and verification tools, and mapping from virtual array to actual architecture.
Sun-Yuan Kung, Jurgen Annevelink, Patrick M. Dewilde, S. C. Lo
ICASSP1
1984 An algorithm basis for systolic/Wavefront array software
abstract
In many signal processing, image processing, and scientific computation applications, there are tremendous demands for large-volume and/or high-speed computations. At the same time, the advent of VLSI offers computing hardware at extremely low cost. These factors combined are bound to have a major effect on the up-grading of future supercomputers. With very large scale integration of systems in mind, a special emphasis is placed on massively parallel, highly modular, and locally interconnected array processors. The algorithmic, architectural, and software principles for the design of VLSI array processors are discussed.
Sun-Yuan Kung
ICASSP1
1984 A state space approach for the 2-D harmonic retrieval problem
abstract
In this paper, a State Space approach to the 2-D harmonic retrieval problem is presented. Under certain assumptions, it is shown that the data and covariance matrices have finite rank and also possess a desirable algebraic structure. Then methods employing the Principal components algorithm are developed to estimate the state space parameters and the sinusoid parameters directly from the data and from the covariance information. Simulation results to support the methods are also provided.
D. V. Bhaskar Rao, Sun-Yuan Kung
ICASSP2
1984 On programming VLSI concurrent array processors
Armin B. Cremers, Sun-Yuan Kung
Integr.2
1983 Highly concurrent Toeplitz eigen-system solver for high resolution spectral estimation
abstract
In this paper, we develop a highly concurrent Toeplitz Eigen-System Solver (TES) for computing the minimum eigenvalue and associated eigenvector of a N by N symmetric Toeplitz matrix. Conventionally, solving an eigen-system will require O(N3) times with sequential machine; or O(N2) time with O(N) processing units. By exploring the Toeplitz structure and adopting a Rayleigh quotient iteration, the TES can solve for the desired minimum eigenvalue in O(KN) time with N processors and K iterations. The development of TES offers a fast algorithm to implement the Pisarenko's high resolution spectral estimation technique.
Yu Hen Hu, Sun-Yuan Kung
ICASSP2
1982 Analysis and implementation of the adaptive notch filter for frequency estimation
abstract
This paper enhances some theoretical and implementation aspects of a constrained autoregressive moving average model, the notch filter model developed in [1] for the estimation of sinusoidal signals in additive, uncorrelated noise, colored or white. This model is shown to approximate the actual signal plus noise model. In addition, the parameter estimates obtained by minimization of the output power of the notch filter approximate the maximum likelihood estimate of the model parameters. The relationship of the notch filtering approach to the existing autoregressive and Pisarenke methods is established. Next, a scheme to combine fast convergence and unbiased estimation is suggested. Lastly, certain implementation aspects of the filter are considered and the method is shown to be amenable to parallel processing.
Sun-Yuan Kung, D. V. Bhaskar Rao
ICASSP1
1982 Wavefront Array Processor: Language, Architecture, and Applications
abstract
This paper describes the development of a wavefront-based language and architecture for a programmable special-purpose multiprocessor array. Based on the notion of computational wavefront, the hardware of the processor array is designed to provide a computing medium that preserves the key properties of the wavefront. In conjunction, a wavefront language (MDFL) is introduced that drastically reduces the complexity of the description of parallel algorithms and simulates the wavefront propagation across the computing network. Together, the hardware and the language lead to a programmable wavefront array processor (WAP). The WAP blends the advantages of the dedicated systolic array and the general-purpose data-flow machine, and provides a powerful tool for the high-speed execution of a large class of matrix operations and related algorithms which have widespread applications.
Sun-Yuan Kung, K. S. Arun, Ron J. Gal-Ezer, D. V. Bhaskar Rao
IEEE Trans. Computers1
1981 Highly parallel architectures for solving linear equations
abstract
An important impact of the fast growing VLSI device technology will be the massive capability of parallel processing which will in turn greatly affect the trend of modern signal processing technology. For the new trend it will be a necessity to have revolutionary architectural design concepts such as topological mapping between algorithm and architecture, simple and regular data flow etc [1]. In this paper, based on a natural topology of the computing structure, a novel "computational wave-front" notion is introduced for describing and and validating data flow in locally connected networks. The parallel architectures include the linear system, with and without pivoting, and the least-square solver using Givens method. We believe that this set of linear system architectures will play a central role in modern signal processing.
Sun-Yuan Kung, D. V. Bhaskar Rao
ICASSP1