Yie-Tarng Chen

dblp:35/4796 · DBLP profile ↗
← Back
39ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0002-7221-1603ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 1 first-author · 6 since 2021Computer networks · 6 · 2 first-authorArtificial intelligence and machine learning · 4 · 1 since 2021
YearPublicationVenuePosition
2025 A Multi-Modal Architecture With Spatio-Temporal-Text Adaptation for Video-Based Traffic Accident Anticipation
abstract
Early and precise accident anticipation is critical for preventing road traffic incidents in advanced traffic systems. This paper presents a Multi-modal Architecture with Spatio-Temporal-Text Adaptation (MASTTA), featuring a Visual Encoder and a Text Encoder within a streamlined end-to-end framework for traffic accident anticipation. Both encoders leverage the CLIP model, pre-trained on large-scale text-image pairs, to utilize visual and textual information effectively. MASTTA captures complex traffic patterns and relationships by fine-tuning only the adapters, reducing retraining demands. In the Visual Encoder, spatio-temporal adaptation is achieved through a novel Temporal Adapter, a novel Spatial Adapter, and an MLP Adapter. The Temporal Adapter enhances temporal consistency in accident-prone areas, while the Spatial Adapter captures spatio-temporal interactions among visual cues. The Text Encoder, equipped with a Text Adapter and an MLP Adapter, aligns latent textual and visual features in a joint embedding space, refining semantic representation. This synergy of text and visual adapters enables MASTTA to model complex spatial interactions across long-range temporal context, improving accident anticipation. We validate MASTTA on DAD and CCD datasets, demonstrating significant improvements in both the earliness and correctness compared to state-of-the-art methods.
Patrik Patera, Yie-Tarng Chen, Wen-Hsien Fang
IEEE Trans. Circuits Syst. Video Technol.2
2024 Spatio-Temporal Adaptation With Dilated Neighbourhood Attention For Accident Anticipation
abstract
Anticipating traffic accidents, which involves predicting potential traffic accidents in advance, is crucial for autonomous vehicles. In this study, we introduce a novel approach that utilises Spatial and Temporal Adapters, specifically designed for image-to-video adaptation through parameter-efficient transfer learning (PEFTL) in the context of traffic accident anticipation. To fully leverage the knowledge from a pretrained CLIP Vision Transformer (CLIP-ViT), the proposed architecture incorporates lightweight Dilated Neighbourhood Attention (DNA) within Adapters. Furthermore, DNA is integrated with a cross-attention mechanism in the Temporal Adapter to capture long-range temporal dependencies. The combination of these Adapters significantly enhances spatio-temporal adaptation, addressing the limitations of existing methods in accurately identifying accident-prone areas while achieving the earliness of accident anticipation in an end-to-end manner. Extensive experiments conducted on two widespread benchmark datasets, DAD and CCD, demonstrate notable performance improvements compared to state-of-the-art works.
Patrik Patera, Yie-Tarng Chen, Wen-Hsien Fang
ICIP2
2022 Learning Spatial-Temporal Graphs with Self-Attention Intensified Conditional Random Field for Video Person Re-identification
abstract
This paper extends the structural graph pooling scheme for video-based person re-identification (re-ID). A temporal-aware feature extractor first employs the short-term temporal correlation of the fine-grained feature maps to generate a set of multi-scale part-based CNN features. Subsequently, a spatio-temporal graph is constructed for these multi-scale part-based features. A structural graph pooling scheme is then used to extract graph features for video re-ID. Specifically, the structured graph pooling is formulated as a node clustering problem based on the structural relationships of the multi-scale part features, addressed by a novel self-attention intensified conditional random field (CRF). Different from the original structural graph pooling approach, CRF is integrated with self-attention to leverage the strength of both schemes to provide long-term structural dependencies. Thereby, it can deal with the deficiency of the existing graph-based approaches on video re-ID in learning the diverse temporal dependency of the multi-scale part features. This enables similar body part information corresponding to the person of interest to be aggregated to diminish the adverse effect of redundant and background information. Simulations on two benchmark datasets showcase the effectiveness of the new method.
Wen-Hsien Fang, Rizard Renanda Adhi Pramono, Yie-Tarng Chen
MMSP3
2022 Spatial-Temporal Action Localization With Hierarchical Self-Attention
abstract
This paper proposes a novel architecture for spatial-temporal action localization in videos. The new architecture first employs a two-stream 3D convolutional neural network (3D-CNN) to provide initial action detection. Next, a new Hierarchical Self-Attention Network (HiSAN), the core of this architecture, is developed to learn the spatial-temporal relationships of key actors. Spatial Gaussian priors (SGP) are also imbued to the bidirectional self-attention to enhance HiSAN in modelling the relationships of neighboring actors. Such a combination of 3D-CNN and SGP augmented HiSAN allows us to effectively extract both of the spatial context information and the long-term temporal dependency to improve action localization accuracy. Afterwards, a new fusion strategy is employed, which first re-scores the bounding boxes to settle the inconsistent detection scores caused by background clutter or occlusion, and then aggregates the motion and appearance information from the two-stream network with the motion saliency to alleviate the impact of camera movement. Finally, a tube association network based on the self-similarity of the actors’ appearance and spatial information across frames is addressed to efficaciously construct the action tubes. Simulations on four widespread datasets reveal the efficacy of the new approach.
Rizard Renanda Adhi Pramono, Yie-Tarng Chen, Wen-Hsien Fang
IEEE Trans. Multim.2
2021 Dance with Self-Attention: A New Look of Conditional Random Fields on Anomaly Detection in Videos
abstract
This paper proposes a novel weakly supervised approach for anomaly detection, which begins with a relation-aware feature extractor to capture the multi-scale convolutional neural network (CNN) features from a video. Afterwards, self-attention is integrated with conditional random fields (CRFs), the core of the network, to make use of the ability of self-attention in capturing the short-range correlations of the features and the ability of CRFs in learning the inter-dependencies of these features. Such a framework can learn not only the spatio-temporal interactions among the actors which are important for detecting complex movements, but also their short- and long-term dependencies across frames. Also, to deal with both local and non-local relationships of the features, a new variant of self-attention is developed by taking into consideration a set of cliques with different temporal localities. Moreover, a contrastive multi-instance learning scheme is considered to broaden the gap between the normal and abnormal instances, resulting in more accurate abnormal discrimination. Simulations reveal that the new method provides superior performance to the state-of-the-art works on the widespread UCF-Crime and Shang-haiTech datasets.
Didik Purwanto, Yie-Tarng Chen, Wen-Hsien Fang
ICCV2
2021 Relational Reasoning for Group Activity Recognition via Self-Attention Augmented Conditional Random Field
abstract
This paper presents a new relational network for group activity recognition. The essence of the network is to integrate conditional random fields (CRFs) with self-attention to infer the temporal dependencies and spatial relationships of the actors. This combination can take advantage of the capability of CRFs in modelling the actors' features that depend on each other and the capability of self-attention in learning the temporal evolution and spatial relational contexts of every actor in videos. Additionally, there are two distinct facets of our CRF and self-attention. First, the pairwise energy of the new CRF relies on both of the temporal self-attention and spatial self-attention, which apply the self-attention mechanism to the features in time and space, respectively. Second, to address both local and non-local relationships in group activities, the spatial self-attention takes account of a collection of cliques with different scales of spatial locality. The associated mean-field inference thereafter can thus be reformulated as a self-attention network to generate the relational contexts of the actors and their individual action labels. Lastly, a bidirectional universal transformer encoder (UTE) is utilized to aggregate the forward and backward temporal context information, scene information and relational contexts for group activity recognition. A new loss function is also employed, consisting of not only the cost for the classification of individual actions and group activities, but also a contrastive loss to address the miscellaneous relational contexts between actors. Simulations show that the new approach can surpass previous works on four commonly used datasets.
Rizard Renanda Adhi Pramono, Wen-Hsien Fang, Yie-Tarng Chen
IEEE Trans. Image Process.3
2020 Empowering Relational Network by Self-attention Augmented Conditional Random Fields for Group Activity Recognition
Rizard Renanda Adhi Pramono, Yie-Tarng Chen, Wen-Hsien Fang
ECCV (1)2
2020 Corrections to "Three-Stream Network With Bidirectional Self-Attention for Action Recognition in Extreme Low Resolution Videos"
abstract
Presents corrections to funding agency information for the above named paper.
Didik Purwanto, Rizard Renanda Adhi Pramono, Yie-Tarng Chen, Wen-Hsien Fang
IEEE Signal Process. Lett.3
2020 CNN-Based Multiple Path Search for Action Tube Detection in Videos
abstract
This paper presents an effective two-stream convolutional neural network (CNN)-based approach to detect multiple spatial-temporal action tubes in videos. A novel video localization refinement (VLR) scheme is first addressed to iteratively rectify the potentially inaccurate bounding boxes by exploiting the temporal consistency between adjacent frames. Then, to provide more faithful detection scores, a new fusion strategy is considered, which combines not only the appearance and the flow information of the two-stream networks but also the motion saliency, the latter of which is included to address the small camera motion. In addition, an efficient multiple path search (MPS) algorithm is developed to simultaneously identify multiple paths in a single run. In the forward message passing of MPS, each node stores information of a prescribed number of connections based on the accumulated scores determined in the previous stages. A backward path tracing is invoked afterward to find all multiple paths at the same time by fully reusing the information generated in the forward pass without repeating the search process. Thus, the complexity incurred can be reduced. The simulation results show that, together with VLR and the new fusion scheme, the proposed MPS, in general, can provide superior performance compared with the state-of-the-art works on four public datasets.
Erick Hendra Putra Alwando, Yie-Tarng Chen, Wen-Hsien Fang
IEEE Trans. Circuits Syst. Video Technol.2
2019 Hierarchical Self-Attention Network for Action Localization in Videos
abstract
This paper presents a novel Hierarchical Self-Attention Network (HISAN) to generate spatial-temporal tubes for action localization in videos. The essence of HISAN is to combine the two-stream convolutional neural network (CNN) with hierarchical bidirectional self-attention mechanism, which comprises of two levels of bidirectional self-attention to efficaciously capture both of the long-term temporal dependency information and spatial context information to render more precise action localization. Also, a sequence rescoring (SR) algorithm is employed to resolve the dilemma of inconsistent detection scores incurred by occlusion or background clutter. Moreover, a new fusion scheme is invoked, which integrates not only the appearance and motion information from the two-stream network, but also the motion saliency to mitigate the effect of camera motion. Simulations reveal that the new approach achieves competitive performance as the state-of-the-art works in terms of action localization and recognition accuracy on the widespread UCF101-24 and J-HMDB datasets.
Rizard Renanda Adhi Pramono, Yie-Tarng Chen, Wen-Hsien Fang
ICCV2
2019 Three-Stream Network With Bidirectional Self-Attention for Action Recognition in Extreme Low Resolution Videos
abstract
This letter presents a novel three-stream network for action recognition in extreme low resolution (LR) videos. In contrast to the existing networks, the new network uses the trajectory-spatial network, which is robust against visual distortion, instead of the pose information to complement the two-stream network. Also, the new three-stream network is combined with the inflated 3D ConvNet (I3D) model pre-trained on kinetics to produce more discriminative spatio-temporal features in blurred LR videos. Moreover, a bidirectional self-attention network is aggregated with the three-stream network to further manifest various temporal dependence among the spatio-temporal features. A new fusion strategy is devised as well to integrate the information from the three different modalities. Simulations show that the new architecture outperforms the main state-of-the-art extreme LR action recognition methods on the HMDB-51 and IXMAS datasets.
Didik Purwanto, Rizard Renanda Adhi Pramono, Yie-Tarng Chen, Wen-Hsien Fang
IEEE Signal Process. Lett.3
2019 First-Person Action Recognition With Temporal Pooling and Hilbert-Huang Transform
abstract
This paper presents a convolutional neural network (CNN)-based approach for first-person action recognition with a combination of temporal pooling and the Hilbert–Huang transform (HHT). The new approach first adaptively performs temporal sub-action localization, treats each channel of the extracted trajectory pooled CNN features as a time series, and summarizes the temporal dynamic information in each sub-action by temporal pooling. The temporal evolution across sub-actions is then modeled by rank pooling. Thereafter, to account for the highly dynamic scene changes in first-person videos, the HHT is employed to decompose the ranked pooling features into finite and often few data-dependent functions, called intrinsic mode functions (IMFs), through empirical mode decomposition. Hilbert spectral analysis is then applied to each IMF component, and four salient descriptors are scrutinized and aggregated into the final video descriptor. Such a framework cannot only precisely acquire both long- and short-term tendencies, but also address the cumbersome significant camera motion in first-person videos to render better accuracy. Furthermore, it works well for complex actions for limited training samples. Simulations show that the proposed approach outperforms the main state-of-the-art methods when applied to four publicly available first-person video datasets.
Didik Purwanto, Yie-Tarng Chen, Wen-Hsien Fang
IEEE Trans. Multim.2
2018 Efficient Weighted Kernel Sharing Convolutional Neural Networks
abstract
To lessen the redundancy of convolutional kernels, this paper proposes a new convolutional structure, i.e., weighted kernel sharing convolution (WKSC), which gathers the inputs with the same kernel, so the inputs in each group can share the same convolutional kernel. Also, an extra weighting is imposed for each input channel before the sharing process to manifest its diversity. As a consequence, the number of kernels can be greatly reduced, leading to a reduction of model parameters and the speedup of inference. Moreover, WKSC can be combined with other existing compression models such as depthwise separable convolutions, resulting in a more compressed architecture. Extensive experiments on CIFAR-100 and ImageNet classification demonstrate the effectiveness of the new approach in both computation cost and the parameters required compared with the state-of-the-art works.
Helong Zhou, Yie-Tarng Chen, Jie Zhang 0071, Wen-Hsien Fang
VCIP2
2018 Improved Object Detection With Iterative Localization Refinement in Convolutional Neural Networks
abstract
To facilitate object localization, the existing convolutional neural network (CNN)-based object detection often requires an object proposal method, which, however, may produce inaccurate region proposals and thus impact the performance. To overcome this setback, this paper presents a novel iterative localization refinement method which, undertaken at a mid-layer of a CNN architecture, progressively refines a subset of region proposals in order to match as much ground-truth as possible. In each iteration, the refinement task is cast into a probabilistic framework based on an ingeniously devised probability function. To expedite the computation of the probability function, a divide-and-conquer paradigm is developed by the theorem of total probability. Moreover, an approximate variant based on a refined sampling strategy is also addressed to further reduce the complexity. The proposed ILR method is not only data-driven and free of learning, but it can also be incorporated with many existing CNN-based object detection algorithms, such as Faster R-CNN to enhance the detection accuracy without changing their configurations. Simulations show that the proposed method can improve the main state-of-the-art works on the PASCAL VOC 2007, 2012 and Youtube-Objects data sets.
Kai-Wen Cheng, Yie-Tarng Chen, Wen-Hsien Fang
IEEE Trans. Circuits Syst. Video Technol.2
2017 Multiple path search for action tube detection in videos
abstract
This paper presents an efficient convolutional neural network (CNN)-based multiple path search (MPS) algorithm to detect multiple spatial-temporal action tubes in videos. With the pass information and the accumulated scores generated by forward message passing, the new algorithm reuses these information to simultaneously find multiple paths in backward path tracing without repeating the search process. Moreover, to rectify the potentially inaccurate bounding boxes, we also propose a video localization refinement scheme to further boost the detection accuracy. Simulations show that the proposed algorithm provides competing performance compared with the main state-of-the-art works on the widespread UCF-101 dataset with yet lower complexity.
Erick Hendra Putra Alwando, Yie-Tarng Chen, Wen-Hsien Fang
ICIP2
2017 Temporal aggregation for first-person action recognition using Hilbert-Huang transform
abstract
This paper presents a new approach for action recognition in the first-person videos which aggregates both of the short- and long-term trends based on the coefficients of the Hilbert-Huang transform (HHT), a renowned time-frequency analysis tool. In contrast to previous works like Pooled Time Series (PoT), the new scheme can extract the salient features of activities based on the non-stationary HHT analysis, which consists of empirical mode decomposition and Hilbert spectral analysis, and can be incorporated with the convolutional neural network (CNN) features such as trajectory pooled CNN features to achieve superior detection accuracy. Conducted simulations show that the proposed method outperforms the main state-of-the-art works on two widespread public first-person datasets.
Didik Purwanto, Yie-Tarng Chen, Wen-Hsien Fang
ICME2
2016 Iterative localization refinement in convolutional neural networks for improved object detection
abstract
Accurate region proposals are of importance to facilitate object localization in the existing convolutional neural network (CNN)-based object detection methods. This paper presents a novel iterative localization refinement (ILR) method which, undertaken at a mid-layer of a CNN architecture, iteratively refines region proposals in order to match as much ground-truth as possible. The search for the desired bounding box in each iteration is first formulated as a statistical hypothesis testing problem and then solved by a divide-and-conquer paradigm. The proposed ILR is not only data-driven, free of learning, but also compatible with a variety of CNNs. Furthermore, to reduce complexity, an approximate variant based on a refined sampling strategy using linear interpolation is addressed. Simulations show that the proposed method improves the main state-of-the-art works on the PASCAL VOC 2007 dataset.
Kai-Wen Cheng, Yie-Tarng Chen, Wen-Hsien Fang
ICIP2
2016 An efficient subsequence search for video anomaly detection and localization
Kai-Wen Cheng, Yie-Tarng Chen, Wen-Hsien Fang
Multim. Tools Appl.2
2016 Importance Sampling-Based Maximum Likelihood Estimation for Multidimensional Harmonic Retrieval
abstract
This letter addresses a maximum likelihood (ML) algorithm for multidimensional (m-D) harmonic retrieval (MHR) problems. The new algorithm iteratively estimates the parameters in a rough to fine manner, intervened with filtering processes to separate the signals into appropriate groups. To facilitate implementations of the ML estimation, a Monte Carlo method, importance sampling (IS), and the theorem of Pincus are utilized to determine the ML estimates. Moreover, the pairing of the estimated parameters is automatically achieved without extra overhead. Conducted simulations demonstrate that the new algorithm outperforms the main state-of-the-art works and can achieve the Cramer-Rao lower bound (CRLB) even in low signal-to-noise ratio (SNR) scenarios.
Wen-Hsien Fang, Yi-Chiao Lee, Yie-Tarng Chen
IEEE Signal Process. Lett.3
2015 Video anomaly detection and localization using hierarchical feature representation and Gaussian process regression
abstract
This paper presents a hierarchical framework for detecting local and global anomalies via hierarchical feature representation and Gaussian process regression. While local anomaly is typically detected as a 3D pattern matching problem, we are more interested in global anomaly that involves multiple normal events interacting in an unusual manner such as car accident. To simultaneously detect local and global anomalies, we formulate the extraction of normal interactions from training video as the problem of efficiently finding the frequent geometric relations of the nearby sparse spatio-temporal interest points. A codebook of interaction templates is then constructed and modeled using Gaussian process regression. A novel inference method for computing the likelihood of an observed interaction is also proposed. As such, our model is robust to slight topological deformations and can handle the noise and data unbalance problems in the training data. Simulations show that our system outperforms the main state-of-the-art methods on this topic and achieves at least 80% detection rates based on three challenging datasets.
Kai-Wen Cheng, Yie-Tarng Chen, Wen-Hsien Fang
CVPR2
2015 Gaussian Process Regression-Based Video Anomaly Detection and Localization With Hierarchical Feature Representation
abstract
This paper presents a hierarchical framework for detecting local and global anomalies via hierarchical feature representation and Gaussian process regression (GPR) which is fully non-parametric and robust to the noisy training data, and supports sparse features. While most research on anomaly detection has focused more on detecting local anomalies, we are more interested in global anomalies that involve multiple normal events interacting in an unusual manner, such as car accidents. To simultaneously detect local and global anomalies, we cast the extraction of normal interactions from the training videos as a problem of finding the frequent geometric relations of the nearby sparse spatio-temporal interest points (STIPs). A codebook of interaction templates is then constructed and modeled using the GPR, based on which a novel inference method for computing the likelihood of an observed interaction is also developed. Thereafter, these local likelihood scores are integrated into globally consistent anomaly masks, from which anomalies can be succinctly identified. To the best of our knowledge, it is the first time GPR is employed to model the relationship of the nearby STIPs for anomaly detection. Simulations based on four widespread datasets show that the new method outperforms the main state-of-the-art methods with lower computational burden.
Kai-Wen Cheng, Yie-Tarng Chen, Wen-Hsien Fang
IEEE Trans. Image Process.2
2012 A Novel Subspace Decomposition-Based Detection Scheme with Soft Interference Cancellation for OFDMA Uplink
abstract
In this paper we propose a novel subspace decomposition-based detection scheme with the assistance of soft interference cancellation in the uplink of interleaved orthogonal frequency division multiple access (OFDMA) systems. By utilizing the inherent data structure, the interference is first separated with the desired symbol and then further decomposed into the one caused by the residues of decision errors and the other one by the undetected symbols in the successive interference cancellation (SIC) process.With such an ingenious interference decomposition along with the soft processing scheme, the new receiver can render more thorough interference cancellation, which in turn entails enhanced system performance. Moreover, for practical implementations, a low-complexity version, which only deal with the principal components of inter-carrier interference (ICI), is also addressed. Conducted simulations show that the developed receiver and its low-complexity implementation can provide superior performance compared with pervious works and is resilient to the presence of carrier-frequency offsets (CFOs). The low complexity implementation, in particular, requires substantially lower computational overhead with only slight performance loss.
Yung-Ping Tu, Wen-Hsien Fang, Yie-Tarng Chen
VTC Spring3
2012 A two-stage receiver with soft interference cancellation for space-time block code and spatial multiplexing combined systems
abstract
Abstract This paper presents a new space–time two‐stage receiver with the assistance of soft information for the Alamouti space–time block code (STBC) and spatially multiplexing (SM) combined multiple‐input multiple‐output (MIMO) systems, which possess both the advantages of high diversity gain and high data rates to entail the next generation wireless communication systems. The first stage of the receiver, utilizing the inherent structure of the STBC, consists of a bank of soft generalized sidelobe canceller (GSC)‐based detectors, each for every STBC block, and intends to yield a more precise initial estimate of the transmitted symbols. In the second stage, the groupwise detection is conducted successively by using the matched filters (MFs) to simultaneously detect the two consecutive symbols in one STBC block with the removal of the soft interferences in between. Since the interferences have been faithfully reproduced and thoroughly annihilated, the new receiver can yield accurate symbol detection even with simple MFs. Moreover, some extreme cases regarding the soft information employed in the new receiver and its extension to the multiuser (MU) MIMO downlink are addressed as well. Conducted simulations show that the developed receiver, with modest computational load, can provide superior performance compared with pervious works, especially in the MU MIMO downlink. Copyright © 2010 John Wiley & Sons, Ltd.
Yung-Ping Tu, Wen-Hsien Fang, Yie-Tarng Chen
Wirel. Commun. Mob. Comput.3
2011 Two-stage power allocation for amplify-andforward cooperative networks with distributed gabba space-time codes
abstract
This study presents a two-stage power allocation for the amplify-and-forward (AF) cooperative networks with distributed generalised ABBA (GABBA) space–time codes. The new power allocation scheme first determines the transmit power between the source node and the relay nodes by maximising the instantaneous rate, and thereafter optimises the power distribution among the relay nodes via the water-filling. Also, a maximum-likelihood detection, which makes use of the encoding structure of the distributed GABBA space–time codes, is addressed to alleviate the computational overhead. Moreover, a performance analysis including the array gain and the diversity gain is also scrutinised to provide further insights into the cooperative networks considered, where the destination is equipped with multiple antennas. Conducted simulations show that the GABBA coded AF cooperative networks incorporated with the proposed two-stage power allocation can attain close or even superior performance compared with previous works but with substantially reduced computational complexity.
Hung-Shiou Chen, Wen-Hsien Fang, Yie-Tarng Chen
IET Commun.3
2011 Joint source and relay power allocation in amplify-and-forward relay networks: a unified geometric programming framework
abstract
This study presents some joint source and relay power allocation algorithms in the amplify-and-forward (AF) relay networks using the efficacious geometric programming (GP). According to the constraints, approximate expressions of the received signal-to-noise ratio are first obtained. Thereafter, the problems are cast into appropriate GP forms according to the constraints. The power allocated to the source(s) and to the relays is then determined iteratively via the single condensation method. The new GP approach is shown to be applicable to a variety of constraints by using all of the relays for assistance and is amenable to asymmetric channels in both the single-user and multi-user AF relay networks. Conducted simulations show that the proposed power allocation schemes can attain superior performance compared with the previous works under various constraints in miscellaneous scenarios.
Wen-Hsien Fang, M.-J. Deng, Yie-Tarng Chen
IET Commun.3
2011 Genetic algorithm-assisted joint quantised precoding and transmit antenna selection in multi-user multi-input multi-output systems
abstract
This study presents a simple and efficient genetic algorithm-assisted approach for joint quantised precoding and transmit antenna selection based on the criterion of maximum capacity. The objective is to alleviate the effect of multi-user interference and to reduce hardware costs, such as the cost of radio frequency chains associated with antennas in the downlink of multi-input multi-output systems with limited feedback. To avoid the enormous search effort required by existing approaches, the authors propose a novel variant of the conventional genetic algorithm, called the hybrid genetic algorithm, in which each chromosome is divided into a bit string for precoding vector selection and an integer string for transmit antenna selection. In addition, new crossover and mutation operations are employed to accommodate these new chromosomes. The results of simulations show that the performance of the proposed approach is close to that of the exhaustive search method, but its computational complexity is substantially lower.
Wen-Hsien Fang, Shen-Chia Huang, Yie-Tarng Chen
IET Commun.3
2010 A hybrid human fall detection scheme
abstract
This paper presents a novel video-based human fall detection system that can detect a human fall in real-time with a high detection rate. This fall detection system is based on an ingenious combination of skeleton feature and human shape variation, which can efficiently distinguish “fall-down” activities from “fall-like” ones. The experimental results indicate that the proposed human fall detection system can achieve a high detection rate and low false alarm rate.
Yie-Tarng Chen, Yu-Ching Lin, Wen-Hsien Fang
ICIP1
2010 Relaying Through Distributed GABBA Space-Time Coded Amplify-and-Forward Cooperative Networks With Two-Stage Power Allocation
abstract
This paper presents a two-stage power allocation for the distributed GABBA space-time coded amplify-and-forward (AF) cooperative networks . The new power allocation scheme first determines the transmit power between the source node and the relay nodes by maximizing the instantaneous rate, and thereafter optimizes the power distribution among the relay nodes via the watering filling. Moreover, a maximum-likelihood (ML) detection, which makes use of the encoding structure of the GABBA code, is addressed to alleviate the computational overhead. Conducted simulations show that the GABBA coded AF cooperative networks incorporated with the proposed twostage power allocation can attain close performance as the opportunistic relaying reported in the literature but with substantially reduced computational complexity.
Hung-Shiou Chen, Wen-Hsien Fang, Yie-Tarng Chen
VTC Spring3
2010 Hybrid Genetic Algorithm for Joint Precoding and Transmit Antenna Selection in Multiuser MIMO Systems with Limited Feedback
abstract
To alleviate the interference while lowering the hardware cost such as the RF chains associated with antennas in the downlink of multiuser multi-input multi-output (MIMO) systems with limited feedback, this paper presents a simple, yet effective approach for joint precoding and transmit antenna selection with the assistance of the genetic algorithm (GA). To overcome the enormous amount of search called for, a novel variant of the conventional GA is addressed, where each chromosome consists of a bit string for the precoding vector selection and an integer string for the transmit antenna selection. A new crossover operation and a new mutation operation are also considered for the new chromosome. Conducted simulations show that the new approach yields close performance as the exhaustive search approach but with substantially reduced computational complexity.
Shen-Chia Huang, Wen-Hsien Fang, Hung-Shiou Chen, Yie-Tarng Chen
VTC Spring4
2010 Joint Direction Finding and Propagation Delay Estimation in the Presence of Mutual Coupling
abstract
In this paper we propose a Multiple-SIgnal-Classification (MUSIC) based algorithm for joint direction finding and delay estimation of the impinging rays in the presence of antenna mutual coupling. We show that with the addition of auxiliary sensors on both sides of the antenna array the algorithm developed earlier can be extended to this scenario. The new algorithm also consists of three stages of the one-dimensional (1-D) MUSIC algorithms which alternatively estimate the impinging directions of arrival (DOAs) and delays in a hierarchical tree structure. Moreover, a constrained temporal filtering process or a constrained spatial beamforming process are employed to partitioned progressively the signals with either close DOAs or close delays into finer subgroups, aiming at higher estimation accuracy and lower computational load. In addition, the developed algorithm proceeds in a tree structure, so the estimated parameters are automatically paired. Conducted simulation results show that the new algorithm provides satisfactory performance but with drastically reduced computations compared with previous work in the presence of mutual coupling.
Chun-Hung Lin, Wen-Hsien Fang, Van-Khang Vu, Yie-Tarng Chen
VTC Spring4
2010 A Two-Stage Receiver with Soft Interference Cancellation for Space-Time Block Code and Spatial Multiplexing Combined Systems
abstract
This paper presents a new space-time two-stage receiver with the assistance of soft information for the Alamouti space-time block code (STBC) and spatially multiplexing (SM) combined systems. The first stage of the receiver consists of a set of soft generalized sidelobe canceller (GSC)-based detectors, each for every STBC group and is intended to yield a more precise initial estimate of the transmitted symbols. In the second stage, the groupwise detection is conducted successively using the matched filters (MF) with the removal of the soft interferences in between. Conducted simulations show that the developed receiver, with modest computational load, can provide superior performance compared with pervious works, especially in multiple user (MU)downlink MIMO scenarios.
Yung-Ping Tu, Wen-Hsien Fang, Tsung-Yu Tsai, Yie-Tarng Chen
VTC Spring4
2010 Joint Carrier Frequency Offset and Direction of Arrival Estimation via Hierarchical ESPRIT for Interleaved OFDMA/SDMA Uplink Systems
abstract
In this paper, we propose an efficient algorithm to jointly estimate the directions of arrival (DOAs) and carrier frequency offsets (CFOs) in interleaved orthogonal frequency division multiple access / space division multiple access (OFDMA/SDMA) uplink networks. The algorithm makes use of the signal structure by estimating the CFOs and DOAs in a hierarchical tree structure, in which two CFO estimations and one DOA estimation are employed alternatively. One special feature in the proposed algorithm is that the algorithm proceeds in a coarse-fine manner with temporal filtering or spatial beamforming being invoked between the parameter estimations to decompose the signals progressively into subgroups so as to enhance the estimation accuracy and lower the computational overhead. Simulations show that the proposed algorithm can provide satisfactory performance with increased channel capacity.
Kuo-Hsiung Wu, Wen-Hsien Fang, Yie-Tarng Chen
VTC Spring3
2009 Antenna-array-assisted frequency offset estimation and data detection in an uplink multiuser MIMO-OFDM interference network
abstract
The high-density deployment of access points (APs) and their serious mutual interference have made both frequency acquisition and data detection even more difficult in wireless local area network (WLAN). In light of this, this paper presents an antenna-array-assisted algorithm to solve above problems in a multiuser multiple-input-multiple-output (MIMO) orthogonal frequency division multiplexing (OFDM) interference network. The algorithm begins with the estimation of the channel parameters, including the frequency offsets, delays, and angle selectivity. To make a good use of the array signal characteristics, these parameters are estimated in an frequency-angle-frequency (FAF) tree structure, in which two frequency estimations and one angle estimation are employed alternatively. One special feature in the FAF tree structure is that temporal filtering or spatial beamforming is invoked between the parameter estimations to decompose the signals so as to enhance the estimation accuracy. Thereafter, based on these parameter estimates, a data detection procedure is developed to mitigate both multiple access interference (MAI) and co-channel interference (CCI). Simulations show that the proposed algorithm can provide satisfactory performance even in networks with MAIs and CCIs sharing the same frequency band.
Kuo-Hsiung Wu, Wen-Hsien Fang, Yie-Tarng Chen, Jiunn-Tsair Chen
PIMRC3
2008 Opportunistic Uplink Retransmission Control with Active-User Estimation in Multi-User Packet CDMA Systems
abstract
This paper presents an uplink retransmission control scheme to improve the system throughput based on the active-user estimation capabilities offered by Kalman filtering in a multiuser single CDMA cell. The improvement is achieved by means of selecting the best subset of users with the power- controlled opportunistic retransmission control (PORC) based on the radio resources reservation and a joint PHY-MAC opportunity function, which also takes the channel information and the waiting time into account. Simulation results show that the proposed strategy exhibits significant improvement in terms of throughput and fairness.
Yie-Tarng Chen, Kuo-Liang Yeh, Wen-Hsien Fang
VTC Spring1
2008 Joint Generalized Antenna Combination and Symbol Detection Based on Minimum Bit Error Rate: A Particle Swarm Optimization Approach
abstract
In order to reduce hardware cost and achieve superior performance in multi-input multi-output (MIMO) systems, this paper proposes a novel scheme for joint antenna combination and symbol detection. More specifically, the new approach simultaneously determines the transformation weighting for antenna combination to lower the RF chains called for and to design the minimum bit error rate (MBER) detector to effectively mitigate the impairment due to interference. The joint decision statistic, however, is highly nonlinear and the particle swarm optimization (PSO) algorithm is employed to reduce the computational overhead. Conducted simulation results show that the new approach yields satisfactory performance with reduced computational overhead compared with pervious works.
Kyar-Chan Huang, Wen-Hsien Fang, Hoang-Yang Lu, Yie-Tarng Chen
VTC Spring4
2008 Iterative Multiuser Detection with Soft Interference Cancellation for Multirate MC-CDMA Systems
abstract
This paper presents an effective multi-rate multiuser detector (MUD) for the uplink of single-input multiple- output (SIMO) multi-carrier code division multiple access (MC- CDMA) systems. The MUD considered is an iterative receiver which utilizes the soft information to refine the estimation of the interference to enhance the interference cancellation capability. More specifically, users with different transmission rates are classified into separate groups and, in each iteration, these groups of users are detected sequentially based on a set of minimum mean-squared error (MMSE) group detectors with the removal of multiple access interferences (MAI) group by group. Furthermore, the estimated interferences in each group, either from the same or the other groups, are refined successively with the assistance of the soft information in the symbol detection process. Conducted simulations show that the proposed MUD, with moderate computational overhead, can effectively suppress the MAI to render superior performance compared with previous works.
Yung-Ping Tu, Wen-Hsien Fang, Hoang-Yang Lu, Yie-Tarng Chen
VTC Spring4
2007 Heterogenous Information Aided Semiblind Group MUD for MIMO MC-CDMA Systems
abstract
This paper presents an effective semiblind group multiuser detector (MUD) for uplink multiple-input multiple- output (MIMO) multi-carrier code division multiple access (MC- CDMA) systems. The MUD considered is a two-stage linear constrained minimum variance (LCMV) detector which uses heterogeneous information (both hard and soft) to enhance the interference cancellation capability. The first stage LCMV detector produces a tentative detection output, while at the second stage a bank of soft LCMV detectors successively refine the estimated interference group by group with the assistance of the heterogeneous information in the symbol detection process. In addition, to avoid the annoying ordering mechanism, a channel- strength criterion is employed to order the groups at the second stage. Conducted simulations show that the proposed MUD can drastically enhance the performance compared with previous works, especially when the hardware cost is at a premium.
Hoang-Yang Lu, Wen-Hsien Fang, Yie-Tarng Chen, Kuo-Liang Yeh
VTC Fall3
2003 An efficient packet classification algorithm for network processors
abstract
The exponential growth in optical link speed has stressed the performance of routers and switches. Consequently, a new breed of microprocessors, called network processors, are designed and fabricated specifically to effectively process packets on switches and routers. Packet classification is a major function in network processors to fit requirements of next-generation Internet. In this paper, we present a hardware-based packet classification algorithm for network processors. The innovative aspect of the proposed algorithm is to use the prior knowledge of rule characteristics to avoid performance fluctuation under different rule characteristics. First, we use divide-and-conquer approach to partition rules into several clusters and perform parallel search in different clusters. Then, we encode each rule into shorter bit string to prune unnecessary search. Finally, we employ level compression scheme to accelerate the lookup time. By running an intensive computer simulation, we show that the performance of the proposed simulation can achieve 8 million packets by 549 KB 10-ns SRAM for 20000 four-dimensional rules. This result demonstrates that the proposed scheme is superior to previous approaches.
Yie-Tarng Chen, Shin-Shian Lee
ICC1
2001 A novel signature-based packet classification
abstract
This paper describes a novel signature-based packet classification that can achieve gigabit speed at limited memory consumption. The innovative aspect of signature-based scheme is to extract the rule into an equivalent signature, a unique variable-length bit string with shorter width. Therefore, only a small fraction of a rule is inspected in search, resulting in considerable saving in lookup time as well as providing an effective solution for high dimensional rule. Moreover, the signature-based packet classification can perform well at high dimensions. By running a simulation model that incorporates the publicly available packet traces, we show that the performance of the signature-based scheme can reach 11 million packets per second even in the worst case, when implemented by 3.96 M 10 ns SRAM for 10000 rule four-dimensional classifier. This result demonstrates signature-based scheme is superior to previous approaches.
Yie-Tarng Chen, Ya-Hsin Yang
GLOBECOM1