VLDB 2026 Research / reviewers in the wild / expert
Yu Hen Hu
dblp:54/4645 · also Yu-Hen Hu, Yuhen Hu
· DBLP profile ↗
182ranked-venue papers
24as first author
15since 2021 · last 2026
0000-0003-3427-0677ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 96 · 17 first-author · 5 since 2021Systems, architecture and hardware · 51 · 6 first-author · 2 since 2021Computer networks · 17 · 2 since 2021Artificial intelligence and machine learning · 10 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Ubiquitous Vision-Augmented Intelligent Communications: A Generalized Loosely Coupled ApproachabstractExisting studies have shown that merging sensory data, especially visual signals, into Radio Frequency (RF) networks can enhance their perception and adaptability to hidden disruptive environmental factors (e.g., blockages). A majority of these methods assume tight integration of the RF and non-RF components within a wireless system, making them rely heavily on prior network deployment knowledge and applicable to only specific deployments. Aiming at ubiquitous physically assisted intelligent communications, this article presents a new LOOSELY COUPLED approach that incorporates the videos from widely installed external surveillance cameras to aid the operation of existing networks. To enable adaptive integration across individual sensing and communication systems in various unknown deployments, we devise a novelobject-centric hierarchical learningscheme that can combine decoupled visual and RF signals to autonomously discover the compositional structure of the blockage scene and analyze its spatiotemporal development for inferring the future blockage condition changes. We demonstrate the application of these structural representations in building a generalized Vision-Perceptive Link Quality Predictor (VP-LQP). To support the training and testing of VP-LQP, we construct a new real-world multi-scenario dataset, and the experimental results validate the superiority of our design. Ming Xia 0005, Ziyang Lin, Jiaquan Jin, Yu Hen Hu, Zhen Cheng 0001, Kaikai Chi |
IEEE Trans. Wirel. Commun. | 4 |
| 2025 | From Prototypes to General Distributions: An Efficient Curriculum for Masked Image ModelingabstractMasked Image Modeling (MIM) has emerged as a powerful self-supervised learning paradigm for visual representation learning, enabling models to acquire rich visual representations by predicting masked portions of images from their visible regions. While this approach has shown promising results, we hypothesize that its effectiveness may be limited by optimization challenges during early training stages, where models are expected to learn complex image distributions from partial observations before developing basic visual processing capabilities. To address this limitation, we propose a prototype-driven curriculum learning framework that structures the learning process to progress from prototypical examples to more complex variations in the dataset. Our approach introduces a temperature-based annealing scheme that gradually expands the training distribution, enabling more stable and efficient learning trajectories. Through extensive experiments on ImageNet-1K, we demonstrate that our curriculum learning strategy significantly improves both training efficiency and representation quality while requiring substantially fewer training epochs compared to standard Masked Auto-Encoding. Our findings suggest that carefully controlling the order of training examples plays a crucial role in self-supervised visual learning, providing a practical solution to the early-stage optimization challenges in MIM. Jinhong Lin, Cheng-En Wu, Huanran Li, Jifan Zhang, Yu Hen Hu, Pedro Morgado 0001 |
CVPR | 5 |
| 2025 | Patch Ranking: Token Pruning as Ranking Prediction for Efficient CLIPabstractContrastive image-text pre-trained models such as CLIP have shown remarkable adaptability to downstream tasks. However, they face challenges due to the high computational requirements of the Vision Transformer (ViT) backbone. Current strategies to boost ViT efficiency focus on pruning patch tokens but fall short in addressing the multimodal nature of CLIP and identifying the optimal subset of tokens for maximum performance. To address this, we propose greedy search methods to establish a “Golden Ranking” and introduce a lightweight predictor specifically trained to approximate this Ranking. To compensate for any performance degradation resulting from token pruning, we incorporate learnable visual tokens that aid in restoring and potentially enhancing the model's performance. Our work presents a comprehensive and systematic investigation of pruning tokens within the ViT backbone of CLIP models. Through our framework, we successfully reduced 40% of patch tokens in CLIP's ViT while only suffering a minimal average accuracy loss of 0.3% across seven datasets. Our study lays the groundwork for building more computationally efficient multimodal models without sacrificing their performance, addressing a key challenge in the application of advanced vision-language models.11Project Page: https://github.com/CEWu/PatchRanking Cheng-En Wu, Jinhong Lin, Yu Hen Hu, Pedro Morgado 0001 |
WACV | 3 |
| 2025 | A Single-Camera Method for Estimating Lift Asymmetry Angles Using Deep Learning Computer Vision AlgorithmsabstractA computer vision (CV) method to automatically measure the revised NIOSH lifting equation asymmetry angle (A) from a single camera is described and tested. A laboratory study involving ten participants performing various lifts was used to estimateAin comparison to ground truth joint coordinates obtained using 3-D motion capture (MoCap). To address challenges, such as obstructed views and limitations in camera placement in real-world scenarios, the CV method utilized video-derived coordinates from a selected set of landmarks. A 2-D pose estimator (HR-Net) detected landmark coordinates in each video frame, and a 3-D algorithm (VideoPose3D) estimated the depth of each 2-D landmark by analyzing its trajectories. The mean absolute precision error for the CV method, compared to MoCap measurements using the same subset of landmarks for estimatingA, was 6.25° (SD = 10.19°, N = 360). The mean absolute accuracy error of the CV method, compared against conventional MoCap landmark markers was 9.45° (SD = 14.01°,N= 360). Zhengyang Lou, Zitong Zhan, Yin Li 0003, Yu Hen Hu, Ming-Lun Lu, Dwight Werren, Robert G. Radwin |
IEEE Trans. Hum. Mach. Syst. | 5 |
| 2024 | AI-Based Automatic System for Assessing Upper-Limb Spasticity of Patients With Stroke Through Voluntary MovementabstractSpasticity is a common complication for patients with stroke, but only few studies investigate the relation between spasticity and voluntary movement. This study proposed a novel automatic system for assessing the severity of spasticity (SS) of four upper-limb joints, including the elbow, wrist, thumb, and fingers, through voluntary movements. A wearable system which combined 19 inertial measurement units and a pressure ball was proposed to collect the kinematic and force information when the participants perform four tasks, namely cone stacking (CS), fast flexion and extension (FFE), slow ball squeezing (SBS), and fast ball squeezing (FBS). Several time and frequency domain features were extracted from the collected data, and two feature selection approaches based on recursive feature elimination were adopted to select the most influential features. The selected features were input into five machine learning techniques for assessing the SS for each joint. The results indicated that using CS task to assess the SS of elbow and fingers and using FBS task to assess the SS of thumb and wrist can reach the highest weighted-average F1-score. Furthermore, the study also concluded that FBS is the optimal task for assessing all the four upper-limb joints. The overall result shown that the proposed automatic system can assess four upper-limb joints through voluntary movements accurately, which is a breakthrough of finding the relation between spasticity and voluntary movement. I-Jung Lee, Yu Hen Hu, Pei-Chi Hsiao, Shu-Yu Yang, Hsin-Te Lin, Yu-Chung Chen, Bor-Shing Lin |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Why Is Prompt Tuning for Vision-Language Models Robust to Noisy Labels?abstractVision-language models such as CLIP [27] learn a generic text-image embedding from large-scale training data. A vision-language model can be adapted to a new classification task through few-shot prompt tuning. We find that such a prompt tuning process is highly robust to label noises. This intrigues us to study the key reasons contributing to the robustness of the prompt tuning paradigm. We conducted extensive experiments to explore this property and find the key factors are: 1) the fixed classname tokens provide a strong regularization to the optimization of the model, reducing gradients induced by the noisy samples; 2) the powerful pre-trained image-text embedding that is learned from diverse and generic web data provides strong prior knowledge for image classification. Further, we demonstrate that noisy zero-shot predictions from CLIP can be used to tune its own prompt, significantly enhancing prediction accuracy in the unsupervised setting. The code is available at https://github.com/CEWu/PTNL. Cheng-En Wu, Haichao Yu, Pedro Morgado 0001, Yu Hen Hu |
ICCV | 6 |
| 2023 | Efficient online real-time video stabilization with a novel least squares formulation and parallel AC-RANSAC
Jianwei Ke, Alex J. Watras, Jae-Jun Kim, Hewei Liu, Hongrui Jiang, Yu Hen Hu |
J. Vis. Commun. Image Represent. | 6 |
| 2023 | Physical-Assisted Routing for Proactive Avoidance of Nomadic Obstacles in IoTabstractWith the broadening of the radio spectrum to higher frequency bands, wireless links are more prone to blockages by nomadic obstacles. However, existing routing schemes mostly follow the network-oriented design principle, which makes it difficult to react quickly to sudden obstruction. This article proposes PAR, a novel physical-assisted routing scheme for the Internet of Things. PAR takes a physical-oriented viewpoint attempting to mitigate unexpected link blockages by leveraging sensor observations of obstacles at individual nodes. It analyzes the signatures of sensor measurements to directly estimate the positions and sizes of obstacles and to proactively infer remedial routing decisions even before the transmission failure occurs. During the network deployment time, PAR lets each node learn and calibrate the geographic distribution of neighbors with respect to its sensor measurements by exchanging the sensor observations of a mobile guidance obstacle. During the runtime, an obstacle bypassing algorithm is then developed to find the shortest detour routes by comparing the local sensor measurements with the calibrated positions of neighbors to immediately resume data forwarding. We evaluate the efficacy of PAR using simulation and a real-world testbed. It is observed that PAR significantly improves the success rate while reducing redundant hops and routing control overhead. Ming Xia 0005, Jiaquan Jin, Biqian Liu, Yu Hen Hu, Xiaoyan Wang 0007, Kaikai Chi |
ACM Trans. Sens. Networks | 4 |
| 2022 | $\text{Edge}^{n}$ AI: Distributed Inference with Local Edge Devices and Minimal LatencyabstractWe propose$\text{Edge}^{n}$AI, a framework to decompose a complex deep neural networks (DNN) over$n$available local edge devices with minimal communication overhead and overall latency. Our framework creates small DNNs (SNNs) from an original DNN by partitioning its classes across the edge devices, while taking into account their available resources. Class-aware pruning is applied to aggressively reduce the size of the SNN on each edge device. The SNNs perform inference in parallel, and are configured to generate a ‘Don't Know’ response when an unassigned class is identified. Our experiments show up to 17X inference speedup compared to a recent work, on devices of at most 150 MB memory when distributing a variant of VGG-16 over 20 parallel edge devices. Maede Hemmat, Azadeh Davoodi, Yu Hen Hu |
ASP-DAC | 3 |
| 2022 | A two-level scheme for multiobjective multidebris active removal mission planning in low Earth orbits
Xiaolei Hou, Yong Liu 0025, Yu Hen Hu, Quan Pan 0001 |
Sci. China Inf. Sci. | 4 |
| 2022 | Video-Based Automatic Wrist Flexion and Extension ClassificationabstractA computer vision method was developed to automatically measure wrist flexion and extension from a 2-D video for occupational health and safety research. Marker-less tracked skeletal joints of the elbow, wrist, and hand estimated the wrist flexion/extension angle between the hand and forearm. Based on the estimated angles, wrist posture was classified as flexion (palmar bending), neutral (no bending), or extension (dorsal bending) for each cycle of hand movement. Applying to a set of laboratory videos of a simulated repetitive motion task, we demonstrated the feasibility of using this algorithm for assessing the state of hand activities during manual work. Tested on 1464 frames from 61 recorded videos for 16 participants, the algorithm achieved an average performance of 72.40% correct, per-class accuracy. The sensitivity and specificity for flexion were 66.16% and 91.47%, respectively. The sensitivity and specificity for extension were 77.12% and 89.72%, respectively. This compared favorably against a previously reported consistency rate of 57% between human analyst estimates and wrist electrogoniometer measured wrist flexion/extension angles. We also applied this technique to 262 video frames of hand flexion instances selected from industrial field video data. For these videos, the average correct per-class accuracy was 76.03% in comparison to human observers. The sensitivity and specificity for flexion were 69.23% and 94.17%, respectively, and the sensitivity and specificity for extension were 91.95% and 80.57%, respectively. Cheng-Hsien Lee, Yu Hen Hu, Stephen Bao, Robert G. Radwin |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2021 | Power Reduction of a Set-Associative Instruction Cache Using a Dynamic Early Tag LookupabstractAn energy-efficient instruction cache lookup technique with low area overheads is proposed. The key concept of this Dynamic Early Tag Lookup (DETL) method is to exploit the presence of instruction fetch-bubble cycles. In a fetch-bubble cycle, the index of the matching cache set can be determined earlier. Hence, the dynamic energy for parallel memory accesses to irrelevant cache banks can be saved. We implemented the proposed DETL algorithm on a 4-way set-associative instruction cache in a RISC-V micro-architecture, and tested its performance using the SPEC CPU2006 benchmark suite. The experiment results showed a 19.38% dynamic power reduction with an area overhead smaller than 0.1 %. Chun-Chang Yu, Yu Hen Hu, Yi-Chang Lu, Charlie Chung-Ping Chen |
DATE | 2 |
| 2021 | Efficient Real-Time Video Stabilization with a Novel Least Squares FormulationabstractWe present a novel video stabilization algorithm (LSstab) that removes unwanted motions in real-time. LSstab is based on a novel least squares formulation of the smoothing cost function to alleviate the undesirable camera jitter. A recursive least square solver is derived to minimize the smoothing cost function with an O(N) computation complexity. LSstab is evaluated using a suite of publicly available videos against the state of the art video stabilization methods. Results show LSstab reaches comparable or better performance, achieving real-time processing speed when a GPU is used. Jianwei Ke, Alex J. Watras, Jae-Jun Kim, Hewei Liu, Hongrui Jiang, Yu Hen Hu |
ICASSP | 6 |
| 2021 | Multi-residual Connection Network for Edge Detection
Lide Wang, Yu Hen Hu |
Neural Process. Lett. | 4 |
| 2021 | Load Asymmetry Angle Estimation Using Multiple-View VideosabstractA robust computer vision-based approach is developed to estimate the load asymmetry angle defined in the revised NIOSH lifting equation (RNLE). The angle of asymmetry enables the computation of a recommended weight limit for repetitive lifting operations in a workplace to prevent lower back injuries. An open-source package OpenPose is applied to estimate the 2D locations of skeletal joints of the worker from two synchronous videos. Combining these joint location estimates, a computer vision correspondence and depth estimation method is developed to estimate the 3D coordinates of skeletal joints during lifting. The angle of asymmetry is then deduced from a subset of these 3D positions. Error analysis reveals unreliable angle estimates due to occlusions of upper limbs. A robust angle estimation method that mitigates this challenge is developed. We propose a method to flag unreliable angle estimates based on the average confidence level of 2D joint estimates provided by OpenPose. An optimal threshold is derived that balances the percentage variance reduction of the estimation error and the percentage of angle estimates flagged. Tested with 360 lifting instances in a NIOSH-provided dataset, the standard deviation of angle estimation error is reduced from 10.13° to 4.99°. To realize this error variance reduction, 34% of estimated angles are flagged and require further validation. Xuan Wang 0022, Yu Hen Hu, Ming-Lun Lu, Robert G. Radwin |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2020 | Towards Real-Time, Multi-View Video StereopsisabstractWe present a real-time, multi-view video stereopsis (RTMVS) algorithm. This algorithm processes five synchronized video streams from cameras of a stationary camera array using a commodity laptop computer equipped with an Nvidia GPU. It provides 3D visualization of a dynamic scene from a chosen viewpoint at the video frame rate. In RTMVS, 3D surfaces are represented as a set of triangles anchored on a sparse set of 3D feature points. The computationally intensive Structure-from-Motion (SfM) algorithm is executed as an initial step. Feature points in each video stream are tracked using a KLT tracker. Triangles will be updated only when at least one vertex moves from its current position. The algorithm will redetect features every X frame. Epipolar geometry and Trifocal tensor are also exploited to accelerate sparse feature point matching and track filtering. Compared to a dense point cloud multi-view stereopsis baseline algorithm, RTMVS reduces the processing time per frame from 30s down to less than 44 ms. Jianwei Ke, Alex J. Watras, Jae-Jun Kim, Hewei Liu, Hongrui Jiang, Yu Hen Hu |
ICASSP | 6 |
| 2020 | An Adaptive Blind Cyclic Calibration Technique for Two-Channel Frequency-Interleaved ADCsabstractA digital adaptive blind calibration technique for a two-channel frequency-interleaved analog-to-digital converter (2 FI-ADC) is proposed. It simultaneously estimates and corrects the inherent circuit errors, including spectral leakage, harmonic folding, jitters, I/Q and aliasing images, in a cyclic manner. The effectiveness of this novel architecture is validated via extensive simulations. Jinpeng Song, Jianping An, Yu Hen Hu, Yu Zhao 0016, Jian Gao 0020 |
ISCAS | 3 |
| 2020 | Two motion models for improving video object tracking performance
Lide Wang, Yu Hen Hu |
Comput. Vis. Image Underst. | 3 |
| 2020 | Efficient Proposals: Scale Estimation for Object Proposals in Pedestrian Detection TasksabstractDue to projective transformation, a great variety of pedestrian sizes appears on the image depending on their depths in the real world. In this letter, we analyze the object proposal generation strategies in existing object detection methods that underperform in some real-world applications owing to ignorance of spatial factors. To this end, we propose a neural network predictor to estimate the sizes of object proposals based on their locations on the image. Furthermore, our model can efficiently generate high-quality proposals using very few training samples with the help of a data augmentation strategy. The proposed size estimation method is compared against several state-of-art object proposal estimation methods by two metrics on a driving dataset and a train station surveillance dataset which shows significant performance advantages. Lide Wang, Yu Hen Hu |
IEEE Signal Process. Lett. | 4 |
| 2020 | Video Salient Object Detection via Robust Seeds Extraction and Multi-Graphs Manifold PropagationabstractVideo salient object detection aims at distinguishing the salient objects from the complex background and highlighting them uniformly in the spatiotemporal domain, which still suffers from the interference of the complicated dynamic background in unconstrained videos. To address this problem, we propose a novel coarse-to-fine spatiotemporal salient object detection method. Specifically, we first model a novel motion energy to exclude the motion noise by exploiting the motion magnitude and motion orientation. Then, a supervoxel-level inter-frame graph model is constructed for each pair of adjacent frames independently, and a robust graph clustering-based saliency seed generation method is proposed to produce a coarse saliency map. Furthermore, the supervoxel-level inter-frame graph model is reconstructed by considering the regional spatiotemporal consistency constraint based on the coarse saliency map. The prior information obtained from pixel clustering is also taken into account to optimize the weight of the inter-frame graph model. Finally, a multi-graphs saliency propagation method is exploited under the manifold regularization framework by fusing the motion energy and appearance feature to refine the coarse saliency map. The extensive experiments on two widely used datasets validate the effectiveness and superiority of the proposed method against 13 state-of-the-art methods in terms of PR-curves, scores of S-measure,$F_{\beta }$, and MAE. Bing Liu 0022, Junbao Li, Yu Hen Hu, Shou Feng |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | PACE: Physically-Assisted Channel EstimationabstractRadio link quality is highly influenced by changes in the physical environment. To sustain reliable and efficient data delivery, link quality estimation is essential for Cyber-Physical Systems (CPSs) or Internet of Things (IoT). Network-based link quality estimation methods estimate the link quality by monitoring data transmissions. In a dynamic environment, the accuracy of link quality so estimated may become degraded because the accuracy must be balanced against the overhead of data transmissions. In this work, we propose to incorporate sensor readings available in a CPS/IoT system to augment existing link quality estimation. We call this a Physical-Assisted Channel Estimator (PACE). By analyzing sensor readings that are highly correlated to the link quality, PACE may detect the change of link quality in real-time. Evaluation conducted on a real intelligent parking system shows that compared to existing network-based methods, PACE reacts to persistent disturbances much more quickly without sacrificing robustness to transient fluctuations, and achieves higher accuracy even under a low data transmission rate. With PACE, the data delivery performance of routing protocols can be significantly improved. We expect PACE to be the first milestone towards Physical-Assisted Cyber Systems (PACSs) for fulfilling the vision of environment-aware computing and communication. Ming Xia 0005, Biqian Liu, Yu Hen Hu, Kaikai Chi, Xiaoyan Wang 0007, Jiajia Liu 0001 |
IEEE Trans. Wirel. Commun. | 3 |
| 2019 | Learn to Detect: Improving the Accuracy of Earthquake DetectionabstractEarthquake early warning system uses high-speed computer network to transmit earthquake information to population center ahead of the arrival of destructive earthquake waves. This short (10 s of seconds) lead time will allow emergency responses such as turning off gas pipeline valves to be activated to mitigate potential disaster and casualties. However, the excessive false alarm rate of such a system imposes heavy cost in terms of loss of services, undue panics, and diminishing credibility of such a warning system. At the current, the decision algorithm to issue an early warning of the onset of an earthquake is often based on empirically chosen features and heuristically set thresholds and suffers from excessive false alarm rate. In this paper, we experimented with three advanced machine learning algorithms, namely, K-nearest neighbor (KNN), classification tree, and support vector machine (SVM) and compared their performance against a traditional criterion-based method. Using the seismic data collected by an experimental strong motion detection network in Taiwan for these experiments, we observed that the machine learning algorithms exhibit higher detection accuracy with much reduced false alarm rate. Tai-Lin Chin, Chin-Ya Huang, Shan-Hsiang Shen, You-Cheng Tsai, Yu Hen Hu, Yih-Min Wu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2019 | Video Saliency Detection via Graph Clustering With Motion Energy and Spatiotemporal ObjectnessabstractWe present a novel, robust estimation method to distinguish salient objects from complicated, dynamic backgrounds in videos. In this method, we propose a novel approach to model motion energy based on motion magnitude, motion orientation, gradient flow field, and spatial gradient of the video frame. Furthermore, an effective spatiotemporal objectness map is also proposed to estimate a compact object-like region in the current video frame leveraging both the objectness proposals and the saliency map of the previous frame. Then the current video frame is oversegmented into the granularity of superpixels using the simple linear iterative clustering algorithm. Each superpixel is designated as a node of a graph. The similarity between adjacent superpixels will be assigned as the weight of an edge that connects these two nodes. The feature values of motion energy and spatiotemporal objectness within each superpixel will be averaged respectively, and used to graphically cluster similar superpixels to form the detected salient object. Extensive experiments comparing this proposed new method against twelve existing salient object detection (SOD) methods have been performed using the benchmark datasets unconstrained video saliency detection and densely annotated video segmentation. Superior performance of this proposed SOD method has been observed through three well-known performance metrics: precision-recall curves, F-measure curves, and the mean absolute error. Bing Liu 0022, Junbao Li, Yu Hen Hu |
IEEE Trans. Multim. | 5 |
| 2018 | Exploring energy and accuracy tradeoff in structure simplification of trained deep neural networksabstractThis paper presents a structure simplification procedure that allows efficient energy and accuracy tradeoffs in implementation of trained deep neural networks (DNNs). This structure simplification procedure identifies and eliminates redundant neurons in any layer of the DNN based on the trained weights connected from these neurons. This procedure may be applied to all layers of a DNN. For each layer different configurations with Pareto-optimal accuracy and energy consumption are realized. Our work is the first to use energy-accuracy trade-offs to guide optimal structure realization of trained DNNs. After redundant neurons are discarded, the weights of remaining neurons will be updated using matrix multiplication without retraining. Yet, retraining may still be applied if desired to further fine tune the performance. In our experiments, we show energy-accuracy tradeoff provides clear guidance to achieve efficient realization of trained DNNs. We also observe significant implementation cost reductions with up to 33X in energy and 12X in memory while the performance (accuracy) loss is negligible. Boyu Zhang 0001, Azadeh Davoodi, Yu Hen Hu |
ASP-DAC | 3 |
| 2018 | Frame-Subsampled, Drift-Resilient Video Object TrackingabstractPerformance-cost trade-offs in video object tracking tasks for long video sequences is investigated. A novel frame-subsampled, drift-resilient (FSDR) video object tracking algorithm is presented that would achieve desired tracking accuracy while dramatically reducing computing time by processing only sub-sampled video frames. A new pattern matching score metric is proposed to estimate the probability of drifting. A drift-recovery procedure is developed to enable the algorithm to recover from a drift situation and resume accurate tracking. Compared against state-of-the-art video object tracking algorithms, dramatic performance (accuracy) enhancement and cost (computing time) reduction are observed. Xuan Wang 0022, Yu Hen Hu, Robert G. Radwin, John D. Lee |
ICASSP | 2 |
| 2018 | Adaptive Visual Target Tracking Based on Label Consistent K-Svd Sparse Coding and Kernel Particle FilterabstractWe propose an adaptive visual target tracking algorithm based on Label-Consistent K -Singular Value Decomposition (LC-KSVD) dictionary learning. To construct target templates, local patch features are sampled from foreground and background of the target. LC-KSVD then is applied to these local patches to simultaneously estimate a set of low-dimension dictionary and classification parameters (CP). To track the target over time, a kernel particle filter (KPF) is proposed that integrates both local and global motion information of the target. An adaptive template updating scheme is also developed to improve the robustness of the tracker. Experimental results demonstrate superior performance of the proposed algorithm over state-of-art visual target tracking algorithms in scenarios that include occlusion, background clutter, illumination change, target rotation and scale changes. Jinlong Yang 0002, Yu Hen Hu |
ICASSP | 3 |
| 2018 | Improving Disparity Map Estimation for Multi-View Noisy ImagesabstractA robust multi-view disparity estimation algorithm for noisy images is presented. The proposed algorithm constructs 3D focus image stacks (3DFIS) by projecting and stacking multi-view images and estimates a disparity map based on the 3DFIS. To make the algorithm robust to noise and occlusion, a texture-based view selection and patch size variation scheme based on texture map is proposed. Experiment results indicate that the proposed algorithm outperforms conventional stereo matching algorithms as well as previously reported multi-view disparity estimation algorithms under noisy conditions. Zhengyang Lou, Yu Hen Hu, Hongrui Jiang |
ICASSP | 3 |
| 2018 | Cross-Layer Design for Network Lifetime Maximization in Underwater Wireless Sensor NetworksabstractThis paper investigates the cross-layer design problem with the goal of maximizing the network lifetime for energy-constrained underwater wireless sensor networks (UWSNs). We first jointly consider link scheduling, transmission power and transmission rate in a proposed optimization problem with the adoption of time division multiple access (TDMA) schedules. Then, we propose an iterative algorithm to solve the optimization problem. It alternates between (1) link scheduling and (2) computation of transmission powers and transmission rates. In fact, the convergence of such iterative algorithm can be mathematically and empirically supported. We evaluate our algorithm for several network topologies. Extensive simulation results demonstrate the superiority of the proposed approach. Yuan Zhou 0006, Yu Hen Hu, Boyu Wang 0004, Sun-Yuan Kung |
ICC | 3 |
| 2018 | Frame-Sub Sampled, Drift-Resilient Long-Term Video Object TrackingabstractA novel frame-subsampled, drift-resilient (FSDR) video object tracking (VOT) algorithm is proposed. Two design goals are sought: to improve the accuracy and to reduce the processing time. The drifting problem is mitigated with a drift detector and accompanying drift recovery mechanism. When a drift is detected, the recovery mechanism provides an opportunity to put the tracking back on track. To gather context-dependent statistics required for these procedures, an initial short segment of the video sequence will be used as a training sequence. Thus, this algorithm is more suitable for video sequences much longer than several minutes. To reduce computing time, a novel frame-subsampling strategy is proposed to process the VOT on small subset of frames. The trajectory of the tracked object on frames that are skipped will be estimated via interpolation. Compared with state of art VOT algorithms, dramatic improvement of performance (accuracy) and orders of magnitude computing time reduction are observed. Xuan Wang 0022, Yu Hen Hu, Robert G. Radwin, John D. Lee |
ICME | 2 |
| 2018 | Multiple view image denoising using 3D focus image stacks
Zhengyang Lou, Yu Hen Hu, Hongrui Jiang |
Comput. Vis. Image Underst. | 3 |
| 2018 | Quantized Kalman Filter Tracking in Directional Sensor NetworksabstractMoving target tracking in directional sensor networks has recently attracted attention by right of special directional sensing features. Unlike omnidirectional sensors, the directional sensor senses the target only in the direction of its orientation. It can provide quantized direction that indicates the presence or absence of the target in the sensing field, rather than just the analog measurement of sensing signal with respect to the detected target. A quantized Kalman filter (QKF) based on both quantized directions and analog ranging measurements is derived in the minimum mean-square error (MMSE) sense. Its performances of mean square estimation error (MSE) and complexity are also analyzed. Then, a reduced-complexity QKF of high-accuracy is pursued to facilitate its implementation. It is proved that the QKF yields a smaller MSE than the traditional extended Kalman filter (EKF) merely based on analog measurements. The posterior Cramér-Rao lower bound (PCRLB) is introduced as the performance measure. The performance advantages of the proposed QKF are demonstrated using Monte Carlo simulations in a target tracking application using ultrasonic ranging sensors. Xiaoqing Hu, Ming Bao, Xiao-Ping Zhang 0002, Sha Wen, Xiaodong Li 0002, Yu Hen Hu |
IEEE Trans. Mob. Comput. | 6 |
| 2017 | Patch-based multiple view image denoising with occlusion handlingabstractA novel patch-based multi-view image denoising algorithm is proposed. This method leverages the 3D focus image stacks structure to exploit self-similarity and image redundancy inherent in multiple view images. Then a depth-guided adaptive window and dynamic view selection criterion is developed to aid proper selection of most consistent patches for the multi-view image denoising. Extensive experiments have been performed. Comparing the outcomes against those of state of the art image denoising algorithms, our proposed algorithm demonstrates significant performance advantage. Yu Hen Hu, Hongrui Jiang |
ICASSP | 2 |
| 2017 | Passive source localization from array covariance matrices via joint sparse representations
Ji-an Luo, Kai Yu 0005, Zhi Wang 0003, Yu Hen Hu |
Neurocomputing | 4 |
| 2016 | Atmospheric lidar imaging and poisson inverse problemsabstractThis paper describes an atmospheric lidar photon-limited imaging problem in which observations are contaminated with Poisson noise. The observations are a nonlinear function of two spatially varying physical parameters. The first parameter, called the transmittance, is known to be a bounded monotonic non-increasing function. The second parameter, called the backscatter cross-section, is non-negative and can be approximated with a piecewise constant function. Current statistical estimators in the lidar community do not take these constraints in account, and at times the estimates violate the physical properties. The standard practice is such that estimation of the parameters are not treated as a statistical inference imaging problem, except that the only imaging technique that is used is averaging over non-overlapping blocks to reduce the noise variance. The proposed method of this paper leverages novel Poisson image reconstruction algorithms with monotonicity constraints. Specifically, an isotonic regression algorithm called PAV is incorporated into total-variation regularized Poisson estimation methods. Simulations demonstrate that the proposed approach - relative to competitor algorithms - consistently yields improved estimates (smaller MSE) of the transmittance parameter, with a negligible deterioration of the backscatter cross-section. Willem J. Marais, Robert E. Holz, Yu Hen Hu, Rebecca Willett |
ICIP | 3 |
| 2016 | DOA Estimation From One-Bit Compressed Array Data via Joint Sparse RepresentationabstractA one-bit joint sparse representation direction of arrival (OBJSR-DOA) estimation approach is proposed in this letter. By exploiting the joint spatial and spectral correlations inherent in acoustic sensor array data, the proposed OBJSR-DOA approach provides reliable DOA estimation from only the sign bit of randomly subsampled acoustic sensor data. The random subsampling and single-bit quantization allow significant reduction of data to be transmitted to the fusion center without additional energy consumption requirement in the source coding/compression operation. Compared with existing compressive sensing-based DOA estimation methods, the superiority of the proposed approach in providing data volume reduction and performance improvement is verified by both simulations and field experiments using a prototype wireless sensor array network platform. Kai Yu 0005, Yimin Zhang 0001, Ming Bao, Yu Hen Hu, Zhi Wang 0003 |
IEEE Signal Process. Lett. | 4 |
| 2014 | Facial image de-identification using identiy subspace decompositionabstractHow to conceal the identity of a human face without covering the facial image? This is the question investigated in this work. Leveraging the high dimensional feature representation of a human face in an Active Appearance Model (AAM), a novel method called the identity subspace decomposition (ISD) method is proposed. Using ISD, the AAM feature space is deposed into an identity sensitive subspace and an identity insensitive subspace. By replacing the feature values in the identity sensitive subspace with the averaged values of k individuals, one may realize a k-anonymity de-identification process on facial images. We developed a heuristic approach to empirically select the AAM features corresponding to the identity sensitive subspace. We showed that after applying k-anonymity de-identification to AAM features in the identity sensitive subspace, the resulting facial images can no longer be distinguished by either human eyes or facial recognition algorithms. Hehua Chi, Yu Hen Hu |
ICASSP | 2 |
| 2014 | CS-Based Framework for Sparse Signal Transmission over Lossy LinkabstractIn this work, compressive sensing (CS) is applied to facilitate efficient wireless information transmission over lossy communication links. Inherently sparse data packets are transmitted without compression or error protection. The packet loss during transmission is modeled as a random sampling process of the transmitted data. The original signal then is reconstructed based on correctly received data packets using CS-based reconstruction method. No computations for source, channel coding or random measurement sampling will be required at the transmitter side. Thus, this method is suitable for applications where transmitters have extreme low power constraints such as wireless sensor networks. Compared with traditional error protection technique(automatic repeat request, data interleaving and interpolation), the proposed method delivers higher quality of sparse signal while significantly reducing energy consumption at transmitter as well as transmission latency. Liantao Wu, Kai Yu 0005, Yu Hen Hu, Zhi Wang 0003 |
MASS | 3 |
| 2014 | Generalised Kalman filter tracking with multiplicative measurement noise in a wireless sensor networkabstractA new generalised Kalman filtering algorithm using a multiplicative measurement noise model is developed for tracking moving targets in a wireless sensor network. This multiplicative error model facilitates more accurate characterisation of the distance dependence measurement errors of range‐estimating sensors. Two new formulations of extended Kalman filter (EKF) and unscented Kalman filter (UKF), called generalised EKF (GEKF) and generalised UKF (GUKF) are derived. Comparing with conventional EKF and UKF formulations, it is shown that GEKF and GUKF can achieve smaller tracking error than traditional EKF and UKF. Simulation results are also reported that demonstrated the superior performance of GEKF and GUKF over existing methods. Xiaoqing Hu, Yu Hen Hu, Bugong Xu |
IET Signal Process. | 2 |
| 2014 | Energy-Balanced Scheduling for Target Tracking in Wireless Sensor NetworksabstractA novel energy-balanced task-scheduling method is proposed that extends the lifespan of wireless sensor networks (WSNs) for collaborative target tracking using an unscented Kalman filter (UKF) algorithm. It is shown that the tracking accuracy is approximately proportional to the number of active sensor nodes participating in collaborative tracking. Excessive sensor nodes thus may be put to sleep mode to conserve energy provided there are a sufficient number of active sensor nodes. It is then shown that the lifespan of a WSN is dictated by the distribution of residue energy of sensor nodes. Specifically, we have shown that an energy-balanced WSN is likely to maximize its lifespan. As such, at each step of the tracking task, the head node must judiciously select active nodes from all sensors within the sensing range to minimize residue energy variations (energy balanced) while achieving desired tracking accuracy. This is formulated as a subset selection problem, which is shown to have a complexity that is NP-hard. Several energy-balanced scheduling for tracking (EBaST) heuristic algorithms are proposed to solve this problem with polynomial execution complexities. Extensive simulations have been conducted to compare EBaST against some state-of-the-art scheduling algorithms. It is observed that EBaST is more capable of significantly extending the WSNs lifespan than competing algorithms while delivering comparable or better tracking accuracy. Xiaoqing Hu, Yu Hen Hu, Bugong Xu |
ACM Trans. Sens. Networks | 2 |
| 2013 | Target tracking with distance-dependent measurement noise in wireless sensor networksabstractA distributed extended Kalman filter (EKF) algorithm is developed for tracking moving targets in a wireless sensor network equipped with distance estimating sensors. In particular, a distance-dependent measurement error of range-estimating sensors is modeled as a multiplicative noise in the observation model. A new formulation of EKF, called generalized EKF (GEKF) based on the multiplicative noise model is developed. Compared to conventional EKF formulation, it is shown that GEKF can achieve smaller estimation error than traditional EKF. Simulation results also demonstrated superior performance of GEKF. Xiaoqing Hu, Bugong Xu, Yu Hen Hu |
ICASSP | 3 |
| 2013 | Cycle efficient bit rate matching for LTE-a with insructions supportabstractEfficient software implementation of Long Term Evolution Advanced (LTE-A) wireless standard over word-based micro-processor architecture is investigated. The very high data rate of LTE-A requires more sophisticated and high throughput bit level algorithms. Due to data format mismatch, traditional word-based microprocessors face great challenge implementing these complex bit-manipulations. In this work, the implementation of bit-level interleaving operation using perfect shuffling word-level instruction is studied. In particular, an efficient implementation of bit-rate matching algorithm is demonstrated on the Texas Instruments c6416 CCS cycle accurate simulator with order of magnitude performance enhancement. Jui-Chieh Lin, Yu Hen Hu |
ICASSP | 2 |
| 2013 | Data centric multi-shift sensor scheduling for wireless sensor networksabstractA multi-shift sensor scheduling method is proposed to extend the operating lifespan of a wireless sensor network. Sensor nodes in the WSN are partitioned into N subnetworks and the operating schedule is partitioned into N shifts of equal duration. Exploiting spatial correlations among sensor nodes, data collected using each subnetwork can well approximate the data collected using original sensor network. Each sub-network also form a connected component to ensure proper data collection. This task is formulated as a NP-hard constrained subset selection problem. A polynomial time heuristic algorithm leveraging breath-first search and subspace approximation is proposed. Simulations using a real world data set demonstrate superior performance and extended lifespan of this proposed method. Yu Hen Hu |
ICASSP | 2 |
| 2013 | Feature extraction in developing an airs cloud maskabstractCloud and clear-sky detection is a crucial part in the analyses of AIRS (Atmospheric InfraRed Sensor) measurements. Currently cloud detection is done using spectral tests, which are based on well understood properties of the atmosphere. This paper gives an account of an investigation in using binary classification and feature extraction techniques to develop an AIRS cloud mask, where CALIOP (Cloud-Aerosol LIDAR with Orthogonal Polarization) observations were used as “oracle” data. The objective was to produce an AIRS cloud mask which is either on par or better than the MODIS (Moderate Resolution Imaging Spectro-radiometer) cloud mask. Willem J. Marais, Yu Hen Hu, Robert E. Holz |
IGARSS | 2 |
| 2013 | Direct Localization of Multiple Sources in Sensor Array Networks: A Joint Sparse Representation of Array Covariance Matrices ApproachabstractA novel sparse representation based multi-source localization method is presented in this work. We envision a wireless network infrastructure containing multiple phase arrays of acoustic sensors. With multiple arrays, direct estimation of a set of source locations is achieved using a new joint sparse representation of array covariance matrices (JSRACM). This representation transforms the source location estimation problem into a spatial sparse signal representation (SSSR) optimization problem. To mitigate the high computation complexity of JSRACM, a novel binary sparse indicative vector (SIV) is introduced to represent the support of joint SSSR of array covariance matrices. As such, the multiple source locations may be estimated by solving an unconstrained optimization problem of the SIV vector using existing FOCUSS-like algorithms. The resulting SIVR-JSRACM algorithm does not require prior information of the number of sources nor initial source location estimates. It promises super-resolution, robustness to noise, and low computing complexity which is independent of the number of sensor phase arrays. Simulation results demonstrate superior performance of the proposed algorithm. Ji-an Luo, Zhi Wang 0003, Yu Hen Hu |
MASS | 3 |
| 2013 | A Direct Wideband Direction of Arrival Estimation under Compressive SensingabstractCompressive Sensing (CS) theory witnesses great breakthrough in signal acquisition and processing. In this paper we focus on the application of direction of arrival (DoA) and proposed a compressive sensing based direct DoA estimation framework (CS-DDoA) for wireless sensor array network. Unlike the former signal reconstruction-processing scheme, CS-DDoA utilize the spatial and spectral sparsity of target, then reformulates them into a global problems. It eliminates the error propagation between this two independent procedure and provides a enhanced DoA performance. Under this framework, array data transmission volume can be greatly reduce without distinct performance decline, which has great potential for wireless sensor array network with limited resources. At last theoretical analysis and prototype system are provided to validate this framework. Kai Yu 0005, Ji-an Luo, Ming Bao, Yu Hen Hu, Zhi Wang 0003 |
MASS | 5 |
| 2013 | How many wireless resources are needed to resolve the hidden terminal problem?
Yu Hen Hu |
Comput. Networks | 2 |
| 2013 | A unified link-layer fault-tolerant architecture for network-based many-core embedded systems
Wen-Chung Tsai, Deng-Yuan Zheng, Yu Hen Hu, Sao-Jie Chen |
J. Syst. Archit. | 3 |
| 2013 | Cycle-Efficient LFSR Implementation on Word-Based MicroarchitectureabstractCycle-efficient implementation of the linear feedback shift register (LFSR) algorithm on a word-based microarchitecture is investigated. This work examines an algorithm transformation method, called term-preserving look-ahead transformation (TePLAT), that transforms the bit-serial LFSR algorithm into a bit parallel format while maintaining the overhead of the original LFSR algorithm. Detailed implementation methodologies as well as extensive simulation results are presented. We apply TePLAT to 25 commonly used LFSRs and test the resulting parallel formulations on two popular word-based microprocessor development platforms: a Texas Instrument C6416 Code Composition Simulator and an ARM-9 Simulator. In all 25 cases, TePLAT transformed LFSR formulations consistently achieve much higher throughput than those of a naïve implementation and a traditional look-ahead transformation-based implementation. Jui-Chieh Lin, Sao-Jie Chen, Yu Hen Hu |
IEEE Trans. Computers | 3 |
| 2012 | Cycle-efficient lineary feedback shift register implementation on word-based micro-architectureabstractA novel algorithm transformation method, called term-preserving look-ahead transformation (TePLAT) is proposed to transform the bit-serial linear feedback shift register (LFSR) algorithm into a bit-parallel formulation which promises order of magnitudes improvement of execution speed compared to the traditional look-ahead algorithm transformation approach. TePLAT is applied to 26 commonly used LFSRs and tested on two popular word-based micro-processor development platforms: a Texas Instrument C6416 Code Composition Simulator and an ARM-9 Simulator. In all 26 cases, TePLAT transformed implementations consistently deliver much higher throughput than those implementations based on traditional look-ahead algorithm transformation. Jui-Chieh Lin, Sao-Jie Chen, Yu Hen Hu |
ICASSP | 3 |
| 2012 | Compressed sensing enhanced random equivalent samplingabstractThe feasibility of compressed sensing (CS) based waveform-reconstruction for data sampled from random equivalent sampling (RES) method is investigated. A novel measurement matrix motivated by the Whittaker-Shannon interpolation formula is proposed for this purpose. Experiments indicate that, for spectrally-sparse signal, the CS reconstructed waveform exhibits significantly higher signal-to-noise ratio (SNR) than that using the traditional time alignment method. A prototype realization of this proposed CS-RES method has been developed using off-the-shelf components. It is able to capture analog waveform at an equivalent sampling rate of 25 GHz while sampled at 100 MHz physically. Yijiu Zhao, Yu Hen Hu, Houjun Wang |
ICASSP | 2 |
| 2012 | A scalable and fault-tolerant network routing scheme for many-core and multi-chip systems
Wen-Chung Tsai, Kuo-Chih Chu, Yu Hen Hu, Sao-Jie Chen |
J. Parallel Distributed Comput. | 3 |
| 2012 | Efficient Thermal Simulation for 3-D IC With Thermal Through-Silicon ViasabstractA novel virtual power source (VPS) method is proposed that can significantly speedup 3-D integrated circuit (IC) thermal simulation with sparely deployed thermal TSVs leveraging Green function-based analytical spectral method. Specifically, it is shown that the impact of TSVs with nonhomogeneous thermal conductivity on thermal simulation may be modeled as VPSs located at mesh grids inside the thermal TSVs. Using this VPS model, the temperature distribution over the entire 3-D IC may be obtained by solving for the thermal distribution of a thermally homogeneous substrate in the presence of these VPSs. As such, the fast spectral method may be applied to accelerate the simulation. Significant (3–100 times) speedup over a baseline finite difference method implementation has been observed in preliminary simulations. Dongkeun Oh, Charlie Chung-Ping Chen, Yu Hen Hu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2011 | A fault-tolerant NoC scheme using bidirectional channelabstractA novel Bidirectional Fault-Tolerant NoC (BFT-NoC) architecture capable of mitigating both static and dynamic channel failures is proposed. In a traditional NoC platform, a faulty data channel will force blocked packets to make costly detours, resulting in significant performance hits. In this work, novel fault-tolerance measures for a bidirectional NoC platform are proposed. The dynamically reconfigurable bidirectional channels of the BFT-NoC offer great flexibility to contain data-link permanent or transient faults while incurring negligible performance loss. Potential performance advantages in terms of failure rate reduction and reliability enhancement of the BFT-NoC architecture are carefully analyzed. Extensive experimental results clearly validate the fault-tolerance performance of BFT-NoC at both synthetic and real world network traffic patterns. Wen-Chung Tsai, Deng-Yuan Zheng, Sao-Jie Chen, Yu Hen Hu |
DAC | 4 |
| 2011 | Optimal Detector Based on Data Fusion for Wireless Sensor NetworksabstractThis paper investigates target detection problem in wireless sensor networks. Sensors carry out sensing operations and make consensus decisions about the presence or absence of a target or event. Most of previous studies for target detection either assume an unrealistic disk model for making detection decision or provide complicated numerical methods to evaluate detection performance. This paper develops the Uniformly Most Powerful(UMP) detector based on likelihood ratio test and derives simple and elegant test rules for target presence and absence. Moreover, detection performance measured by missing rate is also derived analytically. Simulations are conducted to show the performance of the UMP detector compared to a detector developed previously based on value fusion. The results show that the proposed detector dramatically outperforms the value fusion detector even in vulnerable locations. Tai-Lin Chin, Yu Hen Hu |
GLOBECOM | 2 |
| 2011 | A Bidirectional NoC (BiNoC) Architecture With Dynamic Self-Reconfigurable ChannelabstractA bidirectional channel network-on-chip (BiNoC) architecture is proposed to enhance the performance of on-chip communication. In a BiNoC, each communication channel allows to be dynamically self-reconfigured to transmit flits in either direction. This added flexibility promises better bandwidth utilization, lower packet delivery latency, and higher packet consumption rate. Novel on-chip router architecture is developed to support dynamic self-reconfiguration of the bidirectional traffic flow. This area-efficient BiNoC router delivers better performance and requires smaller buffer size than that of a conventional network-on-chip (NoC). The flow direction at each channel is controlled by a channel direction control (CDC) algorithm. Implemented with a pair of finite state machines, this CDC algorithm is shown to be high performance, free of deadlock, and free of starvation. Extensive cycle-accurate simulations using synthetic and real-world traffic patterns have been conducted to evaluate the performance of the BiNoC. These results exhibit consistent and significant performance advantage over conventional NoC equipped with hard-wired unidirectional channels. Ying-Cherng Lan, Yueh-Chi Lin, Shih-Hsin Lo, Yu Hen Hu, Sao-Jie Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2011 | Optimal Layered Video IPTV Multicast Streaming Over Mobile WiMAX SystemsabstractA WiMAX radio resource allocation (WRA) problem is investigated in the context of IPTV broadcasting over mobile WiMAX multicast, broadcast services (MBS) channels. The goal is to maximize quality of services, in terms of number of subscribers served, number of IPTV channels carried, and perceived video qualities of individual viewers, subject to constraints on multicast channel capacities and space-time channel quality variations. We present an efficient heuristic algorithm based on the Pareto principle that achieves near-optimal results in polynomial time complexity. Extensive simulations are conducted to compare this new approach against state-of-the-art heuristic algorithms and consistent superior performance has been observed. Po-Han Wu, Yu Hen Hu |
IEEE Trans. Multim. | 2 |
| 2010 | Runtime temperature-based power estimation for optimizing throughput of thermal-constrained multi-core processorsabstractTechnology scaling has allowed integration of multiple cores into a single die. However, high power consumption of each core leads to very high heat density, limiting the throughput of thermal-constrained multi-core processors. To maximize the throughput, various software-based dynamic thermal management and optimization techniques have been proposed, many of which depend on accurate temperature sensing of each core. However, the decision for dynamic thermal management and throughput optimization only based on the temperature of each core can result in less optimal throughput in certain circumstances according to our investigation. In this paper, we propose 1) a dynamic power estimation method using a single thermal sensor for each core in multi-core processors, 2) a die temperature reconstruction method using the estimated power, and 3) a throughput optimization method based the estimated power instead of the temperature. According to our experiment using 90nm technology, the proposed method results in less than 3% error in estimating power and hot-spot temperature of a multi-core processor. Furthermore, the proposed throughput optimization method based on the estimated power leads to up to 4% higher throughput than a temperature-based optimization method. Dongkeun Oh, Nam Sung Kim, Charlie Chung-Ping Chen, Azadeh Davoodi, Yu Hen Hu |
ASP-DAC | 5 |
| 2010 | Collaborative Sampling in Wireless Sensor NetworksabstractA new paradigm for wireless sensor networks (WSN) information gathering, called Collaborative Sampling is proposed in this paper. Collaborative Sampling is a distributed active sampling regime that requires that each sensor to coordinate the tasks of data sampling and dissemination through succinct collaboration with neighboring sensor nodes. This is made possible by exploiting the inherent correlation among neighboring sensor measurements. The goal is to minimize energy wasted on transmitting redundant information through wireless channels while ensuring the quality of collected data. Collaborative Sampling entails two parts of key operations: at individual sensor nodes, a model of sensor measurements as a function of space time coordinates are computed based on local sensor measurements and those obtained by overhearing neighboring sensor nodes wireless broadcasting. Sensor nodes collaborate by announcing its local measurements to its neighbors to help build a global function approximation of the sensor measurements over the entire sensor field. To reduce redundant information being broadcast, each sensor node will determine whether announcing its own reading to its neighbors will contribute most to achieve the global goal of function approximation. A probabilistic competitive mechanism, based on the medium access control (MAC) layer carrier sensing multiple access (CSMA) protocol, is developed to allow the sensor nodes that offers most contribution to report their value earlier and thereby speedup the convergence to a solution. Theoretical analysis and simulation results showing that the new algorithm performs as well as the best possible greedy result are provided. Minglei Huang, Yu Hen Hu |
GLOBECOM | 2 |
| 2010 | Decentralized Robust Acoustic Source Localization with Wireless Sensor Networks for Heavy-Tail Distributed ObservationsabstractIn this work, an energy based acoustic source localization task in a wireless sensor network (WSN) is considered. Based on data gathered from field experiments, it is revealed that the acoustic energy gathered at sensor nodes exhibits a heavy-tail, non-Gaussian characteristic and should be fitted into a contaminated Gaussian model. This property renders conventional least square and maximum likelihood based location estimation methods ineffective. Leveraging the distributed, in-network processing nature of a WSN, a novel de-centralized robust acoustic source localization (DRASL) algorithm is proposed. With the DRASL, local sensor nodes receive sensor readings broadcast from neighboring sensors and independently compute local location estimates using a light-weight Iterative Nonlinear Reweighted Least Square (INRLS) algorithm. The local location estimate then will be relayed to a fusion center where the final location estimate is obtained as a weighted average of the local estimates. The potential advantage of this algorithm is validated using extensive simulation in a real-world operation scenario. It is show that its performance is superior than existing methods while promising to be more energy efficient. Yong Liu 0025, Yu Hen Hu, Quan Pan 0001 |
GLOBECOM | 2 |
| 2010 | Cycle efficient scrambler implementation for software defined radioabstractThe task of efficient implementation of bit-serial scrambler algorithms on a word-parallel software defined radio micro-architecture is considered. By exploiting properties of the exclusive-OR Boolean logic operator, a novel power-of-two look-ahead recursive algorithm transformation method is developed. Together with loop-unrolling, this new method produces an efficient vector scrambler algorithm formulation that realizes orders of magnitude clock-cycle savings compared to state of the art solutions. Using IEEE 802.11a scrambler algorithm as an example, this new formulation is 9 times faster than previously reported results. Jui-Chieh Lin, Ming-Jung Fan-Chiang, Minja Hsieh, Song-Yen Mao, Sao-Jie Chen, Yu Hen Hu |
ICASSP | 6 |
| 2010 | OFDM Peak to Average Power Ratio reduction using sparse bit-plane codingabstractA novel Peak to Average Power Ratio (PAPR) reduction method leveraging the rare occurrence of large magnitude Orthogonal Frequency Division Multiplexing (OFDM) symbols is proposed. A distinct innovation of this approach is the exploitation of the sparseness structure of a bit-plane binary number representation of time domain OFDM symbols. By encoding the sparsely populated non-zero entries in the binary bit plane matrix, the dynamic range of OFDM symbols to be transmitted can be significantly reduced without increasing bit error rate. Smaller dynamic range of OFDM symbols implies reduced PAPR, and hence better performance. This method incurs no data rate reduction penalty which may occur when side information must be transmitted via some data subcarriers. Simulation results reveal that this proposed new method out-performs current PAPR reduction methods by wide margins. Cheng-Han Sung, Yu Hen Hu |
ICASSP | 2 |
| 2010 | An integral-based curvature estimator and its application in face recognitionabstractThis paper addresses an old, yet challenging issue - curvature estimation from discrete sampling points over a curve. We introduce a novel algorithm based on performing line integrals. The proposed method is computationally more efficient than the previous integration-based methods because of the constant computation time. Qualitative tests on synthesized shapes in the presence of noise are performed, which show the robustness of our approach. Also, the discriminative capability of the estimated curvature is evaluated by conducting experiments on the FRGC (Face Recognition Grand Challenge) v2.0 dataset which contains 4007 3D facial images recorded from 466 subjects. The results show that recognition performance is significantly improved by using our curvature estimation method. This novel approach presents potential for a broad class of multimedia applications. Wei-Yang Lin, Yen-Lin Chiu, Kerry R. Widder, Yu Hen Hu, Nigel Boston |
ICME | 4 |
| 2010 | QoS aware BiNoC architectureabstractA quality-of-service (QoS) aware, bi-directional channel NoC (BiNoC) architecture is proposed to support guarantee-service (GS) traffic while reducing packet delivery latency. By incorporating dynamically self-reconfigured bidirectional communication channels between adjacent routers, BiNoC architecture promises more flexibility for various traffic flow patterns. A novel inter-router communication protocol is proposed that prioritizes bandwidth arbitration in favor of high priority GS traffic flows. Multiple virtual channels with prioritized routing policy are also implemented to facilitate data transmission with QoS considerations. Combining these architectural innovations, the QoS aware BiNoC promises reduced latency of packet delivery and more efficient channel resource utilizations. Cycle-accurate simulations demonstrate significant performance advantage over conventional unidirectional NoC architecture equipped with hard-wired unidirectional channels. Shih-Hsin Lo, Ying-Cherng Lan, Hsin-Hsien Yeh, Wen-Chung Tsai, Yu Hen Hu, Sao-Jie Chen |
IPDPS | 5 |
| 2010 | Perfect shuffling for cycle efficient puncturer and interleaver for software defined radioabstractThis work applies perfect-shuffle on software defined radio's bit-shuffling blocks, such as puncturer and interleaver, which data alignment issue is also investigated. A demonstration of such algorithm on IEEE 802.11a (WiFi) standard is implemented with Texas Instruments®intrinsic perfect-shuffle and inverse perfect-shuffle instructions combined with ANSI C on a TMS320 C64x digital signal processor. Simulation results indicate the performance enhancement using perfect-shuffle technique dramatically improved cycle efficiency comparing to a naive one bit per integer implementation. Jui-Chieh Lin, Minja Hsieh, Ming-Jung Fan-Chiang, Song-Yen Mao, Chu Yu, Sao-Jie Chen, Yu Hen Hu |
ISCAS | 7 |
| 2010 | TM-FAR: Turn-Model based Fully Adaptive Routing for Networks on ChipabstractA novel Turn-Model based Fully-Adaptive-Routing (TM-FAR) algorithm is proposed for Networks-on-Chip (NoC). TM-FAR retains the deadlock-free property of traditional turn-model based routing algorithms (e.g., XY, Odd-Even), while alleviating restrictions on turn and path selections. Just like the current Virtual-Channel based Fully-Adaptive-Routing (VC-FAR) algorithm, TM-FAR allows full exploitation of all available minimal paths, yet TM-FAR does not use virtual channels. This fully adaptive routing capability of TM-FAR promises improved routing adaptivity and enhanced level of fault-tolerance. Preliminary experimental results indicate that the TM-FAR achieves an averaged delay reduction of 30.08% and a throughput rate increase of 4.54% compared to the state-of-the-art NoC routing algorithm based on the Odd-Even turn model. Wen-Chung Tsai, Kuo-Chih Chu, Sao-Jie Chen, Yu Hen Hu |
VLSI-SoC | 4 |
| 2010 | Optimal Layered Video IPTV Multicast Streaming over IEEE 802.16e WiMAX SystemsabstractA wireless resource scheduling problem is formulated in the context of IPTV multicasting streaming services over WiMAX channels. The goal is to maximize quality of services, in terms of number of subscribers served, number of IPTV channels carried, and video qualities of individual channels; subject to constraints of finite amount of multicast channel capacities as well as time and spatially varying wireless channel qualities. While a globally optimal solution is NP-hard, we present an efficient heuristic algorithm that exploits the structure of the problem formulation to yield solutions that are very close to the optimal solution. Po-Han Wu, Yu Hen Hu, Jenq-Neng Hwang |
VTC Spring | 2 |
| 2010 | Optimal Multiple-Bit Huffman DecodingabstractA variable-bit look-ahead Huffman decoding problem is investigated in this paper. The objective is to maximize the decoding throughput rate by exploiting different equivalent state diagrams. The decoding throughput rate is estimated based on the state transition probability of the corresponding Huffman encoding table. We propose an efficient algorithm to search for a heuristic solution that usually yields very good optimization results in polynomial computation time. Ya-Nan Wen, Guang-Huei Lin, Sao-Jie Chen, Yu Hen Hu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2009 | Robust Maximum Likelihood Acoustic Source Localization in Wireless Sensor NetworksabstractSensor measurements in a wireless sensor network (WSN) may significantly deviate from a commonly used Gaussian noise model due to harsh operating conditions, unreliable wireless communication links, or sensor failures. In this work, a mixed Gaussian and impulse noise model is proposed to more accurately model these types of non-Gaussian noise. However, existing maximum likelihood (ML) acoustic energy based source localization algorithms are very sensitive to non-Gaussian noise perturbations. To mitigate this shortcoming, a novel M-estimate based robust estimation formulation is derived. Extensive simulation results demonstrated superior and consistent performance advantage of this robust estimation approach compared to conventional ML estimates over a wide range of practical scenarios. Yong Liu 0025, Yu Hen Hu, Quan Pan 0001 |
GLOBECOM | 2 |
| 2009 | Estimating correspondence between multiple cameras using joint invariantsabstractThe joint invariants of the projective group PSL(3,R) on R2, the five-point volume cross-ratios, are studied to address the problem of correspondence in a camera network. The distribution of cross-ratios over the unit square as well as in a small local-neighbourhood of a reference point are found to have a heavy tail. No cross ratio value is unique but the collection of five point cross ratios generated by taking all possible combination of five points completely prescribes the curve. Sections of the signature submanifold that admit large enough variation of cross ratios are found to be sufficient in providing correspondence across wide perspectives. Such invariant signatures may be collected independently at cameras with different viewpoints and shared, thereby achieving the registration of objects in the image. Experimental results with license plate database are provided. Raman Arora, Yu Hen Hu, Charles R. Dyer |
ICASSP | 2 |
| 2009 | A Novel Cross-layer Data Aggregation Approach for Extreme Values in Wireless Sensor NetworksabstractIn this paper, the task of distributed sensor network data aggregation of extreme (say, maximum) value of sensor measurements is investigated. We exploit the local broadcasting topology of a wireless network to facilitate the development of a novel energy efficient aggregation algorithm. The method models the extreme value aggregation process as a search problem, and devises an approach to accelerate convergence. Through both theoretical analysis and extensive simulation, we characterize the performance of these methods both analytically and experimentally. We observe marked performance advantage of these methods. Minglei Huang, Yu Hen Hu |
ISCAS | 2 |
| 2009 | Summation Invariant Multi-region Fusion ComparisonabstractApplications of summation invariant features to multi-region face recognition are explored in this work. Earlier, we have demonstrated the potential benefits of this approach. In this paper, we provide a systematic, thorough comparison of all the summation invariant features derived to-date, and propose a new multi-feature fusion approach to further improve the overall performance. We also identify summation invariant features that yield superior performance for face recognition applications. Special attention is given to the implementation of 3D summation invariants. Extensive experimental results with the FRGC (Face Recognition Grand Challenge) 2.0 data set confirms the advantage of summation invariant features for 3D face recognition. Kerry R. Widder, Yu Hen Hu, Nigel Boston, Wei-Yang Lin |
ISCAS | 2 |
| 2009 | Adaptive nodes scheduling approach for clustered sensor networksabstractEnergy efficiency is an important issue in wireless sensor networks. One available power saving strategy is having only a portion of nodes work, but this would always compromise data quality as a result. In this paper, we propose an adaptive nodes scheduling approach (ADNS) to conserve energy while maintaining the overall data quality. ADNS selects a subset of nodes to be active and puts the others into sleep mode to save energy. An efficient active node selection (ANS) algorithm is presented. Using the spatial correlation among sensor readings as the prediction model, the values of sleep nodes are predicted by data collected from active nodes to ensure the data quality. In order to maintain the data quality throughout network operation, the prediction errors are validated timely and prediction models are adaptively reconstructed if necessary. We evaluate ADNS on a real-world sensor network data set and validate its effectiveness. Yu Hen Hu, Jingping Bi |
ISCC | 2 |
| 2009 | BiNoC: A bidirectional NoC architecture with dynamic self-reconfigurable channelabstractA Bidirectional channel Network-on-Chip (BiNoC) architecture is proposed to enhance the performance of on-chip communication. The BiNoC allows each communication channel to be dynamically self-configured to transmit flits in either direction in order to better utilize on-chip hardware resources. This added flexibility promises better bandwidth utilization, lower packet delivery latency, and higher packet consumption rate at each on-chip router. In this paper, a novel on-chip router architecture supporting the self-configuring bidirectional channel mechanism is presented. It is shown that the associated hardware overhead is negligible. Cycle-accurate simulation runs on this BiNoC network under synthetic and real-world traffic patterns demonstrate consistent and significant performance advantage over conventional mesh grid NoC architecture equipped with hard-wired unidirectional channels. Ying-Cherng Lan, Shih-Hsin Lo, Yueh-Chi Lin, Yu Hen Hu, Sao-Jie Chen |
NOCS | 4 |
| 2008 | An optimal algorithm for sizing sequential circuits for industrial library based designsabstractIn this paper, we propose an optimal gate sizing and clock skew optimization algorithm for globally sizing synchronous sequential circuits. The number of constraints and variables in our formulation is linear with respect to the number of circuit components and hence our algorithm can efficiently find the optimal solution for industrial scale designs. To the best of our knowledge our method is the first exact gate sizing algorithm that can handle cyclic sequential circuits. Experimental results on industrial cell libraries demonstrate that our algorithm can yield an average of 12.6% improvement in the optimal clock period by combining clock skew optimization with gate sizing. For identical clock period, our algorithm can achieve an average of 11.3% area savings over a popular commercial synthesis tool. Sanghamitra Roy, Yu Hen Hu, Charlie Chung-Ping Chen, Shih-Pin Hung, Tse-Yu Chiang, Jiuan-Guei Tseng |
ASP-DAC | 2 |
| 2008 | Optimal Target Detection with Localized Fusion in Wireless Sensor NetworksabstractDetecting the presence/absence of an object in a region of interest is one of the important applications for sensor networks. A considerable amount of work has been seen in the literature for detecting events or objects using wireless sensor networks. Most of the prior work uses a simple binary detection model or an average signal strength model to make decisions of detection. Such methods are not optimal in terms of detection probability. This paper derives a detection approach which is optimal in the sense of Neyman-Pearson test and shows that the detection performance of the traditional average based method is much lower than the optimal. To reduce power consumption and communication cost, a localized fusion method is also developed by carefully selecting sensors in the vicinity of a target location. The paper shows that the localized fusion can dramatically reduce the number of sensors participating the fusion while maintain high detection performance. Tai-Lin Chin, Yu Hen Hu |
GLOBECOM | 2 |
| 2008 | Planar-projective summation invariant features for camera networksabstractRecently a novel family of geometrically invariant features, called summation invariant, has been developed and applied to object recognition. The range of this family of features is expanded here beyond the Euclidean and afline transformation groups to planar projective transformations. Whereas other methods require small changes in view, or collinear points, this method removes those limitations and allows recognition of general planar objects over wide ranges of viewpoint. The derivation of these new features requires the innovation of deriving the invariants in the homogeneous coordinate space, yet yields results formulated in terms of Cartesian coordinates. Simulations demonstrate the effectiveness of this new approach to object recognition under projective transformations like those encountered in camera networks. Kerry R. Widder, Wei-Yang Lin, Nigel Boston, Yu Hen Hu |
ICASSP | 4 |
| 2008 | Discovering panoramas in web videosabstractWhile methods for stitching panoramas have been successful given proper source images, providing these source images still remains a burden. In this paper, we present a method to discover panoramic source images within widely available web videos. The challenge comes from the fact that many of these videos are not recorded intentionally for stitching panoramas. Our method aims to find segments within a video that work as panorama sources. Specifically, we determine a video segment to be a valid panorama source according to the following three criteria. First, its camera motion should cover a wide field-of-view of the scene. Second, its frames should be "mosaicable", which states that the inter-frame motion should observe the underlying conditions for stitching a panorama. Third, its frames should have good image quality. Based on these criteria, we formulate discovering panoramas in a video as an optimization problem that aims to find an optimal set of video segments as panorama sources. After discovering these panorama sources, we synthesize regular scene panoramas using them. When significant dynamics is detected in the sources, we fuse the dynamics into the scene panoramas to make activity synopses to convey the dynamics. Our experiment of querying panoramas from YouTube confirms the feasibility of using web videos as panorama sources and demonstrates the effectiveness of our method. Feng Liu 0015, Yu Hen Hu, Michael Gleicher |
ACM Multimedia | 2 |
| 2007 | Algorithm Transformation to Improve Data Locality for Multimedia SOCabstractIn this paper we propose a more systematic approach to improve data locality of multimedia algorithm that contain nested loop while still preserve the nested loop program semantic by algorithm transformation. We introduce procedure to evaluate the reuse vector of the data access. The reuse vectors represent the input dependencies of the program in the form of index vector which reveal the information of which loop carrying reuse and thus are used as a guide to select the loops subject to transformation. We use full search block matching motion estimation (FBME) algorithm as a case study because this algorithm has 6 level nested loop and thus is a good showcase of our method. Anne Pratoomtong, Yu Hen Hu |
ICASSP (2) | 2 |
| 2007 | 3D Face Recognition Under Expression Variations using Similarity Metrics FusionabstractWe present a novel 3D face recognition method that incorporates summation invariant features extracted from multiple sub-regions of a facial range images, and optimal fusion of similarity scores between corresponding sub-regions. The key innovation of this paper is the development of the fusion-based face recognition algorithm that delivers significant performance enhancement while requiring very little computation. Experiments on the FRGC (Face Recognition Grand Challenge) version 2 dataset show that our algorithm improves the recognition performance significantly in the presence of facial expressions. Wei-Yang Lin, Kin-Chung Wong, Nigel Boston, Yu Hen Hu |
ICME | 4 |
| 2007 | Event-Based Segmentation of Sports Video Using Motion EntropyabstractAn event-based segmentation method for sports videos is presented. A motion entropy criterion is employed to characterize the level of intensity of relevant object motion in individual frames of a video sequence. The resulting motion entropy curve then is approximated with a piece-wise linear model using a homoscedastic error model based time series change point detection algorithm. It is observed that interesting sports events are correlated with specific patterns of the piece-wise linear model. A set of empirically derived classification rules then is derived based on these observations. Application of these rules to the motion entropy curve leads to this motion entropy curve, one is able to segment the corresponding video sequence into individual sections, each consisting of a semantically relevant event. The proposed method is tested on six hours of sports videos including basketball, soccer and tennis. Excellent experimental results are observed. Chen-Yu Chen, Jia-Ching Wang, Jhing-Fa Wang, Yu Hen Hu |
ISM | 4 |
| 2007 | Fusion of Multiple Facial Regions for Expression-Invariant Face RecognitionabstractIn this paper, we describe a fusion-based face recognition method that is able to compensate for facial expressions even when training samples contain only neutral expression. The similarity metric between two facial images are calculated by combining the similarity scores of the corresponding facial regions, e.g. the similarity between two mouths, the similarity between two noses, etc. In contrast with other approaches where equal weights are assigned on each region, a novel fusion method based on linear discriminant analysis (LDA) is developed to maximize the verification performance. We also conduct a comparative study on various face recognition schemes, including the FRGC baseline algorithm, the fusion of multiple regions by sum rule, and the fusion of multiple regions by LDA. Experiments on the FRGC (Face Recognition Grand Challenge) V2.0 dataset, containing 4007 face images recorded from 266 subjects, show that the proposed method significantly improves the verification performance in the presence of facial expressions. Wei-Yang Lin, Kerry R. Widder, Yu Hen Hu, Nigel Boston |
MMSP | 4 |
| 2007 | Multiscale Integral Invariants For Facial Landmark Detection in 2.5D DataabstractIn this paper, we introduce a novel 3D surface landmark detection method using a 3D integral invariant feature extended from that proposed by Manay et al. for 2D contours. We apply this new feature to detect the nose tips of 2.5D range images of human faces. Using the Face Recognition Grand Challenge 2.0 dataset, our method compares favorably with a recently proposed competing method. Adam Slater, Yu Hen Hu, Nigel Boston |
MMSP | 2 |
| 2007 | Numerically Convex Forms and Their Application in Gate SizingabstractConvex-optimization techniques are very popular in the very large-scale-integration design society due to their guaranteed convergence to a global optimal point. The table data need to be fitted into convex forms to be used in the convex optimization problems. Fitting the tables into polynomials, which are analytically convex under logarithmic transformation, may suffer from the excessive fitting errors as the fitting problem is nonconvex. In this paper, we propose to directly adjust the lookup-table values into a numerically convex lookup table without any explicit analytical form. We show that numerically "convexifying" the lookup-table data with minimum perturbation can be formulated as a convex semidefinite optimization problem, and hence, optimality can be reached in polynomial time. We also propose three algorithms to make the table data smooth to enable faster convergence of the convex optimizer. Results from extensive experiments on industrial cell libraries demonstrate 9.6 improvement in fitting error over a well-developed polynomial-fitting procedure. We illustrate the effectiveness of this model in a convex optimization problem by providing results for using our model in the optimal gate sizing of standard cells. We observe a 5.07% improvement in the delay of International Symposium on Circuits and Systems (ISCAS) benchmark circuits over the polynomial-fitting procedure. Sanghamitra Roy, Weijen Chen, Charlie Chung-Ping Chen, Yu Hen Hu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2007 | Optimal Linear Combination of Facial Regions for Improving Identification PerformanceabstractThis paper presents a novel 3-D multiregion face recognition algorithm that consists of new geometric summation invariant features and an optimal linear feature fusion method. A summation invariant, which captures local characteristics of a facial surface, is extracted from multiple subregions of a 3-D range image as the discriminative features. Similarity scores between two range images are calculated from the selected subregions. A novel fusion method that is based on a linear discriminant analysis is developed to maximize the verification rate by a weighted combination of these similarity scores. Experiments on the Face Recognition Grand Challenge V2.0 dataset show that this new algorithm improves the recognition performance significantly in the presence of facial expressions. Kin-Chung Wong, Wei-Yang Lin, Yu Hen Hu, Nigel Boston |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2006 | Convergence-provable statistical timing analysis with level-sensitive latches and feedback loopsabstractStatistical timing analysis has been widely applied to predict the timing yield of VLSI circuits when process variations become significant. Existing statistical latch timing methods are either having exponential complexity or unable to treat the random variable's self-dependence caused by the coexistence of level-sensitive latches and feedback loops. In this paper, an efficient iterative statistical timing algorithm with provable convergence is proposed for latch-based circuits with feedback loops. Based on a new notion of iteration mean, we prove that the algorithm converges unconditionally. Moreover, we show that the converged value of iteration mean ca be used to predict the circuit yield during design time. Tested by ISCAS'89 benchmark circuits, the proposed algorithm shows a error of 1.1% and speedup of 303 /spl times/ on average when compared with the Monte Carlo simulation. Lizheng Zhang, Jeng-Liang Tsai, Weijen Chen, Yu Hen Hu, Charlie Chung-Ping Chen |
ASP-DAC | 4 |
| 2006 | Fusion of Summation Invariants in 3D Human Face RecognitionabstractA novel family of 2D and 3D geometrically invariant features, called summation invariants is proposed for the recognition of the 3D surface of human faces. Focusing on a rectangular region surrounding the nose of a 3D facial depth map, a subset of the so called semi-local summation invariant features is extracted. Then the similarity between a pair of 3D facial depth maps is computed to determine whether they belong to the same person. Out of many possible combinations of these set of features, we select, through careful experimentation, a subset of features that yields best combined performance. Tested with the 3D facial data from the on-going Face Recognition Grand Challenge v1.0 dataset, the proposed new features exhibit significant performance improvement over the baseline algorithm distributed with the datase Wei-Yang Lin, Kin-Chung Wong, Nigel Boston, Yu Hen Hu |
CVPR (2) | 4 |
| 2006 | Statistical timing analysis with path reconvergence and spatial correlationsabstractState of the art statistical timing analysis (STA) tools often yield less accurate results when timing variables become correlated. Spatial correlation and correlation caused by path reconvergence are among those which are most difficult to deal with. Existing methods treating these correlations will either suffer from high computational complexity or significant errors. In this paper, we present a sensitivity pruning method which significantly reduces the computational cost to consider path reconvergence correlation. We also develop an accurate and efficient model to deal with the spatial correlation. Lizheng Zhang, Yu Hen Hu, Charlie Chung-Ping Chen |
DATE | 2 |
| 2006 | 3D Human Face Recognition Using Summation InvariantsabstractA novel family of geometrically invariant features, called summation invariants are proposed for the recognition of the 3D surface of human faces. In particular, a 2D semi-local summation invariant feature is extracted from each column and each row of a rectangular region surrounding the nose of a 3D facial depth map. Through extensive experimentation, we empirically identify the most efficient 2D summation invariant features. We also investigate the proper pre-processing method for the 2D summation invariant features. Tested with the 3D facial data from the Face Recognition Grand Challenge v1.0 dataset, the proposed new features exhibit significant performance improvement over the baseline algorithm distributed with the dataset. Wei-Yang Lin, Kin-Chung Wong, Nigel Boston, Yu Hen Hu |
ICASSP (2) | 4 |
| 2006 | Statistical static timing analysis with conditional linear MAX/MIN approximation and extended canonical timing modelabstractAn efficient and accurate statistical static timing analysis (SSTA) algorithm is reported in this paper, which features 1) a conditional linear approximation method of the MAX/MIN timing operator, 2) an extended canonical representation of correlated timing variables, and 3) a variation pruning method that facilitates intelligent tradeoff between simulation time and accuracy of simulation result. A special design focus of the proposed algorithm is on the propagation of the statistical correlation among timing variables through nonlinear circuit elements. The proposed algorithm distinguishes itself from existing block-based SSTA algorithms in that it not only deals with correlations due to dependence on global variation factors but also correlations due to signal propagation path reconvergence. Tested with the International Symposium on Circuits and Systems (ISCAS) benchmark suites, the proposed algorithm has demonstrated very satisfactory performance in terms of both accuracy and running time. Compared with Monte-Carlo-based statistical timing simulation, the output probability distribution got from the proposed algorithm is within 1.5% estimation error while a 350 times speed-up is achieved over a circuit with 5355 gates. Lizheng Zhang, Weijen Chen, Yu Hen Hu, Charlie Chung-Ping Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2006 | Correlation-Preserved Statistical Timing With a Quadratic Form of Gaussian VariablesabstractA recent study shows that the existing first-order canonical timing model is not sufficient to represent the dependency of the gate/wire delay on the processing and operational variations when these variations become more and more significant. Due to nonlinear mapping from variation sources to the gate/wire delay, the distribution of the delay will no longer be Gaussian even if variation sources are normally distributed. A novel “quadratic timing model” is proposed to capture the nonlinearity of the dependency of gate/wire delays and arrival times on the variation sources. Systematic methodology is also developed to evaluate the correlation and distribution of the quadratic timing model. Based on these, a statistical static timing analysis algorithm that retains the complete correlation information during timing analysis and has linear computation complexity with respect to both the circuit size and the number of variation sources is proposed. Tested on the ISCAS circuits, the proposed algorithm shows significant accuracy improvement over the existing first-order algorithm with a small amount of computational cost. Lizheng Zhang, Weijen Chen, Yu Hen Hu, John A. Gubner, Charlie Chung-Ping Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2005 | Wave-pipelined on-chip global interconnectabstractA novel wave-pipelined global interconnect system is developed for reliable, high throughput, on-chip data communication. We argue that because there is only a single signal propagation path and a single type of 1-input gate(inverter), a wave-pipelined interconnect will have less stringent timing constraints than a wave-pipelined combinational logic block. A phase-lock loop based clock and data recovery unit architecture, adopted from off-chip high speed digital serial link, is designed for on-chip application so as to minimize power and area cost. Preliminary Monte Carlo simulation indicated that the wave-pipelined global interconnect architecture potentially can offer 18% higher throughput than a flip-flop pipelined global interconnect architecture at about the same level of reliability. While delivering data through long interconnect at the same bit rate, the wave-pipelined architecture consumes less power and requires less chip real estate. Lizheng Zhang, Yu Hen Hu, Charlie Chung-Ping Chen |
ASP-DAC | 2 |
| 2005 | Block based statistical timing analysis with extended canonical timing modelabstractBlock based statistical timing analysis (STA) tools often yield less accurate results when timing variables become correlated due to global source of variations and path reconvergence. To the best of our knowledge, no good solution is available handling both types of correlations simultaneously.In this paper, we present a novel statistical timing algorithm, AMECT (Asymptotic MAX/MIN approximation & Extended Canonical Timing model), that produces accurate timing estimation by handling both types of correlations simultaneously. An extended canonical timing model is developed to evaluate and decompose correlations between arbitrary timing variables. And an intelligent pruning method is designed enabling trade-off runtime with accuracy.Tested with ISCAS benchmark suites, AMECT shows both high accuracy and high performance compared with Monte Carlo simulation results: with distribution estimation error < 1.5% while with around 350X speed up on a circuit with 5355 gates. Lizheng Zhang, Yu Hen Hu, Charlie Chung-Ping Chen |
ASP-DAC | 2 |
| 2005 | Correlation-preserved non-gaussian statistical timing analysis with quadratic timing modelabstractRecent study shows that the existing first order canonical timing model is not sufficient to represent the dependency of the gate delay on the variation sources when processing and operational variations become more and more significant. Due to the nonlinearity of the mapping from variation sources to the gate/wire delay, the distribution of the delay is no longer Gaussian even if the variation sources are normally distributed. Anovelquadratic timing model is proposed to capture the non-linearity of the dependency of gate/wire delays and arrival times on the variation sources. Systematic methodology is also developed to evaluate the correlation and distribution of the quadratic timing model. Based on these, a novel statistical timing analysis algorithm is propose which retains the complete correlation information during timing analysis and has the same computation complexity as the algorithm based on the canonical timing model. Tested on the ISCAS circuits, the proposed algorithm shows 10 × accuracy improvement over the existing first order algorithm while no significant extra runtime is needed. Lizheng Zhang, Weijen Chen, Yu Hen Hu, John A. Gubner, Charlie Chung-Ping Chen |
DAC | 3 |
| 2005 | Statistical Timing Analysis with Extended Pseudo-Canonical Timing ModelabstractState of the art statistical timing analysis (STA) tools often yield less accurate results when timing variables become correlated due to global source of variations and path reconvergence. To the best of our knowledge, no good solution is available for dealing both types of correlations simultaneously. In this paper, we present a novel extended pseudo-canonical timing model to retain and evaluate both types of correlation during statistical timing analysis with minimum computation cost. Also, an intelligent pruning method is introduced to enable trade-off runtime with accuracy. Tested with ISCAS benchmark suites, our method shows both high accuracy and high performance. For example, on the circuit c6288, our distribution estimation error shows 15/spl times/ accuracy improvement compared with previous approaches. Lizheng Zhang, Weijen Chen, Yu Hen Hu, Charlie Chung-Ping Chen |
DATE | 3 |
| 2005 | Summation invariant and its applications to shape recognitionabstractA novel summation invariant of curves under transformation group action is proposed. This new invariant is less sensitive to noise than the differential invariant and does not require an analytical expression for the curve as the integral invariant does. We exploit this summation invariant to define a shape descriptor called a semi-local summation invariant and use it as a new feature for shape recognition. Tested on a database of noisy shapes of fish, it was observed that the summation invariant feature exhibited superior discriminating power compared to that of wavelet-based invariant features. Wei-Yang Lin, Nigel Boston, Yu Hen Hu |
ICASSP (5) | 3 |
| 2005 | On-Chip Cache Algorithm Design for Multimedia SOCabstractIn order to optimally implement real time, high throughput, data intensive multimedia applications, it is crucial to optimize the performance of the memory subsystem to minimize excessive off-chip memory bandwidth subject to the constraint of available on-chip memory cache size. This can be accomplished by customizing algorithm transformation and designing a customized cache address mapping algorithm for a specific class of multimedia applications. In this paper, we propose an algorithm transformation and customized cache mapping to improve the data reusability and reduce address conflict which in turn, reduces the cache miss and memory I/O bandwidth for the block-based full-search motion estimation algorithm. Simulation results using test video sequences demonstrate marked performance improvement. Saengrawee Pratoomtong, Yu Hen Hu |
ICASSP (2) | 2 |
| 2005 | Distributed particle filters for wireless sensor network target trackingabstractWe propose two distributed particle filters to estimate and track the moving targets in a wireless sensor network. The observations by the sensors are divided into a set of disjoint uncorrelated cliques. The first distributed algorithm runs the local particle filters sequentially at each clique. The second distributed algorithm runs the local particle filters in parallel to obtain the local sufficient statistics, and then send these statistics to a centralized location through multi-hops to obtain the final estimates. The two distributed algorithms are both almost surely convergent. In addition, we proposed to use the local Gaussian mixture model (GMM) to approximate the posteriori distribution obtained from the local particle filter. By propagating the GMM parameters rather than belief, we achieve significant bandwidth and power consumption reduction. Very promising simulation results are reported as well. Xiaohong Sheng, Yu Hen Hu |
ICASSP (4) | 2 |
| 2005 | Distributed particle filter with GMM approximation for multiple targets localization and tracking in wireless sensor networkabstractTwo novel distributed particle filters with Gaussian mixer approximation are proposed to localize and track multiple moving targets in a wireless sensor network. The distributed particle filters run on a set of uncorrelated sensor cliques that are dynamically organized based on moving target trajectories. These two algorithms differ in how the distributive computing is performed. In the first algorithm, partial results are updated at each sensor clique sequentially based on partial results forwarded from a neighboring clique and local observations. In the second algorithm, all individual cliques compute partial estimates based only on local observations in parallel, and forward their estimates to a fusion center to obtain final output. In order to conserve bandwidth and power, the local sufficient statistics (belief) is approximated by a low dimensional Gaussian mixture model (GMM) before propagating among sensor cliques. We further prove that the posterior distribution estimated by distributed particle filter convergence almost surely to the posterior distribution estimated from a centralized Bayesian formula. Moreover, a data-adaptive application layer communication protocol is proposed to facilitate sensor self-organization and collaboration. Simulation results show that the proposed DPF with GMM approximation algorithms provide robust localization and tracking performance at much reduced communication overhead. Xiaohong Sheng, Yu Hen Hu, Parameswaran Ramanathan |
IPSN | 2 |
| 2005 | Summation Invariant Features for 3D Face RecognitionabstractA novel summation invariant feature under transformation group action for 3D surface recognition is proposed, and its application to 3D face recognition is investigated. Based on a systematic mathematical procedure called moving frame, we derived the summation invariant feature that is invariant under affine transformation. Compared with classical differential invariants, such as the mean curvature or the Gaussian curvature, summation invariant feature is far less sensitive to observation noise in the data. A further enhancement leads to a new type of invariant 3D surface shape descriptor called a semi-local summation invariant. We demonstrate one important, potential application of this new feature to 3D human face recognition Wei-Yang Lin, Nigel Boston, Yu Hen Hu |
MMSP | 3 |
| 2005 | Single access variable block size motion estimation for multimedia SoCabstractTo obtain an optimum rate-distortion characteristic, it is crucial to have accurate motion and residue information to guild the optimum block size selection process. In this paper, we first modified the full search block matching motion estimation (FBME) algorithm to accommodate the motion vector and sum of absolute different (SAD) calculation of each of the 7 modes when perform FBME on one macroblock without any substantial change in the original program. Then we transform the algorithm to reduce the excessive data access to reference frame. With the reduction of redundant data access and the accuracy of the FBME algorithm, the purpose algorithm provides the most accurate motion and residue information without substantially increasing the complexity and memory requirement of the variable block size motion estimation (VBME) Saengrawee Pratoomtong, Yu Hen Hu |
MMSP | 2 |
| 2004 | Statistical timing analysis in sequential circuit for on-chip global interconnect pipeliningabstractWith deep-sub-micron (DSM) technology, statistical timing analysis becomes increasingly crucial to characterize signal transmission over global interconnect wires. In this paper, a novel statistical timing analysis approach has been developed to analyze the behavior of two important pipelined architectures for multiple clock-cycle global interconnect, namely, the flip-flop inserted global wire and the latch inserted global wire. We present analytical formula that is based on parameters obtained using Monte Carlo simulation. These results enable a global interconnect designer to explore design trade-offs between clock frequency and probability of bit-error during data transmission. Categories and Subject Descriptors Lizheng Zhang, Yu Hen Hu, Charlie Chung-Ping Chen |
DAC | 2 |
| 2004 | Content based blurring coding artifact reduction using patch-based texture synthesis [image/video coding applications]abstractBlurring artifacts occur in low bit-rate image or video coding. This manifests itself as a blurred patch within a textured region. Previously, we proposed a constrained texture synthesis postprocessing algorithm to regenerate the texture of the blurred patch using surrounding texture of the same kind. However, a human operator must manually identify the blurred target, and a valid source region from its surroundings. In this work, we present an effort to automate this process. Specifically, we developed an efficient modified k-means algorithm method to segment and identify potentially blurred patches; and a content-based region selection method to choose the candidate source region. Preliminary experiment results indicate that our algorithm produces results consistent with that produced by a human operator. Rajas A. Sambhare, Yu Hen Hu |
ICASSP (3) | 2 |
| 2004 | Sequential acoustic energy based source localization using particle filter in a distributed sensor networkabstractA sequential source localization method using a particle filter is presented to estimate and track multiple-target locations. This method is designed to make use of an acoustic signal measured at multiple acoustic sensors randomly deployed in a wireless distributed sensor network. By using the particle filter, a non-Gaussian probability density function of the target locations is represented by a discrete set of "particles". The positions of these particles are propagated sequentially using known state transition equation, and updated using new location estimates via the observation equation. Compared to a previously proposed maximum likelihood source localization algorithm, this new approach is computationally effective and more robust to parameter perturbation. Xiaohong Sheng, Yu Hen Hu |
ICASSP (3) | 2 |
| 2004 | Optimal decision fusion with applications to target detection in wireless ad hoc sensor networksabstractDecision fusion is a decentralized decision making process where local decisions are combined to reach a global decision. In this work, we propose a complementary optimal decision fusion (CODF) method to the target detection task that arises in wireless ad hoc sensor network signal processing. We conduct extensive comparative study using real world sensor signal data, and observe superior performance of CODF when compared with state-of-the-art decision fusion methods. In addition to distributed sensor network applications, the proposed CODF algorithm can be applied to numerous multi-modality, multi-agent, multi-media signal processing problems. Marco F. Duarte, Yu Hen Hu |
ICME | 2 |
| 2004 | Content-based image post-processing for blurring artifact reductionabstractBlurring artifacts occur in low bit-rate wavelet images or in video coding. They manifest themselves as blurred patches within textured regions. As such, it is possible to exploit this texture level correlation between adjacent image regions to reconstruct the lost texture. In order to automatically identify the blurred patches, and use valid source regions to draw textures from their surroundings regions, we developed an efficient content-based statistical testing method to choose the candidate source region and to identify potentially blurred patches. Preliminary experiment results indicate that our algorithm produces excellent results that are consistent with that produced by a human operator. Rajas A. Sambhare, Yu Hen Hu |
ICME | 2 |
| 2004 | Optimal decision fusion with applications to target detection in wireless ad hoc sensor networksabstractDecision fusion is a decentralized decision making process where local decisions are combined to reach a global decision. In this work, we propose a complementary optimal decision fusion (CODF) method to the target detection task that arises in wireless ad hoc sensor network signal processing. We conduct extensive comparative study using standard datasets, and observe superior performance of CODF when compared with state-of-the-art decision fusion methods. In addition to distributed sensor network applications, the proposed CODF algorithm can be applied to numerous multi-modality, multi-agent, multi-media signal processing problems. Marco F. Duarte, Yu Hen Hu |
MMSP | 2 |
| 2004 | Vehicle classification in distributed sensor networks
Marco F. Duarte, Yu Hen Hu |
J. Parallel Distributed Comput. | 2 |
| 2004 | A memory-efficient and high-speed sine/cosine generator based on parallel CORDIC rotationsabstractThe sine/cosine function generator is based on parallelization of the original CORDIC algorithm by predicting all the rotation directions directly from the binary bits of the initial input angle. Unlike previous approaches that require complicated circuits or exponentially increased ROM, our proposed architecture has a relatively simple prediction scheme through an efficient angle recoding. The critical path delay is also reduced by utilizing the predicted rotation directions to design an efficient multioperand carry-save addition structure. Shen-Fu Hsiao, Yu Hen Hu, Tso-Bing Juang |
IEEE Signal Process. Lett. | 2 |
| 2003 | Constrained texture synthesis for image post processingabstractA novel constrained texture synthesis approach is proposed to enhance the visual quality of a degraded image by reconstructing its high frequency texture content. In low-bit-rate image and video compression and communication systems, high frequency transformed coefficients are often lost due to aggressive quantization or uneven error protection schemes. However, with the block-based encoding and transmission methods, the amount of high frequency texture loss is uneven between adjacent blocks. As such, it is possible to exploit this texture-level correlation between adjacent image blocks to reconstruct the lost texture. This is accomplished by applying a state of art patch-based quilting texture synthesis algorithm in this paper. Yu Hen Hu, Rajas A. Sambhare |
ICASSP (3) | 1 |
| 2003 | Constrained texture synthesis for image post processingabstractA novel constrained texture synthesis approach is proposed to enhance the visual quality of a degraded image by reconstructing its high frequency texture content. In low-bit-rate image and video compression and communication systems, high frequency transformed coefficients are often lost due to aggressive quantization or un-even error protection schemes. However, with the block-based encoding and transmission methods, the amount of high frequency texture loss is uneven between adjacent blocks. As such, it is possible to exploit this texture-level correlation between adjacent image blocks to reconstruct the lost texture. This is accomplished by applying a state of art patch-based quilting texture synthesis algorithm in this paper. Yu Hen Hu, Rajas A. Sambhare |
ICME | 1 |
| 2003 | Processor Array Synthesis from Shift-Variant Deep Nested Do Loops
Surin Kittitornkun, Yu Hen Hu |
J. Supercomput. | 2 |
| 2003 | Mapping deep nested do-loop DSP algorithms to large scale FPGA array structuresabstractRecently, FPGAs (field programmable gate arrays) technology have made significant advances in both speed and capacity. Millions of logic gates are now available for reconfiguration programming. To fully exploit the potential of so many programmable devices, powerful design methodology must be developed. In this paper, we propose a novel systematic computer-aided design methodology that can efficiently implement deeply nested do-loop algorithms on a FPGA. Specifically, our design methodology maps the loop dependence graph onto a linear array of locally connected processing elements to exploit parallelism. Due to the regular structure of this linear array of processors, it can be easily implemented on a FPGA. While this method is based on conventional systolic array design methodology, our proposed approach exhibits two distinct features that contribute to its superior performance: 1) We developed a novel multiple-order dependence graph representation that is able to efficiently represent distinct, yet correct algorithm execution orders. 2) We developed new FPGA-specific architectural constraints during the mapping process. As such, FPGA implementations based on our approach will utilize much fewer lookup tables while achieving superior performance. Surin Kittitornkun, Yu Hen Hu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2002 | Correlation-Based Web Document Clustering for Adaptive Web Interface Design
Zhong Su, Qiang Yang 0001, HongJiang Zhang, Xiaowei Xu 0001, Yu Hen Hu, Shaoping Ma |
Knowl. Inf. Syst. | 5 |
| 2001 | Automatic Training of a Neural Net for Active Stereo 3D ReconstructionabstractAddresses the problem of recovering 3D geometry using an active stereo vision system. Calibration procedures can be adapted to the active stereo configuration, however, considerable effort is required to accurately model and calibrate the kinematics to avoid poor reconstruction. In the active stereo case there will also be errors due to uncertainty in the kinematics of the system. In addition, data collection needs to be automated because active stereo requires significantly more information for calibration. We present a biologically inspired neural network trained to determine the mapping between 3D geometry and stereo image points. To train the network, we have developed a system to automatically collect accurate calibration data. We compare the reconstructed 3D geometry obtained using a kinematic model based approach with our neural network approach. Jeremiah Neubert, Anthony Hammond, Yongtae Do, Yu Hen Hu, Nils Guse |
ICRA | 4 |
| 2001 | Low bit rate video sequence coding artifact removalabstractThe picture quality of video frames encoded at very low-bit rates often suffers from both blocking and ringing artifacts. We present two post-processing methods to mitigate the visual quality degradation caused by these artifacts. To reduce the blocking artifact of decoded images, we substitute IDCT for the lapped orthogonal transform embedded inverse discrete cosine transform (le-IDCT). On the other hand, we post-process the decoded video frames using a nonlinear robust filter to reduce the ringing artifact. Extensive simulation results indicated significant improvement in both objective and subjective visual qualities. The computation overhead incurred due to these quality enhancement operations is quite moderate, and can be easily optimized to achieve real-time operation. Seungjoon Yang, Surin Kittitornkun, Yu Hen Hu, Truong Q. Nguyen, Damon L. Tull |
MMSP | 3 |
| 2001 | Frame-level pipelined motion estimation array processorabstractA systolic motion estimation processor (MEP) core architecture implementing the full-search block-matching (FSBM) algorithm is presented. A unique feature of this MEP architecture is its support of frame-level pipelined operation. As such, it is possible to process pixels from consecutive frames without any processor idle time. It is designed so that no data broadcasting operations are required, and achieves 100% fully pipelined computation. It compares favorably with existing MEP architectures in terms of both performance and complexity of architecture. Surin Kittitornkun, Yu Hen Hu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2001 | Maximum-likelihood parameter estimation for image ringing-artifact removalabstractAt low bit rates, image compression codecs based on overlapping transforms introduce spurious oscillations known as ringing artifacts in the vicinity of major edges. Unlike previous works, we present a maximum-likelihood approach to the ringing-artifact removal problem. Our approach employs a parameter estimation method based on the k-means algorithm with the number of clusters determined by a cluster-separation measure. The proposed algorithm and its simplified approximation are applied to JPEG2000 compressed images. Our results show effective and efficient removal of ringing artifacts. Seungjoon Yang, Yu Hen Hu, Truong Q. Nguyen, Damon L. Tull |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2000 | Data Partitioning and Reversible Variable Length Codes for Robust Video CommunicationsabstractLow bit-rate multimedia communication over wireless channels has received much attention recently. A key challenge in low bit-rate wireless communication is the very high error rate during transmission. This demands error resilient services that exhibit graceful performance degradation while operating in a highly noisy channel. Among video, audio and other multimedia communication modalities, we focus on the error-resilient transmission of video signals over wireless channels in this paper. Specifically, efforts in developing H.263++ Annex V error resilient data partitioning with reversible variable length code (RVLC) are described in detail. The performance over error-prone channels is analyzed and the effectiveness of the syntax is demonstrated with extensive simulations. Adam H. Li, Surin Kittitornkun, Yu Hen Hu, Dong-Seek Park, John D. Villasenor |
Data Compression Conference | 3 |
| 2000 | Maximum Likelihood Parameter Estimation for Image Ringing Artifact RemovalabstractAt low bit rates, image compression codecs based on overlapping transforms introduce spurious oscillation known as ringing artifacts in the vicinity of major edges. The image quality can be enhanced considerably by removing the artifacts. We present a maximum likelihood approach to the ringing artifact removal problem. Our approach employs a parameter estimation method based on the k-means algorithm with the number of clusters determined by a cluster separation measure. The proposed algorithm and its simplified approximation are applied to JPEG2000 compressed images to demonstrate their effectiveness. Seungjoon Yang, Yu Hen Hu, Damon L. Tull, Truong Q. Nguyen |
ICIP | 2 |
| 2000 | Blocking Artifact Free Inverse Discrete Cosine TransformabstractThis paper presents the generalized lapped biorthogonal transform embedded inverse discrete cosine transform (ge-IDCT) as an alternative to the IDCT. The ge-IDCT with nonlinear weighting in the embedded transform domain can reconstruct the signal with alleviated blockishness. Additional complexity, imposed by the replacement, is trivial thanks to an efficient lattice structure. The proposed ge-IDCT is applied in the JPEG still image compression standard to demonstrate its validity. Seungjoon Yang, Surin Kittitornkun, Yu Hen Hu, Truong Q. Nguyen, Damon L. Tull |
ICIP | 3 |
| 1999 | Minimum initiation interval of multi-module recurrent signal processing algorithm realization with fixed communication delayabstractA novel iterative algorithm is proposed to compute the theoretical minimum initiation interval of a given recurrent algorithm when there is a known, fixed inter-module communication delay. Specifically, for a twin-module implementation problem, a novel representation called necessary initiation interval is introduced to facilitate the development of an iterative algorithm which yields both the minimum initiation interval and the corresponding cut set of the cyclic iterative computational dependence graph (ICDG). The convergence of this iterative algorithm in finite iterations is also proved. Hung-Ying Tyan, Yu Hen Hu |
ICASSP | 2 |
| 1999 | Blocking effect removal using robust statistics and line processabstractThe Gibbs distribution with spatially adaptive clique potential functions is used as the image prior distribution in the coding artifact reduction problem. Unlike previous approaches, the image is divided in regions based on the local statistics and a clique potential function is assigned for each region. A non-differentiable potential function that preserves edge in detailed image regions is introduced. Images recovered by the proposed approach contain reduced coding artifact and improved preservation of details. Seungjoon Yang, Yu Hen Hu, Damon L. Tull |
MMSP | 2 |
| 1999 | Optimal linear spectral unmixingabstractThe optimal estimate of ground cover components of a linearly mixed spectral pixel in remote-sensing imagery is investigated. The problem is formulated as two consecutive constrained least-squares (LS) problems: the first problem concerns the estimation of the end-member spectra (EMS), and the second concerns the estimate, within each mixed pixel, of ground cover class proportions (CCPs) given the estimated EMS. For the EMS estimation problem, the authors propose a total least-squares (TLS) solution as an alternative to the conventional LS approach. The authors pose the CCP estimation problem as a constrained LS optimization problem. Then, they solve for exact solution using a quadratic programming (QP) method, as opposed to the Lagrange multiplier (LM)-based approximated solution proposed by Settle and Drake (1993). Preliminary computer experiments indicated that the TLS-estimated EMS always leads to better estimates of CCP than that of the LS-estimated EMS. Yu Hen Hu, H. B. Lee, F. L. Scarpace |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 1998 | A novel, batch modular learning approach for ECG beat classificationabstractWe investigate a modular architecture for ECG beat classification. The feature space is divided into distinct regions and individual classifiers are developed for each region. We compare different combination strategies, and feature space partition strategies. We also describe a novel, batch modular learning method that can be used to incrementally improve the performance of the modular network. Vijay P. Mani, Yu Hen Hu, Surekha Palreddy |
ICASSP | 2 |
| 1998 | Blocking Effect Removal using Regularization and DitheringabstractImage compression codecs suffer from the various coding artifacts at low bit rate. The coding artifacts can be removed by way of post-processing. Most of the coding artifacts removal algorithms depend on a priori knowledge of the original image, which is about its smoothness. The resulting surface of the restored image is often too smooth to be the one in the natural scene. Dithering is based on the idea of adding controlled noise to the system to achieve the better results. An algorithm based on regularization and dithering is proposed to remove the coding artifacts. The results show that the algorithm removes the coding artifacts successfully and the surface of the restored image is more visually pleasing. Seungjoon Yang, Yu Hen Hu |
ICIP (1) | 2 |
| 1998 | Image coding ringing artifact reduction using morphological post-filteringabstractRinging is an annoying artifact frequently encountered in low bit-rate transform and subband decomposition based compression of different media such as image, intra frame video and graphics. A mathematical morphology based post-processing algorithm is presented in this paper for image ringing artifact suppression. First, we use binary morphological operators to isolate the regions of an image where the ringing artifact is most prominent to the human visual system (HVS) while preserving genuine edges and other (high-frequency) fine details present in the image. Then, a gray-level morphological nonlinear smoothing filter is applied to the unmasked regions of the image under the filtering mask to eliminate ringing within this constraint region. To gauge the effectiveness of this approach, we propose an HVS compatible objective measure of the ringing artifact. Preliminary simulations indicate that the proposed method is capable of significantly reducing the ringing artifact on both subjective and objective basis. Seyfullah H. Oguz, Yu Hen Hu, Truong Q. Nguyen |
MMSP | 2 |
| 1998 | A modular high-throughput architecture for logarithmic search block-matching motion estimationabstractA high-throughput modular architecture for a logarithmic search block-matching algorithm is presented. The design efforts are focused on exploiting the search area data dependencies using special data input ordering constraints. The input bandwidth problem has been solved by a random access on-chip memory, and a simple address generation procedure has been described. Furthermore, this architecture can handle a large search range with unequal horizontal and vertical spans using a technique called pipeline interleaving. Compared to the existing architectures for the three-step search BMA, this architecture delivers a high throughput rate with fewer input lines, and is linearly scalable. Hangu Yeo, Yu Hen Hu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1997 | Committee pattern classifiersabstractMethods which combine outputs of multiple pattern classifiers to enhance the overall performance of pattern classification are presented. Specific attention is given to combination rules which are independent of the input feature vectors. Potentials and pitfalls of this so called stack generalization method are discussed, and experimentation using several machine learning databases are reported. Yu Hen Hu, Jong-Min Park, Thomas Knoblock |
ICASSP | 1 |
| 1997 | On-line learning in pattern classification using active samplingabstractAn adaptive on-line learning method is presented to facilitate pattern classification using active sampling to identify optimal decision boundary for a stochastic oracle with minimum number of training samples. The strategy of sampling at the current estimate of the decision boundary is shown to be optimal in the sense that the probability of convergence toward the true decision boundary at each step is maximized, offering theoretical justification on the popular strategy of category boundary sampling used by many query learning algorithms. Analysis of convergence in distribution is formulated using the Markov chain model. Jong-Min Park, Yu Hen Hu |
ICASSP | 2 |
| 1997 | A motion estimation and image segmentation technique based on the variable block sizeabstractWe discuss our effort to develop a motion estimation algorithm based on variable block size, which reduces the computational complexity dramatically while maintaining a good picture quality as well as a high compression ratio. A key step in this work is to segment each image frame into different regions using a simple binary-level classifier which performs bit-wise comparison. In the second stage, the motion estimation is performed for every block of variable block size within the changed region with a predetermined maximum search range. The proposed technique has been applied to interframe video coding, and it has been shown that this scheme can be a feasible solution for the low bit rate coding application such as video telephony. Hangu Yeo, Yu Hen Hu |
ICASSP | 2 |
| 1997 | Coding Artifact Removal Using Biased Anisotropic DiffusionabstractBiased anisotropic diffusion is applied to the coding artifacts removal of the DCT based codec. It is formulated as a cost minimization problem. The weighting factors of the cost function are controlled such that the solution removes the blocking effect and conceals the block losses. It has an advantage over other postprocessing schemes because it handles the discontinuity of the image, smoothes the image selectively, and takes the visual masking in to account. Features needed for the weighting factors are extracted directly from the DCT coefficients to reduce the computational complexity. Seungjoon Yang, Yu Hen Hu |
ICIP (2) | 2 |
| 1997 | Joint optimization of lattice vector quantizer and entropy coder for a Laplacian sourceabstractThis paper presents a joint optimization algorithm for lattice vector quantization (LVQ) and entropy coding for a Laplacian source at all ranges of bit rates. Entropy-constrained lattice vector quantizers (ECLVQs) are often used in practical coding systems. In order to develop an ECLVQ design algorithm, we derive estimation expressions for both distortion and entropy. From these estimations, we develop an algorithm that jointly optimizes LVQ and the entropy coder pair for a given entropy rate. Compared to previously reported approaches, the approach reported quickly computes a highly accurate optimal ECLVQ at all ranges of bit rates. Since a Laplacian source represents a wide class of subband transformed data, the algorithm can be readily applied as a subband coding method. When the proposed algorithm is applied to a wavelet based image coding, the coding performance surpasses those of any previously reported subband coders, especially at low bit rates. Wonha Kim, Yu Hen Hu, Truong Q. Nguyen |
MMSP | 2 |
| 1996 | A Novel Matching Criterion And Low Power Architecture For Real-Time Block Based Motion EstimationabstractIn recent years, minimizing the power consumption has become a key issue in the design of portable electronic devices. In this paper, low power architecture which can support the real time motion estimation of video signals is presented. The architecture is based on a binary level matching criterion which performs a bit-wise comparison. The processor level design based on simple combinational logic using the binary level matching criterion has been introduced. Compared with the existing architectures, the proposed architecture delivers higher throughput rate, requires fewer input/output lines, and reduces the total power consumption. Hangu Yeo, Yu Hen Hu |
ASAP | 2 |
| 1996 | Synthesis of Real-Time Recursive DSP Algorithms Using Multiple ChipsabstractIn this paper, the problem of synthesizing real-time recursive DSP algorithms with fixed interprocessor communication delay is addressed. The effects of the communication delay to the initiation interval and number of chips are studied. We differentiate our problem from previous work in two parts. First, the DSP algorithms we consider are recurrence. Second, communication delay is considered. By modifying previously proposed scheduling and allocation algorithm, we are able to derive an implementation if it exists under the given real-time and area constraints. Some experiments have been made and results are very promising. Duen-Jeng Wang, Yu Hen Hu |
Great Lakes Symposium on VLSI | 2 |
| 1996 | A Modular Architecture for Real Time HDTV Motion Estimation with Large Search RangeabstractA modular architecture with random access on-chip local memory for real-time motion estimation has been proposed. The random access on-chip local memory with simple address generation has been proposed to overcome the irregular data flow of the three-step search BMA. This architecture features simple interconnection with low memory bandwidth and throughput rate as high as 1/N block per clock cycle for an N/spl times/N block with the search range of d/sub m/=N/2-1 pixels with 100% processor utilization. By using a method called pipeline interleaving, this architecture offers a feasible solution for the Grand Alliance HDTV picture format with large search range. Hangu Yeo, Yu Hen Hu |
Great Lakes Symposium on VLSI | 2 |
| 1996 | A high-throughput modular architecture for three-step search block matching motion estimationabstractThe three-step hierarchical search block matching motion estimation algorithm has played an important role in low bit rate video coding because of its low computation complexity compared to the full search block matching motion estimation algorithm (FBMA). A modular architecture for the three-step hierarchical search BMA is presented, which features a throughput rate, as high as 1/N block per clock cycle, and a low memory bandwidth with random access on-chip local memory. Furthermore, 100% processor utilization has been achieved by using a method called pipeline interleaving. As such, this architecture offers a feasible solution for the Grand Alliance HDTV picture format with a large search range. Hangu Yeo, Yu Hen Hu |
ICASSP | 2 |
| 1996 | On-line learning for active pattern recognitionabstractAn adaptive on-line learning method is presented to facilitate pattern classification using active sampling to identify the optimal decision boundary for a stochastic oracle with a minimum number of training samples. The strategy of sampling at the current estimate of the decision boundary is shown to be optimal compared to random sampling in the sense that the probability of convergence toward the true decision boundary at each step is maximized, offering theoretical justification on the popular strategy of category boundary sampling used by many query learning algorithms. Jong-Min Park, Yu Hen Hu |
IEEE Signal Process. Lett. | 2 |
| 1996 | A Novel Implementation of CORDIC Algorithm Using Backward Angle Recoding (BAR)abstractWe propose a backward angle recoding (BAR) method to eliminate redundant CORDIC elementary rotations and hence expedite the CORDIC rotation computation. We prove that for each of the linear, circular, and hyperbolic CORDIC rotations, the use of BAR guarantees more than 50% reduction of elementary CORDIC rotations provided the scaling factor needs not be kept constant. The proposed BAR algorithm is simple, and amenable to VLSI implementation. Taking practical applications into consideration, we discuss how to incorporate convergence range enhancement procedure with BAR, and how easy it is to devise a constant-scaling-factor BAR algorithm while still enjoying 25% reduction of CORDIC elementary rotations. Yu Hen Hu, Homer H. M. Chern |
IEEE Trans. Computers | 1 |
| 1996 | Analysis of convergence properties of a stochastic evolution algorithmabstractIn this paper, the convergence properties of a stochastic optimization algorithm called the stochastic evolution (SE) algorithm is analyzed. We show that a generic formulation of the SE algorithm can be modeled by an ergodic Markov chain. As such, the global convergence of the SE algorithm is established as the state transition from any initial state to the globally optimal states. We propose a new criterion called the mean first visit time (MFVT) to characterize the convergence rate of the SE algorithm. With MFVT, we are able to show analytically that on average, the SE algorithm converges faster than the random search method to the globally optimal states. This result Is further confirmed using the Monte Carlo simulation. Chi-Yu Mao, Yu Hen Hu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1995 | Breath detection using a fuzzy neural network and sensor fusionabstractWe have developed and trained a fuzzy neural network (FNN) to detect individual breaths using information from multiple independent noninvasive ventilation sensors. We derive input features from simultaneous recordings from impedance and inductance plethysmographs, and a pneumotachometer while healthy adults performed several different combinations of ventilation and motion. We first tested our FNN using membership functions, rules and consequent sets derived using a heuristic approach. Using all features, on 4 subjects we found that the average rate of combined false-positive and false-negative detections was 5.1%. When we trained our FNN using a gradient descent algorithm, the average rate of combined false-positive and false-negative detections was reduced to 2.6%. Kevin P. Cohen, Yu Hen Hu, Willis J. Tompkins, John G. Webster |
ICASSP | 2 |
| 1995 | Optimal VLSI architecture for vector quantizationabstractOptimal VLSI array structure design for the implementation of vector quantization (VQ) are investigated in this paper. After a brief review of the VQ algorithms, the algorithm and architecture design issues will be discussed. This is followed by a brief survey of existing VQ implementation strategies and architecture. Yu Hen Hu |
ICASSP | 1 |
| 1995 | Wavelet packet based optimal subband coderabstractProposes an algorithm to use in designing a subband coder (SBC) constructed by wavelet packet, to achieve minimum distortion for a given bit budget and implementation complexity. The authors map the QMF tree structures onto a binary tree, then formulate the task as an optimization problem including coding bit and implementation complexity constraints. The problem is dissected into two phases. First, they derive the optimal bit allocation strategy which covers the entire range of bit rate, and second, they search for the optimal subband decomposition by using a fast dynamic program. Wonha Kim, Yu Hen Hu |
ICASSP | 2 |
| 1995 | A novel modular systolic array architecture for full-search block matching motion estimationabstractProposes a modular systolic array architecture for the full-search block matching motion estimation algorithm (FBMA). With this novel architecture, the authors are able to generate a motion vector for every reference block in raster scan order while achieving 100% processor utilization and high throughput rate. Furthermore, they devised a scheme to save the pin count (I/O) by sharing memory units. This results in low memory bandwidth. This architecture is scalable in that it can easily be adapted to handle larger search ranges and different block sizes without increasing the effective latency. Hangu Yeo, Yu Hen Hu |
ICASSP | 2 |
| 1995 | Automated entry system for Chinese printed documents
Bor-Shenn Jeng, Tung-Ming Shieh, Char-Shin Miou, Chun-Jen Lee, Bing-Shan Chien, Yu Hen Hu, Gan-How Chang |
Image Vis. Comput. | 6 |
| 1995 | A novel modular systolic array architecture for full-search block matching motion estimationabstractA novel modular systolic array architecture for the full search block matching motion estimation algorithm (FBMA) is presented. The design efforts are focused on matching the array computation to system level input/output constraints. Compared to previously proposed FBMA architectures, this new architecture delivers highest throughput rate, achieves 100% processor utilization, requires much fewer input/output lines (pin count), and is linearly scalable. As such, this architecture offers a feasible solution for progressive-scan HDTV picture format.> Hangu Yeo, Yu Hen Hu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1995 | Multiprocessor implementation of real-time DSP algorithmsabstractIn this paper, we consider multiprocessor implementation of real-time recursive digital signal processing algorithms. The objective is to devise a periodic schedule, and a fully static task assignment scheme to meet the desired throughput rate while minimizing the number of processors. Toward this goal, we propose a notion called cutoff time. We prove that the minimum-processor schedule can be found within a finite time interval bounded by the cutoff time. As such the complexity of the scheduling algorithm and the allocation algorithm can be significantly reduced. Next, taking advantage of the cutoff time, we derive efficient heuristic algorithms which promise better performance and less computation complexity compared to other existing algorithms. Extensive benchmark examples are tested which yield most encouraging results.> Duen-Jeng Wang, Yu Hen Hu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1994 | An efficient multiprocessor implementation scheme for real-time DSP algorithmsabstractAn algorithm to derive minimum-processor implementation for real-time DSP algorithms is proposed. In order to make the number of possible schedules finite and to assure the optimal schedule within the search space, the authors define a novel notion of cutoff time. All the possible schedules can find an equivalent schedule that finishes before cutoff time. Next, they apply all efficient heuristic periodic scheduling and fully static allocation algorithms derived from two generic problem solving heuristics developed in a branch of artificial intelligence research called planning. Extensive benchmarks have been tested and the results are most encouraging.> Yu Hen Hu, Duen-Jeng Wang |
Great Lakes Symposium on VLSI | 1 |
| 1994 | Convergence analyses of simulated evolution algorithmsabstractIn this paper, we show that simulated evolution (SE) can be modeled by an ergodic Markov chain. As such, the global convergence of the SE algorithm is established. Moreover, we propose to use the mean first visit time of an ergodic Markov chain to characterize the convergence time of the SE algorithm such that the fast convergence feature of SE can be assessed theoretically and experimentally.> Chi-Yu Mao, Yu Hen Hu |
Great Lakes Symposium on VLSI | 2 |
| 1994 | Is it Possible to achieve a Teraflop/s on a chip? From High Performance Algorithms to ArchitecturesabstractThe forumnists address the question of high density computations on a single chip. The surface of a chip offers an ideal medium not only to store information or to process data, but also to execute computations. The 1 Giga floating point operations per second per chip mark has been achieved, we are now moving towards the teraflop mark. How is this going to happen, what are the limitations, what are the opportunities-those are central questions.> Francky Catthoor, Ed F. Deprettere, Yu Hen Hu, Jan M. Rabaey, Heinrich Meyr, Lothar Thiele |
ISCAS | 3 |
| 1994 | EDLICS: A New Relaxation-Based Electrical Circuit Simulation TechniqueabstractEvent-Driven Local-Interactive Circuit Simulation (EDLICS) is a new approach to the transient simulation of electrical circuits. Unlike the simulation methods which are based on differential equations, EDLICS directly models the physical behavior of electrical components. The EDLICS approach propagates incremental changes in circuit variables (voltage, current, and charge) as events similar to the behavior of event-driven logic simulation. An electrical component receives an event (either voltage, current, or charge) from a node, converts it to another event and delivers it to a neighboring node. The transient behavior of a circuit is analyzed by evaluating capacitors with current flow for charge buildup. EDLICS is the generalization of relaxation-based simulation, although it originated in a different perspective. The Kirchhoff Current Law, Kirchhoff Voltage Law, and capacitor current integration formula are used interactively in the Gauss-Seidel fashion. Since EDLICS is based on the true behavior modeling of electrical components, it is capable of handling circuit components and circuit configurations which other relaxation-based simulators cannot. The EDLICS technique is applied to a MOS circuit simulator, EDsim. EDsim is one to two orders of magnitude faster than the latest version of SPICE3 and compares favorably with iSPLICE3.> Jai-Cheol Lee, Yu Hen Hu |
ISCAS | 2 |
| 1993 | A fast pipelined CORDIC-based adaptive lattice filter
Dorin Panescu, Yu Hen Hu, Willis J. Tompkins |
ICASSP (3) | 2 |
| 1993 | Optimized code generation for programmable digital signal processors
Kin H. Yu, Yu Hen Hu |
ICASSP (1) | 2 |
| 1993 | Solving Gate-Matrix Layout Problems by Simulated Evolution
Yu Hen Hu, Chi-Yu Mao |
ISCAS | 1 |
| 1993 | An Angle Recoding Method for CORDIC Algorithm ImplementationabstractThe coordinate rotation digital computer (CORDIC), an iterative arithmetic algorithm for computing generalized vector rotations without performing multiplications, is discussed. For applications where the angle of rotation is known in advance, a method to speed up the execution of the CORDIC algorithm by reducing the total number of iterations is presented. This is accomplished by using a technique called angle recoding, which encodes the desired rotation angle as a linear combination of very few elementary rotation angles. Each of these elementary rotation angles takes one CORDIC iteration to compute. The fewer the number of elementary rotation angles, the fewer the number of iterations are required. A greedy algorithm which takes only O(n/sup 2/) operations is developed to perform CORDIC angle recoding. It is proven that this algorithm is able to reduce the total number of required elementary rotation angles by at least 50% without affecting the computational accuracy.> Yu Hen Hu, S. Naganathan |
IEEE Trans. Computers | 1 |
| 1993 | SaPOSM: an optimization method applied to parameter extraction of MOSFET modelsabstractPoints out that SaPOSM integrates an efficient deterministic optimization algorithm, called POSM, with the popular stochastic optimization paradigm simulated annealing (SA). It offers great promise for improving the optimization results significantly while using only a moderate amount of computing time. Tested on a suite of multi-minima optimization benchmark problems, SaPOSM's performance rivals a recently reported fast simulated diffusion method. SaPOSM was used to extract the parameters of state-of-the-art submicron (0.3- mu m channel length) MOSFET transistors, and very favorable results have been obtained. For a second difficult parameter extraction problem (18 parameters, five different channel lengths), simulation results indicate that SaPOSM achieves performance comparable to the SA method. Specifically, both SA and SaPOSM are able to minimize the modeling error to several orders of magnitude smaller than that obtained using POSM alone. At the same time, the computing time taken by SaPOSM is only a very small fraction of that taken by the SA method.> Yu Hen Hu, ShaoWei Pan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1993 | PYFS-a statistical optimization method for integrated circuit yield enhancementabstractAn efficient optimization method for the statistical design of integrated circuits is presented. This method, called the pseudo yield function substitution (PYFS) algorithm, is developed to help a designer select design parameters to maximize the product yield. The design goal of PYFS is to use fewer simulation runs to reach a yield-optimized design. This is accomplished with the development of an improved response surface method for accurate estimation of the circuit response function, and the use of a novel PYFS method for yield maximization.> ShaoWei Pan, Yu Hen Hu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1992 | On systolic mapping of multi-stage algorithmsabstractThe authors present a more general mapping problem called multi-stage systolic mapping which focuses on the computing algorithms containing more than one nested loop constructs to be executed sequentially. Since the emerged interface problem now becomes the dominant factor in performing the mapping, the authors argue that the adjacent stages should have matched interface to reduce the overhead. For this, the conditions of interface matching between two stage's mappings are established. A systematic method to derive the interface matched mapping is also presented. To improve the performance degradation due to the initial and final phases of computation in systolic computing, the inter-stage computation concurrency is explored by overlapping part of the computations in successive stages and thus effectively reduces the computation latency. With these results, the multi-stage systolic mapping tool (MSSM) is developed and several design examples are presented to illustrate the potential use of MSSM.> Yin-Tsung Hwang, Yu Hen Hu |
ASAP | 2 |
| 1992 | Fully static multiprocessor realization for real-time recursive DSP algorithmsabstractA systematic approach to implement a real time recursive digital signal processing algorithm on a dedicated multiprocessor array is presented. First, the authors unfold the algorithm so that its corresponding dependence graph becomes a newly defined generalized perfect rate graph. They prove that the dependence graph of a recursive algorithm admits a desirable rate optimal, full static multiprocessor implementation if and only if it is a generalized perfect rate graph. Based on these results, an efficient heuristic algorithm is presented to perform optimal multi-processor scheduling and task assignment so that the number of processors required is minimized.> Duen-Jeng Wang, Yu Hen Hu |
ASAP | 2 |
| 1991 | Structural simplification of a feed-forward, multilayer perceptron artificial neural networkabstractSeveral methods to reduce the excessive number of neurons and synaptic weights in a feedforward, multilayer perceptron artificial neural network (ANN) are presented. To reduce the synaptic weights, the authors replace the original weight matrix by a product of two smaller matrices so that the number of multiplications required can be reduced. To reduce the hidden units, they exploit the correlation among the outputs of the hidden neurons in the same layer. A method to identify and remove redundant hidden units and update the weights of the remaining neurons is proposed. This approach offers potentially good performance without retraining. When retraining is applied to fine-tune the reduced network, the updated weights become very good initial conditions enabling much faster training compared with training with random initial conditions.> Yu Hen Hu, Qiuzhen Xue, Willis J. Tompkins |
ICASSP | 1 |
| 1991 | Optimal scheduling of linear recurrence equations on a multiprocessor arrayabstractThe authors propose a systematic approach to a rate-optimal, fully static multiprocessor implementation for a real-time recurrence algorithm. The objective is to minimize the number of processors and the number of interprocessor communication links. For any arbitrary algorithm, rate optimal fully static implementation may not be obtained without algorithm transformation. The authors generalize the perfect-rate graphs, which can always achieve their rate-optimal fully static implementations, requiring no algorithm transformation, by relaxing the restriction stated by K. K. Parki and D. G. Messerschmitt (1989). An optimal unfolding factor is introduced to tell at least how many times to unfold a loop to its corresponding generalized perfect-rate graph. A scheme employing the artificial-intelligence planning problem solver is proposed to do scheduling and processor assignment for a generalized perfect-rate graph so that the design can meet the goal.> Duen-Jeng Wang, Yu Hen Hu |
ICASSP | 2 |
| 1990 | An efficient VLSI CORDIC array structure implementation of Toeplitz eigensystem solversabstractA novel, efficient implementation of the Toeplitz eigensystem solver using a doubly pipelined VLSI CORDIC array processor is presented. First, a backward CORDIC angle recoding scheme is proposed which is able to reduce the number of internal CORDIC iterations by at least 50%. It is shown how to apply this scheme to a family of feedforward algorithms for solving general linear systems, especially the Toeplitz systems. This leads to the implementation of a Toeplitz eigensystem solver with CORDIC-based array processors.> Yu Hen Hu, Homer H. M. Chern |
ICASSP | 1 |
| 1990 | Analyses of the hidden units of the multi-layer perceptron and its application in acoustic-to-articulatory mappingabstractAn artificial neural network (ANN) is applied to perform the task of acoustic-to-articulatory inversion. The objective is to model the highly nonlinear mapping from linear predictive coding (LPC) code to corresponding articulatory parameters with a multilayer perceptron ANN structure. Such information will facilitate the study of the relationships between the acoustic signal and the physical vocal tract which produces it. Several novel approaches for devising the ANN structure have been evaluated. Specifically, the performance of two learning algorithms, a backpropagation (BP) algorithm, and a random optimization (RM) algorithms, are compared. To reduce excessive, redundant hidden units in the multilayer perceptron model, a singular value decomposition is applied to either the weight matrix or the output covariance matrix of the hidden units to check their corresponding ranks. In both cases, their ranks are closely related to the number of essential decision regions in the input data.> Qiuzhen Xue, Yu Hen Hu, Paul H. Milenkovic |
ICASSP | 2 |
| 1990 | GM Plan: a gate matrix layout algorithm based on artificial intelligence planning techniquesabstractThe CMOS gate matrix layout problem is formulated and solved as an artificial intelligence planning problem in which a plan (the solution algorithm) is to be generated to achieve a goal (the gate matrix layout). The overall goal consists of many subgoals, each of which corresponds to the placement of a gate to a slot, and to the routing of associated nets connecting to that gate. As different nets compete for track (resource) usage, these subgoals interact (interfere) with each other, rendering suboptimal solutions. Here, such interaction among subgoals is managed with two artificial intelligence planning techniques: hierarchical subgoal organization and domain-independent search control policies. The subgoal hierarchy facilitates an object classification of the subgoals into priority classes according to a proposed distance measure of connectivity. Two search control policies (general problem-solving heuristics)-most-constraint (MC) and least impact (LI)-are used to guide the search process. A planning-based gate matrix layout algorithm, called GM Plan, which combines the gate placement and net routing into a single, incremental, problem-solving loop has been developed using these techniques. Encouraging results have been observed in a number of test examples.> Yu Hen Hu, Sao-Jie Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1989 | Parallel eigenvalue decomposition for Toeplitz and related matricesabstractParallel algorithms for computing the eigenvalues and eigenvectors of real, symmetric Toeplitz and Toeplitz-related low-displacement-rank matrices (e.g. sample covariance matrices) are presented. In particular, the parallel implementation of a class of modified Rayleigh-quotient iteration methods is discussed. Apart from parallel factorization of Toeplitz and Toeplitz-related matrices, other levels of inherent parallelism are exploited, rendering higher efficiency for parallel implementation of these algorithms. Specifically, a parallel multisectioning method is developed on a linear array of locally connected processors.> Yu Hen Hu |
ICASSP | 1 |
| 1988 | A rotation based method for solving covariance and related linear systemsabstractAn effective algorithm is presented for the Cholesky factorization of symmetric linear system equations with low displacement ranks. This proposed method represents an improved implementation of the generalized Schur algorithm (GSA) proposed by T. Kailath et al. (1979). It is shown that the (GSA) can be implemented with a sequence of circular and hyperbolic plane rotations. With careful arrangement, the number of the numerically undesirable hyperbolic rotations can be reduced to one per iteration. Hence the numerical stability of its algorithm is significantly improved. It is also shown that the GSA can be generalized to handle indefinite low-displacement rank liner systems as well. This improvement expands the potential applications of GSA for practical problems.> Yu Hen Hu |
ICASSP | 1 |
| 1988 | The quantization effects of the CORDIC algorithm [coordinate rotation digital computer]abstractCORDIC is a rotation-based arithmetic computing algorithm which has many important signal processing applications. Quantization errors occurring in the CORDIC algorithm are analyzed. Two types of quantization errors in the CORDIC algorithm are identified: one is an approximation error due to a quantized representation of rotation angles; another is the rounding error due to finite-precision arithmetic. Tight error bounds for these two types of errors are derived for both fixed-point and floating-point arithmetic. The effect of scaling (normalization) has been taken into account. These theoretical results are verified by simulation examples. The impact of these results on the architecture of a practical CORDIC processor is discussed.> Yu Hen Hu |
ICASSP | 1 |
| 1988 | Notes on eigenvalue distribution of Toeplitz matricesabstractThe author investigates the asymptotic eigenvalue distributions of degenerate Toeplitz matrices. A Toeplitz matrix is degenerate if its rank remains constant while its dimension increases. This type of Toeplitz matrix can arise as the covariance matrix of a degenerated harmonic process. He shows that asymptotically the individual eigenvalues of the degenerate Toeplitz matrix converge to the amplitude of corresponding harmonic components. This is a stronger result than previously developed.> Yu Hen Hu |
ICASSP | 1 |
| 1988 | New method for time-varying harmonic frequency trackingabstractA tracking method is presented for tracking time-varying harmonic frequencies. The approach is taken in two steps: based on the observation (x(t)), a raw frequency estimate y(t) is first computed using an adaptive Pisarenko harmonic decomposition (PHD) method. In practice, results are quite noisy due to a short data record and frequency changes. The results are then fed into an event-detection tracking unit to extract the trajectory information about the underlying frequencies. The frequency estimates are fed back to the adaptive PHD algorithm to produce successive estimates of frequencies. The method uses an artificial intelligence planning strategy called least commitment, for which the idea is to hold off decision-making further in time.> Binh C. Phan, Yu Hen Hu |
ICASSP | 2 |
| 1988 | Parallel LU factorization for circuit simulation on an MIMD computerabstractDirect method circuit simulation on an MIMD (multiple-instruction, multiple-data-stream) machine is studied. The focus is on the parallel LU (lower-upper) factorization a sparse matrix with a nested bordered-block diagonal (BBD) ordering. A novel computation model for the parallel factorization is proposed, and simulation results conducted on a ten-processor Sequent Balance 21000 parallel computer are reported. It is concluded that nested BBD ordering proves to be a highly concurrent structure for parallel LU factorization in direct method circuit simulation.> Chien-Chih Chen, Yu Hen Hu |
ICCD | 2 |
| 1987 | Function Search from Behavioral Description of a Digital SystemabstractWe present a novel approach for automating the functional design of digital systems. Given a set of behavioral specifications, the objective is to produce an optimal functional design which minimizes certain design criteria. One distinct feature of this approach is adding the step of function minimization. That is, the abstraction of the primitive operations into a set of functions that generates the desired behavior attempts to minimize the cost of that set according to the design criteria. For this purpose, it is important to have a powerful search strategy which will lead to a near-optimal solution in a reasonable time. We have adopted best-first search (A* algorithm) as the general framework, and developed several domain-specific heuristic functions (h') which control the search process. Preliminary experimental results are reported. Jung-Gen Wu, William P.-C. Ho, Yu Hen Hu, David Y. Y. Yun, H. J. Yu |
DAC | 3 |
| 1987 | Parallel VLSI computing array implementation for signal subspace updating algorithmabstractThis paper concerns the parallel VLSI computing array implementaion for a novel signal subspace iteration algorithm (SSIA) proposed by Karasalo. Specifically, by making use of a sparse structure, a Linearly Connected VLSI computing structure is developed for the Singular Value Decomposition (SVD) operation employed in this algorithm. We first show that by making use of a sparse structue matrix the computing time of this algorithm can be reduced from O(N3) to O(N2) with single processors. Then we show that the parallel architecture is able to reduce the overall computing time for SVD from O(N2) to O(N) using O(N) processors. Where N is the dimension of the signal subspace. This makes the total computing time of SSIA from max(O(K2) O(N2K)) with single processors to O(K) with O(N2) processors. Ali H. Abdallah, Yu Hen Hu |
ICASSP | 2 |
| 1987 | A principal component approach for adaptive ARMA model identificationabstractThis paper presents a numerically stable method for adaptive ARMA model parameter identification. Our approach derives the data adaptive formulation of a novel principal component based system identification method proposed by K.S. Arun. It is shown that the new method exhibits significant performance gain. Kambiz Heidarian, Yu Hen Hu |
ICASSP | 2 |
| 1987 | Knowledge-based adaptive signal processingabstractIn this paper, the use of Artificial Intelligence (A.I.) techniques for improving the performance of adaptive signal processing algorithms is considered. Three potential application areas are identified: (1) The selection and adaptation of secondary control parameters and underlying models, (2) intelligent search methods for adaptation algorithms, and (3) the integration of information from various sources. Various A.I. methods will be proposed for each of these problem areas. Yu Hen Hu, Ali Hussein Abdallah |
ICASSP | 1 |
| 1987 | Subspace approximation based algorithms for adaptive high resolution spectrum estimateabstractIn this paper, subspace approximation based algorithms are developed for adaptive high resolution spectrum estimation. Our approach is to adopt adaptive eigen-subspace computation algorithms into subspace approximation methods. Three subspace approximation methods are considered in this paper. They are the Multiple Signal Classification Method (MUSIC), Toeplitz Approximation Method (TAM) and Noise Subspace Approximation Method (NOSSAM). Given an eigen-subspace of a Hermitian covariance matrix, our goal is to update the eigen-subspace estimate when the original covariance matrix is undergone a rank one update. To facilitate real time computation, it is desired to avoid the eigen decomposition on the newly updated covariance matrix. Three algorithms, namely, the Adaptive Block Power method (ABPM), the Adaptive Subspace Iteration method (ASI), and the Adaptive Block Gradient Subspace Iteration method (BGSI) are derived. Among these three algorithms, the adaptive BGSI method stands out due to its superb performance. Sample simulation results will be reported to illustrate the methods presented in this paper. Yu Hen Hu, Pin-Kuan Chou, Ali Hussein Abdallah |
ICASSP | 1 |
| 1986 | Effective adaptive Pisarenko spectrum estimateabstractIn this paper, we present a new method for real time computation of the Pisarenko's spectrum estimates [1]. This method makes use of subspace iteration technique and Rayleigh-Ritz procedure to find the extreme eigenvector of a symmetric matrix. When apply to Pisarenko's high resolution spectrum estimation problem, the proposed method can update the minimum eigenvector adaptively and make real time processing possible. Yu Hen Hu, Pin-Kuan Chou |
ICASSP | 1 |
| 1986 | VLSI Implementation of real-time Kalman filterabstractIn this paper, the problem of parallel implementation of the square-root Kalman filters is addressed. In the system level, our approach is to apply systolic type processor arrays as basic building blocks to speed up the matrix operations required in each iteration. Specifically, by utilizing a sparse matrix structure, we derive a simple systolic array configuration which is able to solve a rotation operation very efficiently. To maximize the parallelism, we also exploit an inter-array pipelining scheme through the overlapping of execution between successive processor arrays. As a result, several modules can be tightly coupled to form a dedicate Kalman Filter processor for real time applications. We estimate that with O(n2) processors, it would take O(4n+3r-3) time units to complete one Kalman filter iteration, where n is number of states and r is number of inputs. Tze-Yun Sung, Yu Hen Hu |
ICASSP | 2 |
| 1986 | Doubly pipelined Cordic array for digital signal processing algorithmsabstractIn this paper, we present a doubly pipelined VLSI Cordic array processor for digital signal processing computations. The basic notion of doubly pipelined CORDIC computation will be introduced first. Then, some potential applications to digital signal processing problems will be discussed. Specifically, we shall demonstrate how a doubly pipelined CORDIC processor array can be applied to compute discrete Fourier transform and Fast Fourier transform, to implement Lattice filters, to solve Toeplitz systems as well as matrix QR factorizations. It is shown that by adopting a secondary pipelining, about one third hardware can be saved, and sometimes the throughput of the entire CORDIC processor array may be doubled. Tze-Yun Sung, Yu Hen Hu, H. J. Yu |
ICASSP | 2 |
| 1986 | A novel implementation of pipelined Toeplitz system solverabstractThis letter describes a novel implementation of the VLSI Toeplitz linear system solver using a lattice filter structure. By representing the triangular factorization with reflection coefficients, only O(N) storage elements are required. This compares favorably with the O(N/sup 2/) storage elements required in the previous implementation. With a linear array of O(N) processors the total computing time will be O(N) units with pipelined computations. I-Chang Jou, Yu Hen Hu, W. S. Feng |
Proc. IEEE | 2 |
| 1985 | Adaptive methods for real time Pisarenko spectrum estimateabstractIn this paper, we discuss new methods for real time computation of the Pisarenko's spectrum estimates [1]. In particular, we shall focus on the data adaptive computation of minimum eigenvector of a covariance matrix and propose several novel methods for doing so. Yu Hen Hu |
ICASSP | 1 |
| 1985 | VLSI Architecture for solving covariance eigen systemabstractIn this paper, we shall first develop parallel architectures for covariance matrix operations. Then we shall discuss how to make use of them to develop parallel covariance eigen system solvers for real time Pisarenko spectrum estimation. Yu Hen Hu |
ICASSP | 1 |
| 1984 | Constrained lattice structures for harmonic retrievalabstractThis paper presents two novel Lattice structures for retrieving single sinusoidal signal from noisy data samples. Our approach is to make use of a two-stage cascaded lattice filter with the second reflection coefficient being set equal to unity. The first reflection coefficient, from which the sinusoidal frequency is estimated, is obtained by minimizing the sum of the "forward" and "backward" prediction error of the output in a procedure similar to that of the Burg's method. At noiseless case, such a structure is able to compute the exact sinusoidal frequency regardless of the initial phase or data record length. With the presence of white noise, however, such an approach will yield biased frequency estimate. For this, we propose a further modification by including a normalization factor at the output, then minimize the resulting forward and backward prediction errors. It is shown that with ideal white noise, this second lattice structure will give unbiased frequency estimate. Yu Hen Hu, Yao-Cheng Ling |
ICASSP | 1 |
| 1983 | Highly concurrent Toeplitz eigen-system solver for high resolution spectral estimationabstractIn this paper, we develop a highly concurrent Toeplitz Eigen-System Solver (TES) for computing the minimum eigenvalue and associated eigenvector of a N by N symmetric Toeplitz matrix. Conventionally, solving an eigen-system will require O(N3) times with sequential machine; or O(N2) time with O(N) processing units. By exploring the Toeplitz structure and adopting a Rayleigh quotient iteration, the TES can solve for the desired minimum eigenvalue in O(KN) time with N processors and K iterations. The development of TES offers a fast algorithm to implement the Pisarenko's high resolution spectral estimation technique. Yu Hen Hu, Sun-Yuan Kung |
ICASSP | 1 |