Yen-Kuang Chen

dblp:44/4065 · DBLP profile ↗
← Back
64ranked-venue papers
11as first author
7since 2021 · last 2026
0000-0003-4546-9497ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 34 · 7 first-author · 4 since 2021Systems, architecture and hardware · 20 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 8 · 3 since 2021Databases, data management, data science and information retrieval · 6Software engineering, systems software and programming languages · 3Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorComputer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Efficient and distributed learning · 78% Image recognition and object detection · 11% Deep learning architectures and training · 8%
Computer architecture, parallel and distributed computing, and storage systems
12 papers
Hardware accelerators and domain-specific architectures · 41% Memory systems · 15% Parallel and multicore computing · 13%
Computer networks
1 paper
Edge and fog computing · 100%
Theoretical computer science
1 paper
Mathematical optimization · 100%
Databases, data mining, and information retrieval
3 papers
Data mining · 100%

Topics — the 30 heaviest of 43, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
2.042022
Effective Model Sparsification by Scheduled Grow-and-Prune Methods · ICLR 2022
CHEX: CHannel EXploration for CNN Model Compression · CVPR 2022
INVITED: Computation on Sparse Neural Networks and its Implications for Future Hardware · DAC 2020
Computer vision › Image recognition and object detection
image classification
0.622022
Learning in the Frequency Domain · CVPR 2020
CHEX: CHannel EXploration for CNN Model Compression · CVPR 2022
Machine learning › Efficient and distributed learning › model compression › pruning › structured pruning
channel pruning
0.612022
CHEX: CHannel EXploration for CNN Model Compression · CVPR 2022
Machine learning › Efficient and distributed learning › model compression › sparsity
model sparsification
0.612022
Effective Model Sparsification by Scheduled Grow-and-Prune Methods · ICLR 2022
Machine learning › Deep learning architectures and training
frequency-domain learning
0.412020
Learning in the Frequency Domain · CVPR 2020
Machine learning › Efficient and distributed learning › model compression
pruning
0.412020
DARB: A Density-Adaptive Regular-Block Pruning for Deep Neural Networks · AAAI 2020
Machine learning › Efficient and distributed learning › model compression
sparse neural network
0.412020
INVITED: Computation on Sparse Neural Networks and its Implications for Future Hardware · DAC 2020
Machine learning › Efficient and distributed learning › model compression › pruning
structured pruning
0.412020
DARB: A Density-Adaptive Regular-Block Pruning for Deep Neural Networks · AAAI 2020
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.412020
INVITED: Computation on Sparse Neural Networks and its Implications for Future Hardware · DAC 2020
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
sparse neural network acceleration
0.412020
INVITED: Computation on Sparse Neural Networks and its Implications for Future Hardware · DAC 2020
Edge and fog computing
task scheduling
0.312018
Scheduling in Visual Fog Computing: NP-Completeness and Practical Efficient Solutions · AAAI 2018
Cloud and datacenter computing › resource management
resource management and scheduling
0.312018
Scheduling in Visual Fog Computing: NP-Completeness and Practical Efficient Solutions · AAAI 2018
Mathematical optimization › combinatorial optimization
scheduling complexity
0.312018
Scheduling in Visual Fog Computing: NP-Completeness and Practical Efficient Solutions · AAAI 2018
Computer vision › Segmentation and scene understanding
instance segmentation
0.112020
Learning in the Frequency Domain · CVPR 2020
Hardware accelerators and domain-specific architectures
efficient inference
0.112020
DARB: A Density-Adaptive Regular-Block Pruning for Deep Neural Networks · AAAI 2020
Data mining › pattern mining
frequent pattern mining
0.122007
Cache-conscious frequent pattern mining on modern and emerging processors · VLDB J. 2007
Cache-conscious Frequent Pattern Mining on a Modern Processor · VLDB 2005
Performance modeling and evaluation
analytical modeling
0.112011
Moguls: a model to explore the memory hierarchy for bandwidth improvements · ISCA 2011
Memory systems › memory bandwidth
memory bandwidth modeling
0.112011
Moguls: a model to explore the memory hierarchy for bandwidth improvements · ISCA 2011
Memory systems › memory hierarchy
memory hierarchy design
0.112011
Moguls: a model to explore the memory hierarchy for bandwidth improvements · ISCA 2011
Memory systems › memory access
atomic memory operation
0.112008
Atomic Vector Operations on Chip Multiprocessors · ISCA 2008
Parallel and multicore computing › parallel algorithms › sorting
parallel sorting
0.112008
Efficient implementation of sorting on multi-core SIMD CPU architecture · Proc. VLDB Endow. 2008
Processor architecture and microarchitecture
SIMD
0.112008
Efficient implementation of sorting on multi-core SIMD CPU architecture · Proc. VLDB Endow. 2008
Parallel and multicore computing › data parallelism
SIMD vectorization
0.112008
Atomic Vector Operations on Chip Multiprocessors · ISCA 2008
Data mining
pattern mining
0.112007
Cache-conscious frequent pattern mining on modern and emerging processors · VLDB J. 2007
Computer animation and physical simulation
parallel simulation
0.112007
Physical simulation for animation and visual effects: parallelization and characterization for chip multiprocessors · ISCA 2007
Parallel and multicore computing
data parallelism
0.112007
ALP: Efficient support for all levels of parallelism for complex media applications · ACM Trans. Archit. Code Optim. 2007
Processor architecture and microarchitecture › chip multiprocessor
multithreaded chip multiprocessors
0.112007
ALP: Efficient support for all levels of parallelism for complex media applications · ACM Trans. Archit. Code Optim. 2007
Processor architecture and microarchitecture
superscalar processor
0.112007
ALP: Efficient support for all levels of parallelism for complex media applications · ACM Trans. Archit. Code Optim. 2007
Machine learning › Optimization for machine learning › regularized risk minimization
support vector machine training
0.112006
Incremental approximate matrix factorization for speeding up support vector machines · KDD 2006
Data mining › pattern mining › graph pattern mining
frequent subgraph mining
0.112006
Adaptive Parallel Graph Mining for CMP Architectures · ICDM 2006

Methods — techniques the papers use, named apart from their topics

optimization · 1.0NP-completeness proof · 1.0sparse training · 0.9density-adaptive regular-block pruning · 0.9block-max weight masking · 0.9scheduling · 0.6grow-and-prune · 0.6column subset selection · 0.6channel exploration · 0.6spectral analysis · 0.4frequency selection · 0.4discrete cosine transform · 0.4design space exploration · 0.1analytical modeling · 0.1cycle-accurate simulation · 0.1architectural support for scatter-gather · 0.1SIMD extension · 0.1parallelization · 0.1
YearPublicationVenuePosition
2026 A Flexible Zero-Shot Approach to Tone Mapping via Structure-Preserving Diffusion Models
abstract
With the prevalence of high dynamic range (HDR) imaging, tone mapping techniques, which convert HDR images to high-quality standard dynamic range (SDR) images for display, have become increasingly important. However, obtaining paired HDR and high-quality SDR images is almost impossible, posing challenges to learning-based tone mapping methods. To address this issue, we propose a zero-shot tone mapping framework without requiring any HDR training samples. Our approach decomposes images into two components: structural information and tonal information. A diffusion-based mapping model taking the structural information as input is first trained in the high-quality SDR domain, then transferred to the HDR domain that has less readily available training data for inference, leveraging the equivalent distribution of the structural information across both domains. To preserve the original image’s structure, we modify the reverse sampling process and explicitly incorporate the original structural information into the intermediate results. To improve the image details, we introduce a dual-control network, enabling different conditional inputs to control different scales of the output. Additionally, we devise a flexible tone adjustment strategy, with a bunch of novel loss functions to modify the trained score function dynamically during reverse sampling, allowing users to customize the style of the generated image according to their preference during testing. Initially designed for tone mapping, our model can be applied to various tasks including image fusion, exposure correction, dehazing, etc., without retraining. Experimental results demonstrate that our approach surpasses previous state-of-the-art methods, indicating that it can serve as an effective, flexible and versatile solution to various tone-mapping tasks. Source code is available at https://github.com/ZSDM-HDR/Zero-Shot-Diffusion-HDR.
Ruoxi Zhu, Shusong Xu, Peiye Liu, Yanheng Lu, Dimin Niu, Hongzhong Zheng, Yen-Kuang Chen, Ming-e Jing, Yibo Fan
IEEE Trans. Circuits Syst. Video Technol.8
2025 A Tightly Coupled AI-ISP Vision Processor
abstract
To achieve high-quality and high-resolution image processing, this work presents a novel vision processor that facilitates deep learning-enhanced image processing pipelines. At the system level, by identifying that a divide-and-conquer approach is essential to synergize both classical image processing and image enhancement networks, we develop a tightly coupled system with strip-tile conversion dataflow to enable fine-grained low-latency data interactions between image signal processors (ISPs) and the deep learning accelerator (DLA). At the architecture level, we design a comprehensive set of 21 efficient image processing modules to construct classical ISP pipelines, a tile-based strip layer fusion DLA specifically optimized for networks, and a programmable pixel pool that seamlessly supports the data access patterns of the ISP and the DLA. At the software and hardware co-design level, we propose a comprehensive optimization framework to address the implementation overhead of networks while maintaining the image quality. Finally, evaluations of the AI-ISP vision processor demonstrate 53.95% external memory access reduction and 35.51% latency reduction, delivering superior image quality with minimal on-chip memory overhead. A throughput of up to 168.5 frames per second facilitates efficient processing of ultra-high definition (UHD) resolution images.
Hao Zhang 0126, Sicheng Li 0001, Yupeng Gui, Zhiyong Li 0016, Shusong Xu, Yanheng Lu, Dimin Niu, Hongzhong Zheng, Yen-Kuang Chen, Yuan Xie 0001, Yibo Fan
IEEE Trans. Circuits Syst. Video Technol.9
2022 CHEX: CHannel EXploration for CNN Model Compression
abstract
Channel pruning has been broadly recognized as an effective technique to reduce the computation and memory cost of deep convolutional neural networks. However, conventional pruning methods have limitations in that: they are restricted to pruning process only, and they require a fully pre-trained large model. Such limitations may lead to sub-optimal model quality as well as excessive memory and training cost. In this paper, we propose a novel Channel Exploration methodology, dubbed as CHEX, to rectify these problems. As opposed to pruning-only strategy, we propose to repeatedly prune and regrow the channels throughout the training process, which reduces the risk of pruning important channels prematurely. More exactly: From intra-Layer's aspect, we tackle the channel pruning problem via a well-known column subset selection (CSS) formulation. From inter-Layer's aspect, our regrowing stages open a path for dynamically re-allocating the number of channels across all the layers under a global channel sparsity constraint. In addition, all the exploration process is done in a single training from scratch without the need of a pre-trained large model. Experimental results demonstrate that CHEX can effectively reduce the FLOPs of diverse CNN architectures on a variety of computer vision tasks, including image classification, object detection, instance segmentation, and 3D vision. For example, our compressed ResNet-50 model on ImageNet dataset achieves 76% top-l accuracy with only 25% FLOPs of the original ResNet-50 model, outperforming previous state-of-the-art channel pruning methods. The checkpoints and code are available at here.
Zejiang Hou, Minghai Qin, Fei Sun 0002, Kun Yuan 0001, Yi Xu 0008, Yen-Kuang Chen, Rong Jin 0001, Yuan Xie 0001, Sun-Yuan Kung
CVPR7
2022 2022 ICCAD CAD Contest Problem C: Microarchitecture Design Space Exploration
abstract
It is vital to select microarchitectures to achieve good trade-offs between performance, power, and area in the chip development cycle. Combining high-level hardware description languages and optimization of electronic design automation tools empowers microarchitecture exploration at the circuit level. Due to the extremely large design space and high runtime cost to evaluate a microarchitecture, ICCAD 2022 CAD Contest Problem C calls for an effective design space exploration algorithm to solve the problem. We formulate the research topic as a contest problem and provide benchmark suites, contest benchmark platforms, etc., for all contestants to innovate and estimate their algorithms.
Sicheng Li 0001, Xuechao Wei, Bizhao Shi, Yen-Kuang Chen, Yuan Xie 0001
ICCAD5
2022 Effective Model Sparsification by Scheduled Grow-and-Prune Methods
Minghai Qin, Fei Sun 0002, Zejiang Hou, Kun Yuan 0001, Yi Xu 0008, Yanzhi Wang 0001, Yen-Kuang Chen, Rong Jin 0001, Yuan Xie 0001
ICLR8
2022 Compact Multi-level Sparse Neural Networks with Input Independent Dynamic Rerouting
abstract
Deep neural networks (DNNs) have shown to provide superb performance in many real life applications, but their large computation cost and storage requirement have prevented them from being deployed to many edge and internet-of-things (IoT) devices. Sparse deep neural networks, whose majority weight parameters are zeros, can substantially reduce the computation complexity and memory consumption of the models. In real-use scenarios, devices may suffer from large fluctuations of the available computation and memory resources under different environment, and the quality of service (QoS) is difficult to maintain due to the long tail inferences with large latency. Facing the real-life challenges, we propose to train a sparse model that supports multiple sparse levels. That is, a hierarchical structure of weights are satisfied such that the locations and the values of the non-zero parameters of the more-sparse sub-model are a subset of the less-sparse sub-model. In this way, one can dynamically select the appropriate sparsity level during inference, while the storage cost is capped by the least sparse sub-model. We have verified our methodologies on a variety of DNN models and tasks, including the ResNet-50, PointNet++, GNMT, and graph attention networks. We obtain sparse sub-models with an average of 13.38% weights and 14.97% FLOPs, while the accuracies are as good as their dense counterparts. More-sparse sub-models with 5.38% weights and 4.47% of FLOPs, which are subsets of the less-sparse ones, can be obtained with only 3.25% relative accuracy loss. In addition, our proposed hierarchical model structure supports the mechanism to inference the first part of the model with less sparsity, and dynamically reroute to the more-sparse level if the real-time latency constraint is estimated to be violated. Preliminary analysis shows that we can improve the QoS by one or two nines depending on the task and the computation-memory resources of the inference engine.
Minghai Qin, Tianyun Zhang, Fei Sun 0002, Yen-Kuang Chen, Makan Fardad, Yanzhi Wang 0001, Yuan Xie 0001
ICTAI4
2022 Learning from the CNN-based Compressed Domain
abstract
Images are transmitted or stored in their compressed form and most of the AI tasks are performed from the re-constructed domain. Convolutional neural network (CNN)-based image compression and reconstruction is growing rapidly and it achieves or surpasses the state-of-the-art heuristic image compression methods, such as JPEG or BPG. A major limitation of the application of the CNN-based image compression is on the computation complexity during compression and reconstruction. Therefore, learning from the compressed domain is desirable to avoid the computation and latency caused by reconstruction. In this paper, we show that learning from the compressed domain can achieve comparative or even better accuracy than from the reconstructed domain. At a high compression rate of 0.098 bpp, for example, the proposed compression-learning system has over 3% absolute accuracy boost over the traditional compression-reconstruction-learning flow. The improvement is achieved by optimizing the compression-learning system targeting original-sized instead of standardized (e.g., 224x224) images, which is crucial in practice since real-world images into the system have different sizes. We also propose an efficient model-free entropy estimation method and a criterion to learn from a selected subset of features in the compressed domain to further re-duce the transmission and computation cost without accuracy degradation.
Minghai Qin, Yen-Kuang Chen
WACV3
2020 DARB: A Density-Adaptive Regular-Block Pruning for Deep Neural Networks
abstract
The rapidly growing parameter volume of deep neural networks (DNNs) hinders the artificial intelligence applications on resource constrained devices, such as mobile and wearable devices. Neural network pruning, as one of the mainstream model compression techniques, is under extensive study to reduce the model size and thus the amount of computation. And thereby, the state-of-the-art DNNs are able to be deployed on those devices with high runtime energy efficiency. In contrast to irregular pruning that incurs high index storage and decoding overhead, structured pruning techniques have been proposed as the promising solutions. However, prior studies on structured pruning tackle the problem mainly from the perspective of facilitating hardware implementation, without diving into the deep to analyze the characteristics of sparse neural networks. The neglect on the study of sparse neural networks causes inefficient trade-off between regularity and pruning ratio. Consequently, the potential of structurally pruning neural networks is not sufficiently mined.In this work, we examine the structural characteristics of the irregularly pruned weight matrices, such as the diverse redundancy of different rows, the sensitivity of different rows to pruning, and the position characteristics of retained weights. By leveraging the gained insights as a guidance, we first propose the novel block-max weight masking (BMWM) method, which can effectively retain the salient weights while imposing high regularity to the weight matrix. As a further optimization, we propose a density-adaptive regular-block (DARB) pruning that can effectively take advantage of the intrinsic characteristics of neural networks, and thereby outperform prior structured pruning work with high pruning ratio and decoding efficiency. Our experimental results show that DARB can achieve 13× to 25× pruning ratio, which are 2.8× to 4.3× improvements than the state-of-the-art counterparts on multiple neural network models and tasks. Moreover, DARB can achieve 14.3× decoding efficiency than block pruning with higher pruning ratio.
Ao Ren, Tao Zhang 0032, Yuhao Wang 0002, Sheng Lin 0001, Peiyan Dong, Yen-Kuang Chen, Yuan Xie 0001, Yanzhi Wang 0001
AAAI6
2020 Learning in the Frequency Domain
abstract
Deep neural networks have achieved remarkable success in computer vision tasks. Existing neural networks mainly operate in the spatial domain with fixed input sizes. For practical applications, images are usually large and have to be downsampled to the predetermined input size of neural networks. Even though the downsampling operations reduce computation and the required communication bandwidth, it removes both redundant and salient information obliviously, which results in accuracy degradation. Inspired by digital signal processing theories, we analyze the spectral bias from the frequency perspective and propose a learning-based frequency selection method to identify the trivial frequency components which can be removed without accuracy loss. The proposed method of learning in the frequency domain leverages identical structures of the well-known neural networks, such as ResNet-50, MobileNetV2, and Mask R-CNN, while accepting the frequency-domain information as the input. Experiment results show that learning in the frequency domain with static channel selection can achieve higher accuracy than the conventional spatial downsampling approach and meanwhile further reduce the input data size. Specifically for ImageNet classification with the same input size, the proposed method achieves 1.60% and 0.63% top-1 accuracy improvements on ResNet-50 and MobileNetV2, respectively. Even with half input size, the proposed method still improves the top-1 accuracy on ResNet-50 by 1.42%. In addition, we observe a 0.8% average precision improvement on Mask R-CNN for instance segmentation on the COCO dataset.
Kai Xu 0007, Minghai Qin, Fei Sun 0002, Yuhao Wang 0002, Yen-Kuang Chen, Fengbo Ren
CVPR5
2020 INVITED: Computation on Sparse Neural Networks and its Implications for Future Hardware
abstract
Neural network models are widely used in solving many challenging problems, such as computer vision, personalized recommendation, and natural language processing. Those models are very computationally intensive and reach the hardware limit of the existing server and IoT devices. Thus, finding better model architectures with much less amount of computation while maximally preserving the accuracy is a popular research topic. Among various mechanisms that aim to reduce the computation complexity, identifying the zero values in the model weights and in the activations to avoid computing them is a promising direction. In this paper, we summarize the current status of the research on the computation of sparse neural networks, from the perspective of the sparse algorithms, the software frameworks, and the hardware accelerations. We observe that the search for the sparse structure can be a general methodology for high-quality model explorations, in addition to a strategy for high-efficiency model execution. We discuss the model accuracy influenced by the number of weight parameters and the structure of the model. The corresponding models are called to be located in the weight dominated and structure dominated regions, respectively. We show that for practically complicated problems, it is more beneficial to search large and sparse models in the weight dominated region. In order to achieve the goal, new approaches are required to search for proper sparse structures, and new sparse training hardware needs to be developed to facilitate fast iterations of sparse models.
Fei Sun 0002, Minghai Qin, Tianyun Zhang, Liu Liu 0017, Yen-Kuang Chen, Yuan Xie 0001
DAC5
2020 Orchestrating Medical Image Compression and Remote Segmentation Networks
Zihao Liu 0015, Sicheng Li 0001, Yen-Kuang Chen, Tao Liu 0023, Qi Liu 0017, Xiaowei Xu 0004, Yiyu Shi 0001, Wujie Wen
MICCAI (4)3
2019 Iotbench: A Benchmark Suite for Intelligent Internet of Things Edge Devices
abstract
IoT devices must and will be more intelligent in the future. However, due to the lack of benchmarks representative to the diverse IoT applications, there are limited architecture performance studies on IoT devices. This paper presents IoTBench, the first benchmark suite targeting at IoT edge-device applications. This suite includes seven representative programs from three major IoT categories. We investigate the computational demand and execution efficiency of these benchmarks running on a popular IoT device platform. We also analyze and characterize the energy consumption of IoTBench using analytic approaches. Overall, IoTBench establishes a foundation for innovative architecture design of IoT edge devices.
Chien-I Lee, Meng-Yao Lin, Chia-Lin Yang, Yen-Kuang Chen
ICIP4
2018 Scheduling in Visual Fog Computing: NP-Completeness and Practical Efficient Solutions
abstract
The visual fog paradigm envisions tens of thousands of heterogeneous, camera-enabled edge devices distributed across the Internet, providing live sensing for a myriad of different visual processing applications. The scale, computational demands, and bandwidth needed for visual computing pipelines necessitates offloading intelligently to distributed computing infrastructure, including the cloud, Internet gateway devices, and the edge devices themselves. This paper focuses on the visual fog scheduling problem of assigning the visual computing tasks to various devices to optimize network utilization. We first prove this problem is NP-complete, and then formulate a practical, efficient solution. We demonstrate sub-minute computation time to optimally schedule 20,000 tasks across over 7,000 devices, and just 7-minute execution time to place 60,000 tasks across 20,000 devices, showing our approach is ready to meet the scale challenges introduced by visual fog.
Hong-Min Chu, Shao-Wen Yang, Padmanabhan Pillai, Yen-Kuang Chen
AAAI4
2017 Distributed video codec with spatiotemporal side information
abstract
In this paper, a distributed video coding (DVC) system with spatiotemporal side information is proposed. The proposed framework addresses the problem of poor compression performance of DVC for high-motion video sequences by integrating temporal and spatial prediction schemes in one framework. Super-resolution techniques are employed for spatially-predicted side information generation, and a support vector machine is trained to adaptively select the coding structure with both spatial and temporal prediction. In addition, an encoder-driven coding mode selection at different granularities, including frame, block and coefficient levels, is adopted to further improve the coding performance for various video conditions. Experimental results show that the average BD rate reduction of the proposed framework is 12.93% compared with the DISCOVER DVC, and the coding gain is significant, especially for high-motion sequences. Moreover, the average computing complexity is only 92.26% of the DISCOVER DVC.
Yueh-Ying Lee, Pin-Hung Kuo, Chia-han Lee, Yen-Kuang Chen, Shao-Yi Chien
ISCAS4
2017 A framework for visual fog computing
abstract
Visual data are rich, which have opened vast analytics opportunities and been widely used in many applications. However, the demanding requirements of computational resources and bandwidth have prevented the data from being useful in an economically efficient manner. A visual fog paradigm is needed for efficient processing of continuous video streams by collaboratively using things in the Internet of Video Things (IoVT), comprising edge devices, intermediate gateways, and servers on premise or in the cloud, as the computing platform. The challenges lying ahead include (1) Reusability-a reusable framework across multiple vertical applications, (2) Efficiency-the intelligence for online distributing and redistributing work-load for optimal system performance, and (3) Configurability-the user interface for (layperson) users to easily analyze the visual data as well as the corresponding metadata. This paper spells out the need of a framework for visual fog computing and suggest promising research directions towards instantiations of a visual fog computing framework.
Shao-Wen Yang, Omesh Tickoo, Yen-Kuang Chen
ISCAS3
2015 Distributed computing in IoT: System-on-a-chip for smart cameras as an example
abstract
There are four major components in application systems with internet-of-things (IoT): sensors, communications, computation and service, where large amount of data are acquired for ultra-big data analysis to discover the context information and knowledge behind signals. To support such large-scale data size and computation tasks, it is not feasible to employ centralized solutions on cloud servers. Thanks for the advances of silicon technology, the cost of computation become lower, and it is possible to distribute computation on every node in IoT. In this paper, we take video sensing network as an example to show the idea of distributed computing in IoT. Existing related works are reviewed and the architecture of a system-on-a-chip solution for distributed smart cameras is proposed with coarse-grained reconfigurable image stream processing architecture. It can accelerate various computer vision algorithms for distributed smart cameras in IoT.
Shao-Yi Chien, Wei-Kai Chan, Yu-Hsiang Tseng, Chia-han Lee, V. Srinivasa Somayazulu, Yen-Kuang Chen
ASP-DAC6
2014 Low complexity on-line video summarization with Gaussian mixture model based clustering
abstract
Techniques of video summarization have attracted significant research interests in the past decade due to the rapid progress in video recording, computation, and communication technologies. However, most of the existing methods analyze the video in an off-line manner, which greatly reduces the flexibility of the system. On-line summarization, which can progressively process video during video recording, is then proposed for a wide range of applications. In this paper, an on-line summarization method using Gaussian mixture model is proposed. As shown in the experiments, the proposed method outperforms other on-line methods in both summarization quality and computational efficiency. It can generate summarization with a shorter latency and much lower computation resource requirements.
Shun-Hsing Ou, Chia-han Lee, V. Srinivasa Somayazulu, Yen-Kuang Chen, Shao-Yi Chien
ICASSP4
2014 Error resilience for key frames in distributed video coding with rate-distortion optimized mode decision
abstract
Distributed video coding (DVC) is a potential solution for distributed video sensors in wireless visual sensor and machine-to-machine (M2M) networks. However, the error resilience schemes have not been fully investigated in literatures. In this paper, we propose several error resilience schemes for key frame transmission of DVC, including resending, refining, and hybrid modes. The resending mode asks the transmitters to resend the lost packets as automatic repeat-request, the refining mode conceals the corrupted key frames by temporal error concealment techniques with considering the characteristics of DVC, and the hybrid mode can adaptively select the better mode packet-by-packet with the proposed rate-distortion optimized mode decision. Experimental results show that the proposed schemes can achieve around 6-dB gain in PSNR compared with a baseline approach with intra error concealment. Furthermore, the hybrid mode can achieve 0.5-dB gain in PSNR for some sequences compared with the better one among the resending and refining modes.
Hsin-Fang Wu, Chia-han Lee, V. Srinivasa Somayazulu, Yen-Kuang Chen, Shao-Yi Chien
ISCAS4
2014 Communication-efficient multi-view keyframe extraction in distributed video sensors
abstract
Video sensors are widely used in many applications such as security monitoring and home care. However, the growth of the number of sensors makes it impractical to stream all videos back to a central server for further processing, due to communication bandwidth and server storage constraints. Multi-view video summarization allows us to discard redundant data in the video streams taken by a group of sensors. All prior multi-view summarization methods, however, process video data in an off-line and centralized manner, which means that all videos are still required to be streamed back to the server before conducting the summarization. This paper proposes an on-line, distributed multi-view summarization system, which integrates the ideas of Maximal Marginal Relevance (MMR) and MS-Wave, a bandwidth-efficient distributed algorithm for finding k-nearest-neighbors and k-farthest-neighbors. Empirical studies show that our proposed system can discard redundant videos and keep important keyframes as effectively as centralized approaches, while transmitting only 1/6 to 1/3 as much data.
Shun-Hsing Ou, Yu-Chen Lu, Jui-Pin Wang, Shao-Yi Chien, Shou-De Lin, Mi-Yen Yeh, Chia-han Lee, Phillip B. Gibbons, V. Srinivasa Somayazulu, Yen-Kuang Chen
VCIP10
2013 Locality-aware task management for unstructured parallelism: a quantitative limit study
abstract
As we increase the number of cores on a processor die, the on-chip cache hierarchies that support these cores are getting larger, deeper, and more complex. As a result, non-uniform memory access effects are now prevalent even on a single chip. To reduce execution time and energy consumption, data access locality should be exploited. This is especially important for task-based programming systems, where a scheduler decides when and where on the chip the code segments, i.e., tasks, should execute. Capturing locality for structured task parallelism has been done effectively, but the more difficult case, unstructured parallelism, remains largely unsolved - little quantitative analysis exists to demonstrate the potential of locality-aware scheduling, and to guide future scheduler implementations in the most fruitful direction.
Richard M. Yoo, Christopher J. Hughes, Changkyu Kim, Yen-Kuang Chen, Christoforos E. Kozyrakis
SPAA4
2012 Challenges and opportunities of internet of things
abstract
To date, most Internet applications focus on providing information, interaction, and entertainment for humans. However, with the widespread deployment of networked, intelligent sensor technologies, an Internet of Things (IoT) is steadily evolving, much like the Internet decades ago. In the future, hundreds of billions of smart sensors and devices will interact with one another without human intervention, on a Machine-to-Machine (M2M) basis. They will generate an enormous amount of data at an unprecedented scale and resolution, providing humans with information and control of events and objects even in remote physical environments. The scale of the M2M Internet will be several orders of magnitude larger than the existing Internet, posing serious research challenges. This paper will provide an overview of challenges and opportunities presented by this new paradigm.
Yen-Kuang Chen
ASP-DAC1
2012 Power optimization of wireless video sensor nodes in M2M networks
abstract
Low-power wireless video sensor nodes play important roles for applications in machine-to-machine (M2M) network. Several design issues to optimize the power consumption of a video sensor node are addressed in this paper. For the video coding engine selection, the comparison between conventional video coding system and distributed video coding (DVC) system shows that although the rate-distortion performance of existing DVC codec still has room to improve, it can provide lower power consumption with a noisy transmission channel. Furthermore, it also demonstrated that video analysis unit can help to filter out video contents without event-of-interest to reduce transmission power. Finally, several future research directions are addressed, and the trade-off between the video analysis unit, video coding unit, and data transmission should be further studied to design wireless video sensors with optimized power consumption.
Shao-Yi Chien, Teng-Yuan Cheng, Chieh-Chuan Chiu, Pei-Kuei Tsung, Chia-han Lee, V. Srinivasa Somayazulu, Yen-Kuang Chen
ASP-DAC7
2012 Hybrid distributed video coding with frame level coding mode selection
abstract
Distributed video coding (DVC), a new video coding paradigm based on Slepian-Wolf and Wyner-Ziv theories, is a promising solution for implementing low-power and low-cost distributed wireless video sensors since most of the computation load is moved from the encoder to the decoder. It has been showed that there is still room to improve the coding efficiency of the current DVC codec. In this paper, we propose a hybrid coding structure with frame-level coding mode selection (CMS) to allow the DVC encoder to flexibly choose channel coding or entropy coding to code each band. The experimental results show the significant improvement on the rate-distortion (R-D) performance with only slight increase in the encoding complexity compared to the DISCOVER codec. The proposed DVC system performs comparably to H.264 No Motion with much lower encoding complexity.
Chieh-Chuan Chiu, Shao-Yi Chien, Chia-han Lee, V. Srinivasa Somayazulu, Yen-Kuang Chen
ICIP5
2011 Moguls: a model to explore the memory hierarchy for bandwidth improvements
abstract
In recent years, the increasing number of processor cores and limited increases in main memory bandwidth have led to the problem of the bandwidth wall, where memory bandwidth is becoming a performance bottleneck. This is especially true for emerging latency-insensitive, bandwidth-sensitive applications. Designing the memory hierarchy for a platform with an emphasis on maximizing bandwidth within a fixed power budget becomes one of the key challenges. To facilitate architects to quickly explore the design space of memory hierarchies, we propose an analytical performance model called Moguls. The Moguls model estimates the performance of an application on a system, using the bandwidth demand of the application for a range of cache capacities and the bandwidth provided by the system with those capacities. We show how to extend this model with appropriate approximations to optimize a cache hierarchy under a power constraint. The results show how many levels of cache should be designed, and what the capacity, bandwidth, and technology of each level should be. In addition, we study memory hierarchy design with hybrid memory technologies, which shows the benefits of using multiple technologies for future computing systems.
Guangyu Sun 0003, Christopher J. Hughes, Changkyu Kim, Jishen Zhao, Cong Xu 0002, Yuan Xie 0001, Yen-Kuang Chen
ISCA7
2011 Emerging applications for multi/many-core processors
abstract
There has always been a strong relationship between computer applications and computer architectures. Advances in computer architecture enable new usage models; new usage models challenge new architectures. For many decades, the interplays between applications and architectures have resulted in significant progress in the computer technologies. Recently, the computer industry adopts the multi/many-core architecture for better performance, energy efficiency and reliability. This industry-wide movement towards multi/many-core architectures opens up many opportunities for developing new classes of applications. This paper uses a few emerging applications as examples to illustrate the interplay between applications and architecture. It further discusses the key characteristics of emerging applications and how multi/many-core architecture addresses the needs of the new applications. The paper closes with recommendations for new application developers to allow better utilization of new multi/many-core processors.
Victor W. Lee, Yen-Kuang Chen, Pradeep Debuy
ISCAS2
2011 Distributed video coding: A promising solution for distributed wireless video sensors or not?
abstract
Low-power and low-cost distributed wireless video sensors play important roles for applications in machine-to-machine (M2M) and wireless sensor networks. Distributed video coding (DVC), an emerging coding technology based on Wyner-Ziv theory, seems to be a possible solution for implementing low-power video sensors since most of the computational complexity is moved from the encoder to the decoder. In this paper, existing works on DVC are discussed with rate-distortion and power consumption analyses compared with H.264/AVC-based approaches. We show that, since more transmission power is required for compensating the lower rate-distortion performance, the power consumption of sensor nodes using DVC is just similar to that of using H.264/AVC with zero motion vectors. Therefore, there is still a room for improvement to make DVC applicable for distributed wireless video sensors. Based on our analysis results, several possible research directions, such as studies on the trade- off between hardware cost and system power consumption, are also addressed in this paper under a unified DVC framework.
Chieh-Chuan Chiu, Shao-Yi Chien, Chia-han Lee, V. Srinivasa Somayazulu, Yen-Kuang Chen
VCIP5
2009 Challenges and opportunities of obtaining performance from multi-core CPUs and many-core GPUs
abstract
Multi-core processors represent a major development in computing technology. For example, IntelregCoretrade 2 Quad processors, IBM Cell processors, and Nvidia GeForce 9800 GX2, are widely used. However, most applications struggle to make the best use of the power provided by many-core processors. Easy-to-use software tools are hard to find. Furthermore, it's not clear what changes need to be made to algorithms to fully utilize many-core CPUs or GPUs. In this paper, we try to offer a bird's eye view of the opportunities lying ahead in two folds: (1) software tools and (2) workload analysis. With good software tools and insightful workload analysis, software and algorithm developers can not only harness the power of many computing cores, but also innovate new algorithms that best utilize the many computing cores. New algorithms and applications are thus made possible with the computing power not available before.
Trista Pei-Chun Chen, Yen-Kuang Chen
ICASSP2
2009 Special Issue: Algorithm/Architecture Co-Exploration of Visual Computing on Emerging Platforms
abstract
Provides notice of upcoming special issue(s) of interest to practitioners and researchers.
Yen-Kuang Chen, Gwang-Gook Lee, Marco Mattavelli, Euee S. Jang
IEEE Trans. Circuits Syst. Video Technol.1
2009 Algorithm/Architecture Co-Exploration of Visual Computing on Emergent Platforms: Overview and Future Prospects
abstract
Concurrently exploring both algorithmic and architectural optimizations is a new design paradigm. This survey paper addresses the latest research and future perspectives on the simultaneous development of video coding, processing, and computing algorithms with emerging platforms that have multiple cores and reconfigurable architecture. As the algorithms in forthcoming visual systems become increasingly complex, many applications must have different profiles with different levels of performance. Hence, with expectations that the visual experience in the future will become continuously better, it is critical that advanced platforms provide higher performance, better flexibility, and lower power consumption. To achieve these goals, algorithm and architecture co-design is significant for characterizing the algorithmic complexity used to optimize targeted architecture. This paper shows that seamless weaving of the development of previously autonomous visual computing algorithms and multicore or reconfigurable architectures will unavoidably become the leading trend in the future of video technology.
Gwo Giun Lee, Yen-Kuang Chen, Marco Mattavelli, E. S. Jang
IEEE Trans. Circuits Syst. Video Technol.2
2008 Novel parallel Hough Transform on multi-core processors
abstract
After analyzing the performance bottlenecks of the Hough transform on multi-core processors, this paper proposes a new Hough transform implementation. The performance of microprocessors improves significantly because of the introduction of multiple cores. To harness the computation power of such multi-core processors, we must effectively execute many threads at the same time. This paper first studies a coarse-grain and a fine-grain parallelization of a straightforward Hough transform implementation on an 8-core machine. Due to parallelization overheads and memory requirements, these schemes do not fully utilize computation power. After that, we propose a new Hough transform implementation for parallelization. Experimental data shows that the new Hough transform exposes a significant amount of concurrency and pretty good data locality. On the 8-core machine, the new implementation has 25% better performance than the old ones.
Yen-Kuang Chen, Wenlong Li 0003, Tao Wang 0003
ICASSP1
2008 Atomic Vector Operations on Chip Multiprocessors
abstract
The current trend is for processors to deliver dramatic improvements in parallel performance while only modestly improving serial performance. Parallel performance is harvested through vector/SIMD instructions as well as multithreading (through both multithreaded cores and chip multiprocessors). Vector parallelism can be more efficiently supported than multithreading, but is often harder for software to exploit. In particular, code with sparse data access patterns cannot easily utilize the vector/SIMD instructions of mainstream processors. Hardware to scatter and gather sparse data has previously been proposed to enable vector execution for these codes. However, on multithreaded architectures, a number of applications spend significant time on atomic operations (e.g., parallel reductions), which cannot be vectorized using previously proposed schemes. This paper proposes architectural support for atomic vector operations (referred to as GLSC) that addresses this limitation. GLSC extends scatter-gather hardware to support atomic memory operations. Our experiments show that the GLSC provides an average performance improvement on a set of important RMS kernels of 54% for 4-wide SIMD.
Daehyun Kim 0001, Mikhail Smelyanskiy, Yen-Kuang Chen, Jatin Chhugani, Christopher J. Hughes, Changkyu Kim, Victor W. Lee, Anthony D. Nguyen
ISCA4
2008 Memory hierarchy performance measurement of commercial dual-core desktop processors
Lu Peng 0001, Jih-Kwon Peir, Tribuvan K. Prakash, Carl Staelin, Yen-Kuang Chen, David M. Koppelman
J. Syst. Archit.5
2008 Convergence of Recognition, Mining, and Synthesis Workloads and Its Implications
abstract
This paper examines the growing need for a general-purpose ldquoanalytics enginerdquo that can enable next-generation processing platforms to effectively model events, objects, and concepts based on end-user input, and accessible datasets, along with an ability to iteratively refine the model in real-time. We find such processing needs at the heart of many emerging applications and services. This processing is further decomposed in terms of an integration of three fundamental compute capabilities-recognition, mining, and synthesis (RMS). The set of RMS workloads is examined next in terms of usage, mathematical models, numerical algorithms, and underlying data structures. Our analysis suggests a workload convergence that is analyzed next for its platform implications. In summary, a diverse set of emerging RMS applications from market segments like graphics, gaming, media-mining, unstructured information management, financial analytics, and interactive virtual communities presents a relatively focused, highly overlapping set of common platform challenges. A general-purpose processing platform designed to address these challenges has the potential for significantly enhancing users' experience and programmer productivity.
Yen-Kuang Chen, Jatin Chhugani, Pradeep Dubey, Christopher J. Hughes, Daehyun Kim 0001, Victor W. Lee, Anthony D. Nguyen, Mikhail Smelyanskiy
Proc. IEEE1
2008 Efficient implementation of sorting on multi-core SIMD CPU architecture
abstract
Sorting a list of input numbers is one of the most fundamental problems in the field of computer science in general and high-throughput database applications in particular. Although literature abounds with various flavors of sorting algorithms, different architectures call for customized implementations to achieve faster sorting times. This paper presents an efficient implementation and detailed analysis of MergeSort on current CPU architectures. Our SIMD implementation with 128-bit SSE is 3.3X faster than the scalar version. In addition, our algorithm performs an efficient multiway merge, and is not constrained by the memory bandwidth. Our multi-threaded, SIMD implementation sorts 64 million floating point numbers in less than 0.5 seconds on a commodity 4-core Intel processor. This measured performance compares favorably with all previously published results. Additionally, the paper demonstrates performance scalability of the proposed sorting algorithm with respect to certain salient architectural features of modern chip multiprocessor (CMP) architectures, including SIMD width and core-count. Based on our analytical models of various architectural configurations, we see excellent scalability of our implementation with SIMD width scaling up to 16X wider than current SSE width of 128-bits, and CMP core-count scaling well beyond 32 cores. Cycle-accurate simulation of Intel's upcoming x86 many-core Larrabee architecture confirms scalability of our proposed algorithm.
Jatin Chhugani, Anthony D. Nguyen, Victor W. Lee, William Macy, Mostafa Hagog, Yen-Kuang Chen, Akram Baransi, Pradeep Dubey
Proc. VLDB Endow.6
2007 Computer Vision on Multi-Core Processors: Articulated Body Tracking
abstract
The recent emergence of multi-core processors enables a new trend in the usage of computers. Computer vision applications, which require heavy computation and lots of bandwidth, usually cannot run in real-time. Recent multi-core processors can potentially serve the needs of such workloads. In addition, more advanced algorithms can be developed utilizing the new computation paradigm. In this paper, we study the performance of an articulated body tracker on multi-core processors. The articulated body tracking workload encapsulates most of the important aspects of a computer vision workload. It takes multiple camera inputs of a scene with a single human object, extracts useful features, and performs statistical inference to find the body pose. We show the importance of properly parallelizing the workload in order to achieve great performance: speedups of 26 on 32 cores. We conclude that: (1) data-domain parallelization is better than function-domain parallelization for computer vision applications; (2) data-domain parallelism by image regions and particles is very effective; (3) reducing serial code in edge detection brings significant performance improvements; (4) domain knowledge about low/mid/high level of vision computation is helpful in parallelizing the workload.
Trista Pei-Chun Chen, Dmitry Budnikov, Christopher J. Hughes, Yen-Kuang Chen
ICME4
2007 Memory Performance and Scalability of Intel's and AMD's Dual-Core Processors: A Case Study
abstract
As chip multiprocessor (CMP) has become the mainstream in processor architectures, Intel and AMD have introduced their dual-core processors to the PC market. In this paper, performance studies on an Intel Core 2 Duo, an Intel Pentium D and an AMD Athlon 64times2 processor are reported. According to the design specifications, key derivations exist in the critical memory hierarchy architecture among these dual-core processors. In addition to the overall execution time and throughput measurement using both multiprogrammed and multi-threaded workloads, this paper provides detailed analysis on the memory hierarchy performance and on the performance scalability between single and dual cores. Our results indicate that for the best performance and scalability, it is important to have (1) fast cache-to-cache communication, (2) large L2 or shared capacity, (3) fast L2 to core latency, and (4) fair cache resource sharing. Three dual-core processors that we studied have shown benefits of some of these factors, but not all of them. Core 2 Duo has the best performance for most of the workloads because of its microarchitecture features such as shared L2 cache. Pentium D shows the worst performance in many aspects due to its technology-remap of Pentium 4.
Lu Peng 0001, Jih-Kwon Peir, Tribuvan K. Prakash, Yen-Kuang Chen, David M. Koppelman
IPCCC4
2007 Physical simulation for animation and visual effects: parallelization and characterization for chip multiprocessors
abstract
We explore the emerging application area of physics-based simulation for computer animation and visual special effects. In particular, we examine its parallelization potential and characterize its behavior on a chip multiprocessor (CMP). Applications in this domain model and simulate natural phenomena, and often direct visual components of motion pictures. We study a set of three workloads that exemplify the span and complexity of physical simulation applications used in a production environment: fluid dynamics, facial animation, and cloth simulation. They are computationally demanding, requiring from a few seconds to several minutes to simulate a single frame; therefore, they can benefit greatly from the acceleration possible with large scale CMPs.
Christopher J. Hughes, Radek Grzeszczuk, Eftychios Sifakis, Daehyun Kim 0001, Andrew Selle, Jatin Chhugani, Matthew J. Holliman, Yen-Kuang Chen
ISCA9
2007 ALP: Efficient support for all levels of parallelism for complex media applications
abstract
The real-time execution of contemporary complex media applications requires energy-efficient processing capabilities beyond those of current superscalar processors. We observe that the complexity of contemporary media applications requires support for multiple forms of parallelism, including ILP, TLP, and various forms of DLP, such as subword SIMD, short vectors, and streams. Based on our observations, we propose an architecture, called ALP, that efficiently integrates all of these forms of parallelism with evolutionary changes to the programming model and hardware. The novel part of ALP is a DLP technique called SIMD vectors and streams (SVectors/SStreams) , which is integrated within a conventional superscalar-based CMP/SMT architecture with subword SIMD. This technique lies between subword SIMD and vectors, providing significant benefits over the former at a lower cost than the latter. Our evaluations show that each form of parallelism supported by ALP is important. Specifically, SVectors/SStreams are effective, compared to a system with the other enhancements in ALP. They give speedups of 1.1 to 3.4X and energy-delay product improvements of 1.1 to 5.1X for applications with DLP.
Ruchira Sasanka, Man-Lap Li, Sarita V. Adve, Yen-Kuang Chen, Eric Debes
ACM Trans. Archit. Code Optim.4
2007 Cache-conscious frequent pattern mining on modern and emerging processors
Amol Ghoting, Gregory Buehrer, Srinivasan Parthasarathy 0001, Daehyun Kim 0001, Anthony D. Nguyen, Yen-Kuang Chen, Pradeep Dubey
VLDB J.6
2006 Adaptive Parallel Graph Mining for CMP Architectures
abstract
Mining graph data is an increasingly popular challenge, which has practical applications in many areas, including molecular substructure discovery, Web link analysis, fraud detection, and social network analysis. The problem statement is to enumerate all subgraphs occurring in at least sigma graphs of a database, where sigma is a user specified parameter. Chip multiprocessors (CMPs) provide true parallel processing, and are expected to become the de facto standard for commodity computing. In this work, building on the state-of-the-art, we propose an efficient approach to parallelize such algorithms for CMPs. We show that an algorithm which adapts its behavior based on the runtime state of the system can improve system utilization and lower execution times. Most notably, we incorporate dynamic state management to allow memory consumption to vary based on availability. We evaluate our techniques on current day shared memory systems (SMPs) and expect similar performance for CMPs. We demonstrate excellent speedup, 27-fold on 32 processors for several real world datasets. Additionally, we show our dynamic techniques afford this scalability while consuming up to 35% less memory than static techniques.
Gregory Buehrer, Srinivasan Parthasarathy 0001, Yen-Kuang Chen
ICDM3
2006 Coterminous locality and coterminous group data prefetching on chip-multiprocessors
abstract
Due to shared cache contentions and interconnect delays, data prefetching is more critical in alleviating penalties from increasing memory latencies and demands on chip-multiprocessors (CMPs). Through deep analysis of SPEC2000 applications, we find that a part of the nearby data memory references often exhibit highly-repeated patterns with long, but equal block reuse distance. These references can form a coterminous group (CG). Coterminous locality is introduced as that when a member in a CG is referenced, the remaining members will likely be referenced in the near future. Based on the coterminous locality behavior, we implement a novel CG data prefetcher on CMPs. Performance evaluations show that the proposed prefetcher can accurately cover up to 40-50% of the total misses, and result in 50-60% of potential performance improvement for several selected workload mixes
Xudong Shi 0003, Jih-Kwon Peir, Lu Peng 0001, Yen-Kuang Chen, Victor W. Lee, B. Liang
IPDPS5
2006 Incremental approximate matrix factorization for speeding up support vector machines
abstract
Traditional decomposition-based solutions to Support Vector Machines (SVMs) suffer from the widely-known scalability problem. For example, given a one-million training set, it takes about six days for SVMLight to run on a Pentium-4 sever with 8G-byte memory. In this paper, we propose an incremental algorithm, which performs approximate matrix-factorization operations, to speed up SVMs. Two approximate factorization schemes, Kronecker and incomplete Cholesky, are utilized in the primal-dual interior-point method (IPM) to directly solve the quadratic optimization problem in SVMs. We found out that a coarse approximate algorithm enjoys good speedup performance but may suffer from poor training accuracy. Conversely, a fine-grained approximate algorithm enjoys good training quality but may suffer from long training time. We subsequently propose an incremental training algorithm, which uses the approximate IPM solution of a coarse factorization to initialize the IPM of a fine-grained factorization. Extensive empirical studies show that our proposed incremental algorithm with approximate factorizations substantially speeds up SVM training while maintaining high training accuracy. In addition, we show that our proposed algorithm is highly parallelizable on an Intel dual-coreprocessor.
Gang Wu 0005, Edward Y. Chang, Yen-Kuang Chen, Christopher J. Hughes
KDD3
2006 Implementation of H.264 encoder and decoder on personal computers
Yen-Kuang Chen, Eric Q. Li, Xiaosong Zhou, Steven Ge
J. Vis. Commun. Image Represent.1
2005 A Characterization of Data Mining Workloads on a Modern Processor
Amol Ghoting, Gregory Buehrer, Srinivasan Parthasarathy 0001, Daehyun Kim 0001, Anthony D. Nguyen, Yen-Kuang Chen, Pradeep Dubey
DaMoN6
2005 Cache-conscious Frequent Pattern Mining on a Modern Processor
Amol Ghoting, Gregory Buehrer, Srinivasan Parthasarathy 0001, Daehyun Kim 0001, Anthony D. Nguyen, Yen-Kuang Chen, Pradeep Dubey
VLDB6
2005 A compiler for exploiting nested parallelism in OpenMP programs
Xinmin Tian, Jay P. Hoeflinger, Grant Haab, Yen-Kuang Chen, Milind Girkar, Sanjiv Shah
Parallel Comput.4
2004 Fast multi-frame motion estimation algorithm with adaptive search strategies in H.264
abstract
In the new H.264/AVC video coding standard, motion estimation takes up a significant encoding time, especially when using the straightforward full search algorithm (FS). A fast flexible multi-frame motion estimation algorithm with adaptive search strategies (FMASS) is presented. With special considerations on the multiple reference frames and block modes, several techniques, i.e., adaptive search strategies for single frame and flexible multi-frame selection, have been utilized to improve significantly the speed-up performance in H.264. Extensive simulations show that it can minimize the matching points by more than 1190 times compared with FS. In addition, the output quality of the encoded sequences loses only 0.053 dB in terms of PSNR on average. Fast speed-up performance and unnoticeable quality losses make the proposed algorithm outperform most of the other well-known algorithms proposed recently, such as ARPS3, MVFAST and UMHexagonS, of which the latter two have been already accepted by MPEG-4 and JVT respectively.
Xiang Li 0003, Eric Q. Li, Yen-Kuang Chen
ICASSP (3)3
2004 The energy efficiency of CMP vs. SMT for multimedia workloads
abstract
This paper compares the energy efficiency of chip multiprocessing (CMP) and simultaneous multithreading (SMT) on modern out-of-order processors for the increasingly important multimedia applications. Since performance is an important metric for real-time multimedia applications, we compare configurations at equal performance. We perform this comparison for a large number of performance points derived using different processor architectures and frequencies/voltages.We find that for the design space explored, for each workload, at each performance point, CMP is more energy efficient than SMT. The difference is small for two thread systems, but large (18% to 44%) for four thread systems. We also find that the best SMT and the best CMP configuration for a given performance target have different architecture and frequency/voltage. Therefore, their relative energy efficiency depends on a subtle interplay between various factors such as capacitance, voltage, IPC, frequency, and the level of clock gating, as well as workload features. We perform a detailed analysis considering these factors and develop a mathematical model to explain these results.Although CMP shows a clear energy advantage for four-thread (and higher) workloads, it comes at the cost of increased silicon area. We therefore investigate a hybrid solution where a CMP is built out of SMT cores, and find it to be an effective compromise. Finally, we find that we can reduce energy further for CMP with a straightforward application of previously proposed techniques of adaptive architectures and dynamic voltage/frequency scaling.
Ruchira Sasanka, Sarita V. Adve, Yen-Kuang Chen, Eric Debes
ICS3
2004 Towards Efficient Multi-Level Threading of H.264 Encoder on Intel Hyper-Threading Architectures
abstract
Summary form only given. Exploiting thread-level parallelism is a promising way to improve the performance of multimedia applications that are running on multithreading general-purpose processors. We describe the work in developing our threaded H.264 encoder. We parallelize the H.264 encoder using the OpenMP programming model, which allows us to leverage the advanced compiler technologies in the Intel/spl reg/ C++ compiler for Intel hyper-threading architectures. After we present our design considerations in the parallelization process, we describe two efficient methods for multilevel data partitioning, which can improve the performance of our multithreaded H.264 encoder. Furthermore, we exploit different options in the OpenMP programming. While one implementation that uses the task queuing model is slightly slower than the other implementation, it is easier to be read than the other one. The results have shown good speedups ranging from 3.74x to 4.53x over the well-optimized sequential code performance on a system of 4 Intel Xeon/spl trade/processors with hyper-threading technology.
Yen-Kuang Chen, Xinmin Tian, Steven Ge, Milind Girkar
IPDPS1
2004 Implementation of H.264 encoder on general-purpose processors with hyper-threading technology
abstract
H.264 is the emerging video coding standard, which aims at compressing high-quality video contents at low bit-rates. While its new encoding and decoding processes are similar to many previous standards, the new standard includes a number of new features and thus requires much more computation than most existing standards do. The complexity of H.264 standard poses a large amount of challenges to implementing the encoder/decoder in real-time via software on personal computers. Even after 2~3x performance improvement with media instruction on modern general-purpose processors and another 2~4x improvement from algorithmic optimization, the H.264 encoder is still too complicated to be implemented in real-time on a single processor. Based on the detailed analysis of the possibilities of parallelism in H.264 encoder, we proposed an efficient multithreading implementation of the H.264 video encoder. In order to guarantee enough concurrency of the whole system, an elaborate macroblock and inter-frame parallel scheduling scheme is presented. In addition, our macroblock-based multithreading scheme achieves almost no video quality losses in contrast to other parallelization schemes. Our results show that the multithreaded encoder can obtain another 3.96x speed-up on a four-processor system or 4.6x speed-up on a four-processor system with Hyper-Treading Technology. The techniques demonstrated in this work can be applied not only to H.264, but also to other video/image coding/decoding applications on personal computers.
Eric Q. Li, Yen-Kuang Chen
VCIP2
2003 Quality-delay-and-computation trade-off analysis of acoustic echo cancellation on general-purpose CPU
abstract
While many previous studies have examined acoustic echo cancellation (AEC) in terms of quality, computation complexity, and implementation issues on DSP processors, this work evaluates quality-delay-computation trade-off of the unconstrained frequency-domain recursive-least-square AEC algorithm on general purpose microprocessors. Specially, the trade-off among echo cancellation quality, sampling delay, and computation time on Intel Pentium 4 systems is analyzed. Our quantitative analysis shows that the effectiveness of echo cancellation does not depend on availability of CPU as long as CPU can provide sufficient computational power for online real-time processing. Today's general-purpose microprocessor-based AEC can deliver satisfactory echo cancellation quality at a computationally acceptable price (less than 5% CPU usage). On the other hand, the effectiveness depends on sampling delay. And no matter how fast a microprocessor would be, it is unlikely to guarantee both smaller sampling delay and larger echo-return-loss-enhancement (ERLE) at the same time. Finally, considering possible application of general-purpose processor-based AEC in laptop, office and meeting room environments, we analyzed the acoustic channel delay's influence on both ERLE and CPU computation, showing that the general-purpose microprocessor AEC's outstanding ability in tolerating various computing environments. Our experimental results can be used to design good configuration to meet specific quality requirements in terms of quality and sampling delay.
Justin J. Song, Yen-Kuang Chen
ICASSP (2)3
2003 Quality-delay-and-computation trade-off analysis of acoustic echo cancellation on general-purpose CPU
abstract
While many previous studies have examined acoustic echo cancellation (AEC) in terms of quality, computation complexity, and implementation issues on DSP processors, this work evaluates quality-delay-computation trade-off of unconstrained frequency-domain recursive-least-square AEC algorithm on general purpose microprocessors. Specially, trade-off among echo cancellation quality, sampling delay, and computation time on Intel Pentium 4 systems is analyzed. Our quantitative analysis shows that the effectiveness of echo cancellation does not depend on availability of CPU as long as CPU can provide sufficient computational power for online real-time processing. Today's general-purpose microprocessor-based AEC can deliver satisfactory echo cancellation quality at a computationally acceptable price (less than 5% CPU usage). On other hand, the effectiveness depends on sampling delay. And no matter how fast a microprocessor would be, it is unlikely to guarantee both smaller sampling delay and larger echo-return-loss-enhancement (ERLE) at the same time. Finally, considering possible application of general-purpose processor-based AEC in laptop, office and meeting room environments, we analyzed acoustic channel delay's influence on both ERLE and CPU computation, showing that general- purpose microprocessor AEC's outstanding ability in tolerating various computing environments. Our experimental results can be used to design good configuration to meet specific quality requirements in terms of quality and sampling delay.
Justin J. Song, Yen-Kuang Chen
ICME3
2002 Error drifting reduction in enhanced fine granularity scalability
abstract
We incorporate fading and reset mechanisms in an enhanced fine granularity scalability algorithm to reduce the drifting error at low bit rate while still maintaining 1.5dB PSNR gain at high bit rate over the current MPEG-4 fine granularity scalability. Many previous works use enhancement layers to predict enhancement layers so as to increase the compression efficiency. Drifting error occurs because the enhancement layer, the predictor, is not received as expected. Our fading mechanism linearly combines the current reconstructed base layer and previously reconstructed enhancement layer with fading factors between 0 and 1. Our reset mechanism sets the reference frame for prediction to be the base layer periodically. Our theoretical formulation and experimental results show that drifting error can be distributed more uniformly and maximum accumulated mismatch error is significantly reduced while our mechanisms are turned on. Around 1dB can be improved at low bit rate comparing to the one without any drifting reduction mechanism.
Wen-Hsiao Peng, Yen-Kuang Chen
ICIP (2)2
2002 Video applications on hyper-threading technology
abstract
This paper characterizes selected workloads of multimedia applications on current superscalar architectures, and then it characterizes the same workloads on Intel hyper-threading technology. This technology enables multiple threads to run in parallel on a processor, by interleaving instructions from different threads in the pipeline. The workloads, including video encoding, decoding, and watermark detection, are optimized for the Intel Pentium 4 processor. Even if the workloads are very well optimized for the Pentium 4 processor, most of the modules in these well-optimized workloads cannot fully utilize all the execution units available in the microprocessor, due to the inherently sequential constitution of the algorithms. Some of the modules are memory-bounded, while some are computation-bounded. Therefore, hyper-threading technology is a promising architecture feature that allows more CPU resources to be used at a given moment. Our goal is to provide a better explanation of the performance improvements that are possible in multimedia applications using hyper-threading technology. We demonstrate different task partition/scheduling schemes and discuss their trade-offs so that the reader can understand how to develop efficient applications on processors with hyper-threading technology.
Yen-Kuang Chen, Matthew J. Holliman, Eric Debes
ICME (2)1
2002 Characterizing multimedia kernels on general-purpose processors
abstract
Media applications have been driving microprocessor development for more than a decade. Future multimedia applications will have much higher computational requirements. Indeed, tomorrow's PC experiences will be even richer in audio-visual effects, easier to use, and, more importantly, computing will merge with communications. Such application characteristics will differ from current application characteristics and thus, a new detailed analysis of such kernels is required. This paper proposes a selection of key kernels we believe will be prominent across applications running on future processors. An extensive study of these applications on current general-purpose processor architectures makes it possible to determine future media and communication characteristics and to identify bottlenecks that are not well addressed in current designs.
Eric Debes, William Macy, Yen-Kuang Chen, Minerva M. Yeung
ICME (2)3
2002 Evaluating and Improving Performance of Multimedia Applications on Simultaneous Multi-Threading
abstract
This paper presents the study and results of running several core multimedia applications on a simultaneous multithreading (SMT) architecture, including some detailed analysis ranging from memory-bounded kernels to computational-bounded functions. A performance metric to evaluate effective SMT performance gain is introduced, and compared to similar metrics on symmetric multiprocessor (SMP) systems. In addition, we analyze and compare SMT versus SMP systems, and highlight the advantages in the studied applications. The results indicate that sharing the cache in SMT processors can provide better cache locality and thus better performance although sharing the cache can introduce cache conflicts and reduce the actual cache size available for each logical processor. We also propose "mutual prefetching" -a technique to schedule threads so that they prefetch data for each other in order to reduce cache miss penalty.
Yen-Kuang Chen, Eric Debes, Rainer Lienhart, Matthew J. Holliman, Minerva M. Yeung
ICPADS1
2001 Mode-adaptive fine granularity scalability
abstract
We propose a new algorithm, which utilizes the enhancement layer prediction to further improve the coding efficiency of current fine granularity scalability (FGS) defined in MPEG-4. The proposed algorithm adaptively uses (1) the previously reconstructed enhancement layer macroblock after motion compensation along with (2) current reconstructed base layer macroblock and (3) the combination of both to form the predicted macroblock for the current enhancement layer. The new algorithm is designed so that other error drifting reduction methods can be used to avoid the drifting problem due to prediction from the enhancement layer. In addition, the proposed algorithm can re-use the implemented B-frame hardware to form the enhancement layer predicted frame. Simulation results show that about 1 dB gain in PSNR can be achieved compared to the current FGS algorithm at moderate to high bit rates. Thus, the proposed algorithm is a cost efficient solution to improve the coding efficiency of fine granularity scalability.
Wen-Shiaw Peng, Yen-Kuang Chen
ICIP (2)2
1999 A fast rate-optimized motion estimation algorithm for low-bit-rate video coding
abstract
Motion estimation is known to be the main bottleneck in real-time encoding applications, and the search for an effective motion estimation algorithm (in terms of computational complexity and compression efficiency) has been a challenging problem for years. This paper describes a new block-matching algorithm that is much faster than the full search algorithm and occasionally even produces better rate-distortion curves than the full search algorithms. We observe that a piecewise continuous motion field reduces the bit rate for differentially encoded motion vectors. Our motion estimation algorithm exploits the spatial correlations of motion vectors effectively in the sense of producing better rate-distortion curves. Furthermore, we incorporate such correlations in a multiresolution framework to reduce the computational complexity. Simulation shows that this method is successful because of the homogeneous and reliable estimation of the displacement vectors. In nine out of our ten benchmark simulations, the performance of the full search algorithm and that of our subblock multiresolution method is about the same. In one out of our ten benchmark simulations, our method has improvement.
John C.-H. Ju, Yen-Kuang Chen, Sun-Yuan Kung
IEEE Trans. Circuits Syst. Video Technol.2
1998 Frame-rate up-conversion using transmitted true motion vectors
abstract
In this paper, we present a video frame-rate up-conversion scheme that uses transmitted true motion vectors for motion-compensated interpolation. In a past work, we demonstrated that a neighborhood-relaxation motion tracker can provide more accurate true motion information than a conventional minimal-residue block-matching algorithm. Although the technique to estimate the true motion vectors is a novelty in its own right, the strength of this technique can be further demonstrated through various spatio-temporal interpolation applications. In this work, we focus on the particular problem of frame-rate up-conversion. In the proposed scheme, the true motion field is derived by the encoder and transmitted by normal means (e.g., MPEG or H.263 encoding). Then, it is recovered by the decoder and is used not only for motion compensated predictions but also used to reconstruct missing data. It is shown that the use of our neighborhood-relaxation motion estimation provides a method of constructing high quality image sequences in a practical manner.
Yen-Kuang Chen, Anthony Vetro, Huifang Sun, Sun-Yuan Kung
MMSP1
1997 Rate optimization by true motion estimation
abstract
We propose a rate-optimized motion estimation based on a "true" motion tracker. We observe that the piecewise continuous motion field reduces the bit rate for differentially encoded motion vectors. Hence, a neighborhood relaxation method is proposed. In addition, in current MPEG-4 video VM, each video-object-plane (VOP) is individually coded by a block-based approach. The bit rate can be further improved by the removal of redundancy among the block motion vectors within the same VOP. Therefore, we also propose an object-and-block hybrid coding.
Yen-Kuang Chen, Sun-Yuan Kung
MMSP1
1997 On architectural styles for multimedia signal processors
abstract
After presenting several possible multimedia signal processor (MSP) architecture styles, we propose an architecture style which could provide high performance and high flexibilities, and require less external memory accesses and I/O operations. It is a hierarchical and scalable architecture style which facilitates the hardware-software co-design of MSP circuits and systems.
Sun-Yuan Kung, Yen-Kuang Chen
MMSP2
1996 Motion-based segmentation by principal singular vector (PSV) clustering method
abstract
Motion-based segmentation has attracted a lot of attention. The task of identifying independent objects is called segmentation. Motion-based segmentation has a broad video application domain. An approach based on principal singular vectors (PSVs) of the image measurement matrix was proposed for separating independent moving objects in Kung and Yun-Ting Lin (1995). After applying SVD (singular value decomposition), feature blocks with different object-based motions tend to form separate clusters on the PSV space. Therefore, a frame can be divided into regions each with consistent motion. Our approach offers several additional features: (1) a multi-candidate feature tracker is adopted. (2) Multiple frames are utilized to facilitate motion-based separation. (3) We would like to achieve not only accurate motion estimation, but also the object regions should retain some neighborhood property (to save the bits for the coding boundary). For this, a neighborhood sensitivity parameter /spl delta/ is introduced. One application of motion-based segmentation is low-bit-rate video compression. In very low bit-rate video coding, only motion vectors of finite regions and the region boundary (coded in prediction error) need to be transmitted. Yet simulations yield quite respectable compensated frames.
Sun-Yuan Kung, Yun-Ting Lin, Yen-Kuang Chen
ICASSP3
1996 A feature tracking algorithm using neighborhood relaxation with multi-candidate pre-screening
abstract
Tracking of features in video sequences has many applications. Conventionally, the minimum displaced frame difference (referred to as DFD or residue) of a block of pixels is used as the criterion for tracking in block-matching algorithms (BMA). However, such a criterion often misses the true motion vectors, due to many practical factors, e.g. affine warping, image noise, object occlusion, lighting variation, and existence of multiple minimal DFD. Our goal is to find motion vectors of the features for object-based motion tracking, in which (1) any region of an object contains a good number of blocks, whose motion vectors exhibit certain consistency; and (2) only true motion vectors for a few blocks per region are needed. Hence, we propose a new tracking method. (1) At the outset, we disqualify some of the reference blocks which are considered to be unreliable to track. (2) We adopt a multi-candidate pre-screening to provide some robustness in selecting motion candidates. (3) Assuming the true motion field is piecewise continuous, we determine the motion of a feature block by consulting all its neighboring blocks' directions. This allows for the chance that a singular and erroneous motion vector may be corrected by its surrounding motion vectors (just like median filtering). Our method is also designed for tracking more flexible affine-type motions, such as rotation, zooming, sheering, etc. Finally, the performance improvement over other existing methods is demonstrated.
Yen-Kuang Chen, Yun-Ting Lin, Sun-Yuan Kung
ICIP (2)1
1996 Object-based scene segmentation combining motion and image cues
abstract
This paper presents an object-based scene segmentation algorithm which combines the temporal information (e.g. motion) from video and image cues from individual frame. First a motion-based segmentation is decided based on the hierarchical principal component split (HPCS) algorithm for multi-moving-object motion classification. HPCS is a binary-tree-structured recursive procedure which clusters the feature blocks according to their principal component of the feature track matrix. Tracking of feature blocks from multiple frames (/spl ges/2) can be effectively processed and this results in a more accurate rigid motion classification. Experimental result shows that by using motion alone, some mostly homogeneous blocks may fit well to more than one motion classes so that ambiguity occurs. Such blocks are categorized into the so-called "undetermined" region (or U-region) for further processing. An image segmentation scheme using local pixel statistics of blocks in the U-region (U-blocks) is applied to find "valid voting regions" (VVRs). A VVR has a mostly homogeneous interior and is surrounded by a closed contour consisting of relatively high gradient points, which can offer the needed discriminating power for classifying each VVR to its belonging object class by motion voting. By combining the motion-based segmentation with the classification result of VVRs, the final object-based scene segmentation is determined. Simulation results are presented.
Yun-Ting Lin, Yen-Kuang Chen, Sun-Yuan Kung
ICIP (1)2